Catwalk Default Metrics

Evaluates language models across multiple dataset categories using standardized metrics attached to datasets rather than model implementations. Probes classification/entailment accuracy, question-answering fidelity, and language modeling fluency/probability calibration. Use when the user has predictions and gold and needs to compute Accuracy & relative improvement over random baseline, SQuAD metric.

qhjqhj00 f431918 3.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/catwalk-default-metrics commit f431918faf

Frequently asked questions

npx skillmds add qhjqhj00/catwalk-default-metrics