Disco Eval

This protocol evaluates how well a metamodel can predict the benchmark performance of unseen models using a highly condensed subset of test samples. It probes the trade-off between evaluation cost reduction and the fidelity of accuracy estimation and model ranking preservation across language and vision benchmarks. Use when the user wants to benchmark on MMLU, HellaSwag, Winogrande, ARC, ImageNet-1k, or asks about evaluating this task. Reports MAE, Spearman rank correlation.

qhjqhj00 cca6849 3.2 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/disco-eval commit cca684953a

Frequently asked questions

npx skillmds add qhjqhj00/disco-eval