Benchmark Diversity Stability Eval

Evaluates the inherent trade-off between diversity (agreement of model rankings across tasks) and stability (sensitivity of final rankings to label noise) in multi-task machine learning benchmarks. It quantifies how much a benchmark's leaderboard ranking changes when trivial label noise is injected, and how diverse the rankings are across its constituent tasks. Use when the user wants to benchmark on GLUE, SuperGLUE, MTEB, BigBenchHard, MMLU, OpenLLM, VTAB, ImageNet, or asks about evaluating this task. Reports Kendall's τ.

qhjqhj00 e1e8d63 4.2 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/benchmark-diversity-stability-eval commit e1e8d6357c

Frequently asked questions

npx skillmds add qhjqhj00/benchmark-diversity-stability-eval