Simba Benchmark Analysis Eval

Evaluates a framework for analyzing language model performance matrices by identifying dataset-model correlations, discovering minimal representative dataset subsets, and predicting held-out model performance while preserving model rankings. Use when the user wants to benchmark on HELM, MMLU, BigBenchLite, or asks about evaluating this task. Reports coverage ($\eta$).

qhjqhj00 02e1eb5 3.4 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/simba-benchmark-analysis-eval commit 02e1eb5fa2

Frequently asked questions

npx skillmds add qhjqhj00/simba-benchmark-analysis-eval