Tinybenchmarks Sampling Eval

Evaluates the efficiency and accuracy of LLM benchmarking by testing how well a small, strategically selected subset of examples predicts overall model performance on standard evaluation scenarios. Use when the user wants to benchmark on HELM, MMLU, AlpacaEval 2.0, Open LLM Leaderboard, or asks about evaluating this task. Reports estimation error.

qhjqhj00 44da8b6 2.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/tinybenchmarks-sampling-eval commit 44da8b6dba

Frequently asked questions

npx skillmds add qhjqhj00/tinybenchmarks-sampling-eval