Bbh Mmlu Predictability Eval

This protocol evaluates how well aggregate and per-task benchmark performance can be predicted from scaled compute using scaling law fits. It probes the monotonicity and predictability of LLM capabilities across compute scaling, distinguishing between stable scaling trends and emergent or non-monotonic behaviors. Use when the user wants to benchmark on BIG-Bench Hard (BBH), MMLU, or asks about evaluating this task. Reports mean absolute error.

qhjqhj00 6bd3d00 2.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/bbh-mmlu-predictability-eval commit 6bd3d001c1

Frequently asked questions

npx skillmds add qhjqhj00/bbh-mmlu-predictability-eval