Big Bench Predictability Eval

Evaluates the predictability of LLM performance across diverse BIG-bench tasks using an MLP-based predictor. It probes how well model scale, task type, and in-context examples correlate with actual benchmark scores, and tests robustness under different holdout strategies. Use when the user wants to benchmark on BIG-bench, or asks about evaluating this task. Reports R².

qhjqhj00 f8fafcf 2.6 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/big-bench-predictability-eval commit f8fafcf21b

Frequently asked questions

npx skillmds add qhjqhj00/big-bench-predictability-eval