Mmlu Bbh Ifeval Eval

Evaluates instruction-tuned language models on factual knowledge, complex reasoning, and instruction-following capabilities. The protocol measures how different data selection methods and model sizes impact performance under strict compute budgets. Use when the user wants to benchmark on MMLU, BBH, IFEval, or asks about evaluating this task. Reports 5-shot accuracy, 3-shot exact match score, 0-shot accuracy.

qhjqhj00 4318aff 2.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mmlu-bbh-ifeval-eval commit 4318aff2b9

Frequently asked questions

npx skillmds add qhjqhj00/mmlu-bbh-ifeval-eval