Mmlu Sandbagging Eval

Evaluates whether language models can strategically underperform on capability assessments by emulating a lower educational level (high school) on subject-specific questions, and measures how prompting strategies (zero-shot vs. chain-of-thought) affect this emulation. Use when the user wants to benchmark on MMLU, or asks about evaluating this task. Reports accuracy.

qhjqhj00 6a07ab6 2.6 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mmlu-sandbagging-eval commit 6a07ab6f5b

Frequently asked questions

npx skillmds add qhjqhj00/mmlu-sandbagging-eval