Capbencher Eval

Evaluates LLMs on standard benchmarks modified with randomized answers to measure performance tracking and detect data contamination via a Bayes accuracy ceiling. The protocol compares model accuracy against a predefined theoretical maximum (Bayes accuracy) to identify overfitting or memorization. It also assesses robustness to reverse-engineering attacks and cross-lingual contamination. Use when the user wants to benchmark on GSM8K, ARC-Challenge, GPQA, MathQA, MMLU, HLE-MC, MMLU-ProX, BoolQ, GPQA (diamond), MMLU-Pro, MATH-500, HumanEval, or asks about evaluating this task. Reports accuracy.

qhjqhj00 5b1ce22 4.0 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/capbencher-eval commit 5b1ce22b3e

Frequently asked questions

npx skillmds add qhjqhj00/capbencher-eval