Math General Reasoning Eval

Evaluates large language models on mathematical and general reasoning capabilities using a standardized suite of benchmarks. It measures the model's problem-solving accuracy under self-play training conditions, tracking sustained performance gains across multiple evolution iterations. Use when the user wants to benchmark on AMC, Minerva, MATH, GSM8K, Olympiad, AIME25, AIME24, SuperGPQA, MMLU-Pro, BBEH, or asks about evaluating this task. Reports pass@1 accuracy.

qhjqhj00 666d941 3.4 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/math-general-reasoning-eval commit 666d9412d3

Frequently asked questions

npx skillmds add qhjqhj00/math-general-reasoning-eval