Math Reasoning Eval

Evaluates the mathematical reasoning capabilities of language models across multiple challenging benchmarks. It measures whether models can correctly solve math problems and follow structured reasoning processes aligned with a teacher model's trace. Use when the user wants to benchmark on MATH-500, MINERVA, OlympiadBench, LiveMathBench, KSAT2025, AIME 2024, AIME 2025, or asks about evaluating this task. Reports Pass@1.

qhjqhj00 a102f8d 2.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/math-reasoning-eval commit a102f8dc97

Frequently asked questions

npx skillmds add qhjqhj00/math-reasoning-eval