Math Benchmarks Eval

Evaluates the mathematical reasoning and problem-solving capabilities of language models across a spectrum of difficulties, from grade-school arithmetic to advanced competition-level mathematics. Use when the user wants to benchmark on GSM8K, MATH, AMC 2023, AIME 2024, Omni-MATH, or asks about evaluating this task. Reports accuracy.

qhjqhj00 cdb8f48 3.2 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/math-benchmarks-eval commit cdb8f4860b

Frequently asked questions

npx skillmds add qhjqhj00/math-benchmarks-eval