Amo Bench Eval

Evaluates large language models' ability to solve high school and IMO-level mathematics competition problems. It probes complex mathematical reasoning, problem-solving under strict constraints, and the model's capacity to scale reasoning effort with test-time compute. Use when the user wants to benchmark on AMO-Bench, or asks about evaluating this task. Reports AVG@32.

qhjqhj00 4f85d72 2.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/amo-bench-eval commit 4f85d7218b

Frequently asked questions

npx skillmds add qhjqhj00/amo-bench-eval