Omni Math Eval

Evaluates large language models on rigorous, Olympiad-level mathematical reasoning across diverse domains and difficulty levels. It probes the model's ability to perform complex logical deduction, multi-step problem solving, and handle non-standard answer formats without relying on trivial or non-mathematical content. Use when the user wants to benchmark on Omni-MATH, or asks about evaluating this task. Reports accuracy.

qhjqhj00 21725cd 2.5 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/omni-math-eval commit 21725cdba5

Frequently asked questions

npx skillmds add qhjqhj00/omni-math-eval