Deepmath 103k Eval

Evaluates mathematical reasoning capabilities on a curated, decontaminated dataset of challenging problems, measuring performance across standardized math competitions and academic benchmarks. Use when the user wants to benchmark on DeepMath-103K, or asks about evaluating this task. Reports accuracy.

qhjqhj00 e547122 2.1 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/deepmath-103k-eval commit e547122896

Frequently asked questions

npx skillmds add qhjqhj00/deepmath-103k-eval