Frontiermath Eval

Evaluates advanced mathematical reasoning and experimental problem-solving. It tests whether models can iteratively write and execute Python code to verify hypotheses, refine strategies, and derive correct solutions to expert-level, unsolved math problems. Use when the user wants to benchmark on FrontierMath, or asks about evaluating this task. Reports accuracy.

qhjqhj00 3b421ac 2.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/frontiermath-eval commit 3b421aca32

Frequently asked questions

npx skillmds add qhjqhj00/frontiermath-eval