Spice Reasoning Eval

This evaluation probes a model's ability to solve challenging mathematical and general reasoning tasks, both from standard benchmarks and document-grounded self-play generated questions. It measures how well the model can extract information, perform multi-step logical deduction, and produce verifiable answers across diverse academic and competition-level datasets. Use when the user wants to benchmark on MATH-500, OlympiadBench, Minerva Math, GSM8K, AMC, AIME'24, AIME'25, SuperGPQA, GPQA-Diamond, MMLU-Pro, BBEH, or asks about evaluating this task. Reports pass rate.

qhjqhj00 63459bb 3.4 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/spice-reasoning-eval commit 63459bbfeb

Frequently asked questions

npx skillmds add qhjqhj00/spice-reasoning-eval