Booststep Eval

This evaluation probes the mathematical reasoning capability of large language models, specifically focusing on single-step reasoning and the effectiveness of step-aligned in-context learning. It measures how well models can solve challenging math problems across text and multi-modal domains when provided with fine-grained, step-level examples. Use when the user wants to benchmark on MATH500, AQuA, OlympiadBench-TO, MATHBench, AMC-10, AMC-12, MathVision, MathVerse, AIME, or asks about evaluating this task. Reports accuracy.

qhjqhj00 4444894 3.5 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/booststep-eval commit 4444894e06

Frequently asked questions

npx skillmds add qhjqhj00/booststep-eval