Math Reward Eval

Evaluates multimodal and language-only models on mathematical reasoning, self-judgment/reward accuracy, and general multimodal capabilities. It measures how well a model can solve complex problems, verify its own answers, and generalize across diverse domains without external reward models. Use when the user wants to benchmark on MathVista, GSM8k, RewardBench2, VL-RewardBench, MMBench, MMStar, or asks about evaluating this task. Reports accuracy.

qhjqhj00 01e70fe 2.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/math-reward-eval commit 01e70fe989

Frequently asked questions

npx skillmds add qhjqhj00/math-reward-eval