Mathreal Eval

Evaluates the ability of multimodal large language models to perform K-12 mathematical reasoning on real-world, mobile-captured images. It probes robustness to visual degradation (blur, rotation, handwritten annotations) and perspective variations, measuring how well models extract text and figures to solve math problems under imperfect conditions. Use when the user wants to benchmark on MathReal, or asks about evaluating this task. Reports Loose Accuracy (Acc).

qhjqhj00 50cef25 3.4 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mathreal-eval commit 50cef25615

Frequently asked questions

npx skillmds add qhjqhj00/mathreal-eval