Mathvista Eval

This benchmark evaluates the mathematical reasoning capabilities of foundation models (LLMs and LMMs) when processing visual contexts. It probes abilities such as figure interpretation, algebraic and geometric reasoning, and the integration of multimodal inputs like images, OCR text, and captions into mathematical problem-solving. Use when the user wants to benchmark on MATHVISTA, or asks about evaluating this task. Reports accuracy.

qhjqhj00 469df88 3.2 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mathvista-eval commit 469df881fa

Frequently asked questions

npx skillmds add qhjqhj00/mathvista-eval