Mm Upt Eval

Evaluates the multi-modal mathematical reasoning capabilities of MLLMs on diverse visual math problems including geometry, charts, and tables. It tests the model's ability to solve multiple-choice and fill-in-the-blank questions using both human-created and synthetically generated unlabeled data. Use when the user wants to benchmark on MathVision, MathVerse, MathVista, We-Math, or asks about evaluating this task. Reports accuracy.

qhjqhj00 c89e198 2.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mm-upt-eval commit c89e198d7e

Frequently asked questions

npx skillmds add qhjqhj00/mm-upt-eval