Mmfinereason Eval

Evaluates multimodal reasoning capabilities across STEM, puzzles, general VQA, and document understanding domains. Probes how well vision-language models perform on complex visual reasoning tasks under strict greedy decoding and high-resolution inference settings. Use when the user wants to benchmark on MMMU_val, MathVista_mini, MathVision_test, MathVerse_mini, Dynamath, LogicVista, VisuLogic, ScienceQA, RealWorldQA, MMBench-EN, MMStar_test, AI2D_test, CharXiv_reas, CharXiv_desc, or asks about evaluating this task. Reports accuracy.

qhjqhj00 c01bdef 3.3 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mmfinereason-eval commit c01bdefe38

Frequently asked questions

npx skillmds add qhjqhj00/mmfinereason-eval