Multimodal Reasoning Eval

This evaluation probes the ability of multimodal large language models to perform complex reasoning across diverse domains (mathematics, science, diagram comprehension, and creative tasks) by requiring them to explicitly ground their reasoning in visual and textual evidence before producing a final answer. Use when the user wants to benchmark on MMMU, MathVista, AI2D, EMMA, Creation-MMBench, Creation-MMBench-TO, or asks about evaluating this task. Reports accuracy.

qhjqhj00 1001252 3.0 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/multimodal-reasoning-eval commit 1001252cd2

Frequently asked questions

npx skillmds add qhjqhj00/multimodal-reasoning-eval