Emma Multimodal Reasoning Eval

This benchmark evaluates multimodal large language models on their ability to perform integrated visual-textual reasoning across mathematics, physics, chemistry, and coding. It probes capabilities such as fine-grained spatial simulation, multi-hop visual inference, and cross-modal problem solving under both direct and chain-of-thought prompting conditions. Use when the user wants to benchmark on EMMA-mini, or asks about evaluating this task. Reports accuracy.

qhjqhj00 9f96b65 3.4 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/emma-multimodal-reasoning-eval commit 9f96b65e0d

Frequently asked questions

npx skillmds add qhjqhj00/emma-multimodal-reasoning-eval