Mmmu Pro Eval

This benchmark evaluates multimodal models' ability to perform robust, multi-discipline reasoning by forcing them to integrate visual and textual information without relying on shortcuts. It specifically probes resistance to guessing strategies through augmented multiple-choice options and tests true vision-text integration by embedding questions directly within images. Use when the user wants to benchmark on MMMU-Pro, or asks about evaluating this task. Reports accuracy.

qhjqhj00 b618651 3.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mmmu-pro-eval commit b618651947

Frequently asked questions

npx skillmds add qhjqhj00/mmmu-pro-eval