Mmerealworld Eval

This benchmark evaluates multimodal large language models on high-resolution real-world image perception and complex reasoning tasks. It probes the models' ability to extract fine-grained details from large images and perform logical inference across diverse domains like autonomous driving, remote sensing, and document understanding. Use when the user wants to benchmark on MME-RealWorld, or asks about evaluating this task. Reports accuracy.

qhjqhj00 6f85530 3.0 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mmerealworld-eval commit 6f855303db

Frequently asked questions

npx skillmds add qhjqhj00/mmerealworld-eval