Mme Realworld Vbench Eval

This evaluation probes a model's ability to perform high-resolution visual reasoning and fine-grained grounding on complex, real-world images. It specifically tests whether the model can accurately localize relevant visual regions and correctly answer multiple-choice questions without explicit grounding supervision. Use when the user wants to benchmark on MME-Realworld, V* Bench, or asks about evaluating this task. Reports accuracy.

qhjqhj00 2c780ae 2.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mme-realworld-vbench-eval commit 2c780aec8c

Frequently asked questions

npx skillmds add qhjqhj00/mme-realworld-vbench-eval