Vista Multimodal Eval

Evaluates cross-modal vision-text alignment in Multimodal Large Language Models (MLLMs) across high-level semantic VQA, general multimodal understanding, and fine-grained visual perception/retrieval tasks. Use when the user wants to benchmark on VQAv2, OK-VQA, GQA, TextVQA, RealWorldQA, DocVQA, MMBench, SEED, AI2D, MMMU, MMStar, MME, RefCOCO, RefCOCO+, RefCOCOg, or asks about evaluating this task. Reports performance.

qhjqhj00 a185e05 3.2 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/vista-multimodal-eval commit a185e05817

Frequently asked questions

npx skillmds add qhjqhj00/vista-multimodal-eval