Mmvp Eval

Evaluates the visual reasoning and grounding capabilities of multimodal large language models (MLLMs) on basic visual patterns such as orientation, counting, viewpoint, and feature presence. It specifically probes whether models fail due to limitations in their visual encoders (e.g., CLIP) rather than language model hallucinations. Use when the user wants to benchmark on MMVP, or asks about evaluating this task. Reports accuracy.

qhjqhj00 cc5c950 2.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mmvp-eval commit cc5c9506c4

Frequently asked questions

npx skillmds add qhjqhj00/mmvp-eval