Mmstar Eval

Evaluates Large Vision-Language Models (LVLMs) on six core capabilities (coarse perception, fine-grained perception, instance reasoning, logical reasoning, science & technology, and mathematics) using a human-curated benchmark designed to enforce strict visual dependency and minimize data leakage. Use when the user wants to benchmark on MMStar, or asks about evaluating this task. Reports accuracy.

qhjqhj00 f719e28 3.3 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mmstar-eval commit f719e286eb

Frequently asked questions

npx skillmds add qhjqhj00/mmstar-eval