Pixel Reasoner Eval

Evaluates multimodal models' ability to perform fine-grained visual reasoning, object counting, temporal video understanding, and complex infographic parsing. It specifically probes whether models can effectively leverage pixel-space operations (e.g., zooming, frame selection) rather than defaulting to text-only reasoning pathways. Use when the user wants to benchmark on V* (V-Star), TallyQA, MVBench, InfographicVQA, or asks about evaluating this task. Reports Acc.

qhjqhj00 93743c8 3.1 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/pixel-reasoner-eval commit 93743c81e4

Frequently asked questions

npx skillmds add qhjqhj00/pixel-reasoner-eval