Vision R1 Eval

Evaluates Large Vision-Language Models on their ability to detect, localize, and ground objects in images across diverse and challenging scenarios, including in-domain dense detection, out-of-domain real-world settings, and generalization to unseen categories or scenes. Use when the user wants to benchmark on MSCOCO Val2017, ODINW-13, or asks about evaluating this task. Reports mAP.

qhjqhj00 9d0b3f9 3.4 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/vision-r1-eval commit 9d0b3f9f6f

Frequently asked questions

npx skillmds add qhjqhj00/vision-r1-eval