Padt Unified Vision Eval

Evaluates a multimodal large language model's ability to perform visual grounding, segmentation, open-vocabulary detection, and referring image captioning by predicting structured visual outputs directly from interleaved visual reference tokens and text. Use when the user wants to benchmark on RefCOCO/+/g, COCO 2017, RIC, or asks about evaluating this task. Reports IoU@0.5 accuracy.

qhjqhj00 48207af 4.2 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/padt-unified-vision-eval commit 48207aff04

Frequently asked questions

npx skillmds add qhjqhj00/padt-unified-vision-eval