Multimodal Grounding Eval

Evaluates a model's ability to localize text phrases and referring expressions within images by generating bounding box coordinates, and conversely to generate text descriptions from given bounding boxes. It probes spatial reasoning, text-image alignment, and zero-shot generalization across different expression types. Use when the user wants to benchmark on Flickr30k Entities, RefCOCO, RefCOCO+, RefCOCOg, or asks about evaluating this task. Reports R@1.

qhjqhj00 2c7d07d 3.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/multimodal-grounding-eval commit 2c7d07ddcb

Frequently asked questions

npx skillmds add qhjqhj00/multimodal-grounding-eval