Detailed Localized Captioning Eval

Evaluates a model's ability to generate detailed, region-specific descriptions for images and videos, ranging from keywords to multi-sentence captions. It probes fine-grained visual grounding, attribute recognition, and hallucination resistance by comparing generated text against reference captions or using attribute-level positive/negative judgments. Use when the user wants to benchmark on DLC-Bench, LVIS, PACO, Flickr30k Entities, Ref-L4, HC-STVG, VideoRefer-Bench-D, or asks about evaluating this task. Reports positive accuracy.

qhjqhj00 8e1b743 5.0 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/detailed-localized-captioning-eval commit 8e1b743b87

Frequently asked questions

npx skillmds add qhjqhj00/detailed-localized-captioning-eval