Image Captioning Retrieval Eval

Evaluates vision-language models on image captioning and image-text retrieval tasks to measure zero-shot and fine-tuned generalization on long-tail visual concepts and out-of-domain data. Use when the user wants to benchmark on nocaps, COCO Captions, Flickr30K, LocNar Flickr30K, or asks about evaluating this task. Reports CIDEr.

qhjqhj00 89c642a 3.4 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/image-captioning-retrieval-eval commit 89c642aa96

Frequently asked questions

npx skillmds add qhjqhj00/image-captioning-retrieval-eval