Vqa Captioning Eval

Evaluates multimodal language models on visual question answering and image captioning tasks, probing their zero-shot and few-shot in-context learning capabilities with interleaved image-text inputs. Use when the user wants to benchmark on OKVQA, TextVQA, COCO, Flickr30k, VQAv2, VizWiz, or asks about evaluating this task. Reports accuracy.

qhjqhj00 2d5e112 3.0 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/vqa-captioning-eval commit 2d5e11237c

Frequently asked questions

npx skillmds add qhjqhj00/vqa-captioning-eval