Captionqa Eval

This benchmark evaluates whether model-generated image captions retain sufficient visual information to answer domain-specific multiple-choice questions without access to the original image. It measures caption utility for downstream reasoning by testing if a text-only QA model can reliably select correct answers or explicitly acknowledge missing information when prompted only with the caption. Use when the user wants to benchmark on CaptionQA, or asks about evaluating this task. Reports Caption Utility Score (avg s).

qhjqhj00 dafb596 4.4 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/captionqa-eval commit dafb596aae

Frequently asked questions

npx skillmds add qhjqhj00/captionqa-eval