Vqa Gen Eval

This benchmark evaluates a model's ability to generalize in Visual Question Answering under coordinated visual and textual distribution shifts. It probes robustness to image corruptions, style transfers, and linguistic variations by measuring in-domain and cross-domain accuracy. Use when the user wants to benchmark on VQA-GEN, or asks about evaluating this task. Reports Accuracy.

qhjqhj00 8a1ead1 4.1 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/vqa-gen-eval commit 8a1ead14fa

Frequently asked questions

npx skillmds add qhjqhj00/vqa-gen-eval