Multiple Choice Vqa Eval

This evaluation probes the true multimodal reasoning capability of vision-language models on multiple-choice question answering tasks. It specifically measures whether models rely on actual question understanding or exploit visual relevance imbalances between correct answers and distractors. Performance is assessed under both standard (vision, question, options) and question-omitted (vision, options) settings to detect easy-option bias. Use when the user wants to benchmark on NExT-QA, MMStar, or asks about evaluating this task. Reports accuracy.

qhjqhj00 7a96d0a 3.2 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/multiple-choice-vqa-eval commit 7a96d0a897

Frequently asked questions

npx skillmds add qhjqhj00/multiple-choice-vqa-eval