Autovivqa Eval

This benchmark evaluates Vietnamese vision-language models on visual question answering, probing their ability to ground textual queries in images and generate semantically accurate, linguistically fluent responses. It specifically tests reasoning complexity across five levels (recognition, relational, compositional, causal, text-in-image) and measures both exact-match accuracy and generation quality. Use when the user wants to benchmark on AutoViVQA, or asks about evaluating this task. Reports Accuracy.

qhjqhj00 ee095f9 4.1 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/autovivqa-eval commit ee095f9124

Frequently asked questions

npx skillmds add qhjqhj00/autovivqa-eval