E3vqa Eval

E3VQA evaluates a model's ability to perform multi-view visual question answering using synchronized egocentric and exocentric image pairs. It specifically probes whether models can identify relevant regions across views, filter redundant information, and integrate complementary visual cues to answer multiple-choice questions. Use when the user wants to benchmark on E3VQA, or asks about evaluating this task. Reports accuracy.

qhjqhj00 6d528d6 2.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/e3vqa-eval commit 6d528d6e17

Frequently asked questions

npx skillmds add qhjqhj00/e3vqa-eval