Mudaif Vl Eval

Evaluates a decoder-only vision-language model's ability to perform visual question answering, image captioning, and multimodal reasoning. It measures cross-modal alignment, computational efficiency, and robustness to input variations like resolution and noise. Use when the user wants to benchmark on VQA-v2, GQA, VizWiz, SEED, MM-Vet, or asks about evaluating this task. Reports accuracy.

qhjqhj00 e1875da 2.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mudaif-vl-eval commit e1875da788

Frequently asked questions

npx skillmds add qhjqhj00/mudaif-vl-eval