Kvasir Vqa X1 Eval

Evaluates multimodal vision-language models on gastrointestinal endoscopy image understanding and clinical question answering. It probes factual recall, multi-step clinical reasoning across varying complexity levels, and robustness to realistic visual perturbations like motion blur and color shifts. Use when the user wants to benchmark on Kvasir-VQA-x1, or asks about evaluating this task. Reports BERT-F1.

qhjqhj00 5755621 4.0 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/kvasir-vqa-x1-eval commit 5755621ab8

Frequently asked questions

npx skillmds add qhjqhj00/kvasir-vqa-x1-eval