Medthinkvqa Eval

Evaluates vision-language models' ability to interpret multiple medical images, integrate cross-view evidence, and perform stepwise clinical reasoning for differential diagnosis. It probes visual grounding, evidence alignment, and reasoning depth beyond simple answer matching. Use when the user wants to benchmark on MedThinkVQA, or asks about evaluating this task. Reports Stepwise Reasoning Evaluation.

qhjqhj00 fc38d2a 3.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/medthinkvqa-eval commit fc38d2aca6

Frequently asked questions

npx skillmds add qhjqhj00/medthinkvqa-eval