Mediconfusion Eval

Probes the visual reasoning reliability and robustness of multimodal medical foundation models by presenting pairs of visually distinct but semantically confused medical images. It measures whether models can correctly answer questions about each image individually and consistently across the pair, revealing shortcut learning and hallucination tendencies. Use when the user wants to benchmark on MediConfusion, or asks about evaluating this task. Reports Set accuracy.

qhjqhj00 af2189c 4.4 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mediconfusion-eval commit af2189c628

Frequently asked questions

npx skillmds add qhjqhj00/mediconfusion-eval