Medsyn Eval

Evaluates multimodal large language models on their ability to generate differential diagnoses (DDx) and select final diagnoses (FDx) for complex clinical cases. It probes cross-modal evidence calibration, testing how models weigh textual versus visual clinical evidence, and measures their sensitivity to specific evidence types. Use when the user wants to benchmark on MEDSYN, or asks about evaluating this task. Reports FDx SelectionAcc. (%).

qhjqhj00 659d6aa 3.2 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/medsyn-eval commit 659d6aae13

Frequently asked questions

npx skillmds add qhjqhj00/medsyn-eval