Multimodal Medical Stress Test Eval

This evaluation probes the robustness and genuine multimodal reasoning capabilities of large language models in clinical settings. It measures how model accuracy degrades when visual inputs are removed, answer options are perturbed, or distractors are replaced, revealing reliance on textual shortcuts and memorization rather than true visual-textual integration. Use when the user wants to benchmark on NEJM, JAMA, VQA-RAD, OmniMedVQA, or asks about evaluating this task. Reports accuracy.

qhjqhj00 f56d5db 3.3 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/multimodal-medical-stress-test-eval commit f56d5db916

Frequently asked questions

npx skillmds add qhjqhj00/multimodal-medical-stress-test-eval