Medrcube Eval

Probes multimodal large language models' fine-grained capabilities in medical imaging across anatomical regions, imaging modalities, and cognitive hierarchies. It assesses reasoning reliability, shortcut behavior, and foundational perceptual skills to reveal how models handle clinical VQA beyond aggregate performance. Use when the user wants to benchmark on MedRCube, or asks about evaluating this task. Reports MedRCube Score.

qhjqhj00 b9a525c 3.2 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/medrcube-eval commit b9a525cde3

Frequently asked questions

npx skillmds add qhjqhj00/medrcube-eval