Neural Medbench Eval

Neural-MedBench probes the clinical reasoning and multimodal synthesis capabilities of vision-language models in neurology diagnostics. It specifically tests whether models can move beyond superficial classification to perform uncertainty resolution, generate clinically justified rationales, and maintain logical coherence when interpreting patient histories and medical imaging. Use when the user wants to benchmark on Neural-MedBench, or asks about evaluating this task. Reports Diagnostic Accuracy (pass@1).

qhjqhj00 46f54aa 4.2 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/neural-medbench-eval commit 46f54aa53a

Frequently asked questions

npx skillmds add qhjqhj00/neural-medbench-eval