Mm Neuroonco Eval

This benchmark evaluates the multimodal diagnostic reasoning capabilities of large vision-language models on brain tumor MRI scans. It probes whether models can integrate subtle visual cues with structured anatomical knowledge to produce accurate diagnoses, while also measuring their ability to recognize uncertainty through explicit rejection options. Use when the user wants to benchmark on MM-NeuroOnco-Bench, or asks about evaluating this task. Reports Accuracy.

qhjqhj00 4e6ad6c 3.3 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mm-neuroonco-eval commit 4e6ad6c61a

Frequently asked questions

npx skillmds add qhjqhj00/mm-neuroonco-eval