brats-mri-vqa-eval
Performance of GPT-5 in Brain Tumor MRI Reasoning — Safari et al. (2025) (arXiv:2508.10865, 2025)
What this evaluates
Evaluates multi-modal medical reasoning and visual question answering capabilities on brain tumor MRI scans. It probes the model's ability to parse clinical features and answer structured questions across three distinct tumor subtypes: metastases, glioblastoma, and meningioma.
Datasets
- BraTS (MET, GLI, MEN cohorts) — total ?; splits: test (-1)
Metrics
accuracy(primary) — range: percent- Percentage of correct answers per cohort, calculated as (number of correct predictions / total number of questions) * 100. A macro-average is computed as the unweighted mean of accuracy across the three cohorts (MET, GLI, MEN).
Input / output format
Input: MRI brain tumor images (triplanar mosaic imaging) paired with structured visual questions.
Output: Textual response to the visual question.
Scoring recipe
def compute_accuracy(predictions, gold):
correct = sum(1 for p, g in zip(predictions, gold) if p.strip().lower() == g.strip().lower())
return (correct / len(gold)) * 100
def compute_macro_average(accuracies):
return sum(accuracies) / len(accuracies)
Common pitfalls
- The macro-average is unweighted, so cohort size imbalances do not affect the final score.
- Accuracy is calculated at the cohort level first, then averaged, rather than globally across all questions combined.
- Zero-shot chain-of-thought prompting was used, which may influence scores compared to direct answering.
Evidence (verbatim from paper)
Table 1: Accuracy (%) across the three BraTS tumor cohorts including brain metastases (MET), glioblastoma (GLI), and meningioma (MEN) and the unweighted macro-average over cohorts.
Citation
@misc{safari2025performance,
title={Performance of GPT-5 in Brain Tumor MRI Reasoning},
author={Safari et al. (2025)},
year={2025},
note={arXiv:2508.10865}
}
- arXiv: 2508.10865