# Brats Mri Vqa Eval

> brats-mri-vqa-eval

- Skill: `qhjqhj00/brats-mri-vqa-eval` (Agent Skill)
- Install (CLI): `npx skillmds@latest add qhjqhj00/brats-mri-vqa-eval`
- Raw SKILL.md: https://api.skillmd.com/api/skills/qhjqhj00/brats-mri-vqa-eval/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: qhjqhj00 (https://skillmd.com/u/qhjqhj00)
- Updated: 2026-09-21
- Page: https://skillmd.com/skills/qhjqhj00/brats-mri-vqa-eval

---


# brats-mri-vqa-eval

> Performance of GPT-5 in Brain Tumor MRI Reasoning — Safari et al. (2025) (arXiv:2508.10865, 2025)

## What this evaluates

Evaluates multi-modal medical reasoning and visual question answering capabilities on brain tumor MRI scans. It probes the model's ability to parse clinical features and answer structured questions across three distinct tumor subtypes: metastases, glioblastoma, and meningioma.

## Datasets

- **BraTS (MET, GLI, MEN cohorts)** — total ?; splits: test (-1)

## Metrics

- `accuracy` **(primary)** — range: percent
  - Percentage of correct answers per cohort, calculated as (number of correct predictions / total number of questions) * 100. A macro-average is computed as the unweighted mean of accuracy across the three cohorts (MET, GLI, MEN).

## Input / output format

**Input**: MRI brain tumor images (triplanar mosaic imaging) paired with structured visual questions.

**Output**: Textual response to the visual question.

## Scoring recipe

```python
def compute_accuracy(predictions, gold):
    correct = sum(1 for p, g in zip(predictions, gold) if p.strip().lower() == g.strip().lower())
    return (correct / len(gold)) * 100

def compute_macro_average(accuracies):
    return sum(accuracies) / len(accuracies)
```

## Common pitfalls

- The macro-average is unweighted, so cohort size imbalances do not affect the final score.
- Accuracy is calculated at the cohort level first, then averaged, rather than globally across all questions combined.
- Zero-shot chain-of-thought prompting was used, which may influence scores compared to direct answering.

## Evidence (verbatim from paper)

> Table 1: Accuracy (%) across the three BraTS tumor cohorts including brain metastases (MET), glioblastoma (GLI), and meningioma (MEN) and the unweighted macro-average over cohorts.

## Citation

```bibtex
@misc{safari2025performance,
  title={Performance of GPT-5 in Brain Tumor MRI Reasoning},
  author={Safari et al. (2025)},
  year={2025},
  note={arXiv:2508.10865}
}
```

- arXiv: 2508.10865

