# Chestagentbench Eval

> chestagentbench-eval

- Skill: `qhjqhj00/chestagentbench-eval` (Agent Skill)
- Install (CLI): `npx skillmds@latest add qhjqhj00/chestagentbench-eval`
- Raw SKILL.md: https://api.skillmd.com/api/skills/qhjqhj00/chestagentbench-eval/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: qhjqhj00 (https://skillmd.com/u/qhjqhj00)
- Updated: 2026-09-21
- Page: https://skillmd.com/skills/qhjqhj00/chestagentbench-eval

---


# chestagentbench-eval

> MedRAX: Medical Reasoning Agent for Chest X-ray — Fallahpour et al. (2025) (arXiv:2502.02673, 2025)

## What this evaluates

Evaluates an AI agent's ability to perform multi-step medical reasoning and tool orchestration for chest X-ray interpretation. It probes capabilities across seven clinically relevant categories: detection, classification, localization, comparison, relationship, diagnosis, and characterization.

## Datasets

- **ChestAgentBench** — total 2500; splits: test (2500); repo https://github.com/bowang-lab/MedRAX

## Metrics

- `Accuracy (%)` **(primary)** — range: percent
  - Calculated as the percentage of correctly answered questions out of the total 2,500 items. A response is considered correct if the extracted letter choice matches the ground truth label.

## Input / output format

**Input**: A chest X-ray image paired with a multiple-choice question containing six options (A-F) that requires multi-step clinical reasoning.

**Output**: A text response containing a single letter choice (A-F). The evaluation pipeline uses regex to extract the choice; invalid, unclear, or timeout responses are retried up to three times before being marked incorrect.

## Scoring recipe

```python
def compute_accuracy(predictions, gold_labels):
    correct = 0
    for pred, gold in zip(predictions, gold_labels):
        extracted = extract_letter_via_regex(pred)
        if extracted is None:
            extracted = retry_extraction(pred, max_retries=3)
        if extracted == gold:
            correct += 1
    return (correct / len(gold_labels)) * 100
```

## Common pitfalls

- Regex-based extraction often fails on verbose chain-of-thought outputs, requiring a retry mechanism that can artificially inflate or deflate accuracy if not standardized.
- The six-choice format (16.7% random baseline) makes models prone to guessing on fine-grained distinctions like localization or characterization without explicit tool grounding.
- Timeouts and API errors during tool execution are counted as incorrect, which heavily penalizes latency-sensitive agents compared to direct VLMs.

## Evidence (verbatim from paper)

> We evaluate models on four complementary benchmarks: (1) ChestAgentBench, our proposed benchmark described in Section [4], which assesses comprehensive CXR reasoning through 2,500 six-choice questions across seven categories: detection, classification, localization, comparison, relationship, characterization, and diagnosis. Model performance is measured by accuracy across all questions. We process model responses using regex to extract letter choices. For unclear responses, errors, or timeouts, we retry up to three times. Responses that remain invalid or do not choose a single choice are marked incorrect.

## Citation

```bibtex
@misc{fallahpour2025medrax,
  title={MedRAX: Medical Reasoning Agent for Chest X-ray},
  author={Fallahpour et al. (2025)},
  year={2025},
  note={arXiv:2502.02673}
}
```

- arXiv: 2502.02673

