# Chexgenbench Eval

> chexgenbench-eval

- Skill: `qhjqhj00/chexgenbench-eval` (Agent Skill)
- Install (CLI): `npx skillmds@latest add qhjqhj00/chexgenbench-eval`
- Raw SKILL.md: https://api.skillmd.com/api/skills/qhjqhj00/chexgenbench-eval/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: qhjqhj00 (https://skillmd.com/u/qhjqhj00)
- Updated: 2026-09-21
- Page: https://skillmd.com/skills/qhjqhj00/chexgenbench-eval

---


# chexgenbench-eval

> CheXGenBench: A Unified Benchmark For Fidelity, Privacy and Utility of Synthetic Chest Radiographs — Raman Dutt et al. (arXiv:2505.10496, 2025)

## What this evaluates

Evaluates text-to-image generative models for synthetic chest radiograph generation across three dimensions: fidelity (image quality and mode coverage), privacy (memorization and re-identification risks), and utility (downstream clinical performance for classification and report generation). It probes whether synthetic medical images can match real data in clinical tasks while preserving patient privacy.

## Datasets

- **MIMIC-CXR** — total 242422; splits: train (237388), test (5034)

## Metrics

- `FID (RadDino)` **(primary)** — range: [0, ∞)
  - Fréchet Inception Distance computed using features extracted by the RadDino vision encoder. Lower values indicate better distribution alignment between real and synthetic images.
- `KID (RadDino)` — range: [0, ∞)
  - Kernel Inception Distance using RadDino features. Measures the distance between feature distributions of real and generated images. Lower is better.
- `Re-ID Score` — range: [0, 1]
  - Similarity score (e.g., cosine similarity or classifier confidence) between real and synthetic image pairs. Higher scores indicate greater memorization and re-identification risk.
- `AUC` — range: [0, 1]
  - Area Under the Receiver Operating Characteristic curve for binary pathology classification trained on synthetic vs. real data. Higher indicates better clinical utility.
- `BLEU-4` — range: [0, 100]
  - 4-gram precision score comparing generated radiology reports to reference reports. Higher indicates better fluency and lexical overlap.
- `F1-RadGraph` — range: [0, 1]
  - F1 score for clinical entity extraction and relation prediction using the RadGraph framework. Higher indicates better semantic accuracy.

## Input / output format

**Input**: Text prompts describing chest radiograph pathologies/conditions for generation; synthetic images or generated reports for downstream evaluation tasks.

**Output**: Synthetic chest radiograph images; classification labels or generated radiology reports.

## Scoring recipe

```python
def compute_metrics(generated_images, real_images, reports_gold, reports_pred):
    fid = frechet_distance(rad_dino_features(real_images), rad_dino_features(generated_images))
    kid = kernel_inception_distance(rad_dino_features(real_images), rad_dino_features(generated_images))
    reid_scores = [cosine_similarity(r, g) for r, g in zip(real_images, generated_images)]
    auc = compute_roc_auc(classifier(real_images), classifier(generated_images))
    bleu4 = nltk.bleu_score.sentence_bleu(reports_gold, reports_pred, weights=(0,0,0,1)) * 100
    f1_rg = radgraph_f1(reports_gold, reports_pred)
    return {'FID': fid, 'KID': kid, 'Re-ID': np.mean(reid_scores), 'AUC': auc, 'BLEU-4': bleu4, 'F1-RadGraph': f1_rg}
```

## Common pitfalls

- Relying solely on micro-averaged FID scores without assessing pathology-specific performance or mode coverage.
- Ignoring privacy risks (e.g., re-identification scores) when evaluating generation fidelity.
- Assuming that larger model sizes or LoRA fine-tuning automatically yield better clinical utility or balanced pathology generation.

## Evidence (verbatim from paper)

> The results are presented in Tab.[1], where we showcase both fidelity and mode coverage metrics. Sana *[xie2025sana]* delivers superior overall performance across key metrics, achieving the lowest FID and KID scores, indicating exceptional generation fidelity, while simultaneously attaining the highest Recall and Coverage, demonstrating its capacity to capture diverse modes (distributions) throughout the dataset.

## Citation

```bibtex
@misc{dutt2025chexgenbench,
  title={CheXGenBench: A Unified Benchmark For Fidelity, Privacy and Utility of Synthetic Chest Radiographs},
  author={Raman Dutt et al.},
  year={2025},
  note={arXiv:2505.10496}
}
```

- arXiv: 2505.10496

