# Medrgb Eval

> medrgb-eval

- Skill: `qhjqhj00/medrgb-eval` (Agent Skill)
- Install (CLI): `npx skillmds@latest add qhjqhj00/medrgb-eval`
- Raw SKILL.md: https://api.skillmd.com/api/skills/qhjqhj00/medrgb-eval/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: qhjqhj00 (https://skillmd.com/u/qhjqhj00)
- Updated: 2026-09-21
- Page: https://skillmd.com/skills/qhjqhj00/medrgb-eval

---


# medrgb-eval

> Comprehensive and Practical Evaluation of Retrieval-Augmented Generation Systems for Medical Question Answering — Ngo et al. (2024) (arXiv:2411.09213, 2024)

## What this evaluates

Evaluates LLMs' ability to perform medical question answering under realistic retrieval-augmented generation (RAG) conditions. It probes four key capabilities: handling insufficient or noisy context, integrating multi-source information via sub-questions, detecting factual errors in retrieved documents, and standard RAG performance.

## Datasets

- **MedRGB** — total ?; splits: test (-1)

## Metrics

- `accuracy` **(primary)** — range: [0, 1]
  - Proportion of correctly answered main questions out of the total number of questions evaluated.
- `exact-match accuracy` — range: [0, 1]
  - Strict string matching between the model's generated sub-question answer and the ground truth answer.
- `GPT-based score` — range: [0, 1]
  - Lenient evaluation of sub-question answers using a GPT-based prompt to reward relatively accurate responses that aid main answer inference, rather than requiring exact matches.
- `factual error detection rate` — range: [0, 1]
  - Proportion of factually incorrect or adversarial documents correctly identified as misinformation by the model.

## Input / output format

**Input**: A main medical question, optionally accompanied by sub-questions, along with a set of retrieved documents. The documents vary in quantity (5 or 20) and signal-to-noise ratio (parameter $p$), including scenarios with insufficient information or adversarial misinformation.

**Output**: A final answer to the main question. For integration and robustness tests, the model also generates answers to sub-questions and/or flags insufficient information or detects factual errors.

## Scoring recipe

```python
def compute_metrics(predictions, gold, sub_preds=None, sub_gold=None, error_flags=None):
    main_acc = sum(1 for p, g in zip(predictions, gold) if p == g) / len(predictions)
    if sub_preds and sub_gold:
        exact_match = sum(1 for p, g in zip(sub_preds, sub_gold) if p == g) / len(sub_preds)
        gpt_score = evaluate_with_gpt_prompt(sub_preds, sub_gold)
    if error_flags:
        detection_rate = sum(error_flags) / len(error_flags)
    return {'accuracy': main_acc, 'exact_match': exact_match, 'gpt_score': gpt_score, 'detection_rate': detection_rate}
```

## Common pitfalls

- Models often ignore retrieved context when $p=0$ and rely on internal knowledge, leading to false 'insufficient information' responses.
- Higher document counts do not guarantee better performance; they can introduce noise or exceed context limits, degrading accuracy.
- GPT-based sub-question scores remain high (>80%) but do not correlate with main task accuracy, making them misleading for optimization.
- Models frequently accept misinformation as truth, resulting in high false-positive rates for factual error detection.

## Evidence (verbatim from paper)

> Additionally, we measure sub-question accuracy in Fig. [14] with two metrics. One of them is the strict exact-match score for extractive QA task, and the other is a more lenient GPT-based score using the prompt from Fig. [10]. The intuition for this metric is that, since these are sub-questions, their exact accuracy is not as important. Sub-answers that are relatively accurate and help infer the main answer should also be rewarded.

## Citation

```bibtex
@misc{ngo2024medrgb,
  title={Comprehensive and Practical Evaluation of Retrieval-Augmented Generation Systems for Medical Question Answering},
  author={Ngo et al. (2024)},
  year={2024},
  note={arXiv:2411.09213}
}
```

- arXiv: 2411.09213

