# Clinical Decision Metrics

> clinical-decision-metrics

- Skill: `qhjqhj00/clinical-decision-metrics` (Agent Skill)
- Install (CLI): `npx skillmds@latest add qhjqhj00/clinical-decision-metrics`
- Raw SKILL.md: https://api.skillmd.com/api/skills/qhjqhj00/clinical-decision-metrics/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: qhjqhj00 (https://skillmd.com/u/qhjqhj00)
- Updated: 2026-09-21
- Page: https://skillmd.com/skills/qhjqhj00/clinical-decision-metrics

---


# clinical-decision-metrics

> Beyond MedQA: Towards Real-world Clinical Decision Making in the Era of LLMs — Xiao et al. (2025) (arXiv:2510.20001, 2025)

## What this evaluates

Evaluates LLMs on real-world clinical decision-making across three dimensions: effectiveness (accuracy on closed-ended questions), efficiency (proportion of correct, non-redundant reasoning steps), and explainability (quality of clinical reasoning/answers via human or LLM judges).

## Datasets

- **Clinical Decision-Making Evaluation Framework** — total ?; splits: test (-1)

## Metrics

- `Accuracy` **(primary)** — range: [0, 1]
  - Calculated as the ratio of correctly predicted samples to the total number of samples: Accuracy = a_c / a_t. It measures match or multiple-choice score on known-answer questions.
- `Efficiency` — range: [0, 1]
  - Calculated as the average of binary indicators across N reasoning steps: Efficiency = (1/N) * sum(e_i), where e_i=1 if the step provides additional insight and does not repeat previous steps, else 0.
- `Explainability` — range: other
  - Evaluated via manual scoring (e.g., Likert scale averaged across raters) or LLM-as-a-judge scoring, where a judge model evaluates the output against context/prompt templates or step-wise rubrics.

## Input / output format

**Input**: Clinical background (varying from precise to incomplete information) paired with a clinical question (true/false, multiple-choice, short-answer, or open-ended).

**Output**: Model prediction (e.g., selected option, short text, or open-ended response) and optionally a step-by-step reasoning process or explanation.

## Scoring recipe

```python
def compute_accuracy(predictions, golds):
    correct = sum(1 for p, g in zip(predictions, golds) if p == g)
    return correct / len(golds)

def compute_efficiency(reasoning_steps):
    N = len(reasoning_steps)
    e_i = [1 if step_provides_insight_and_is_non_redundant(step) else 0 for step in reasoning_steps]
    return sum(e_i) / N

def compute_explainability(output, rubric_or_judge_llm):
    # Manual: average Likert scores from multiple raters
    # LLM-as-judge: P_LLM(x ⊕ C) where x=output, C=scoring context/rubric
    return judge_llm.evaluate(output, rubric_or_judge_llm)
```

## Common pitfalls

- Accuracy is only suitable for closed-ended questions and fails to capture performance on open-ended or incomplete-information clinical scenarios.
- Efficiency and Explainability rely heavily on LLM-as-a-judge or manual scoring, which can introduce subjectivity and inconsistency across studies.
- The paper notes that even advanced models struggle with complex multiple-choice questions (accuracy <50%), highlighting that high benchmark scores do not guarantee clinical readiness.

## Evidence (verbatim from paper)

> For closed-ended questions, accuracy is an easy-to-use metric to measure effectiveness, which is calculated as follows: $Accuracy\=a_{c}/a_{t}$ ... $a_{c}$ is the number of samples predicted correctly, $a_{t}$ is the number of all samples.

## Citation

```bibtex
@misc{xiao2025beyondmedqa,
  title={Beyond MedQA: Towards Real-world Clinical Decision Making in the Era of LLMs},
  author={Xiao et al. (2025)},
  year={2025},
  note={arXiv:2510.20001}
}
```

- arXiv: 2510.20001

