clinical-decision-metrics
Beyond MedQA: Towards Real-world Clinical Decision Making in the Era of LLMs — Xiao et al. (2025) (arXiv:2510.20001, 2025)
What this evaluates
Evaluates LLMs on real-world clinical decision-making across three dimensions: effectiveness (accuracy on closed-ended questions), efficiency (proportion of correct, non-redundant reasoning steps), and explainability (quality of clinical reasoning/answers via human or LLM judges).
Datasets
- Clinical Decision-Making Evaluation Framework — total ?; splits: test (-1)
Metrics
Accuracy(primary) — range: [0, 1]- Calculated as the ratio of correctly predicted samples to the total number of samples: Accuracy = a_c / a_t. It measures match or multiple-choice score on known-answer questions.
Efficiency— range: [0, 1]- Calculated as the average of binary indicators across N reasoning steps: Efficiency = (1/N) * sum(e_i), where e_i=1 if the step provides additional insight and does not repeat previous steps, else 0.
Explainability— range: other- Evaluated via manual scoring (e.g., Likert scale averaged across raters) or LLM-as-a-judge scoring, where a judge model evaluates the output against context/prompt templates or step-wise rubrics.
Input / output format
Input: Clinical background (varying from precise to incomplete information) paired with a clinical question (true/false, multiple-choice, short-answer, or open-ended).
Output: Model prediction (e.g., selected option, short text, or open-ended response) and optionally a step-by-step reasoning process or explanation.
Scoring recipe
def compute_accuracy(predictions, golds):
correct = sum(1 for p, g in zip(predictions, golds) if p == g)
return correct / len(golds)
def compute_efficiency(reasoning_steps):
N = len(reasoning_steps)
e_i = [1 if step_provides_insight_and_is_non_redundant(step) else 0 for step in reasoning_steps]
return sum(e_i) / N
def compute_explainability(output, rubric_or_judge_llm):
# Manual: average Likert scores from multiple raters
# LLM-as-judge: P_LLM(x ⊕ C) where x=output, C=scoring context/rubric
return judge_llm.evaluate(output, rubric_or_judge_llm)
Common pitfalls
- Accuracy is only suitable for closed-ended questions and fails to capture performance on open-ended or incomplete-information clinical scenarios.
- Efficiency and Explainability rely heavily on LLM-as-a-judge or manual scoring, which can introduce subjectivity and inconsistency across studies.
- The paper notes that even advanced models struggle with complex multiple-choice questions (accuracy <50%), highlighting that high benchmark scores do not guarantee clinical readiness.
Evidence (verbatim from paper)
For closed-ended questions, accuracy is an easy-to-use metric to measure effectiveness, which is calculated as follows: $Accuracy=a_{c}/a_{t}$ ... $a_{c}$ is the number of samples predicted correctly, $a_{t}$ is the number of all samples.
Citation
@misc{xiao2025beyondmedqa,
title={Beyond MedQA: Towards Real-world Clinical Decision Making in the Era of LLMs},
author={Xiao et al. (2025)},
year={2025},
note={arXiv:2510.20001}
}
- arXiv: 2510.20001