Er Reason Eval

Evaluates LLMs on longitudinal clinical reasoning across five emergency room workflow stages, including acuity assessment, case summarization, treatment planning, final diagnosis, and patient disposition. It probes the models' ability to integrate sparse clinical notes, perform rule-out differential diagnosis, and align outputs with real-world clinical decision-making and safety constraints. Use when the user wants to benchmark on ER-Reason, or asks about evaluating this task. Reports Accuracy, ROUGE-F1, cTAKES CUI Overlap Ratio.

qhjqhj00 3a7c9d8 4.1 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/er-reason-eval commit 3a7c9d81fd

Frequently asked questions

npx skillmds add qhjqhj00/er-reason-eval