Clibench Eval

CliBench evaluates large language models on real-world clinical decision-making tasks, including diagnosis, procedure recommendation, lab test ordering, and medication prescribing. It probes the models' ability to process complex patient records, generate structured medical codes, and maintain coherence across multi-step clinical workflows in a zero-shot setting. Use when the user wants to benchmark on CliBench (MIMIC-IV derived), or asks about evaluating this task. Reports micro F1.

qhjqhj00 1369874 4.3 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/clibench-eval commit 1369874847

Frequently asked questions

npx skillmds add qhjqhj00/clibench-eval