Mimic Cdm Eval

Evaluates large language models on clinical decision-making tasks by predicting diagnoses from structured patient evidence. It probes whether models possess latent domain-specific reasoning capabilities that are masked by unfamiliarity with benchmark input formats and task definitions. Use when the user wants to benchmark on MIMIC-CDM, or asks about evaluating this task. Reports accuracy.

qhjqhj00 a219972 2.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mimic-cdm-eval commit a219972528

Frequently asked questions

npx skillmds add qhjqhj00/mimic-cdm-eval