Medcasereasoning Eval

Evaluates large language models' ability to perform clinical diagnostic reasoning and arrive at correct final diagnoses based on patient case reports. It specifically probes whether models can align their step-by-step reasoning processes with clinician-authored diagnostic traces, rather than just guessing the final answer. Use when the user wants to benchmark on MedCaseReasoning, or asks about evaluating this task. Reports Diagnostic Accuracy.

qhjqhj00 69068e1 3.6 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/medcasereasoning-eval commit 69068e13b8

Frequently asked questions

npx skillmds add qhjqhj00/medcasereasoning-eval