Medical QA Explanation Eval

Evaluates large language models' ability to answer challenging medical multiple-choice questions and generate step-by-step clinical reasoning explanations. It probes both factual accuracy in clinical decision-making and the quality of model-generated rationales compared to expert-written references. Use when the user wants to benchmark on JAMA Clinical Challenge, Medbullets, or asks about evaluating this task. Reports accuracy.

qhjqhj00 543aa9a 3.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/medical-qa-explanation-eval commit 543aa9a3ae

Frequently asked questions

npx skillmds add qhjqhj00/medical-qa-explanation-eval