Medical Reasoning Eval

Evaluates large language models on complex medical reasoning and knowledge retrieval across multiple-choice and open-ended clinical questions. It probes the model's ability to apply domain-specific knowledge, perform multi-step clinical reasoning, and handle challenging benchmarks that require more than simple fact recall. Use when the user wants to benchmark on MedQA (USMLE), MedMCQA, PubMedQA, MMLU-Pro, GPQA, or asks about evaluating this task. Reports accuracy.

qhjqhj00 ef46b63 2.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/medical-reasoning-eval commit ef46b63023

Frequently asked questions

npx skillmds add qhjqhj00/medical-reasoning-eval