Medical QA Eval

This benchmark evaluates the medical reasoning and question-answering capabilities of language models across multiple-choice and open-ended clinical tasks. It probes the model's ability to retrieve relevant medical knowledge, perform stepwise reasoning, and select or generate correct answers based on clinical guidelines and literature. Use when the user wants to benchmark on MedQA, MedMCQA, MMLU-Med, DDXPlus, AgentClinicNEJM, AgentClinicMedQA, or asks about evaluating this task. Reports accuracy.

qhjqhj00 b752c9c 3.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/medical-qa-eval commit b752c9c7b3

Frequently asked questions

npx skillmds add qhjqhj00/medical-qa-eval