Medical Reasoning Benchmarks Eval

Evaluates large language models' ability to perform medical reasoning across multiple-choice clinical questions, specialist-level board exams, and general-domain medical subsets. It probes factual knowledge integration, diagnostic accuracy, and reasoning under uncertainty in safety-critical settings. Use when the user wants to benchmark on MedQA (USMLE), MedMCQA (Validation), PubMedQA, GPQA, JMED, ReDis-QA, MedXpertQA, MMLU-Pro, or asks about evaluating this task. Reports accuracy.

qhjqhj00 2a2cb91 3.3 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/medical-reasoning-benchmarks-eval commit 2a2cb91ec2

Frequently asked questions

npx skillmds add qhjqhj00/medical-reasoning-benchmarks-eval