Medqa Usmle Eval

Evaluates medical multiple-choice question answering capability of 4B-parameter LLMs, specifically comparing the impact of domain fine-tuning versus retrieval-augmented generation (RAG) on accuracy. Use when the user wants to benchmark on MedQA-USMLE, or asks about evaluating this task. Reports Majority-vote accuracy.

qhjqhj00 248c843 2.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/medqa-usmle-eval commit 248c843003

Frequently asked questions

npx skillmds add qhjqhj00/medqa-usmle-eval