Medexqa Eval

Evaluates medical language models on multiple-choice question answering and the generation of clinically relevant explanations. It probes the model's ability to perform clinical reasoning, avoid hallucinations, and produce coherent, accurate rationales aligned with medical domain knowledge. Use when the user wants to benchmark on MedExQA, or asks about evaluating this task. Reports Classification Accuracy.

qhjqhj00 885d6f3 4.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/medexqa-eval commit 885d6f31dd

Frequently asked questions

npx skillmds add qhjqhj00/medexqa-eval