Medagents Medical QA Eval

Evaluates zero-shot medical reasoning and multiple-choice question answering capabilities of LLMs using a training-free multi-agent collaboration framework. It probes the model's ability to simulate domain expert role-playing and reach consensus without retrieval-augmented generation. Use when the user wants to benchmark on MedQA, MedMCQA, PubMedQA, MMLU Anatomy, MMLU Clinical Knowledge, MMLU College Medicine, MMLU Medical Genetics, MMLU Professional Medicine, MMLU College Biology, or asks about evaluating this task. Reports accuracy.

qhjqhj00 b3cb002 3.1 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/medagents-medical-qa-eval commit b3cb00236e

Frequently asked questions

npx skillmds add qhjqhj00/medagents-medical-qa-eval