Diagnosisarena Eval

Clinical diagnostic reasoning capability of LLMs, requiring them to generate plausible diagnoses from patient case descriptions and imaging/symptom details. It probes the model's ability to perform complex, multi-step medical deduction and generalize across 28 clinical specialties. Use when the user wants to benchmark on DiagnosisArena, or asks about evaluating this task. Reports accuracy.

qhjqhj00 b6bc44a 4.0 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/diagnosisarena-eval commit b6bc44a5bd

Frequently asked questions

npx skillmds add qhjqhj00/diagnosisarena-eval