Mentalbench Eval

Evaluates large language models' ability to perform psychiatric diagnostic decision-making using DSM-5 criteria. It probes their capacity to handle information incompleteness, perform differential diagnosis among overlapping disorders, and calibrate diagnostic commitment under varying prompt constraints. Use when the user wants to benchmark on MentalBench, or asks about evaluating this task. Reports accuracy (exact match).

qhjqhj00 dc517fe 2.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mentalbench-eval commit dc517fe8ca

Frequently asked questions

npx skillmds add qhjqhj00/mentalbench-eval