Medquad Behavioural Eval

Evaluates a medical AI system's clinical reasoning behavior, focusing on uncertainty handling, deferral, and safety rather than raw answer accuracy. It probes the model's ability to maintain clinician-aligned reasoning, avoid speculative completions, and preserve context across diverse medical queries. Use when the user wants to benchmark on MedQuAD benchmark, or asks about evaluating this task. Reports Benchmark Completion Rate.

qhjqhj00 276d092 4.0 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/medquad-behavioural-eval commit 276d0922a3

Frequently asked questions

npx skillmds add qhjqhj00/medquad-behavioural-eval