Healthbench Eval

Evaluates LLM responses to realistic clinical queries using a fine-grained, rubric-based scoring system. It measures medical accuracy, instruction following, completeness, context awareness, and safety by assigning positive or negative points to specific behavioral criteria, then normalizing the total to a [0, 1] scale. Use when the user wants to benchmark on HealthBench, or asks about evaluating this task. Reports HealthBench Score.

qhjqhj00 e3686ee 3.1 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/healthbench-eval commit e3686ee09c

Frequently asked questions

npx skillmds add qhjqhj00/healthbench-eval