Clinconsensus Eval

Evaluates Chinese medical LLMs on their ability to generate clinically usable, consistent, and safe responses across diverse specialties and difficulty levels. It probes reasoning depth, evidence integration, and longitudinal follow-up rather than raw factual accuracy. Use when the user wants to benchmark on ClinConsensus, or asks about evaluating this task. Reports CACS@7.

qhjqhj00 63fc4ac 2.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/clinconsensus-eval commit 63fc4acce5

Frequently asked questions

npx skillmds add qhjqhj00/clinconsensus-eval