Tcm Best4sdt Eval

This benchmark evaluates large language models' capabilities in Traditional Chinese Medicine (TCM) clinical reasoning, specifically focusing on syndrome differentiation and treatment decision-making. It probes the model's ability to accurately diagnose pathological patterns, formulate appropriate herbal prescriptions, and adhere to medical ethics and safety guidelines across 27 dimensions. Use when the user wants to benchmark on TCM-BEST4SDT, or asks about evaluating this task. Reports selected-response evaluation.

qhjqhj00 0b84cc7 4.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/tcm_best4sdt-eval commit 0b84cc7944

Frequently asked questions

npx skillmds add qhjqhj00/tcm-best4sdt-eval