Medec Eval

Evaluates large language models' ability to detect and correct medical errors in clinical text. It probes the model's sensitivity to clinical inaccuracies, its precision in localizing erroneous sentences, and its capacity to generate semantically and lexically accurate corrections using different prompting strategies. Use when the user wants to benchmark on MEDEC, or asks about evaluating this task. Reports AggScore.

qhjqhj00 781eb4f 4.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/medec-eval commit 781eb4feb5

Frequently asked questions

npx skillmds add qhjqhj00/medec-eval