Med Critics Eval

This benchmark evaluates a model's ability to verify factual accuracy in long-form medical texts by recursively decomposing claims into a verification tree. It probes fine-grained fact-checking across six medical domains, requiring the model to distinguish between factual and deliberately falsified claims while accounting for contextual dependencies. Use when the user wants to benchmark on Med-Critics, or asks about evaluating this task. Reports accuracy.

qhjqhj00 a427252 3.5 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/med-critics-eval commit a427252616

Frequently asked questions

npx skillmds add qhjqhj00/med-critics-eval