Medhal Eval

Evaluates AI models' ability to detect factual inconsistencies (hallucinations) in medical text and generate grounded explanations for why statements are non-factual. It probes domain-specific factual consistency reasoning and binary classification under clinical constraints. Use when the user wants to benchmark on MedHal, MedNLI, Hegselmann et al. (2024a) Hallucination Dataset, or asks about evaluating this task. Reports F1-score.

qhjqhj00 bec148d 4.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/medhal-eval commit bec148dd6d

Frequently asked questions

npx skillmds add qhjqhj00/medhal-eval