Selfcheckgpt Eval

Evaluates a model's ability to detect hallucinated versus factual content in generated text using zero-resource consistency metrics across stochastic samples. It probes whether factual knowledge yields coherent, consistent outputs while hallucinated content exhibits divergence across multiple generations. Use when the user wants to benchmark on SelfCheckGPT dataset, or asks about evaluating this task. Reports AUC-PR.

qhjqhj00 f3c9805 3.6 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/selfcheckgpt-eval commit f3c9805d0e

Frequently asked questions

npx skillmds add qhjqhj00/selfcheckgpt-eval