Refact Eval

This benchmark evaluates large language models' ability to detect, localize, and correct scientific confabulations in generated answers. It probes fine-grained factuality awareness, span-level error identification, and factual restoration capabilities under domain-specific scrutiny. Use when the user wants to benchmark on ReFACT, or asks about evaluating this task. Reports accuracy.

qhjqhj00 008097a 4.2 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/refact-eval commit 008097a7a7

Frequently asked questions

npx skillmds add qhjqhj00/refact-eval