Hc3 Human Eval

Evaluates the ability to distinguish AI-generated responses from human expert answers and assesses perceived helpfulness across multiple domains. It probes linguistic realism, factual reliability, and stylistic alignment with human communication. Use when the user wants to benchmark on HC3, or asks about evaluating this task. Reports detection accuracy.

qhjqhj00 a75b8de 2.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/hc3-human-eval commit a75b8dee7c

Frequently asked questions

npx skillmds add qhjqhj00/hc3-human-eval