Authenhallu Eval

Evaluates an LLM's capability to detect and categorize hallucinations in authentic, real-world human-LLM dialogues. It specifically probes whether models can identify input-conflicting, context-conflicting, and fact-conflicting errors in query-response pairs. Use when the user wants to benchmark on AuthenHallu, or asks about evaluating this task. Reports accuracy.

qhjqhj00 7bd16c8 2.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/authenhallu-eval commit 7bd16c8f17

Frequently asked questions

npx skillmds add qhjqhj00/authenhallu-eval