Averitec Eval

Evaluates a model's ability to verify real-world claims by retrieving web evidence, generating supporting questions, predicting veracity stance, and producing textual justifications. It probes retrieval quality, stance detection, and justification generation under realistic conditions with temporal and context constraints. Use when the user wants to benchmark on AVeriTeC, or asks about evaluating this task. Reports Macro-F1.

qhjqhj00 721716c 3.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/averitec-eval commit 721716c626

Frequently asked questions

npx skillmds add qhjqhj00/averitec-eval