Fact Checking Arena Eval

Evaluates LLMs on multi-hop fact-checking by measuring claim extraction, evidence retrieval, and justification quality. Uses an arena-style pairwise comparison framework with LLM judges to rank models across multiple reasoning dimensions. Use when the user wants to benchmark on HOVER, FEVERIOUS, or asks about evaluating this task. Reports Accuracy (%).

qhjqhj00 9410859 3.3 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/fact-checking-arena-eval commit 9410859bc1

Frequently asked questions

npx skillmds add qhjqhj00/fact-checking-arena-eval