QA Benchmarks Eval

Evaluates the capability of retrieval-augmented generation systems to answer complex, multi-hop, and long-form questions by iteratively retrieving, structuring, and accumulating evidence from documents. Use when the user wants to benchmark on StrategyQA, ASQA, NQ, 2WikiMultiHopQA, HotpotQA, or asks about evaluating this task. Reports EM, F1, ACC.

qhjqhj00 9aab08c 2.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/qa-benchmarks-eval commit 9aab08cc31

Frequently asked questions

npx skillmds add qhjqhj00/qa-benchmarks-eval