Hotpotqa Eval

This benchmark evaluates a model's ability to perform multi-hop question answering by reasoning across multiple documents. It specifically probes explainability through supporting fact prediction and tests robustness against distractor paragraphs and large-scale retrieval contexts. Use when the user wants to benchmark on HotpotQA, or asks about evaluating this task. Reports F1.

qhjqhj00 c077025 4.2 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/hotpotqa-eval commit c0770250a5

Frequently asked questions

npx skillmds add qhjqhj00/hotpotqa-eval