Multi Hop QA Eval

Evaluates multi-hop question answering capabilities across diverse reasoning types, including implicit commonsense/arithmetic reasoning, explicit composition/comparison, and fact verification. It tests the model's ability to synthesize information from retrieved evidence and generate step-by-step explanations. Use when the user wants to benchmark on STRATEGYQA, FERMI, QUARTZ, HOTPOTQA, 2WIKIMQA, BAMBOOGLE, FEVEROUS, or asks about evaluating this task. Reports F1.

qhjqhj00 8b2ad98 3.3 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/multi-hop-qa-eval commit 8b2ad98924

Frequently asked questions

npx skillmds add qhjqhj00/multi-hop-qa-eval