Ott QA Eval

This benchmark evaluates a model's ability to perform open-domain question answering by retrieving and fusing evidence from both tabular and textual sources. It specifically probes multi-hop reasoning capabilities where answers require bridging information across separate table segments and text passages. Use when the user wants to benchmark on OTT-QA, or asks about evaluating this task. Reports EM.

qhjqhj00 dd85d7b 3.1 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/ott-qa-eval commit dd85d7ba1e

Frequently asked questions

npx skillmds add qhjqhj00/ott-qa-eval