Ambigqa Eval

This benchmark evaluates a model's ability to identify ambiguous open-domain questions, generate multiple plausible answer spans, and produce disambiguated question rewrites that distinguish between different interpretations of the same query. Use when the user wants to benchmark on AMBIGNQ, or asks about evaluating this task. Reports F1ans.

qhjqhj00 dce5e9b 3.5 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/ambigqa-eval commit dce5e9b502

Frequently asked questions

npx skillmds add qhjqhj00/ambigqa-eval