Mkqa Eval

Evaluates multilingual open-domain question answering models on their ability to generate or extract short, factoid answers across 26 typologically diverse languages. It specifically probes cross-lingual transfer, handling of unanswerable queries, and robustness to language-specific normalization and threshold tuning for abstention. Use when the user wants to benchmark on MKQA, or asks about evaluating this task. Reports token overlap F1.

qhjqhj00 7f0149d 4.1 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mkqa-eval commit 7f0149d687

Frequently asked questions

npx skillmds add qhjqhj00/mkqa-eval