Mrqa 2019 Shared Task Eval

Evaluates out-of-domain generalization in extractive reading comprehension by testing models on held-out datasets from diverse domains (crowdsourced, synthetic, domain experts, Wikipedia, education, etc.) that were not seen during training. Use when the user wants to benchmark on MRQA 2019 Shared Task, or asks about evaluating this task. Reports F1.

qhjqhj00 d4cbb4f 2.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mrqa-2019-shared-task-eval commit d4cbb4fc3a

Frequently asked questions

npx skillmds add qhjqhj00/mrqa-2019-shared-task-eval