Finsquad Eval

Evaluates the quality of a machine-translated extractive QA dataset (FinSQuAD) by training and testing QA models on it, comparing performance against other translated SQuAD datasets and the original English version. It also assesses translation fidelity through backtranslation and manual error analysis. Use when the user wants to benchmark on Finnish SQuAD2.0, SQuAD2.0, or asks about evaluating this task. Reports exact match (EM).

qhjqhj00 b0fc0e4 3.6 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/finsquad-eval commit b0fc0e4c47

Frequently asked questions

npx skillmds add qhjqhj00/finsquad-eval