Uquad1.0 Eval

This benchmark evaluates Machine Reading Comprehension (MRC) capabilities in Urdu by testing a model's ability to extract correct answer spans from context paragraphs in response to questions. It probes span prediction accuracy, handling of multiple valid answers, and performance across different question types and named entities. Use when the user wants to benchmark on UQuAD1.0, or asks about evaluating this task. Reports F1.

qhjqhj00 7d20a65 3.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/uquad1.0-eval commit 7d20a65f6c

Frequently asked questions

npx skillmds add qhjqhj00/uquad1-0-eval