Wqe Metric Eval

Evaluates the ability of unsupervised and supervised metrics to identify word-level translation errors by comparing their continuous scores against human-annotated error spans and multi-annotator agreement rates. Use when the user has predictions and gold and needs to compute Average Precision (AP).

qhjqhj00 59aa960 3.2 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/wqe-metric-eval commit 59aa960b30

Frequently asked questions

npx skillmds add qhjqhj00/wqe-metric-eval