Mt Data Filtering Eval

This evaluation protocol assesses how effectively Quality Estimation (QE) metrics can filter low-quality or noisy sentence pairs from large parallel corpora. It measures whether retaining only the top 50% of high-scoring pairs improves downstream Neural Machine Translation (NMT) performance compared to using the full corpus or alternative filtering baselines like BICLEANER. Use when the user wants to benchmark on WMT & IWSLT Evaluation Campaigns, or asks about evaluating this task. Reports COMET22.

qhjqhj00 4053dbc 3.1 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mt-data-filtering-eval commit 4053dbc8c4

Frequently asked questions

npx skillmds add qhjqhj00/mt-data-filtering-eval