Wmt Mt Eval

This protocol evaluates the machine translation quality of large language models across multiple language pairs. It measures translation accuracy and fluency by comparing model outputs against gold references and state-of-the-art baselines using neural quality estimation metrics. The benchmark probes the model's ability to generalize across diverse language directions and avoid generating near-perfect but flawed translations. Use when the user wants to benchmark on WMT'21 Test Set, WMT'22 Test Set, WMT'23 Test Set, or asks about evaluating this task. Reports KIWI-XXL.

qhjqhj00 73aa256 4.1 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/wmt-mt-eval commit 73aa256bc6

Frequently asked questions

npx skillmds add qhjqhj00/wmt-mt-eval