Mt Incremental Eval

This benchmark evaluates how well automatic machine translation metrics track quality improvements in commercial systems over time. It probes whether metrics consistently rank newer systems higher than older ones, and how their reliability changes as system quality improves or when synthetic references are used. Use when the user wants to benchmark on Commercial MT Systems Corpus, or asks about evaluating this task. Reports Accuracy.

qhjqhj00 f9e63c7 3.7 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mt-incremental-eval commit f9e63c766c

Frequently asked questions

npx skillmds add qhjqhj00/mt-incremental-eval