Comet Mt Eval

Probes the semantic fidelity and grammatical correctness of automated translation pipelines when converting English benchmarks into low-resource languages. It measures how well translation methods preserve task structure and downstream model performance consistency. Use when the user wants to benchmark on FLORES, WMT24++, MMLU, or asks about evaluating this task. Reports COMET.

qhjqhj00 5979b47 3.6 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/comet-mt-eval commit 5979b47d36

Frequently asked questions

npx skillmds add qhjqhj00/comet-mt-eval