Codet Eval

This benchmark probes the robustness of machine translation systems to dialectal variations by measuring how consistently they translate semantically similar sentences in standard vs. dialectal forms. It evaluates whether models maintain translation quality and coherence when exposed to lexical and morphosyntactic variations across multiple languages. Use when the user wants to benchmark on CODET, or asks about evaluating this task. Reports COMET.

qhjqhj00 2f440e9 3.3 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/codet-eval commit 2f440e9d93

Frequently asked questions

npx skillmds add qhjqhj00/codet-eval