Mt Reasoning Scaling Eval

This evaluation probes the test-time scaling properties of reasoning models across diverse machine translation tasks. It measures how varying reasoning budgets and iterative self-correction workflows impact translation quality across literary, biomedical, cultural, and commonsense domains. Use when the user wants to benchmark on WMT24-Literary, MetaphorTrans, LitEval-Corpus, WMT24-Biomedical, WMT23-Biomedical, CAMT, Commonsense-MT, RTT, RAGTrans, or asks about evaluating this task. Reports COMET-22.

qhjqhj00 aec8d0e 4.2 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/mt-reasoning-scaling-eval commit aec8d0eeb2

Frequently asked questions

npx skillmds add qhjqhj00/mt-reasoning-scaling-eval