transbench-eval
TransBench: Benchmarking Machine Translation for Industrial-Scale Applications — Li et al. (2025) (arXiv:2505.14244, 2025)
What this evaluates
TransBench evaluates machine translation models across three industrial capability levels: basic linguistic quality and robustness, domain-specific proficiency (e-commerce/finance), and cultural adaptation (taboo words and honorifics). It probes whether models maintain translation fidelity under input perturbations, adhere to domain terminology, and correctly handle culturally sensitive expressions without omission or over-translation.
Datasets
- TransBench — total 17000; splits: test (-1)
Metrics
BLEU (primary) — range: [0, 1]
- Precision-oriented n-gram co-occurrence metric incorporating a brevity penalty to penalize overly short translations.
TER — range: [0, 1]
- Edit-distance-based metric measuring the minimum number of insertions, deletions, substitutions, and shifts required to transform a candidate into a reference, normalized by reference length.
chrF — range: [0, 1]
- Character n-gram F-score computing a weighted harmonic mean of character-level precision and recall, emphasizing morphological and surface-form similarity.
COMET-XXL — range: [0, 1]
- Model-based metric leveraging multilingual pre-trained models (e.g., XLM-RoBERTa) to model the relationship between human ratings and vector space alignment via regression.
Hallucination Rate (HR) — range: [0, 1]
- HR = sum(F(S_H|S) for S in H_data) / |H_data|, where F is a binary classifier determining if the translation exhibits hallucination (repetition, omission, language mismatch, or length ratio violation).
Marco-MOS — range: [0, 5]
- Domain-tailored Quality Estimation model (fine-tuned Qwen2.5) predicting human Mean Opinion Scores on a 0-5 scale for financial and e-commerce datasets.
Taboo Accuracy (ACC_taboo) — range: [0, 1]
- Fraction of translations containing zero taboo words: sum(1 if no taboo words in translation else 0) / |T_data|.
Honorific Accuracy (ACC_hon) — range: [0, 1]
- Fraction of translations containing all expected honorific units: sum(F(S_Hon|S)) / |HO_data|, where F returns 1 only if all specified honorific tokens are present.
Input / output format
Input: Source sentence (optionally perturbed at sentence, character, or word level for robustness testing) and target language specification.
Output: Generated target-language translation string.
Scoring recipe
def score_translations(dataset, refs, taboo_list, honorific_units):
bleu = compute_bleu(refs, dataset.hyp)
ter = compute_ter(refs, dataset.hyp)
comet = comet_xxl_score(dataset.src, dataset.hyp)
hr = sum(hallucination_classifier(s, h) for s, h in dataset) / len(dataset)
taboo_acc = sum(1 for h in dataset if not any(t in h for t in taboo_list)) / len(dataset)
hon_acc = sum(1 for s, h in dataset if all(u in h for u in honorific_units)) / len(dataset)
marco_mos = marco_mos_model.predict(dataset.src, dataset.hyp)
return {'BLEU': bleu, 'TER': ter, 'chrF': chrF, 'COMET-XXL': comet, 'HR': hr, 'ACC_taboo': taboo_acc, 'ACC_hon': hon_acc, 'Marco-MOS': marco_mos}
Common pitfalls
- Robustness is evaluated by measuring BLEU drop on perturbed sources while keeping references unaltered, which can conflate translation degradation with perturbation sensitivity rather than true model robustness.
- Cultural fidelity metrics use strict exact-match accuracy for taboo words and honorifics, ignoring partial correctness, contextual nuance, or acceptable paraphrasing.
- Hallucination detection relies on heuristic thresholds (e.g., vector distance, length ratios, language detection) and a binary classifier that may not align with human judgments of omission or over-translation.
Evidence (verbatim from paper)
We utilize a set of established automatic metrics to measure fundamental linguistic quality and reliability, which include popular N-gram based metrics such as BLEU and TER, character-based metrics like chrF, and model-based metrics like COMET-XXL which leverage large language models to assess translation quality.
Citation
@misc{li2025transbench,
title={TransBench: Benchmarking Machine Translation for Industrial-Scale Applications},
author={Li et al. (2025)},
year={2025},
note={arXiv:2505.14244}
}
1---2name: transbench-eval3description: transbench-eval4---56# transbench-eval78> TransBench: Benchmarking Machine Translation for Industrial-Scale Applications — Li et al. (2025) (arXiv:2505.14244, 2025)910## What this evaluates1112TransBench evaluates machine translation models across three industrial capability levels: basic linguistic quality and robustness, domain-specific proficiency (e-commerce/finance), and cultural adaptation (taboo words and honorifics). It probes whether models maintain translation fidelity under input perturbations, adhere to domain terminology, and correctly handle culturally sensitive expressions without omission or over-translation.1314## Datasets1516- **TransBench** — total 17000; splits: test (-1)1718## Metrics1920- `BLEU` **(primary)** — range: [0, 1]21 - Precision-oriented n-gram co-occurrence metric incorporating a brevity penalty to penalize overly short translations.22- `TER` — range: [0, 1]23 - Edit-distance-based metric measuring the minimum number of insertions, deletions, substitutions, and shifts required to transform a candidate into a reference, normalized by reference length.24- `chrF` — range: [0, 1]25 - Character n-gram F-score computing a weighted harmonic mean of character-level precision and recall, emphasizing morphological and surface-form similarity.26- `COMET-XXL` — range: [0, 1]27 - Model-based metric leveraging multilingual pre-trained models (e.g., XLM-RoBERTa) to model the relationship between human ratings and vector space alignment via regression.28- `Hallucination Rate (HR)` — range: [0, 1]29 - HR = sum(F(S_H|S) for S in H_data) / |H_data|, where F is a binary classifier determining if the translation exhibits hallucination (repetition, omission, language mismatch, or length ratio violation).30- `Marco-MOS` — range: [0, 5]31 - Domain-tailored Quality Estimation model (fine-tuned Qwen2.5) predicting human Mean Opinion Scores on a 0-5 scale for financial and e-commerce datasets.32- `Taboo Accuracy (ACC_taboo)` — range: [0, 1]33 - Fraction of translations containing zero taboo words: sum(1 if no taboo words in translation else 0) / |T_data|.34- `Honorific Accuracy (ACC_hon)` — range: [0, 1]35 - Fraction of translations containing all expected honorific units: sum(F(S_Hon|S)) / |HO_data|, where F returns 1 only if all specified honorific tokens are present.3637## Input / output format3839**Input**: Source sentence (optionally perturbed at sentence, character, or word level for robustness testing) and target language specification.4041**Output**: Generated target-language translation string.4243## Scoring recipe4445```python46def score_translations(dataset, refs, taboo_list, honorific_units):47 bleu = compute_bleu(refs, dataset.hyp)48 ter = compute_ter(refs, dataset.hyp)49 comet = comet_xxl_score(dataset.src, dataset.hyp)50 51 hr = sum(hallucination_classifier(s, h) for s, h in dataset) / len(dataset)52 53 taboo_acc = sum(1 for h in dataset if not any(t in h for t in taboo_list)) / len(dataset)54 55 hon_acc = sum(1 for s, h in dataset if all(u in h for u in honorific_units)) / len(dataset)56 57 marco_mos = marco_mos_model.predict(dataset.src, dataset.hyp)58 return {'BLEU': bleu, 'TER': ter, 'chrF': chrF, 'COMET-XXL': comet, 'HR': hr, 'ACC_taboo': taboo_acc, 'ACC_hon': hon_acc, 'Marco-MOS': marco_mos}59```6061## Common pitfalls6263- Robustness is evaluated by measuring BLEU drop on perturbed sources while keeping references unaltered, which can conflate translation degradation with perturbation sensitivity rather than true model robustness.64- Cultural fidelity metrics use strict exact-match accuracy for taboo words and honorifics, ignoring partial correctness, contextual nuance, or acceptable paraphrasing.65- Hallucination detection relies on heuristic thresholds (e.g., vector distance, length ratios, language detection) and a binary classifier that may not align with human judgments of omission or over-translation.6667## Evidence (verbatim from paper)6869> We utilize a set of established automatic metrics to measure fundamental linguistic quality and reliability, which include popular N-gram based metrics such as BLEU and TER, character-based metrics like chrF, and model-based metrics like COMET-XXL which leverage large language models to assess translation quality.7071## Citation7273```bibtex74@misc{li2025transbench,75 title={TransBench: Benchmarking Machine Translation for Industrial-Scale Applications},76 author={Li et al. (2025)},77 year={2025},78 note={arXiv:2505.14244}79}80```8182- arXiv: 2505.14244