dip-benchmark-eval
Can Your Context-Aware MT System Pass the DiP Benchmark Tests? : Evaluation Benchmarks for Discourse Phenomena in Machine Translation — Jwalapuram et al. (2020) (arXiv:2004.14607, 2020)
What this evaluates
Evaluates machine translation systems on four discourse phenomena: anaphora resolution, lexical consistency, coherence/readability, and discourse connectives. It probes whether context-aware models can maintain discourse-level quality and consistency across different language pairs beyond standard n-gram overlap.
Datasets
- DiP Benchmark — total ?; splits: test (-1)
Metrics
BLEU(primary) — range: percent- Standard n-gram overlap metric between reference and hypothesis translations. Computed as geometric mean of modified n-gram precisions (typically 1-4 grams) with a brevity penalty applied if the hypothesis is shorter than the reference. The paper uses it as a baseline that often fails to capture discourse-level deficiencies.
Input / output format
Input: Source sentence paired with preceding context (automatically extracted from WMT system outputs), target language.
Output: Translated target sentence.
Scoring recipe
def compute_bleu(predictions, references):
# Standard sentence-level or corpus-level BLEU calculation
# Counts matching n-grams (typically up to 4-grams) between hypothesis and reference
# Applies brevity penalty if hypothesis length < reference length
# Returns score in [0, 100] percent
return standard_bleu_score(predictions, references)
Common pitfalls
- Test sets are automatically generated from existing WMT system errors, which may bias evaluation towards specific error types rather than general discourse quality.
- BLEU scores may not capture discourse-level deficiencies, leading to misleading performance assessments for context-aware models.
- French test sets are smaller and based on older data (2013-2015) due to WMT discontinuation, making cross-lingual comparisons uneven.
Evidence (verbatim from paper)
We include French since Fr-En is a popular translation pair that results in some of the highest BLEU scores. Other source languages that are part of WMT can be extracted as needed; the testsets can also be expanded if older data were to be considered.
Citation
@misc{jwalapuram2020dipbenchmark,
title={Can Your Context-Aware MT System Pass the DiP Benchmark Tests? : Evaluation Benchmarks for Discourse Phenomena in Machine Translation},
author={Jwalapuram et al. (2020)},
year={2020},
note={arXiv:2004.14607}
}
- arXiv: 2004.14607