# Xq Meval Eval

> Evaluates automatic machine translation metrics by measuring their correlation with human quality judgments across multiple language pairs. It probes whether metrics exhibit cross-lingual scoring bias and how reliably they rank translation systems or quality triplets relative to human assessments. Use when the user wants to benchmark on XQ-MEval, or asks about evaluating this task. Reports Kendall-τ.

- Skill: `qhjqhj00/xq-meval-eval` (Agent Skill)
- Install (CLI): `npx skillmds add qhjqhj00/xq-meval-eval`
- Raw SKILL.md: https://api.skillmd.com/api/skills/qhjqhj00/xq-meval-eval/raw
- Safety review: pending
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Productivity
- Author: qhjqhj00 (https://skillmd.com/u/qhjqhj00)
- Updated: 2026-09-08
- Page: https://skillmd.com/skills/qhjqhj00/xq-meval-eval

---


# xq-meval-eval

> XQ-MEval: A Dataset with Cross-lingual Parallel Quality for Benchmarking Translation Metrics — Liu et al. (2026) (arXiv:2604.14934, 2026)

## What this evaluates

Evaluates automatic machine translation metrics by measuring their correlation with human quality judgments across multiple language pairs. It probes whether metrics exhibit cross-lingual scoring bias and how reliably they rank translation systems or quality triplets relative to human assessments.

## Datasets

- **XQ-MEval** — total ?; splits: test (-1); repo https://github.com/zhiqu22/XQ-MEval

## Metrics

- `Kendall-τ` **(primary)** — range: [-1, 1]
  - Measures the ordinal correlation between automatic metric scores and human quality judgments. Computed separately at the system level (averaging scores per system) and triplet level (ranking triplets of translations).

## Input / output format

**Input**: Source sentence, reference translation, and hypothesis translation (or system output) for each instance; evaluated across nine translation directions involving languages such as zh, lo, de, id, ja, si, and en.

**Output**: A continuous quality score for the hypothesis, or a ranked ordering of systems/triplets based on the metric's prediction.

## Scoring recipe

```python
def compute_kendall_tau(pred_scores, human_scores):
    n = len(pred_scores)
    concordant = discordant = 0
    for i in range(n):
        for j in range(i + 1, n):
            diff_pred = pred_scores[i] - pred_scores[j]
            diff_gold = human_scores[i] - human_scores[j]
            if diff_pred * diff_gold > 0:
                concordant += 1
            elif diff_pred * diff_gold < 0:
                discordant += 1
    return (concordant - discordant) / (n * (n - 1) / 2)
```

## Common pitfalls

- Averaging metric scores across languages without normalization introduces unfair evaluation due to systematic cross-lingual scoring bias.
- System-level and triplet-level Kendall-τ capture different ranking granularities and should be reported separately.
- Reference-free metrics (e.g., KIWI22, MX-qe) are evaluated on the same dataset but require input formatting without reference sentences.

## Evidence (verbatim from paper)

> Table 5: Results showing the system-level and triplet-level Kendall-τ correlation between averaged metric scores and human judgments on pseudo systems. Num. of Lang. denotes the number of involved languages. In this setting, Num. of 3 means that the system is sampled from zh, lo, and de; Num. of 6 means that the system is sampled from zh, lo, de, id, ja, and si; Num. of 9 means that the system is sampled from all languages. The abbreviations of metric are as follows: BLEURT, COMET, xCOMET, MX-reg, KIWI22, KIWI23, and MX-qe.

## Citation

```bibtex
@misc{liu2026xqmeval,
  title={XQ-MEval: A Dataset with Cross-lingual Parallel Quality for Benchmarking Translation Metrics},
  author={Liu et al. (2026)},
  year={2026},
  note={arXiv:2604.14934}
}
```

- arXiv: 2604.14934

