sinhala-script-benchmark-eval
A Comprehensive Benchmark of Language Models on Unicode and Romanized Sinhala — Rajapakse et al. (2026) (arXiv:2601.14958, 2026)
What this evaluates
Evaluates language models' ability to process and generate text in two distinct Sinhala writing systems: standard Unicode and Romanized transliteration. It probes script-specific perplexity, coherence, and grammatical correctness, highlighting morphological understanding and training data biases.
Datasets
- Sinhala Unicode & Romanized — total ?; splits: test (-1)
Metrics
perplexity(primary) — range: scalar (lower is better)- Standard language modeling perplexity: PPL = exp(-1/N * sum(log P(x_i))). Lower values indicate better fluency and probability calibration.
Coherence— range: [1, 3] (lower is better)- Human/automated rating on a 1-3 scale (1=Excellent, 2=Acceptable, 3=Poor). Averaged across the dataset.
Grammar/Readability— range: [1, 3] (lower is better)- Human/automated rating on a 1-3 scale (1=Excellent, 2=Acceptable, 3=Poor). Averaged across the dataset.
Input / output format
Input: Prompt text in either Sinhala Unicode or Romanized script (e.g., 'mama kalin…' or 'monawada meke karanna…').
Output: Text completion in the corresponding script (Sinhala Unicode or Romanized).
Scoring recipe
# Perplexity
ppl = exp(-mean(log(model.log_prob(tokens))))
# Qualitative (averaged across dataset)
coherence = mean([1 if excellent else 2 if acceptable else 3 for c in completions])
grammar = mean([1 if excellent else 2 if acceptable else 3 for c in completions])
Common pitfalls
- Lower numerical scores indicate better performance for qualitative ratings (1=Excellent, 3=Poor).
- Models exhibit strong script bias; performance on Unicode does not generalize to Romanized Sinhala and vice versa.
- Subject-verb agreement and morphological endings are frequent failure points across models.
Evidence (verbatim from paper)
The perplexity scores for the open-source models are presented in Table[I]. The results indicate that the Mistral-Nemo-Base-2407 achieved the lowest perplexity for Unicode scripts (2.19) and Mistral-7B-v0.3 achieved the lowest perplexity for Romanized scripts (74.76), outperforming significantly larger models.
Citation
@misc{rajapakse2026comprehensive,
title={A Comprehensive Benchmark of Language Models on Unicode and Romanized Sinhala},
author={Rajapakse et al. (2026)},
year={2026},
note={arXiv:2601.14958}
}
- arXiv: 2601.14958