# Sinhala Script Benchmark Eval

> sinhala-script-benchmark-eval

- Skill: `qhjqhj00/sinhala-script-benchmark-eval` (Agent Skill)
- Install (CLI): `npx skillmds@latest add qhjqhj00/sinhala-script-benchmark-eval`
- Raw SKILL.md: https://api.skillmd.com/api/skills/qhjqhj00/sinhala-script-benchmark-eval/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: qhjqhj00 (https://skillmd.com/u/qhjqhj00)
- Updated: 2026-09-21
- Page: https://skillmd.com/skills/qhjqhj00/sinhala-script-benchmark-eval

---


# sinhala-script-benchmark-eval

> A Comprehensive Benchmark of Language Models on Unicode and Romanized Sinhala — Rajapakse et al. (2026) (arXiv:2601.14958, 2026)

## What this evaluates

Evaluates language models' ability to process and generate text in two distinct Sinhala writing systems: standard Unicode and Romanized transliteration. It probes script-specific perplexity, coherence, and grammatical correctness, highlighting morphological understanding and training data biases.

## Datasets

- **Sinhala Unicode & Romanized** — total ?; splits: test (-1)

## Metrics

- `perplexity` **(primary)** — range: scalar (lower is better)
  - Standard language modeling perplexity: PPL = exp(-1/N * sum(log P(x_i))). Lower values indicate better fluency and probability calibration.
- `Coherence` — range: [1, 3] (lower is better)
  - Human/automated rating on a 1-3 scale (1=Excellent, 2=Acceptable, 3=Poor). Averaged across the dataset.
- `Grammar/Readability` — range: [1, 3] (lower is better)
  - Human/automated rating on a 1-3 scale (1=Excellent, 2=Acceptable, 3=Poor). Averaged across the dataset.

## Input / output format

**Input**: Prompt text in either Sinhala Unicode or Romanized script (e.g., 'mama kalin…' or 'monawada meke karanna…').

**Output**: Text completion in the corresponding script (Sinhala Unicode or Romanized).

## Scoring recipe

```python
# Perplexity
ppl = exp(-mean(log(model.log_prob(tokens))))

# Qualitative (averaged across dataset)
coherence = mean([1 if excellent else 2 if acceptable else 3 for c in completions])
grammar = mean([1 if excellent else 2 if acceptable else 3 for c in completions])
```

## Common pitfalls

- Lower numerical scores indicate better performance for qualitative ratings (1=Excellent, 3=Poor).
- Models exhibit strong script bias; performance on Unicode does not generalize to Romanized Sinhala and vice versa.
- Subject-verb agreement and morphological endings are frequent failure points across models.

## Evidence (verbatim from paper)

> The perplexity scores for the open-source models are presented in Table[I]. The results indicate that the Mistral-Nemo-Base-2407 achieved the lowest perplexity for Unicode scripts (2.19) and Mistral-7B-v0.3 achieved the lowest perplexity for Romanized scripts (74.76), outperforming significantly larger models.

## Citation

```bibtex
@misc{rajapakse2026comprehensive,
  title={A Comprehensive Benchmark of Language Models on Unicode and Romanized Sinhala},
  author={Rajapakse et al. (2026)},
  year={2026},
  note={arXiv:2601.14958}
}
```

- arXiv: 2601.14958

