# Emergent Tts Eval

> emergent-tts-eval

- Skill: `qhjqhj00/emergent-tts-eval` (Agent Skill)
- Install (CLI): `npx skillmds@latest add qhjqhj00/emergent-tts-eval`
- Raw SKILL.md: https://api.skillmd.com/api/skills/qhjqhj00/emergent-tts-eval/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: qhjqhj00 (https://skillmd.com/u/qhjqhj00)
- Updated: 2026-09-21
- Page: https://skillmd.com/skills/qhjqhj00/emergent-tts-eval

---


# emergent-tts-eval

> EmergentTTS-Eval: Evaluating TTS Models on Complex Prosodic, Expressiveness, and Linguistic Challenges Using Model-as-a-Judge — Manku et al. (2025) (arXiv:2505.23009, 2025)

## What this evaluates

Evaluates text-to-speech models on six complex linguistic and prosodic dimensions: emotions, paralinguistics, syntactic complexity, questions, foreign words, and complex pronunciation. It uses a model-as-a-judge framework to measure pairwise preference against a baseline, alongside automated speech quality metrics.

## Datasets

- **EmergentTTS-Eval** — total 1645; splits: test (1645); repo https://github.com/boson-ai/EmergentTTS-Eval-public

## Metrics

- `win-rate` **(primary)** — range: percent
  - Percentage of pairwise comparisons where the evaluated model's audio output is preferred over the baseline (gpt-4o-mini-tts, Alloy voice) by a judge LALM. Computed per category and overall.
- `WER` — range: percent
  - Word Error Rate computed using Whisper-v3-large to measure transcription accuracy of the generated audio.
- `MOS` — range: other
  - Mean Opinion Score estimated using a fine-tuned wav2vec2.0 model to predict human-like quality ratings.

## Input / output format

**Input**: Text utterance. For 'Strong Prompting', input is augmented with category-specific instructions (e.g., 'be emotionally expressive') passed via style descriptors or user messages depending on the model type.

**Output**: Audio waveform (TTS output).

## Scoring recipe

```python
wins = 0
total = 0
for audio_gen, audio_baseline in zip(generated_audios, baseline_audios):
    judge_response = judge_lalm.compare(audio_gen, audio_baseline)
    if judge_response == "preferred_gen":
        wins += 1
    total += 1
win_rate = (wins / total) * 100
wer = whisper_v3_large.transcribe(audio_gen).word_error_rate
mos = wav2vec2_mos_model.predict(audio_gen)
```

## Common pitfalls

- Win-rate scores are highly sensitive to the specific voice used by the evaluated TTS model; results can vary significantly across different voice clones.
- Judge parsing failures (due to incorrect JSON formatting or token limits in reasoning loops) must be filtered out, as they can artificially deflate win-rates.
- Performance gains from 'Strong Prompting' are substantial for some models, so comparisons must explicitly state whether basic or strong prompting was used.

## Evidence (verbatim from paper)

> In addition to the win-rate as described in Section 3.2, we follow standard practice by computing WER using Whisper-v3-large [27], and MOS scores are calculated using a fine-tuned wav2vec2.0 model [7].

## Citation

```bibtex
@misc{manku2025emergentts,
  title={EmergentTTS-Eval: Evaluating TTS Models on Complex Prosodic, Expressiveness, and Linguistic Challenges Using Model-as-a-Judge},
  author={Manku et al. (2025)},
  year={2025},
  note={arXiv:2505.23009}
}
```

- arXiv: 2505.23009

