Clip Tts Eval

Evaluates the naturalness and quality of synthesized speech across single-speaker, multi-speaker, and multi-emotion TTS models. It measures how closely generated audio matches human ground truth in terms of overall speech quality and emotional similarity. Use when the user wants to benchmark on Baker, AISHELL3, LJSpeech, LibriTTS, Emotional Speech Dataset (ESD), or asks about evaluating this task. Reports MOS.

qhjqhj00 f00436f 3.6 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/clip-tts-eval commit f00436f05a

Frequently asked questions

npx skillmds add qhjqhj00/clip-tts-eval