Nvtts Eval

This evaluation protocol assesses the capability of zero-shot text-to-speech models to synthesize nonverbal vocalizations (NVs) like breathing, laughter, coughing, and sighs alongside emotional speech. It measures speech intelligibility, speaker and emotion fidelity, acoustic quality, and the precise alignment of generated NVs with reference audio. Use when the user wants to benchmark on NVTTS, or asks about evaluating this task. Reports WER.

qhjqhj00 3883178 4.4 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/nvtts-eval commit 38831786aa

Frequently asked questions

npx skillmds add qhjqhj00/nvtts-eval