Capspeech Eval

This benchmark evaluates text-to-speech models on generating high-fidelity, intelligible speech conditioned on free-form natural language style captions. It probes the model's ability to control intrinsic speaker traits, expressive styles, accents, emotions, and integrate non-verbal sound events across diverse real-world scenarios. Use when the user wants to benchmark on CapTTS, EmoCapTTS, AccCapTTS, CapTTS-SE, AgentTTS, or asks about evaluating this task. Reports binary_correctness.

qhjqhj00 bc18ea9 3.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/capspeech-eval commit bc18ea9f68

Frequently asked questions

npx skillmds add qhjqhj00/capspeech-eval