Create Audio
Generate high-quality audio from text using local TTS models that run on your CPU.
Quick Start
Basic usage:
# Generate with default voice
scripts/create_audio.py --text "Hello world" --output hello.wav
# Use specific voice
scripts/create_audio.py --voice marius --text "Bonjour" --output greeting.wav
# From stdin
echo "Hello from stdin" | scripts/create_audio.py --output output.wav
Available Providers
Pocket TTS (Default)
- Provider ID:
pocket-tts
- Model: Kyutai Labs Pocket TTS (100M parameters)
- Performance: ~6x real-time on MacBook Air M4, ~200ms first chunk
- Requirements: CPU only (no GPU needed)
- Language: English
- Installation:
pip install pocket-tts or uvx pocket-tts
Built-in Voices:
alba (default) - Female voice
marius - Male voice
javert, jean, fantine, cosette, eponine, azelma - Character voices
MLX-Audio Providers (Apple Silicon Optimized)
MLX-Audio provides 7 different TTS models optimized for Apple Silicon (M1/M2/M3/M4). Each model is registered as a separate provider.
Installation: pip install mlx-audio
Requirements: Apple Silicon Mac (M1/M2/M3/M4), Python 3.9+
1. Kokoro (mlx-audio-kokoro)
- Languages: English, Japanese, Chinese, French, Spanish, Italian, Portuguese, Hindi
- Voices: 10 built-in voices + voice cloning
- American:
af_heart, af_bella, af_nova, af_sky, am_adam, am_echo
- British:
bf_alice, bf_emma, bm_daniel, bm_george
- Best for: Multilingual content, fast generation
- Features: Speed control, voice cloning
2. CSM (mlx-audio-csm)
- Languages: English
- Features: Conversational Speech Model with voice cloning
- Best for: Natural conversations, voice cloning
3. Dia (mlx-audio-dia)
- Languages: English
- Best for: Dialogue-focused content
4. OuteTTS (mlx-audio-oute)
- Languages: English
- Best for: Efficient, fast generation
5. Spark (mlx-audio-spark)
- Languages: English, Chinese
- Best for: Bilingual content
6. Chatterbox (mlx-audio-chatterbox)
- Languages: 16 languages (en, es, fr, de, it, pt, pl, tr, ru, nl, cs, ar, zh, ja, hu, ko)
- Best for: Expressive multilingual content
7. Soprano (mlx-audio-soprano)
- Languages: English
- Best for: High-quality English TTS
ElevenLabs (Cloud API)
Premium cloud-based TTS with the highest quality voices and massive voice library.
Provider ID: elevenlabs
- Languages: 32 languages (en, es, fr, de, it, pt, pl, uk, nl, sv, da, fi, no, cs, sk, el, ro, bg, hr, sr, mk, lv, lt, et, sl, hu, tr, vi, ar, hi, bn, ta, ko, zh, ja, etc.)
- Voices: 10,000+ voices in Voice Library
- Features: Professional voice cloning, voice design, emotional control
- Requirements: API key (ELEVEN_API_KEY environment variable)
- Pricing: Free tier available, paid plans for production
- Best for: Production-quality voiceovers, professional content
- Installation:
pip install elevenlabs
Popular Voices:
rachel, drew, clyde, paul, domi - English
antoni, thomas, charlie, george - Various accents
emily, elli, charlotte, alice, matilda - Female voices
- Or use any custom voice_id from the Voice Library
Coqui TTS Providers (Open Source)
Coqui TTS provides 4 different models for various use cases - all free and open source.
Installation: pip install TTS
Requirements: Python 3.10-3.14, CPU or GPU
1. XTTS v2 (coqui-xtts_v2)
- Languages: 17 languages (en, es, fr, de, it, pt, pl, tr, ru, nl, cs, ar, zh-cn, ja, hu, ko, hi)
- Features: Best quality, voice cloning with 3-10 seconds of audio
- Best for: Multilingual projects with voice cloning needs
- Streaming: <200ms latency
2. VITS (coqui-vits)
- Languages: English
- Features: Fast, single speaker
- Best for: Quick English TTS without voice customization
3. YourTTS (coqui-yourtts)
- Languages: English, French, Portuguese
- Features: Multilingual voice cloning, multi-speaker
- Best for: Voice cloning in multiple languages
4. Bark (coqui-bark)
- Languages: 13 languages (en, de, es, fr, hi, it, ja, ko, pl, pt, ru, tr, zh)
- Features: Highly expressive, multi-speaker
- Best for: Expressive, natural-sounding speech
Core Capabilities
1. Basic Text-to-Speech
Generate audio from any text input:
scripts/create_audio.py \
--text "Your text here" \
--output audio.wav
2. Voice Selection
Choose from built-in voices:
scripts/create_audio.py \
--voice marius \
--text "Hello in a different voice" \
--output voice_demo.wav
3. Voice Cloning
Clone any voice by providing a WAV file:
scripts/create_audio.py \
--voice /path/to/reference_voice.wav \
--text "This will sound like the reference" \
--output cloned.wav
4. Provider Selection
Switch between TTS providers (extensible architecture):
# List all available providers
scripts/create_audio.py --list
# Use specific provider
scripts/create_audio.py \
--provider pocket-tts \
--text "Hello" \
--output audio.wav
5. Pipeline Integration
Integrate with other tools via stdin/stdout:
# From file
cat script.txt | scripts/create_audio.py --output narration.wav
# From command output
echo "This is generated text" | scripts/create_audio.py -o result.wav
# Chain with other tools
cat content.md | sed 's/#//g' | scripts/create_audio.py -o doc_audio.wav
Common Use Cases
Voiceover Generation
# Pocket TTS voiceover
scripts/create_audio.py \
--voice alba \
--text "Welcome to this tutorial about..." \
--output intro_voiceover.wav
# MLX-Audio Kokoro (multilingual)
scripts/create_audio.py \
--provider mlx-audio-kokoro \
--voice af_heart \
--text "Welcome to this tutorial" \
--output intro_mlx.wav
Multi-Voice Content
# Character 1 (Pocket TTS)
scripts/create_audio.py --voice jean --text "First character speaks" -o char1.wav
# Character 2 (MLX-Audio)
scripts/create_audio.py --provider mlx-audio-kokoro --voice bm_george --text "Second character responds" -o char2.wav
Multilingual Content
# English
scripts/create_audio.py \
--provider mlx-audio-kokoro \
--voice af_heart \
--params '{"lang": "en"}' \
--text "Hello world" \
--output en.wav
# Japanese
scripts/create_audio.py \
--provider mlx-audio-kokoro \
--params '{"lang": "ja"}' \
--text "こんにちは世界" \
--output ja.wav
# French
scripts/create_audio.py \
--provider mlx-audio-kokoro \
--params '{"lang": "fr"}' \
--text "Bonjour le monde" \
--output fr.wav
Custom Voice Branding
# Pocket TTS voice cloning
scripts/create_audio.py \
--voice /path/to/your_voice_sample.wav \
--text "All content in my voice" \
--output branded_audio.wav
# MLX-Audio CSM voice cloning
scripts/create_audio.py \
--provider mlx-audio-csm \
--voice custom \
--params '{"custom_voice_path": "/path/to/voice.wav"}' \
--text "Cloned voice with CSM" \
--output csm_cloned.wav
Speed Control
# Faster speech (1.5x)
scripts/create_audio.py \
--provider mlx-audio-kokoro \
--voice af_nova \
--params '{"speed": 1.5}' \
--text "This will be faster" \
--output fast.wav
# Slower speech (0.8x)
scripts/create_audio.py \
--provider mlx-audio-kokoro \
--voice bf_emma \
--params '{"speed": 0.8}' \
--text "This will be slower" \
--output slow.wav
ElevenLabs Production Quality
# Set API key first
export ELEVEN_API_KEY="your_api_key_here"
# Use default voice
scripts/create_audio.py \
--provider elevenlabs \
--voice rachel \
--text "Professional quality voiceover" \
--output professional.wav
# Use custom voice from Voice Library
scripts/create_audio.py \
--provider elevenlabs \
--voice custom \
--params '{"voice_id": "your_voice_id_here"}' \
--text "Custom voice from library" \
--output custom.wav
# Advanced settings (stability, similarity, style)
scripts/create_audio.py \
--provider elevenlabs \
--voice rachel \
--params '{"stability": 60, "similarity_boost": 80, "style": 20, "speed": 1.1}' \
--text "Fine-tuned voice settings" \
--output tuned.wav
# Multilingual with language enforcement
scripts/create_audio.py \
--provider elevenlabs \
--voice antoni \
--params '{"language_code": "es"}' \
--text "Hola mundo" \
--output spanish.wav
Coqui TTS Voice Cloning
# XTTS v2 voice cloning (best quality)
scripts/create_audio.py \
--provider coqui-xtts_v2 \
--voice custom \
--params '{"speaker_wav": "/path/to/reference.wav", "language": "en"}' \
--text "Cloned voice with XTTS v2" \
--output xtts_cloned.wav
# YourTTS multilingual cloning
scripts/create_audio.py \
--provider coqui-yourtts \
--voice custom \
--params '{"speaker_wav": "/path/to/voice.wav", "language": "fr"}' \
--text "Bonjour le monde" \
--output yourtts_fr.wav
# VITS fast English
scripts/create_audio.py \
--provider coqui-vits \
--text "Fast English synthesis" \
--output vits.wav
# Bark expressive speech
scripts/create_audio.py \
--provider coqui-bark \
--voice custom \
--params '{"speaker_idx": 0}' \
--text "Very expressive and natural sounding" \
--output bark.wav
Adding New TTS Providers
The skill uses an extensible provider architecture. To add a new provider:
- Create provider class in
scripts/<provider>_provider.py:
from tts_provider import TTSProvider
class NewProvider(TTSProvider):
@property
def name(self) -> str:
return "new-provider"
@property
def supported_voices(self) -> List[str]:
return ["voice1", "voice2"]
# Implement other required methods...
- Register provider in
scripts/provider_registry.py:
from new_provider import NewProvider
ProviderRegistry.register(NewProvider)
- Use immediately:
scripts/create_audio.py --provider new-provider --voice voice1 --text "Test"
See scripts/pocket_tts_provider.py for a complete implementation example.
Advanced Usage
Provider-Specific Parameters
Pass provider-specific parameters as JSON:
scripts/create_audio.py \
--voice custom \
--params '{"custom_voice_path": "/path/to/voice.wav"}' \
--text "Using custom voice" \
--output result.wav
Batch Processing
Generate multiple audio files:
# Simple loop
for voice in alba marius jean; do
scripts/create_audio.py \
--voice $voice \
--text "Sample text" \
--output "${voice}_sample.wav"
done
Installation Requirements
Pocket TTS
# Via pip
pip install pocket-tts
# Or use uvx (no installation needed)
uvx pocket-tts generate --text "Test"
Dependencies:
- Python 3.10+ (supports up to 3.14)
- PyTorch 2.5+ (CPU version)
- scipy (for WAV file writing)
Troubleshooting
Model not loading?
- First run takes longer (downloads model)
- Model is cached for subsequent uses
- Check internet connection for initial download
Import errors?
pip install pocket-tts scipy
Voice file not found?
- Ensure WAV files are valid audio files
- Use absolute paths for custom voices
- Check file permissions
Technical Details
Architecture
create_audio.py (CLI)
↓
provider_registry.py (Provider management)
↓
tts_provider.py (Base interface)
↓
[pocket_tts_provider.py, future_provider.py, ...]
Provider Interface
All providers implement:
name: Provider identifier
supported_voices: List of available voices
default_voice: Fallback voice
generate_audio(): Core synthesis method
get_provider_params(): Provider-specific options
Output Format
- Format: WAV (uncompressed)
- Sample Rate: Provider-specific (Pocket TTS: 16kHz)
- Channels: Mono
- Bit Depth: 16-bit PCM
Resources
Scripts
create_audio.py - Main CLI tool
tts_provider.py - Base provider interface
pocket_tts_provider.py - Pocket TTS implementation
provider_registry.py - Provider management system
References
Pocket TTS:
MLX-Audio:
ElevenLabs:
Coqui TTS:
1---2name: create-audio3description: Generate audio from text using 13 TTS providers (local + cloud). Use when user wants to create audio files, convert text to speech, generate voiceovers, create audio with different voices, use voice cloning, multilingual TTS, or mentions /create-audio command. Supports Pocket TTS (CPU, 8 voices), MLX-Audio (Apple Silicon, 7 models, 50+ voices), ElevenLabs (cloud API, 32 languages, 10k+ voices), and Coqui TTS (open source, 4 models, voice cloning). Includes 32+ languages, voice cloning, speed control, and both local and cloud options.4---56# Create Audio78Generate high-quality audio from text using local TTS models that run on your CPU.910## Quick Start1112**Basic usage:**13```bash14# Generate with default voice15scripts/create_audio.py --text "Hello world" --output hello.wav1617# Use specific voice18scripts/create_audio.py --voice marius --text "Bonjour" --output greeting.wav1920# From stdin21echo "Hello from stdin" | scripts/create_audio.py --output output.wav22```2324## Available Providers2526### Pocket TTS (Default)27- **Provider ID:** `pocket-tts`28- **Model:** Kyutai Labs Pocket TTS (100M parameters)29- **Performance:** ~6x real-time on MacBook Air M4, ~200ms first chunk30- **Requirements:** CPU only (no GPU needed)31- **Language:** English32- **Installation:** `pip install pocket-tts` or `uvx pocket-tts`3334**Built-in Voices:**35- `alba` (default) - Female voice36- `marius` - Male voice37- `javert`, `jean`, `fantine`, `cosette`, `eponine`, `azelma` - Character voices3839### MLX-Audio Providers (Apple Silicon Optimized)4041MLX-Audio provides **7 different TTS models** optimized for Apple Silicon (M1/M2/M3/M4). Each model is registered as a separate provider.4243**Installation:** `pip install mlx-audio`44**Requirements:** Apple Silicon Mac (M1/M2/M3/M4), Python 3.9+4546#### 1. Kokoro (`mlx-audio-kokoro`)47- **Languages:** English, Japanese, Chinese, French, Spanish, Italian, Portuguese, Hindi48- **Voices:** 10 built-in voices + voice cloning49 - American: `af_heart`, `af_bella`, `af_nova`, `af_sky`, `am_adam`, `am_echo`50 - British: `bf_alice`, `bf_emma`, `bm_daniel`, `bm_george`51- **Best for:** Multilingual content, fast generation52- **Features:** Speed control, voice cloning5354#### 2. CSM (`mlx-audio-csm`)55- **Languages:** English56- **Features:** Conversational Speech Model with voice cloning57- **Best for:** Natural conversations, voice cloning5859#### 3. Dia (`mlx-audio-dia`)60- **Languages:** English61- **Best for:** Dialogue-focused content6263#### 4. OuteTTS (`mlx-audio-oute`)64- **Languages:** English65- **Best for:** Efficient, fast generation6667#### 5. Spark (`mlx-audio-spark`)68- **Languages:** English, Chinese69- **Best for:** Bilingual content7071#### 6. Chatterbox (`mlx-audio-chatterbox`)72- **Languages:** 16 languages (en, es, fr, de, it, pt, pl, tr, ru, nl, cs, ar, zh, ja, hu, ko)73- **Best for:** Expressive multilingual content7475#### 7. Soprano (`mlx-audio-soprano`)76- **Languages:** English77- **Best for:** High-quality English TTS7879### ElevenLabs (Cloud API)8081Premium cloud-based TTS with the highest quality voices and massive voice library.8283**Provider ID:** `elevenlabs`8485- **Languages:** 32 languages (en, es, fr, de, it, pt, pl, uk, nl, sv, da, fi, no, cs, sk, el, ro, bg, hr, sr, mk, lv, lt, et, sl, hu, tr, vi, ar, hi, bn, ta, ko, zh, ja, etc.)86- **Voices:** 10,000+ voices in Voice Library87- **Features:** Professional voice cloning, voice design, emotional control88- **Requirements:** API key (ELEVEN_API_KEY environment variable)89- **Pricing:** Free tier available, paid plans for production90- **Best for:** Production-quality voiceovers, professional content91- **Installation:** `pip install elevenlabs`9293**Popular Voices:**94- `rachel`, `drew`, `clyde`, `paul`, `domi` - English95- `antoni`, `thomas`, `charlie`, `george` - Various accents96- `emily`, `elli`, `charlotte`, `alice`, `matilda` - Female voices97- Or use any custom voice_id from the Voice Library9899### Coqui TTS Providers (Open Source)100101Coqui TTS provides 4 different models for various use cases - all free and open source.102103**Installation:** `pip install TTS`104**Requirements:** Python 3.10-3.14, CPU or GPU105106#### 1. XTTS v2 (`coqui-xtts_v2`)107- **Languages:** 17 languages (en, es, fr, de, it, pt, pl, tr, ru, nl, cs, ar, zh-cn, ja, hu, ko, hi)108- **Features:** Best quality, voice cloning with 3-10 seconds of audio109- **Best for:** Multilingual projects with voice cloning needs110- **Streaming:** <200ms latency111112#### 2. VITS (`coqui-vits`)113- **Languages:** English114- **Features:** Fast, single speaker115- **Best for:** Quick English TTS without voice customization116117#### 3. YourTTS (`coqui-yourtts`)118- **Languages:** English, French, Portuguese119- **Features:** Multilingual voice cloning, multi-speaker120- **Best for:** Voice cloning in multiple languages121122#### 4. Bark (`coqui-bark`)123- **Languages:** 13 languages (en, de, es, fr, hi, it, ja, ko, pl, pt, ru, tr, zh)124- **Features:** Highly expressive, multi-speaker125- **Best for:** Expressive, natural-sounding speech126127## Core Capabilities128129### 1. Basic Text-to-Speech130131Generate audio from any text input:132133```bash134scripts/create_audio.py \135 --text "Your text here" \136 --output audio.wav137```138139### 2. Voice Selection140141Choose from built-in voices:142143```bash144scripts/create_audio.py \145 --voice marius \146 --text "Hello in a different voice" \147 --output voice_demo.wav148```149150### 3. Voice Cloning151152Clone any voice by providing a WAV file:153154```bash155scripts/create_audio.py \156 --voice /path/to/reference_voice.wav \157 --text "This will sound like the reference" \158 --output cloned.wav159```160161### 4. Provider Selection162163Switch between TTS providers (extensible architecture):164165```bash166# List all available providers167scripts/create_audio.py --list168169# Use specific provider170scripts/create_audio.py \171 --provider pocket-tts \172 --text "Hello" \173 --output audio.wav174```175176### 5. Pipeline Integration177178Integrate with other tools via stdin/stdout:179180```bash181# From file182cat script.txt | scripts/create_audio.py --output narration.wav183184# From command output185echo "This is generated text" | scripts/create_audio.py -o result.wav186187# Chain with other tools188cat content.md | sed 's/#//g' | scripts/create_audio.py -o doc_audio.wav189```190191## Common Use Cases192193### Voiceover Generation194```bash195# Pocket TTS voiceover196scripts/create_audio.py \197 --voice alba \198 --text "Welcome to this tutorial about..." \199 --output intro_voiceover.wav200201# MLX-Audio Kokoro (multilingual)202scripts/create_audio.py \203 --provider mlx-audio-kokoro \204 --voice af_heart \205 --text "Welcome to this tutorial" \206 --output intro_mlx.wav207```208209### Multi-Voice Content210```bash211# Character 1 (Pocket TTS)212scripts/create_audio.py --voice jean --text "First character speaks" -o char1.wav213214# Character 2 (MLX-Audio)215scripts/create_audio.py --provider mlx-audio-kokoro --voice bm_george --text "Second character responds" -o char2.wav216```217218### Multilingual Content219```bash220# English221scripts/create_audio.py \222 --provider mlx-audio-kokoro \223 --voice af_heart \224 --params '{"lang": "en"}' \225 --text "Hello world" \226 --output en.wav227228# Japanese229scripts/create_audio.py \230 --provider mlx-audio-kokoro \231 --params '{"lang": "ja"}' \232 --text "こんにちは世界" \233 --output ja.wav234235# French236scripts/create_audio.py \237 --provider mlx-audio-kokoro \238 --params '{"lang": "fr"}' \239 --text "Bonjour le monde" \240 --output fr.wav241```242243### Custom Voice Branding244```bash245# Pocket TTS voice cloning246scripts/create_audio.py \247 --voice /path/to/your_voice_sample.wav \248 --text "All content in my voice" \249 --output branded_audio.wav250251# MLX-Audio CSM voice cloning252scripts/create_audio.py \253 --provider mlx-audio-csm \254 --voice custom \255 --params '{"custom_voice_path": "/path/to/voice.wav"}' \256 --text "Cloned voice with CSM" \257 --output csm_cloned.wav258```259260### Speed Control261```bash262# Faster speech (1.5x)263scripts/create_audio.py \264 --provider mlx-audio-kokoro \265 --voice af_nova \266 --params '{"speed": 1.5}' \267 --text "This will be faster" \268 --output fast.wav269270# Slower speech (0.8x)271scripts/create_audio.py \272 --provider mlx-audio-kokoro \273 --voice bf_emma \274 --params '{"speed": 0.8}' \275 --text "This will be slower" \276 --output slow.wav277```278279### ElevenLabs Production Quality280```bash281# Set API key first282export ELEVEN_API_KEY="your_api_key_here"283284# Use default voice285scripts/create_audio.py \286 --provider elevenlabs \287 --voice rachel \288 --text "Professional quality voiceover" \289 --output professional.wav290291# Use custom voice from Voice Library292scripts/create_audio.py \293 --provider elevenlabs \294 --voice custom \295 --params '{"voice_id": "your_voice_id_here"}' \296 --text "Custom voice from library" \297 --output custom.wav298299# Advanced settings (stability, similarity, style)300scripts/create_audio.py \301 --provider elevenlabs \302 --voice rachel \303 --params '{"stability": 60, "similarity_boost": 80, "style": 20, "speed": 1.1}' \304 --text "Fine-tuned voice settings" \305 --output tuned.wav306307# Multilingual with language enforcement308scripts/create_audio.py \309 --provider elevenlabs \310 --voice antoni \311 --params '{"language_code": "es"}' \312 --text "Hola mundo" \313 --output spanish.wav314```315316### Coqui TTS Voice Cloning317```bash318# XTTS v2 voice cloning (best quality)319scripts/create_audio.py \320 --provider coqui-xtts_v2 \321 --voice custom \322 --params '{"speaker_wav": "/path/to/reference.wav", "language": "en"}' \323 --text "Cloned voice with XTTS v2" \324 --output xtts_cloned.wav325326# YourTTS multilingual cloning327scripts/create_audio.py \328 --provider coqui-yourtts \329 --voice custom \330 --params '{"speaker_wav": "/path/to/voice.wav", "language": "fr"}' \331 --text "Bonjour le monde" \332 --output yourtts_fr.wav333334# VITS fast English335scripts/create_audio.py \336 --provider coqui-vits \337 --text "Fast English synthesis" \338 --output vits.wav339340# Bark expressive speech341scripts/create_audio.py \342 --provider coqui-bark \343 --voice custom \344 --params '{"speaker_idx": 0}' \345 --text "Very expressive and natural sounding" \346 --output bark.wav347```348349## Adding New TTS Providers350351The skill uses an extensible provider architecture. To add a new provider:3523531. **Create provider class** in `scripts/<provider>_provider.py`:354```python355from tts_provider import TTSProvider356357class NewProvider(TTSProvider):358 @property359 def name(self) -> str:360 return "new-provider"361362 @property363 def supported_voices(self) -> List[str]:364 return ["voice1", "voice2"]365366 # Implement other required methods...367```3683692. **Register provider** in `scripts/provider_registry.py`:370```python371from new_provider import NewProvider372ProviderRegistry.register(NewProvider)373```3743753. **Use immediately:**376```bash377scripts/create_audio.py --provider new-provider --voice voice1 --text "Test"378```379380See `scripts/pocket_tts_provider.py` for a complete implementation example.381382## Advanced Usage383384### Provider-Specific Parameters385386Pass provider-specific parameters as JSON:387388```bash389scripts/create_audio.py \390 --voice custom \391 --params '{"custom_voice_path": "/path/to/voice.wav"}' \392 --text "Using custom voice" \393 --output result.wav394```395396### Batch Processing397398Generate multiple audio files:399400```bash401# Simple loop402for voice in alba marius jean; do403 scripts/create_audio.py \404 --voice $voice \405 --text "Sample text" \406 --output "${voice}_sample.wav"407done408```409410## Installation Requirements411412### Pocket TTS413```bash414# Via pip415pip install pocket-tts416417# Or use uvx (no installation needed)418uvx pocket-tts generate --text "Test"419```420421**Dependencies:**422- Python 3.10+ (supports up to 3.14)423- PyTorch 2.5+ (CPU version)424- scipy (for WAV file writing)425426## Troubleshooting427428**Model not loading?**429- First run takes longer (downloads model)430- Model is cached for subsequent uses431- Check internet connection for initial download432433**Import errors?**434```bash435pip install pocket-tts scipy436```437438**Voice file not found?**439- Ensure WAV files are valid audio files440- Use absolute paths for custom voices441- Check file permissions442443## Technical Details444445### Architecture446447```448create_audio.py (CLI)449 ↓450provider_registry.py (Provider management)451 ↓452tts_provider.py (Base interface)453 ↓454[pocket_tts_provider.py, future_provider.py, ...]455```456457### Provider Interface458459All providers implement:460- `name`: Provider identifier461- `supported_voices`: List of available voices462- `default_voice`: Fallback voice463- `generate_audio()`: Core synthesis method464- `get_provider_params()`: Provider-specific options465466### Output Format467468- **Format:** WAV (uncompressed)469- **Sample Rate:** Provider-specific (Pocket TTS: 16kHz)470- **Channels:** Mono471- **Bit Depth:** 16-bit PCM472473## Resources474475### Scripts476- `create_audio.py` - Main CLI tool477- `tts_provider.py` - Base provider interface478- `pocket_tts_provider.py` - Pocket TTS implementation479- `provider_registry.py` - Provider management system480481### References482483**Pocket TTS:**484- [Pocket TTS GitHub](https://github.com/kyutai-labs/pocket-tts)485- [Pocket TTS Blog](https://kyutai.org/blog/2026-01-13-pocket-tts)486- [Kyutai Labs](https://kyutai.org/tts)487488**MLX-Audio:**489- [MLX-Audio GitHub](https://github.com/Blaizzy/mlx-audio)490- [MLX-Audio PyPI](https://pypi.org/project/mlx-audio/)491- [MLX-Audio Tutorial](https://blog.johnys.io/local-text-to-speech-tts-and-voice-cloning-with-mlx-audio/)492493**ElevenLabs:**494- [ElevenLabs API Documentation](https://elevenlabs.io/docs/api-reference/text-to-speech/convert)495- [ElevenLabs Voice Library](https://elevenlabs.io/app/voice-library)496- [Get API Key](https://elevenlabs.io/app/settings/api-keys)497498**Coqui TTS:**499- [Coqui TTS GitHub](https://github.com/coqui-ai/TTS)500- [Coqui TTS Documentation](https://docs.coqui.ai/)501- [Coqui TTS PyPI](https://pypi.org/project/TTS/)