Gemini Text-to-Speech
Generate natural-sounding speech from text using Gemini's TTS models through executable scripts with support for multiple voices and multi-speaker conversations.
When to Use This Skill
Use this skill when you need to:
- Convert text to natural speech
- Create audio for podcasts, audiobooks, or videos
- Generate multi-speaker conversations
- Stream audio for long content
- Choose from multiple voice options
- Create accessible audio content
- Generate voiceovers for presentations
- Batch convert text to audio files
Available Scripts
scripts/tts.py
Purpose: Convert text to speech using Gemini TTS models
When to use:
- Any text-to-speech conversion
- Multi-speaker conversation generation
- Streaming audio for long texts
- Voiceovers for content creation
- Accessible audio generation
Key parameters:
| Parameter |
Description |
Example |
text |
Text to convert (required) |
"Hello, world!" |
--voice, -v |
Voice name |
Kore |
--output, -o |
Base name for output file |
welcome |
--output-dir |
Output directory for audio |
audio/ |
--no-timestamp |
Disable auto timestamp |
Flag |
--model, -m |
TTS model |
gemini-2.5-flash-preview-tts |
--stream, -s |
Enable streaming |
Flag |
--speakers |
Multi-speaker mapping |
"Joe:Kore,Jane:Puck" |
Output: WAV audio file path
Workflows
Workflow 1: Basic Text-to-Speech
python scripts/tts.py "Hello, world! Have a wonderful day."
- Best for: Quick audio generation, simple messages
- Voice:
Kore (default, clear and professional)
- Output:
audio/tts_output_YYYYMMDD_HHMMSS.wav (auto timestamp)
Workflow 2: Choose Different Voice
python scripts/tts.py "Welcome to our podcast about technology trends" --voice Puck --output welcome
- Best for: Friendly, conversational content
- Voice options: Kore, Puck, Charon, Fenrir, Aoede, Zephyr, Sulafat
- Output:
audio/welcome_YYYYMMDD_HHMMSS.wav
Workflow 3: Multi-Speaker Conversation
python scripts/tts.py "TTS the following conversation:
Joe: How's it going today?
Jane: Not too bad, how about you?
Joe: I'm working on a new project.
Jane: Sounds exciting, tell me more!" --speakers "Joe:Kore,Jane:Puck" --output conversation
- Best for: Dialogues, interviews, role-playing content
- Format: Marked conversation with speaker names
- Script automatically routes text to appropriate voices
- Output:
audio/conversation_YYYYMMDD_HHMMSS.wav
Workflow 4: Long Content with Streaming
python scripts/tts.py "This is a very long text that would benefit from streaming..." --stream --output long-form
- Best for: Podcasts, audiobooks, long articles
- Streaming: Processes audio in chunks for long texts
- Output:
audio/long-form_YYYYMMDD_HHMMSS.wav
Workflow 5: Professional Voiceover
python scripts/tts.py "Welcome to our quarterly earnings presentation. Today we'll discuss our growth metrics and future plans." --voice Charon --output voiceover
- Best for: Corporate content, presentations, formal announcements
- Voice:
Charon (deep, authoritative)
- Use when: Professional, serious tone required
Workflow 6: Custom Output Directory
python scripts/tts.py "Save to specific folder." --output-dir ./my-projects/podcasts/ --output episode1
- Best for: Organized project structures
- Directory created automatically if it doesn't exist
- Output:
./my-projects/podcasts/episode1_YYYYMMDD_HHMMSS.wav
Workflow 7: Content Creation Pipeline (Text → Audio)
# 1. Generate script (gemini-text skill)
python skills/gemini-text/scripts/generate.py "Write a 2-minute podcast intro about sustainable energy"
# 2. Generate audio (this skill)
python scripts/tts.py "[Paste generated script]" --voice Fenrir --output podcast-intro
# 3. Use in video or podcast
- Best for: Podcasts, audiobooks, video narration
- Combines with: gemini-text for script generation
Workflow 8: Accessible Content
python scripts/tts.py "Welcome to our accessible website. This audio describes our main navigation options." --voice Aoede --output accessibility
- Best for: Web accessibility, screen reader alternatives
- Voice:
Aoede (melodic, pleasant)
- Use when: Making content accessible to visually impaired users
Workflow 9: Educational Content
python scripts/tts.py "Chapter 1: Introduction to Quantum Computing. Let's explore the fundamental principles..." --voice Zephyr --output chapter1
- Best for: Educational materials, tutorials, e-learning
- Voice:
Zephyr (light, airy)
- Combines well with: gemini-text for content generation
Workflow 10: Disable Timestamp
python scripts/tts.py "Fixed filename." --output my-audio --no-timestamp
- Best for: When you want complete control over filename
- Output:
audio/my-audio.wav (no timestamp)
- Use when: Generating files for specific naming schemes
Parameters Reference
Model Selection
| Model |
Quality |
Speed |
Best For |
gemini-2.5-flash-preview-tts |
Good |
Fast |
General use, high volume |
gemini-2.5-pro-preview-tts |
Higher |
Slower |
Premium content, voiceovers |
Voice Selection
| Voice |
Characteristics |
Best For |
| Kore |
Clear, professional |
Announcements, general purpose (default) |
| Puck |
Friendly, conversational |
Casual content, interviews |
| Charon |
Deep, authoritative |
Corporate, serious content |
| Fenrir |
Warm, expressive |
Storytelling, narratives |
| Aoede |
Melodic, pleasant |
Educational, accessibility |
| Zephyr |
Light, airy |
Gentle content, tutorials |
| Sulafat |
Neutral, balanced |
Documentaries, factual content |
Audio Format
| Specification |
Value |
| Format |
WAV (PCM) |
| Sample rate |
24000 Hz |
| Channels |
1 (mono) |
| Bit depth |
16-bit |
Token Limits
| Limit |
Type |
Description |
| 8,192 |
Input |
Maximum input text tokens |
| 16,384 |
Output |
Maximum output audio tokens |
Output Interpretation
Audio File
- Format: WAV (compatible with most players)
- Mono channel (single audio track)
- Sample rate: 24000 Hz (broadcast quality)
- Can be converted to MP3/AAC if needed
Multi-Speaker Files
- Single WAV file with multiple voices
- Voices separated by timing within file
- Use
--speakers parameter to map speakers to voices
Streaming Output
- Audio processed in chunks during generation
- Script shows "Streaming audio..." message
- Useful for very long texts or real-time applications
Common Issues
"google-genai not installed"
pip install google-genai
"Voice name not found"
- Check voice name spelling
- Use available voices: Kore, Puck, Charon, Fenrir, Aoede, Zephyr, Sulafat
- Voice names are case-sensitive
"No audio generated"
- Check text is not empty
- Verify text doesn't exceed token limit (8,192)
- Try shorter text segments
- Check API quota limits
"Multi-speaker format error"
- Format:
SpeakerName:VoiceName,Speaker2:Voice2
- Separate speakers with commas
- Use colon between speaker and voice
- Example:
"Joe:Kore,Jane:Puck,Host:Charon"
"Output file already exists"
- Script will overwrite existing files
- Change
--output filename to avoid conflicts
- Use unique names for batch generation
Audio quality issues
- Check input text for unusual characters
- Try different voice for better pronunciation
- Consider splitting long text into smaller segments
- Verify audio playback software compatibility
Best Practices
Voice Selection
- Kore: General purpose, clear articulation
- Puck: Conversational, engaging tone
- Charon: Professional, authoritative
- Fenrir: Emotional, storytelling
- Aoede: Soft, gentle for accessibility
- Zephyr: Educational, clear explanations
Text Preparation
- Use natural language and punctuation
- Include pauses with commas and periods
- Spell out difficult words if needed
- Break very long text into logical segments
- Add speaker labels for multi-speaker content
Performance Optimization
- Use streaming for very long texts
- Generate shorter segments for better control
- Use flash model for faster generation
- Batch process multiple files for efficiency
Quality Tips
- Test different voices for your content type
- Use appropriate pacing with punctuation
- Consider context when selecting voice
- Listen to output before final use
- Multi-speaker requires clear speaker labeling
Use Cases by Voice
| Voice |
Ideal Use Cases |
| Kore |
Announcements, navigation, general info |
| Puck |
Podcasts, interviews, casual content |
| Charon |
Corporate, news, formal presentations |
| Fenrir |
Audiobooks, stories, emotional content |
| Aoede |
Accessibility, educational, gentle content |
| Zephyr |
Tutorials, explanations, guides |
| Sulafat |
Documentaries, factual presentations |
Related Skills
- gemini-text: Generate scripts and text for TTS
- gemini-image: Create visuals to accompany audio
- gemini-batch: Process multiple TTS requests efficiently
- gemini-files: Upload audio files for processing
Quick Reference
# Basic
python scripts/tts.py "Your text here"
# Custom voice
python scripts/tts.py "Your text" --voice Puck --output audio.wav
# Multi-speaker
python scripts/tts.py "Joe: Hi. Jane: Hello!" --speakers "Joe:Kore,Jane:Puck"
# Streaming
python scripts/tts.py "Long text..." --stream --output long.wav
# Professional
python scripts/tts.py "Corporate announcement" --voice Charon
Reference
1---2name: gemini-tts3description: Generate speech from text using Google Gemini TTS models via scripts/. Use for text-to-speech, audio generation, voice synthesis, multi-speaker conversations, and creating audio content. Supports multiple voices and streaming. Triggers on "text to speech", "TTS", "generate audio", "voice synthesis", "speak this text".4license: MIT5---6
7# Gemini Text-to-Speech
8
9Generate natural-sounding speech from text using Gemini's TTS models through executable scripts with support for multiple voices and multi-speaker conversations.
10
11## When to Use This Skill
12
13Use this skill when you need to:
14- Convert text to natural speech
15- Create audio for podcasts, audiobooks, or videos
16- Generate multi-speaker conversations
17- Stream audio for long content
18- Choose from multiple voice options
19- Create accessible audio content
20- Generate voiceovers for presentations
21- Batch convert text to audio files
22
23## Available Scripts
24
25### scripts/tts.py
26**Purpose**: Convert text to speech using Gemini TTS models
27
28**When to use**:
29- Any text-to-speech conversion
30- Multi-speaker conversation generation
31- Streaming audio for long texts
32- Voiceovers for content creation
33- Accessible audio generation
34
35**Key parameters**:
36| Parameter | Description | Example |
37|-----------|-------------|---------|
38| `text` | Text to convert (required) | `"Hello, world!"` |
39| `--voice`, `-v` | Voice name | `Kore` |
40| `--output`, `-o` | Base name for output file | `welcome` |
41| `--output-dir` | Output directory for audio | `audio/` |
42| `--no-timestamp` | Disable auto timestamp | Flag |
43| `--model`, `-m` | TTS model | `gemini-2.5-flash-preview-tts` |
44| `--stream`, `-s` | Enable streaming | Flag |
45| `--speakers` | Multi-speaker mapping | `"Joe:Kore,Jane:Puck"` |
46
47**Output**: WAV audio file path
48
49## Workflows
50
51### Workflow 1: Basic Text-to-Speech
52```bash
53python scripts/tts.py "Hello, world! Have a wonderful day."
54```
55- Best for: Quick audio generation, simple messages
56- Voice: `Kore` (default, clear and professional)
57- Output: `audio/tts_output_YYYYMMDD_HHMMSS.wav` (auto timestamp)
58
59### Workflow 2: Choose Different Voice
60```bash
61python scripts/tts.py "Welcome to our podcast about technology trends" --voice Puck --output welcome
62```
63- Best for: Friendly, conversational content
64- Voice options: Kore, Puck, Charon, Fenrir, Aoede, Zephyr, Sulafat
65- Output: `audio/welcome_YYYYMMDD_HHMMSS.wav`
66
67### Workflow 3: Multi-Speaker Conversation
68```bash
69python scripts/tts.py "TTS the following conversation:
70Joe: How's it going today?
71Jane: Not too bad, how about you?
72Joe: I'm working on a new project.
73Jane: Sounds exciting, tell me more!" --speakers "Joe:Kore,Jane:Puck" --output conversation
74```
75- Best for: Dialogues, interviews, role-playing content
76- Format: Marked conversation with speaker names
77- Script automatically routes text to appropriate voices
78- Output: `audio/conversation_YYYYMMDD_HHMMSS.wav`
79
80### Workflow 4: Long Content with Streaming
81```bash
82python scripts/tts.py "This is a very long text that would benefit from streaming..." --stream --output long-form
83```
84- Best for: Podcasts, audiobooks, long articles
85- Streaming: Processes audio in chunks for long texts
86- Output: `audio/long-form_YYYYMMDD_HHMMSS.wav`
87
88### Workflow 5: Professional Voiceover
89```bash
90python scripts/tts.py "Welcome to our quarterly earnings presentation. Today we'll discuss our growth metrics and future plans." --voice Charon --output voiceover
91```
92- Best for: Corporate content, presentations, formal announcements
93- Voice: `Charon` (deep, authoritative)
94- Use when: Professional, serious tone required
95
96### Workflow 6: Custom Output Directory
97```bash
98python scripts/tts.py "Save to specific folder." --output-dir ./my-projects/podcasts/ --output episode1
99```
100- Best for: Organized project structures
101- Directory created automatically if it doesn't exist
102- Output: `./my-projects/podcasts/episode1_YYYYMMDD_HHMMSS.wav`
103
104### Workflow 7: Content Creation Pipeline (Text → Audio)
105```bash
106# 1. Generate script (gemini-text skill)
107python skills/gemini-text/scripts/generate.py "Write a 2-minute podcast intro about sustainable energy"
108
109# 2. Generate audio (this skill)
110python scripts/tts.py "[Paste generated script]" --voice Fenrir --output podcast-intro
111
112# 3. Use in video or podcast
113```
114- Best for: Podcasts, audiobooks, video narration
115- Combines with: gemini-text for script generation
116
117### Workflow 8: Accessible Content
118```bash
119python scripts/tts.py "Welcome to our accessible website. This audio describes our main navigation options." --voice Aoede --output accessibility
120```
121- Best for: Web accessibility, screen reader alternatives
122- Voice: `Aoede` (melodic, pleasant)
123- Use when: Making content accessible to visually impaired users
124
125### Workflow 9: Educational Content
126```bash
127python scripts/tts.py "Chapter 1: Introduction to Quantum Computing. Let's explore the fundamental principles..." --voice Zephyr --output chapter1
128```
129- Best for: Educational materials, tutorials, e-learning
130- Voice: `Zephyr` (light, airy)
131- Combines well with: gemini-text for content generation
132
133### Workflow 10: Disable Timestamp
134```bash
135python scripts/tts.py "Fixed filename." --output my-audio --no-timestamp
136```
137- Best for: When you want complete control over filename
138- Output: `audio/my-audio.wav` (no timestamp)
139- Use when: Generating files for specific naming schemes
140
141## Parameters Reference
142
143### Model Selection
144
145| Model | Quality | Speed | Best For |
146|-------|---------|-------|----------|
147| `gemini-2.5-flash-preview-tts` | Good | Fast | General use, high volume |
148| `gemini-2.5-pro-preview-tts` | Higher | Slower | Premium content, voiceovers |
149
150### Voice Selection
151
152| Voice | Characteristics | Best For |
153|-------|----------------|----------|
154| **Kore** | Clear, professional | Announcements, general purpose (default) |
155| **Puck** | Friendly, conversational | Casual content, interviews |
156| **Charon** | Deep, authoritative | Corporate, serious content |
157| **Fenrir** | Warm, expressive | Storytelling, narratives |
158| **Aoede** | Melodic, pleasant | Educational, accessibility |
159| **Zephyr** | Light, airy | Gentle content, tutorials |
160| **Sulafat** | Neutral, balanced | Documentaries, factual content |
161
162### Audio Format
163
164| Specification | Value |
165|--------------|-------|
166| Format | WAV (PCM) |
167| Sample rate | 24000 Hz |
168| Channels | 1 (mono) |
169| Bit depth | 16-bit |
170
171### Token Limits
172
173| Limit | Type | Description |
174|-------|------|-------------|
175| 8,192 | Input | Maximum input text tokens |
176| 16,384 | Output | Maximum output audio tokens |
177
178## Output Interpretation
179
180### Audio File
181- Format: WAV (compatible with most players)
182- Mono channel (single audio track)
183- Sample rate: 24000 Hz (broadcast quality)
184- Can be converted to MP3/AAC if needed
185
186### Multi-Speaker Files
187- Single WAV file with multiple voices
188- Voices separated by timing within file
189- Use `--speakers` parameter to map speakers to voices
190
191### Streaming Output
192- Audio processed in chunks during generation
193- Script shows "Streaming audio..." message
194- Useful for very long texts or real-time applications
195
196## Common Issues
197
198### "google-genai not installed"
199```bash
200pip install google-genai
201```
202
203### "Voice name not found"
204- Check voice name spelling
205- Use available voices: Kore, Puck, Charon, Fenrir, Aoede, Zephyr, Sulafat
206- Voice names are case-sensitive
207
208### "No audio generated"
209- Check text is not empty
210- Verify text doesn't exceed token limit (8,192)
211- Try shorter text segments
212- Check API quota limits
213
214### "Multi-speaker format error"
215- Format: `SpeakerName:VoiceName,Speaker2:Voice2`
216- Separate speakers with commas
217- Use colon between speaker and voice
218- Example: `"Joe:Kore,Jane:Puck,Host:Charon"`
219
220### "Output file already exists"
221- Script will overwrite existing files
222- Change `--output` filename to avoid conflicts
223- Use unique names for batch generation
224
225### Audio quality issues
226- Check input text for unusual characters
227- Try different voice for better pronunciation
228- Consider splitting long text into smaller segments
229- Verify audio playback software compatibility
230
231## Best Practices
232
233### Voice Selection
234- **Kore**: General purpose, clear articulation
235- **Puck**: Conversational, engaging tone
236- **Charon**: Professional, authoritative
237- **Fenrir**: Emotional, storytelling
238- **Aoede**: Soft, gentle for accessibility
239- **Zephyr**: Educational, clear explanations
240
241### Text Preparation
242- Use natural language and punctuation
243- Include pauses with commas and periods
244- Spell out difficult words if needed
245- Break very long text into logical segments
246- Add speaker labels for multi-speaker content
247
248### Performance Optimization
249- Use streaming for very long texts
250- Generate shorter segments for better control
251- Use flash model for faster generation
252- Batch process multiple files for efficiency
253
254### Quality Tips
255- Test different voices for your content type
256- Use appropriate pacing with punctuation
257- Consider context when selecting voice
258- Listen to output before final use
259- Multi-speaker requires clear speaker labeling
260
261### Use Cases by Voice
262
263| Voice | Ideal Use Cases |
264|-------|-----------------|
265| Kore | Announcements, navigation, general info |
266| Puck | Podcasts, interviews, casual content |
267| Charon | Corporate, news, formal presentations |
268| Fenrir | Audiobooks, stories, emotional content |
269| Aoede | Accessibility, educational, gentle content |
270| Zephyr | Tutorials, explanations, guides |
271| Sulafat | Documentaries, factual presentations |
272
273## Related Skills
274
275- **gemini-text**: Generate scripts and text for TTS
276- **gemini-image**: Create visuals to accompany audio
277- **gemini-batch**: Process multiple TTS requests efficiently
278- **gemini-files**: Upload audio files for processing
279
280## Quick Reference
281
282```bash
283# Basic
284python scripts/tts.py "Your text here"
285
286# Custom voice
287python scripts/tts.py "Your text" --voice Puck --output audio.wav
288
289# Multi-speaker
290python scripts/tts.py "Joe: Hi. Jane: Hello!" --speakers "Joe:Kore,Jane:Puck"
291
292# Streaming
293python scripts/tts.py "Long text..." --stream --output long.wav
294
295# Professional
296python scripts/tts.py "Corporate announcement" --voice Charon
297```
298
299## Reference
300
301- See `references/voices.md` for complete voice documentation
302- Get API key: https://aistudio.google.com/apikey
303- Documentation: https://ai.google.dev/gemini-api/docs/text-to-speech
304- Sample rate: 24000 Hz standard for most applications