Voice Generation Skill
Generate realistic speech using AI (Google Gemini TTS, ElevenLabs, OpenAI TTS).
Prerequisites
At least one API key is required:
GOOGLE_API_KEY - For Google Gemini TTS (same key as video/image/music) ✅
ELEVENLABS_API_KEY - For ElevenLabs high-quality voice synthesis
OPENAI_API_KEY - For OpenAI TTS voices
Available APIs
Google Gemini TTS (Recommended - Same API Key)
- Best for: Podcasts, dialogues, audiobooks with style control
- Voices: 30 voices with natural language style control
- Multi-speaker: Up to 2 speakers for dialogues ✅
- Languages: 24 languages (auto-detected)
- Features: Control style, accent, pace via prompts
- Output: 24kHz WAV
- API Key: Same
GOOGLE_API_KEY as video/image/music ✅
ElevenLabs (Best Quality)
- Best for: Natural-sounding voices, voice cloning, long-form content
- Voices: 100+ pre-made voices + custom voice cloning
- Languages: 29+ languages
- Models: Eleven Multilingual v2, Eleven Turbo v2
OpenAI TTS (Simplest)
- Best for: Quick, reliable text-to-speech with consistent quality
- Voices: alloy, echo, fable, onyx, nova, shimmer
- Models: tts-1 (fast), tts-1-hd (high quality)
- Output: MP3, Opus, AAC, FLAC
Workflow
Step 1: Understand the Request
Parse the user's voice request for:
- Text content: What should be spoken?
- Voice type: Male, female, specific character?
- Tone: Professional, casual, dramatic, cheerful?
- Use case: Narration, voiceover, audiobook, notification?
- Language: English, Spanish, other?
- Speed: Normal, slow, fast?
Step 2: Select Voice and API
Choose based on requirements:
| Use Case |
Recommended API |
Reason |
| Default / Same key as video |
Gemini TTS |
Same GOOGLE_API_KEY ✅ |
| Multi-speaker dialogue |
Gemini TTS |
Up to 2 speakers built-in |
| Style/accent control |
Gemini TTS |
Natural language prompts |
| Voice cloning |
ElevenLabs |
Only API with cloning |
| 100+ voice options |
ElevenLabs |
Widest selection |
| Audiobook/podcast |
ElevenLabs or Gemini |
Both excellent for long content |
| Quick narration |
OpenAI TTS |
Fast, reliable |
| Budget-conscious |
OpenAI TTS |
Lower cost |
Step 3: Prepare the Text
Optimize text for speech:
- Add pauses: Use commas, periods for natural rhythm
- Spell out numbers: "1,234" → "one thousand two hundred thirty-four" (if needed)
- Handle acronyms: "NASA" vs "N.A.S.A." depending on pronunciation
- Mark emphasis: Some APIs support emphasis markers
Example transformation:
- Original: "The Q4 2024 results show a 15% YoY increase."
- Optimized: "The Q4 2024 results show a fifteen percent year-over-year increase."
Step 4: Generate the Audio
Execute the appropriate script from ${CLAUDE_PLUGIN_ROOT}/skills/voice-generation/scripts/:
For Google Gemini TTS (single speaker):
python3 ${CLAUDE_PLUGIN_ROOT}/skills/voice-generation/scripts/gemini_tts.py \
--text "Welcome to our podcast!" \
--voice "Charon"
Gemini TTS with style direction:
python3 ${CLAUDE_PLUGIN_ROOT}/skills/voice-generation/scripts/gemini_tts.py \
--text "Have a wonderful day!" \
--voice "Puck" \
--style "Say cheerfully with a British accent:"
Gemini TTS multi-speaker (dialogue):
python3 ${CLAUDE_PLUGIN_ROOT}/skills/voice-generation/scripts/gemini_tts.py \
--multi \
--speaker "Host:Charon" \
--speaker "Guest:Aoede" \
--text "Host: Welcome to the show!
Guest: Thanks for having me!"
For ElevenLabs:
python3 ${CLAUDE_PLUGIN_ROOT}/skills/voice-generation/scripts/elevenlabs.py \
--text "Your text here" \
--voice "Rachel" \
--model "eleven_multilingual_v2"
For OpenAI TTS:
python3 ${CLAUDE_PLUGIN_ROOT}/skills/voice-generation/scripts/openai_tts.py \
--text "Your text here" \
--voice "nova" \
--model "tts-1-hd"
List Gemini voices:
python3 ${CLAUDE_PLUGIN_ROOT}/skills/voice-generation/scripts/gemini_tts.py --list-voices
Step 5: Deliver the Result
- Provide the generated audio file path
- Mention the voice and settings used
- Offer to:
- Try a different voice
- Adjust speed or tone
- Use a different API
- Generate in a different format
Error Handling
Missing API key: Inform the user which key is needed:
Gemini TTS requires google-genai package: pip install google-genai
Text too long: Split into chunks and concatenate, or suggest shorter text.
Rate limit: Suggest waiting or trying a different API.
Unsupported language: Suggest an alternative API that supports the language.
Multi-speaker limit: Gemini TTS supports max 2 speakers. For more, use ElevenLabs with multiple calls.
Voice Selection Guide
Google Gemini TTS Voices (30 voices)
| Style |
Voices |
Best For |
| Bright/Upbeat |
Zephyr, Puck, Aoede, Laomedeia |
Marketing, cheerful content |
| Firm/Informative |
Charon, Kore, Orus, Rasalgethi |
News, tutorials, professional |
| Soft/Warm |
Achernar, Sulafat, Vindemiatrix |
Meditation, gentle narration |
| Smooth |
Algieba, Despina, Callirrhoe |
Audiobooks, storytelling |
| Clear |
Erinome, Iapetus, Pulcherrima |
Instructions, clarity |
| Character |
Fenrir (excitable), Enceladus (breathy), Algenib (gravelly), Gacrux (mature) |
Character voices, drama |
| Friendly |
Achird, Zubenelgenubi (casual) |
Casual, conversational |
Gemini TTS Style Tips:
- Use natural language:
--style "Say angrily:" or --style "Whisper mysteriously:"
- Specify accents:
--style "Speak with a British accent from London:"
- Control pace:
--style "Speak slowly and deliberately:"
- Combine:
--style "Say excitedly with a Southern US accent:"
OpenAI TTS Voices
| Voice |
Description |
Best For |
| alloy |
Neutral, balanced |
General purpose |
| echo |
Warm, conversational |
Podcasts, casual |
| fable |
Expressive, British |
Storytelling |
| onyx |
Deep, authoritative |
Narration, professional |
| nova |
Friendly, upbeat |
Marketing, tutorials |
| shimmer |
Soft, gentle |
Meditation, ASMR |
ElevenLabs Popular Voices
| Voice |
Description |
Best For |
| Rachel |
Young female, American |
Narration, audiobooks |
| Domi |
Young female, energetic |
Marketing, ads |
| Bella |
Young female, soft |
Storytelling |
| Antoni |
Young male, well-rounded |
Narration |
| Josh |
Young male, deep |
Audiobooks |
| Arnold |
Mature male, authoritative |
Documentary |
| Adam |
Middle-aged male, deep |
Narration |
| Sam |
Young male, raspy |
Character voices |
Best Practices
For Narration
- Use a consistent voice throughout
- Add natural pauses between paragraphs
- Consider pacing for the content type
For Dialogue
- Use different voices for different characters
- Match voice characteristics to character descriptions
- Adjust speed for emotional scenes
For Accessibility
- Use clear, well-paced speech
- Avoid overly stylized voices
- Test with screen readers if applicable
API Comparison
| Feature |
Gemini TTS |
ElevenLabs |
OpenAI TTS |
| API Key |
GOOGLE_API_KEY ✅ |
ELEVENLABS_API_KEY |
OPENAI_API_KEY |
| Voice quality |
Excellent |
Excellent |
Very good |
| Voice variety |
30 voices |
100+ voices |
6 voices |
| Multi-speaker |
✅ Up to 2 |
❌ No |
❌ No |
| Style control |
✅ Natural language |
Limited |
❌ No |
| Voice cloning |
❌ No |
✅ Yes |
❌ No |
| Languages |
24 |
29+ |
50+ |
| Speed control |
Via prompts |
Yes |
Yes (0.25-4x) |
| Max length |
32k tokens |
5,000 chars |
4,096 chars |
| Output format |
WAV (24kHz) |
MP3, WAV |
MP3, Opus, AAC, FLAC |
| Same key as video/image |
✅ Yes |
❌ No |
❌ No |
1---2name: voice-generation3description: Voice Generation Skill4---5# Voice Generation Skill67Generate realistic speech using AI (Google Gemini TTS, ElevenLabs, OpenAI TTS).89## Prerequisites1011At least one API key is required:1213- `GOOGLE_API_KEY` - For Google Gemini TTS (same key as video/image/music) ✅14- `ELEVENLABS_API_KEY` - For ElevenLabs high-quality voice synthesis15- `OPENAI_API_KEY` - For OpenAI TTS voices1617## Available APIs1819### Google Gemini TTS (Recommended - Same API Key)20- **Best for**: Podcasts, dialogues, audiobooks with style control21- **Voices**: 30 voices with natural language style control22- **Multi-speaker**: Up to 2 speakers for dialogues ✅23- **Languages**: 24 languages (auto-detected)24- **Features**: Control style, accent, pace via prompts25- **Output**: 24kHz WAV26- **API Key**: Same `GOOGLE_API_KEY` as video/image/music ✅2728### ElevenLabs (Best Quality)29- **Best for**: Natural-sounding voices, voice cloning, long-form content30- **Voices**: 100+ pre-made voices + custom voice cloning31- **Languages**: 29+ languages32- **Models**: Eleven Multilingual v2, Eleven Turbo v23334### OpenAI TTS (Simplest)35- **Best for**: Quick, reliable text-to-speech with consistent quality36- **Voices**: alloy, echo, fable, onyx, nova, shimmer37- **Models**: tts-1 (fast), tts-1-hd (high quality)38- **Output**: MP3, Opus, AAC, FLAC3940## Workflow4142### Step 1: Understand the Request4344Parse the user's voice request for:45- **Text content**: What should be spoken?46- **Voice type**: Male, female, specific character?47- **Tone**: Professional, casual, dramatic, cheerful?48- **Use case**: Narration, voiceover, audiobook, notification?49- **Language**: English, Spanish, other?50- **Speed**: Normal, slow, fast?5152### Step 2: Select Voice and API5354Choose based on requirements:5556| Use Case | Recommended API | Reason |57|----------|----------------|--------|58| **Default / Same key as video** | Gemini TTS | Same `GOOGLE_API_KEY` ✅ |59| **Multi-speaker dialogue** | Gemini TTS | Up to 2 speakers built-in |60| **Style/accent control** | Gemini TTS | Natural language prompts |61| **Voice cloning** | ElevenLabs | Only API with cloning |62| **100+ voice options** | ElevenLabs | Widest selection |63| **Audiobook/podcast** | ElevenLabs or Gemini | Both excellent for long content |64| **Quick narration** | OpenAI TTS | Fast, reliable |65| **Budget-conscious** | OpenAI TTS | Lower cost |6667### Step 3: Prepare the Text6869Optimize text for speech:70711. **Add pauses**: Use commas, periods for natural rhythm722. **Spell out numbers**: "1,234" → "one thousand two hundred thirty-four" (if needed)733. **Handle acronyms**: "NASA" vs "N.A.S.A." depending on pronunciation744. **Mark emphasis**: Some APIs support emphasis markers7576**Example transformation:**77- Original: "The Q4 2024 results show a 15% YoY increase."78- Optimized: "The Q4 2024 results show a fifteen percent year-over-year increase."7980### Step 4: Generate the Audio8182Execute the appropriate script from `${CLAUDE_PLUGIN_ROOT}/skills/voice-generation/scripts/`:8384**For Google Gemini TTS (single speaker):**85```bash86python3 ${CLAUDE_PLUGIN_ROOT}/skills/voice-generation/scripts/gemini_tts.py \87 --text "Welcome to our podcast!" \88 --voice "Charon"89```9091**Gemini TTS with style direction:**92```bash93python3 ${CLAUDE_PLUGIN_ROOT}/skills/voice-generation/scripts/gemini_tts.py \94 --text "Have a wonderful day!" \95 --voice "Puck" \96 --style "Say cheerfully with a British accent:"97```9899**Gemini TTS multi-speaker (dialogue):**100```bash101python3 ${CLAUDE_PLUGIN_ROOT}/skills/voice-generation/scripts/gemini_tts.py \102 --multi \103 --speaker "Host:Charon" \104 --speaker "Guest:Aoede" \105 --text "Host: Welcome to the show!106Guest: Thanks for having me!"107```108109**For ElevenLabs:**110```bash111python3 ${CLAUDE_PLUGIN_ROOT}/skills/voice-generation/scripts/elevenlabs.py \112 --text "Your text here" \113 --voice "Rachel" \114 --model "eleven_multilingual_v2"115```116117**For OpenAI TTS:**118```bash119python3 ${CLAUDE_PLUGIN_ROOT}/skills/voice-generation/scripts/openai_tts.py \120 --text "Your text here" \121 --voice "nova" \122 --model "tts-1-hd"123```124125**List Gemini voices:**126```bash127python3 ${CLAUDE_PLUGIN_ROOT}/skills/voice-generation/scripts/gemini_tts.py --list-voices128```129130### Step 5: Deliver the Result1311321. Provide the generated audio file path1332. Mention the voice and settings used1343. Offer to:135 - Try a different voice136 - Adjust speed or tone137 - Use a different API138 - Generate in a different format139140## Error Handling141142**Missing API key**: Inform the user which key is needed:143- Gemini TTS: Same `GOOGLE_API_KEY` as video/image - https://aistudio.google.com/apikey144- ElevenLabs: https://elevenlabs.io145- OpenAI: https://platform.openai.com/api-keys146147**Gemini TTS requires google-genai package**: `pip install google-genai`148149**Text too long**: Split into chunks and concatenate, or suggest shorter text.150151**Rate limit**: Suggest waiting or trying a different API.152153**Unsupported language**: Suggest an alternative API that supports the language.154155**Multi-speaker limit**: Gemini TTS supports max 2 speakers. For more, use ElevenLabs with multiple calls.156157## Voice Selection Guide158159### Google Gemini TTS Voices (30 voices)160| Style | Voices | Best For |161|-------|--------|----------|162| Bright/Upbeat | Zephyr, Puck, Aoede, Laomedeia | Marketing, cheerful content |163| Firm/Informative | Charon, Kore, Orus, Rasalgethi | News, tutorials, professional |164| Soft/Warm | Achernar, Sulafat, Vindemiatrix | Meditation, gentle narration |165| Smooth | Algieba, Despina, Callirrhoe | Audiobooks, storytelling |166| Clear | Erinome, Iapetus, Pulcherrima | Instructions, clarity |167| Character | Fenrir (excitable), Enceladus (breathy), Algenib (gravelly), Gacrux (mature) | Character voices, drama |168| Friendly | Achird, Zubenelgenubi (casual) | Casual, conversational |169170**Gemini TTS Style Tips:**171- Use natural language: `--style "Say angrily:"` or `--style "Whisper mysteriously:"`172- Specify accents: `--style "Speak with a British accent from London:"`173- Control pace: `--style "Speak slowly and deliberately:"`174- Combine: `--style "Say excitedly with a Southern US accent:"`175176### OpenAI TTS Voices177| Voice | Description | Best For |178|-------|-------------|----------|179| alloy | Neutral, balanced | General purpose |180| echo | Warm, conversational | Podcasts, casual |181| fable | Expressive, British | Storytelling |182| onyx | Deep, authoritative | Narration, professional |183| nova | Friendly, upbeat | Marketing, tutorials |184| shimmer | Soft, gentle | Meditation, ASMR |185186### ElevenLabs Popular Voices187| Voice | Description | Best For |188|-------|-------------|----------|189| Rachel | Young female, American | Narration, audiobooks |190| Domi | Young female, energetic | Marketing, ads |191| Bella | Young female, soft | Storytelling |192| Antoni | Young male, well-rounded | Narration |193| Josh | Young male, deep | Audiobooks |194| Arnold | Mature male, authoritative | Documentary |195| Adam | Middle-aged male, deep | Narration |196| Sam | Young male, raspy | Character voices |197198## Best Practices199200### For Narration201- Use a consistent voice throughout202- Add natural pauses between paragraphs203- Consider pacing for the content type204205### For Dialogue206- Use different voices for different characters207- Match voice characteristics to character descriptions208- Adjust speed for emotional scenes209210### For Accessibility211- Use clear, well-paced speech212- Avoid overly stylized voices213- Test with screen readers if applicable214215## API Comparison216217| Feature | Gemini TTS | ElevenLabs | OpenAI TTS |218|---------|------------|------------|------------|219| API Key | `GOOGLE_API_KEY` ✅ | `ELEVENLABS_API_KEY` | `OPENAI_API_KEY` |220| Voice quality | Excellent | Excellent | Very good |221| Voice variety | 30 voices | 100+ voices | 6 voices |222| Multi-speaker | ✅ Up to 2 | ❌ No | ❌ No |223| Style control | ✅ Natural language | Limited | ❌ No |224| Voice cloning | ❌ No | ✅ Yes | ❌ No |225| Languages | 24 | 29+ | 50+ |226| Speed control | Via prompts | Yes | Yes (0.25-4x) |227| Max length | 32k tokens | 5,000 chars | 4,096 chars |228| Output format | WAV (24kHz) | MP3, WAV | MP3, Opus, AAC, FLAC |229| Same key as video/image | ✅ Yes | ❌ No | ❌ No |