Voicemode Troubleshooting
Audio Issues
User being cut off mid-sentence
Problem: Silence detection stops recording too early.
Solutions:
Increase
listen_duration_min:converse("What's on your mind?", listen_duration_min=5.0)Decrease VAD aggressiveness:
converse("Tell me more", vad_aggressiveness=0)Disable silence detection entirely:
converse("Please describe in detail", disable_silence_detection=True)
Background noise triggering false starts
Problem: VAD detects non-speech as speech.
Solutions:
Increase VAD aggressiveness:
converse("Can you hear me?", vad_aggressiveness=3)Check environment noise levels
Use better quality microphone
Audio chimes getting cut off
Problem: Bluetooth or audio system delays.
Solutions:
Add leading silence:
converse("Hello", chime_leading_silence=1.0)Add trailing silence:
converse("Hello", chime_trailing_silence=0.5)Add both:
converse("Hello", chime_leading_silence=1.0, chime_trailing_silence=0.5)
No audio output
Problem: TTS not playing.
Solutions:
- Check if
skip_ttsis enabled - Verify VOICEMODE_SKIP_TTS env var
- Force TTS:
converse("Test message", skip_tts=False) - Check audio output device settings
- Verify TTS service endpoint
Voice Activity Detection (VAD)
Understanding VAD levels
- 0 (Least aggressive): Captures everything, including background noise
- 1 (Low): Slightly stricter, good for quiet environments
- 2 (Balanced - default): Good for normal home/office
- 3 (Most aggressive): Strict speech detection, filters most noise
When to adjust VAD
| Environment | Recommended Setting | Reason |
|---|---|---|
| Silent room | 0-1 | Don't miss soft speech |
| Home office | 2 | Balanced (default) |
| Busy office | 2-3 | Filter typing, conversations |
| Cafe/public | 3 | Filter heavy background noise |
| Outdoors | 3 | Filter wind, traffic |
| Dictation mode | 0-1 + high listen_duration_min | Allow thinking pauses |
Connection Issues
STT/TTS endpoint errors
Problem: Cannot connect to speech services.
Check:
Services expose OpenAI-compatible endpoints:
/v1/audio/transcriptions(STT)/v1/audio/speech(TTS)
Environment variables are set correctly:
OPENAI_API_KEY(if using OpenAI)- Service URLs for Whisper/Kokoro
Network connectivity to services
Service logs for errors
Transport issues
Problem: LiveKit or local transport failing.
Solutions:
Try different transport:
# Force local transport converse("Test", transport="local") # Force LiveKit converse("Test", transport="livekit")Check LiveKit room configuration
Verify microphone permissions
Timeout issues
Problem: Operations timing out.
Solutions:
Increase listen duration:
converse("Please elaborate", listen_duration_max=300)Check network latency to services
Verify services are responding
Voice Quality Issues
Incorrect pronunciation
Problem: Words mispronounced, especially for non-English.
Solutions:
For non-English, use Kokoro with appropriate voice:
converse("Bonjour", voice="ff_siwis", tts_provider="kokoro")See
voicemode-languagesresource for language-specific voices
Robotic/unnatural voice
Problem: Voice sounds too mechanical.
Solutions:
Try different voice:
converse("Hello", voice="nova") # OpenAI converse("Hello", voice="af_sky", tts_provider="kokoro")Use HD model for better quality:
converse("Hello", tts_model="tts-1-hd")Add emotional context:
converse( "I'm excited to help!", tts_model="gpt-4o-mini-tts", tts_instructions="Sound warm and friendly" )
Speech too fast/slow
Problem: Default speed doesn't match user preference.
Solution: Adjust speed:
# Slower
converse("Complex information", speed=0.8)
# Faster
converse("Quick update", speed=1.5)
Recognition Issues
STT not recognizing speech
Problem: Speech not being transcribed.
Solutions:
- Check microphone is working
- Verify microphone permissions
- Increase recording duration:
converse("What do you think?", listen_duration_min=5.0) - Disable silence detection to see if it's a VAD issue:
converse("Testing", disable_silence_detection=True, listen_duration_max=10)
Incorrect transcriptions
Problem: Speech transcribed wrong.
Solutions:
- Speak more clearly
- Reduce background noise
- Adjust VAD for your environment
- Use better quality microphone
- Check STT service configuration
Specific words consistently misrecognized
Problem: Certain words or names are repeatedly transcribed incorrectly, even when spoken clearly.
Symptoms:
- Same word always comes out wrong (e.g., "tmux" → "T-Mux" or "T marks")
- Names are consistently misspelled (e.g., "Tali" → "Talley" or "Tolly")
- Technical terms get mangled (e.g., "kubectl" → "cube control")
Solution: Use vocabulary biasing with VOICEMODE_STT_PROMPT.
Set the environment variable with words you frequently use:
# In ~/.voicemode/voicemode.env
VOICEMODE_STT_PROMPT="tmux, Tali, kubectl, pytest, VoiceMode"
This "primes" Whisper to recognize these specific terms correctly.
See: Parameters - Vocabulary Biasing for detailed configuration options.
Performance Issues
Slow response times
Problem: Long delays between speaking and response.
Causes:
- Network latency to STT/TTS services
- Heavy service load
- Large audio files
Solutions:
- Use lower quality audio format if possible
- Check service response times
- Consider local STT/TTS services
- Use
skip_tts=Truefor development:converse("Quick test", skip_tts=True)
Common Mistakes
Using coral voice
❌ Don't: voice="coral" - Not supported
✅ Do: Use supported voices (nova, shimmer, af_sky, etc.)
Not specifying voice for non-English
❌ Don't: converse("Bonjour")
✅ Do: converse("Bonjour", voice="ff_siwis", tts_provider="kokoro")
Setting listen_duration_max too low
❌ Don't: listen_duration_max=5 for complex questions
✅ Do: Use default (120) or higher for long responses
Overriding defaults unnecessarily
❌ Don't: Specify voice, tts_provider, tts_model without reason
✅ Do: Let system auto-select unless specific need
Getting More Help
If issues persist:
- Check service logs for errors
- Verify environment configuration
- Test with minimal parameters first
- Add parameters one at a time to isolate issue
See Also
- Parameters - Full parameter reference (includes vocabulary biasing)
- Patterns - Best practices
- Languages - Language-specific configuration