Text-to-Speech (TTS) Feature
Overview
The Text-to-Speech feature enables users to have AI assistant messages read aloud using Azure Speech Service with high-quality DragonHD voices. This creates a more immersive conversational experience and allows for hands-free interaction with the AI.
Version Implemented
Version: 0.234.189
Dependencies
- Azure Speech Service (Azure Cognitive Services)
- Azure Speech SDK for Python (
azure-cognitiveservices-speech)
- Backend: Flask routes for TTS synthesis
- Frontend: JavaScript module for audio playback and controls
Features
1. Voice Selection
- 27 DragonHD Latest Neural Voices across multiple languages
- Voice preview functionality in user profile
- Customizable speech speed (0.5x - 2.0x)
- Grouped by language for easy selection
2. Playback Controls
- Inline "Listen" button on each AI assistant message
- Play/Pause/Stop controls
- Visual feedback during playback (message highlighting)
- Audio state management (prevents multiple simultaneous playbacks)
3. Auto-play Mode
- Optional automatic playback of AI responses
- Configurable per-user in profile settings
- Automatically disables streaming when auto-play is enabled (for complete message synthesis)
4. Streaming Integration
- Streaming is disabled when TTS auto-play is enabled
- Users can manually play any message regardless of streaming state
- Clear UI feedback about streaming/TTS state
Configuration
Admin Settings (Settings Page)
Located in Search & Extract tab:
- Enable Text-to-Speech AI Responses toggle
- Speech Service Settings (shared with Speech-to-Text):
- Endpoint:
https://<region>.cognitiveservices.azure.com/
- Location: Azure region (e.g.,
eastus2)
- API Key: Azure Speech Service key
User Settings (Profile Page)
New Text-to-Speech Settings section includes:
Voice Selection dropdown
- 27 DragonHD voices organized by language
- Gender indicators (♂ Male, ♀ Female, ◉ Multi)
- Status badges (GA vs Preview)
Speech Speed slider
- Range: 0.5x (slower) to 2.0x (faster)
- Default: 1.0x (normal speed)
Voice Preview button
- Test selected voice with sample phrase
- Shows voice name in sample text
Auto-play Toggle
- Enable automatic playback of AI responses
- Warning: Disables streaming when enabled
Available Voices
English (US) - 14 voices
- en-US-Andrew:DragonHDLatestNeural (Male, GA) - Default
- en-US-Andrew2:DragonHDLatestNeural (Male, GA) - Optimized for conversational content
- en-US-Andrew3:DragonHDLatestNeural (Male, Preview) - Optimized for podcast content
- en-US-Adam:DragonHDLatestNeural (Male, GA)
- en-US-Alloy:DragonHDLatestNeural (Male, Preview)
- en-US-Brian:DragonHDLatestNeural (Male, GA)
- en-US-Davis:DragonHDLatestNeural (Male, GA)
- en-US-Steffan:DragonHDLatestNeural (Male, GA)
- en-US-Ava:DragonHDLatestNeural (Female, GA)
- en-US-Ava3:DragonHDLatestNeural (Female, Preview) - Optimized for podcast content
- en-US-Emma:DragonHDLatestNeural (Female, GA)
- en-US-Emma2:DragonHDLatestNeural (Female, GA) - Optimized for conversational content
- en-US-Aria:DragonHDLatestNeural (Female, Preview)
- en-US-Jenny:DragonHDLatestNeural (Female, Preview)
- en-US-Nova:DragonHDLatestNeural (Female, Preview)
- en-US-Phoebe:DragonHDLatestNeural (Female, Preview)
- en-US-Serena:DragonHDLatestNeural (Female, Preview)
- en-US-MultiTalker-Ava-Andrew:DragonHDLatestNeural (Multi, Preview)
German (DE) - 2 voices
- de-DE-Florian:DragonHDLatestNeural (Male, GA)
- de-DE-Seraphina:DragonHDLatestNeural (Female, GA)
Spanish (ES) - 2 voices
- es-ES-Tristan:DragonHDLatestNeural (Male, GA)
- es-ES-Ximena:DragonHDLatestNeural (Female, GA)
French (FR) - 2 voices
- fr-FR-Remy:DragonHDLatestNeural (Male, GA)
- fr-FR-Vivienne:DragonHDLatestNeural (Female, GA)
Japanese (JP) - 2 voices
- ja-JP-Masaru:DragonHDLatestNeural (Male, GA)
- ja-JP-Nanami:DragonHDLatestNeural (Female, GA)
Chinese (CN) - 2 voices
- zh-CN-Xiaochen:DragonHDLatestNeural (Female, GA)
- zh-CN-Yunfan:DragonHDLatestNeural (Male, GA)
API Endpoints
POST /api/chat/tts
Synthesize text to speech.
Request Body:
{
"text": "Text to synthesize",
"voice": "en-US-Andrew:DragonHDLatestNeural",
"speed": 1.0
}
Response:
- Success: Audio stream (MP3, 48kHz, 192kbps)
- Error: JSON with error message
GET /api/chat/tts/voices
Get list of available DragonHD voices.
Response:
{
"voices": [
{
"name": "en-US-Andrew:DragonHDLatestNeural",
"gender": "Male",
"language": "English (US)",
"status": "GA",
"note": ""
},
...
]
}
Usage Instructions
For Administrators
- Navigate to Admin Settings → Search & Extract tab
- Enable Text-to-Speech AI Responses toggle
- Configure Speech Service Settings:
- Enter Azure Speech Service endpoint
- Enter region/location
- Enter API key
- Save settings
For Users
- Navigate to Profile page
- Scroll to Text-to-Speech Settings section
- Select preferred voice from dropdown
- Adjust speech speed if desired
- Click Play Sample to preview voice
- (Optional) Enable Auto-play for AI Responses
- Click Save Voice Settings
In Chat
Manual Playback:
- Click the speaker icon (🔊) on any AI assistant message
- Click again to pause
- Message will highlight during playback
Auto-play Mode:
- Enable in profile settings
- AI responses will automatically play when complete
- Streaming will be disabled automatically
Technical Specifications
Architecture
- Backend: Python Flask with Azure Speech SDK
- Frontend: ES6 JavaScript modules
- Audio Format: MP3, 48kHz, 192kbps mono
- Synthesis: On-demand (no caching)
File Structure
Backend:
├── route_backend_tts.py # TTS API endpoints
├── config.py # Version and settings
└── functions_settings.py # TTS enable flag
Frontend:
├── static/js/chat/
│ ├── chat-tts.js # TTS module
│ ├── chat-messages.js # Message rendering with TTS button
│ └── chat-streaming.js # Streaming/TTS integration
├── static/css/chats.css # TTS button styles
└── templates/
├── chats.html # TTS script import
├── profile.html # TTS settings UI
└── admin_settings.html # Admin TTS toggle
User Settings Schema
{
"settings": {
"ttsEnabled": true,
"ttsVoice": "en-US-Andrew:DragonHDLatestNeural",
"ttsSpeed": 1.0,
"ttsAutoplay": false
}
}
Testing and Validation
Functional Test
Location: functional_tests/test_tts_integration.py
Tests:
- TTS API endpoint availability
- Voice list retrieval
- Audio synthesis with different voices
- Speed adjustment
- Error handling
Manual Testing Checklist
Performance Considerations
No Caching Strategy
- Audio is generated on-demand for each request
- Reduces storage requirements
- Ensures fresh synthesis with latest settings
- Trade-off: Slightly higher API latency (~1-2 seconds)
API Rate Limits
- Azure Speech Service has per-second request limits
- Consider implementing client-side rate limiting for heavy usage
- Monitor Azure Speech Service quotas
Known Limitations
Streaming Incompatibility
- TTS auto-play requires complete message text
- Streaming must be disabled for auto-play
- Manual playback works on any message
Browser Support
- Requires modern browser with Audio API support
- Chrome, Edge, Firefox, Safari supported
Language Detection
- Voice selection is manual
- No automatic language detection
Message Length
- Azure Speech Service has character limits per request
- Very long messages may be truncated
- Consider implementing chunking for long messages
Future Enhancements
Potential Features
Voice Auto-selection
- Detect message language
- Auto-select appropriate voice
Playback Queue
- Queue multiple messages
- Continuous playback mode
Word Highlighting
- Real-time word-by-word highlighting
- Requires SSML with word boundaries
Audio Caching
- Optional server-side caching
- Reduce API calls for repeated content
Custom Voices
- Support for custom neural voices
- Voice training integration
Emotion and Emphasis
- SSML-based emotion control
- Adjust prosody based on message sentiment
Troubleshooting
TTS Not Working
- Check admin settings:
- TTS feature enabled
- Speech Service configured correctly
- Check user settings:
- Voice selected
- TTS enabled in profile
- Check browser console for errors
- Verify Azure Speech Service credentials
No Audio Output
- Check browser audio permissions
- Verify system volume not muted
- Test with voice preview in profile
- Check network connectivity to Azure
Streaming Not Disabling
- Verify TTS auto-play is enabled
- Check localStorage for cached settings
- Refresh page to reload settings
Security Considerations
Settings Sanitization
- Admin settings sanitized before sending to frontend
- API keys never exposed in client-side code
- User settings validated server-side
API Key Management
- Store in Azure Key Vault (if enabled)
- Use managed identity when possible
- Rotate keys regularly
Rate Limiting
- Implement per-user rate limiting
- Monitor for abuse
- Set reasonable quotas
Related Documentation
Support and Feedback
For issues, feature requests, or feedback, please contact the development team or file an issue in the project repository.
1---2name: text-to-speech-tts-feature3description: The Text-to-Speech feature enables users to have AI assistant messages read aloud using Azure Speech Service with high-quality DragonHD voices.4---5# Text-to-Speech (TTS) Feature67## Overview8The Text-to-Speech feature enables users to have AI assistant messages read aloud using Azure Speech Service with high-quality DragonHD voices. This creates a more immersive conversational experience and allows for hands-free interaction with the AI.910## Version Implemented11**Version:** 0.234.1891213## Dependencies14- **Azure Speech Service** (Azure Cognitive Services)15- **Azure Speech SDK for Python** (`azure-cognitiveservices-speech`)16- **Backend:** Flask routes for TTS synthesis17- **Frontend:** JavaScript module for audio playback and controls1819## Features2021### 1. Voice Selection22- **27 DragonHD Latest Neural Voices** across multiple languages23- Voice preview functionality in user profile24- Customizable speech speed (0.5x - 2.0x)25- Grouped by language for easy selection2627### 2. Playback Controls28- **Inline "Listen" button** on each AI assistant message29- **Play/Pause/Stop** controls30- Visual feedback during playback (message highlighting)31- Audio state management (prevents multiple simultaneous playbacks)3233### 3. Auto-play Mode34- Optional automatic playback of AI responses35- Configurable per-user in profile settings36- **Automatically disables streaming** when auto-play is enabled (for complete message synthesis)3738### 4. Streaming Integration39- Streaming is disabled when TTS auto-play is enabled40- Users can manually play any message regardless of streaming state41- Clear UI feedback about streaming/TTS state4243## Configuration4445### Admin Settings (Settings Page)46Located in **Search & Extract** tab:47481. **Enable Text-to-Speech AI Responses** toggle492. **Speech Service Settings** (shared with Speech-to-Text):50 - Endpoint: `https://<region>.cognitiveservices.azure.com/`51 - Location: Azure region (e.g., `eastus2`)52 - API Key: Azure Speech Service key5354### User Settings (Profile Page)55New **Text-to-Speech Settings** section includes:56571. **Voice Selection** dropdown58 - 27 DragonHD voices organized by language59 - Gender indicators (♂ Male, ♀ Female, ◉ Multi)60 - Status badges (GA vs Preview)61622. **Speech Speed** slider63 - Range: 0.5x (slower) to 2.0x (faster)64 - Default: 1.0x (normal speed)65663. **Voice Preview** button67 - Test selected voice with sample phrase68 - Shows voice name in sample text69704. **Auto-play Toggle**71 - Enable automatic playback of AI responses72 - Warning: Disables streaming when enabled7374## Available Voices7576### English (US) - 14 voices77- **en-US-Andrew:DragonHDLatestNeural** (Male, GA) - *Default*78- **en-US-Andrew2:DragonHDLatestNeural** (Male, GA) - Optimized for conversational content79- **en-US-Andrew3:DragonHDLatestNeural** (Male, Preview) - Optimized for podcast content80- **en-US-Adam:DragonHDLatestNeural** (Male, GA)81- **en-US-Alloy:DragonHDLatestNeural** (Male, Preview)82- **en-US-Brian:DragonHDLatestNeural** (Male, GA)83- **en-US-Davis:DragonHDLatestNeural** (Male, GA)84- **en-US-Steffan:DragonHDLatestNeural** (Male, GA)85- **en-US-Ava:DragonHDLatestNeural** (Female, GA)86- **en-US-Ava3:DragonHDLatestNeural** (Female, Preview) - Optimized for podcast content87- **en-US-Emma:DragonHDLatestNeural** (Female, GA)88- **en-US-Emma2:DragonHDLatestNeural** (Female, GA) - Optimized for conversational content89- **en-US-Aria:DragonHDLatestNeural** (Female, Preview)90- **en-US-Jenny:DragonHDLatestNeural** (Female, Preview)91- **en-US-Nova:DragonHDLatestNeural** (Female, Preview)92- **en-US-Phoebe:DragonHDLatestNeural** (Female, Preview)93- **en-US-Serena:DragonHDLatestNeural** (Female, Preview)94- **en-US-MultiTalker-Ava-Andrew:DragonHDLatestNeural** (Multi, Preview)9596### German (DE) - 2 voices97- **de-DE-Florian:DragonHDLatestNeural** (Male, GA)98- **de-DE-Seraphina:DragonHDLatestNeural** (Female, GA)99100### Spanish (ES) - 2 voices101- **es-ES-Tristan:DragonHDLatestNeural** (Male, GA)102- **es-ES-Ximena:DragonHDLatestNeural** (Female, GA)103104### French (FR) - 2 voices105- **fr-FR-Remy:DragonHDLatestNeural** (Male, GA)106- **fr-FR-Vivienne:DragonHDLatestNeural** (Female, GA)107108### Japanese (JP) - 2 voices109- **ja-JP-Masaru:DragonHDLatestNeural** (Male, GA)110- **ja-JP-Nanami:DragonHDLatestNeural** (Female, GA)111112### Chinese (CN) - 2 voices113- **zh-CN-Xiaochen:DragonHDLatestNeural** (Female, GA)114- **zh-CN-Yunfan:DragonHDLatestNeural** (Male, GA)115116## API Endpoints117118### POST `/api/chat/tts`119Synthesize text to speech.120121**Request Body:**122```json123{124 "text": "Text to synthesize",125 "voice": "en-US-Andrew:DragonHDLatestNeural",126 "speed": 1.0127}128```129130**Response:**131- **Success:** Audio stream (MP3, 48kHz, 192kbps)132- **Error:** JSON with error message133134### GET `/api/chat/tts/voices`135Get list of available DragonHD voices.136137**Response:**138```json139{140 "voices": [141 {142 "name": "en-US-Andrew:DragonHDLatestNeural",143 "gender": "Male",144 "language": "English (US)",145 "status": "GA",146 "note": ""147 },148 ...149 ]150}151```152153## Usage Instructions154155### For Administrators1561. Navigate to **Admin Settings → Search & Extract** tab1572. Enable **Text-to-Speech AI Responses** toggle1583. Configure **Speech Service Settings**:159 - Enter Azure Speech Service endpoint160 - Enter region/location161 - Enter API key1624. Save settings163164### For Users1651. Navigate to **Profile** page1662. Scroll to **Text-to-Speech Settings** section1673. Select preferred voice from dropdown1684. Adjust speech speed if desired1695. Click **Play Sample** to preview voice1706. (Optional) Enable **Auto-play for AI Responses**1717. Click **Save Voice Settings**172173### In Chat1741. **Manual Playback:**175 - Click the speaker icon (🔊) on any AI assistant message176 - Click again to pause177 - Message will highlight during playback1781792. **Auto-play Mode:**180 - Enable in profile settings181 - AI responses will automatically play when complete182 - Streaming will be disabled automatically183184## Technical Specifications185186### Architecture187- **Backend:** Python Flask with Azure Speech SDK188- **Frontend:** ES6 JavaScript modules189- **Audio Format:** MP3, 48kHz, 192kbps mono190- **Synthesis:** On-demand (no caching)191192### File Structure193```194Backend:195├── route_backend_tts.py # TTS API endpoints196├── config.py # Version and settings197└── functions_settings.py # TTS enable flag198199Frontend:200├── static/js/chat/201│ ├── chat-tts.js # TTS module202│ ├── chat-messages.js # Message rendering with TTS button203│ └── chat-streaming.js # Streaming/TTS integration204├── static/css/chats.css # TTS button styles205└── templates/206 ├── chats.html # TTS script import207 ├── profile.html # TTS settings UI208 └── admin_settings.html # Admin TTS toggle209```210211### User Settings Schema212```json213{214 "settings": {215 "ttsEnabled": true,216 "ttsVoice": "en-US-Andrew:DragonHDLatestNeural",217 "ttsSpeed": 1.0,218 "ttsAutoplay": false219 }220}221```222223## Testing and Validation224225### Functional Test226Location: `functional_tests/test_tts_integration.py`227228Tests:229- TTS API endpoint availability230- Voice list retrieval231- Audio synthesis with different voices232- Speed adjustment233- Error handling234235### Manual Testing Checklist236- [ ] Admin can enable/disable TTS feature237- [ ] User can select voice in profile238- [ ] Voice preview plays sample audio239- [ ] Speech speed slider affects playback240- [ ] "Listen" button appears on AI messages241- [ ] Audio plays when clicking "Listen" button242- [ ] Message highlights during playback243- [ ] Only one message plays at a time244- [ ] Auto-play works when enabled245- [ ] Streaming disables when auto-play enabled246- [ ] Streaming button shows correct tooltip247248## Performance Considerations249250### No Caching Strategy251- Audio is generated on-demand for each request252- Reduces storage requirements253- Ensures fresh synthesis with latest settings254- Trade-off: Slightly higher API latency (~1-2 seconds)255256### API Rate Limits257- Azure Speech Service has per-second request limits258- Consider implementing client-side rate limiting for heavy usage259- Monitor Azure Speech Service quotas260261## Known Limitations2622631. **Streaming Incompatibility**264 - TTS auto-play requires complete message text265 - Streaming must be disabled for auto-play266 - Manual playback works on any message2672682. **Browser Support**269 - Requires modern browser with Audio API support270 - Chrome, Edge, Firefox, Safari supported2712723. **Language Detection**273 - Voice selection is manual274 - No automatic language detection2752764. **Message Length**277 - Azure Speech Service has character limits per request278 - Very long messages may be truncated279 - Consider implementing chunking for long messages280281## Future Enhancements282283### Potential Features2841. **Voice Auto-selection**285 - Detect message language286 - Auto-select appropriate voice2872882. **Playback Queue**289 - Queue multiple messages290 - Continuous playback mode2912923. **Word Highlighting**293 - Real-time word-by-word highlighting294 - Requires SSML with word boundaries2952964. **Audio Caching**297 - Optional server-side caching298 - Reduce API calls for repeated content2993005. **Custom Voices**301 - Support for custom neural voices302 - Voice training integration3033046. **Emotion and Emphasis**305 - SSML-based emotion control306 - Adjust prosody based on message sentiment307308## Troubleshooting309310### TTS Not Working3111. Check admin settings:312 - TTS feature enabled313 - Speech Service configured correctly3142. Check user settings:315 - Voice selected316 - TTS enabled in profile3173. Check browser console for errors3184. Verify Azure Speech Service credentials319320### No Audio Output3211. Check browser audio permissions3222. Verify system volume not muted3233. Test with voice preview in profile3244. Check network connectivity to Azure325326### Streaming Not Disabling3271. Verify TTS auto-play is enabled3282. Check localStorage for cached settings3293. Refresh page to reload settings330331## Security Considerations332333### Settings Sanitization334- Admin settings sanitized before sending to frontend335- API keys never exposed in client-side code336- User settings validated server-side337338### API Key Management339- Store in Azure Key Vault (if enabled)340- Use managed identity when possible341- Rotate keys regularly342343### Rate Limiting344- Implement per-user rate limiting345- Monitor for abuse346- Set reasonable quotas347348## Related Documentation349- [Azure Speech Service Documentation](https://docs.microsoft.com/azure/cognitive-services/speech-service/)350- [Speech-to-Text Feature](./SPEECH_TO_TEXT.md)351- [Agent Integration](./AGENT_ORCHESTRATION.md)352353## Support and Feedback354For issues, feature requests, or feedback, please contact the development team or file an issue in the project repository.