Provider Base URL Lists Specification
Overview
This specification describes a system for configuring multiple TTS and STT base URLs as comma-separated lists, with automatic discovery, failover, and provider auto-detection.
Design Principles
- OpenAI API Compatibility: All endpoints must be OpenAI API-compatible
- Graceful Degradation: Handle missing endpoints gracefully
- Priority-Based Selection: Use URLs in the order specified by the user
- Transparent to LLM: The LLM doesn't need to know which provider is being used
Environment Variables
Core Configuration (No Backward Compatibility)
# Comma-separated list of TTS base URLs (tried in order)
VOICEMODE_TTS_BASE_URLS=http://127.0.0.1:8880/v1,https://api.openai.com/v1
# Comma-separated list of STT base URLs (tried in order)
VOICEMODE_STT_BASE_URLS=http://127.0.0.1:2022/v1,https://api.openai.com/v1
# Comma-separated list of preferred TTS voices (tried in order of availability)
VOICEMODE_VOICES=af_sky,nova,alloy
# Comma-separated list of preferred TTS models (optional)
VOICEMODE_TTS_MODELS=tts-1,gpt-4o-mini-tts
# API key for authentication (required)
OPENAI_API_KEY=sk-...
Discovery Process
On Startup
- Iterate through each base URL in
VOICEMODE_TTS_BASE_URLSandVOICEMODE_STT_BASE_URLS - Health Check: Verify endpoint is reachable
- Model Discovery: Query
/v1/modelsendpoint - Voice Discovery (TTS only):
- If URL contains "openai.com" → assume OpenAI voices:
["alloy", "echo", "fable", "nova", "onyx", "shimmer"] - Otherwise → try
/v1/audio/voices(Kokoro endpoint) - If voices endpoint fails but health check passes → assume OpenAI voices
- If URL contains "openai.com" → assume OpenAI voices:
- Build Registry: Store discovered capabilities for runtime use
Voice Discovery Logic
async def discover_voices(base_url: str, client: AsyncOpenAI) -> List[str]:
"""Discover available voices for a TTS endpoint."""
# OpenAI doesn't have a voices endpoint, use known list
if "openai.com" in base_url:
return ["alloy", "echo", "fable", "nova", "onyx", "shimmer"]
# Try Kokoro-style voices endpoint
try:
response = await client.get("/v1/audio/voices")
return response.json()["voices"]
except:
# If endpoint doesn't exist but server is healthy, assume OpenAI voices
return ["alloy", "echo", "fable", "nova", "onyx", "shimmer"]
Registry Structure
The registry stores discovered capabilities for each base URL:
{
"tts": {
"http://127.0.0.1:8880/v1": {
"healthy": true,
"models": ["tts-1"],
"voices": ["af_sky", "af_sarah", "am_adam", "af_nicole", "am_michael"],
"last_health_check": "2024-01-20T10:30:00Z",
"response_time_ms": 45
},
"https://api.openai.com/v1": {
"healthy": true,
"models": ["tts-1", "tts-1-hd", "gpt-4o-mini-tts"],
"voices": ["alloy", "echo", "fable", "nova", "onyx", "shimmer"],
"last_health_check": "2024-01-20T10:30:00Z",
"response_time_ms": 120
}
},
"stt": {
"http://127.0.0.1:2022/v1": {
"healthy": true,
"models": ["whisper-1"],
"last_health_check": "2024-01-20T10:30:00Z",
"response_time_ms": 30
}
}
}
Selection Algorithm
When a TTS request is made:
- Iterate through healthy endpoints in the order specified by
VOICEMODE_TTS_BASE_URLS - Find first endpoint that supports the requested voice (or first preferred voice)
- Use that endpoint for the request
Selection Priority
- User-specified voice/model/provider (if provided)
- First available voice from
VOICEMODE_VOICES - First available model from
VOICEMODE_TTS_MODELS - First healthy endpoint from
VOICEMODE_TTS_BASE_URLS
Example Selection
Given:
VOICEMODE_TTS_BASE_URLS=http://127.0.0.1:8880/v1,https://api.openai.com/v1
VOICEMODE_VOICES=af_sky,nova,alloy
If 127.0.0.1:8880 is healthy and has af_sky, use it. Otherwise, check if OpenAI has nova or alloy.
Registry Updates
When to Update
- On startup: Full discovery of all endpoints
- On request failure: Health check the failed endpoint
- Manual refresh: Via MCP tool/command
- No periodic refresh: Not needed for typical use
Failure Handling
When a request fails:
- Mark endpoint as unhealthy in registry
- Retry with next available endpoint
- Run health check on failed endpoint
- Update registry based on health check result
LLM Integration
The LLM can query the registry to see available options:
async def get_voice_registry() -> Dict:
"""Return the current provider registry for LLM inspection."""
return {
"tts": {
url: {
"healthy": info["healthy"],
"models": info["models"],
"voices": info["voices"],
"response_time_ms": info["response_time_ms"]
}
for url, info in registry["tts"].items()
},
"stt": {
url: {
"healthy": info["healthy"],
"models": info["models"],
"response_time_ms": info["response_time_ms"]
}
for url, info in registry["stt"].items()
}
}
Configuration Examples
Minimal Configuration
# Only API key required - defaults to OpenAI
OPENAI_API_KEY=sk-...
Local Development
VOICEMODE_TTS_BASE_URLS=http://127.0.0.1:8880/v1,https://api.openai.com/v1
VOICEMODE_STT_BASE_URLS=http://127.0.0.1:2022/v1,https://api.openai.com/v1
VOICEMODE_VOICES=af_sky,nova,alloy
OPENAI_API_KEY=sk-...
Production with Fallback
VOICEMODE_TTS_BASE_URLS=http://tts-prod.internal/v1,http://tts-backup.internal/v1,https://api.openai.com/v1
VOICEMODE_STT_BASE_URLS=http://stt-prod.internal/v1,https://api.openai.com/v1
VOICEMODE_VOICES=nova,alloy,echo
VOICEMODE_TTS_MODELS=gpt-4o-mini-tts,tts-1-hd,tts-1
OPENAI_API_KEY=sk-...
Implementation Notes
- Remove all legacy environment variables (TTS_BASE_URL, STT_BASE_URL, etc.)
- No provider-specific code - everything uses OpenAI API
- Graceful fallback - if primary fails, try next URL
- Fast selection - use pre-discovered registry, no discovery during requests
- Simple configuration - just list URLs and preferences