# Oatda Generate Speech

> Text-to-speech (TTS) and AI voice generation through OATDA's unified audio API gateway. Triggers when the user wants to convert text to speech, synthesize voice, create narration, voiceovers, audiobooks, podcast audio, accessibility audio, or generate spoken audio from text. Supports OpenAI, xAI Grok, Google Gemini and other TTS models via a single API key.

- Skill: `devcsde/oatda-generate-speech-2` (Agent Skill)
- Install (CLI): `npx skillmds@latest add devcsde/oatda-generate-speech-2`
- Raw SKILL.md: https://api.skillmd.com/api/skills/devcsde/oatda-generate-speech-2/raw
- Safety review: pending
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Web & Frontend
- Author: devcsde (https://skillmd.com/u/devcsde)
- Updated: 2026-09-22
- Page: https://skillmd.com/skills/devcsde/oatda-generate-speech-2

---


# OATDA Speech Generation

Generate spoken audio from text through OATDA's unified audio API.

## API Key Resolution

All commands need the OATDA API key. Resolve it inline for each `exec` call:

```bash
export OATDA_API_KEY="${OATDA_API_KEY:-$(cat ~/.oatda/credentials.json 2>/dev/null | jq -r '.profiles[.defaultProfile].apiKey' 2>/dev/null)}"
```

If the key is empty or `null`, tell the user to get one at https://oatda.com and configure it.

**Security**: Never print the full API key. Only verify existence or show first 8 chars.

## Model Resolution

> **⚠️ Model availability changes over time.** Always call `oatda-list-models` with `?type=audio` to verify the exact model ID, available voices, and supported parameters before generating speech.

If the user provides `provider/model` format directly (for example `openai/gpt-4o-mini-tts`), split on `/`.

Use the `oatda-list-models` results to determine:
- Available voices (from `supported_params.voice.values`)
- Supported response formats
- Optional parameters like `language` or `instructions`

If the user does not specify a model, query `oatda-list-models` first and offer a choice from the currently available TTS models.

## Discovering Audio Model Parameters

Query available audio models and inspect `supported_params` before sending optional fields:

```bash
export OATDA_API_KEY="${OATDA_API_KEY:-$(cat ~/.oatda/credentials.json 2>/dev/null | jq -r '.profiles[.defaultProfile].apiKey' 2>/dev/null)}" && \
curl -s -X GET "https://oatda.com/api/v1/llm/models?type=audio" \
  -H "Authorization: Bearer $OATDA_API_KEY" | jq '.audio_models[] | {id, supported_params}'
```

Look for:
- `audio_modes` containing `tts`
- supported `voice` values
- allowed `response_format` values
- optional fields like `instructions` or `language`

## API Call

The speech endpoint returns **binary audio**, not JSON. Always save the response to a file.

```bash
export OATDA_API_KEY="${OATDA_API_KEY:-$(cat ~/.oatda/credentials.json 2>/dev/null | jq -r '.profiles[.defaultProfile].apiKey' 2>/dev/null)}" && \
curl -s -X POST "https://oatda.com/api/v1/llm/speech" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $OATDA_API_KEY" \
  -d '{
    "provider": "<PROVIDER>",
    "model": "<MODEL>",
    "input": "<TEXT_TO_SPEAK>",
    "voice": "alloy",
    "response_format": "mp3",
    "speed": 1.0
  }' \
  --output speech.mp3
```

### Common Parameters

- `input`: Text to convert to speech, max 15000 characters
- `voice`: Voice name, e.g. `alloy`, `nova`, `shimmer`
- `response_format`: `mp3`, `opus`, `aac`, `flac`, `wav`, `pcm`, `mulaw`, or `alaw`
- `speed`: 0.25 to 4.0, default 1.0
- `instructions`: Optional tone/style guidance for supported models
- `language`: Optional language code for supported models

## Success Handling

If the request succeeds, tell the user where the file was saved, for example:

> Speech generated successfully: `speech.mp3`

If headers matter, use `curl -D headers.txt` while still saving the audio body with `--output`.

## Error Handling

| HTTP Status | Meaning | Action |
|-------------|---------|--------|
| 401 | Invalid API key | Tell user to check their key |
| 402 | Insufficient credits | Tell user to check balance |
| 400 | Bad request / model not supported | Check model format and query `oatda-list-models` with `type=audio` |
| 429 | Rate limited or monthly cap | Wait briefly and retry once |
| 500 | Provider error | Show the error message if returned |

## Example

```bash
export OATDA_API_KEY="${OATDA_API_KEY:-$(cat ~/.oatda/credentials.json 2>/dev/null | jq -r '.profiles[.defaultProfile].apiKey' 2>/dev/null)}" && \
curl -s -X POST "https://oatda.com/api/v1/llm/speech" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $OATDA_API_KEY" \
  -d '{
    "provider": "openai",
    "model": "gpt-4o-mini-tts",
    "input": "Welcome to OATDA, one API to direct all.",
    "voice": "alloy",
    "response_format": "mp3",
    "speed": 1.0
  }' \
  --output speech.mp3
```

> **Note:** The model ID above is an example. Always verify the current model ID via `oatda-list-models` before use.

## Notes

- Endpoint: `/api/v1/llm/speech`
- Use `input`, not `prompt`, for TTS requests
- Always save the response with `--output`
- Use `oatda-list-models` to discover available audio models
- Equivalent capability name: `generate_speech`
- Related skills: `oatda-list-models`, `oatda-transcribe-audio`, `oatda-translate-audio`

