Inworld AI
Text-to-Speech platform with voice cloning, audio markups, and timestamp alignment.
Quick Navigation
| Topic |
Reference |
| Installation |
installation.md |
| Voice Cloning |
cloning.md |
| Voice Control |
voice-control.md |
| API Reference |
api.md |
When to Use
- Text-to-speech audio generation
- Voice cloning from 5-15 seconds of audio
- Emotion-controlled speech (
[happy], [sad], etc.)
- Word/phoneme timestamps for lip sync
- Custom pronunciation with IPA
Models
| Model |
ID |
Latency |
Price |
| TTS 1.5 Max |
inworld-tts-1.5-max |
~200ms |
$10/1M chars |
| TTS 1.5 Mini |
inworld-tts-1.5-mini |
~120ms |
$5/1M chars |
Minimal Example
import requests, base64, os
response = requests.post(
"https://api.inworld.ai/tts/v1/voice",
headers={"Authorization": f"Basic {os.getenv('INWORLD_API_KEY')}"},
json={"text": "Hello!", "voiceId": "Ashley", "modelId": "inworld-tts-1.5-max"}
)
audio = base64.b64decode(response.json()['audioContent'])
Key Features
- 15 languages — en, zh, ja, ko, ru, it, es, pt, fr, de, pl, nl, hi, he, ar
- Instant cloning — 5-15 seconds audio, no training
- Audio markups —
[happy], [laughing], [sigh] (English only)
- Timestamps — word, phoneme, viseme timing for lip sync
- Streaming —
/voice:stream endpoint
Prohibitions
- Audio markups work only in English
- Use ONE emotion markup at text beginning
- Match voice language to text language
- Instant cloning may not work for children's voices or unique accents
Links
1---2name: inworld3description: Inworld TTS API. Covers voice cloning, audio markups, timestamps. Keywords: text-to-speech, visemes.4---5
6# Inworld AI
7
8Text-to-Speech platform with voice cloning, audio markups, and timestamp alignment.
9
10## Quick Navigation
11
12| Topic | Reference |
13| ------------- | ----------------------------------------------- |
14| Installation | [installation.md](references/installation.md) |
15| Voice Cloning | [cloning.md](references/cloning.md) |
16| Voice Control | [voice-control.md](references/voice-control.md) |
17| API Reference | [api.md](references/api.md) |
18
19## When to Use
20
21- Text-to-speech audio generation
22- Voice cloning from 5-15 seconds of audio
23- Emotion-controlled speech (`[happy]`, `[sad]`, etc.)
24- Word/phoneme timestamps for lip sync
25- Custom pronunciation with IPA
26
27## Models
28
29| Model | ID | Latency | Price |
30| ------------ | ---------------------- | ------- | ------------ |
31| TTS 1.5 Max | `inworld-tts-1.5-max` | ~200ms | $10/1M chars |
32| TTS 1.5 Mini | `inworld-tts-1.5-mini` | ~120ms | $5/1M chars |
33
34## Minimal Example
35
36```python
37import requests, base64, os
38
39response = requests.post(
40 "https://api.inworld.ai/tts/v1/voice",
41 headers={"Authorization": f"Basic {os.getenv('INWORLD_API_KEY')}"},
42 json={"text": "Hello!", "voiceId": "Ashley", "modelId": "inworld-tts-1.5-max"}
43)
44audio = base64.b64decode(response.json()['audioContent'])
45```
46
47## Key Features
48
49- **15 languages** — en, zh, ja, ko, ru, it, es, pt, fr, de, pl, nl, hi, he, ar
50- **Instant cloning** — 5-15 seconds audio, no training
51- **Audio markups** — `[happy]`, `[laughing]`, `[sigh]` (English only)
52- **Timestamps** — word, phoneme, viseme timing for lip sync
53- **Streaming** — `/voice:stream` endpoint
54
55## Prohibitions
56
57- Audio markups work **only in English**
58- Use **ONE** emotion markup at text **beginning**
59- Match voice language to text language
60- Instant cloning may not work for children's voices or unique accents
61
62## Links
63
64- Docs: https://docs.inworld.ai/docs/tts/tts
65- Platform: https://platform.inworld.ai/