happy-audio-gen
Turns text into speech across 6 providers through one CLI. All providers are synchronous (TTS is fast — typically under 10 seconds) except Bailian's voice-design flow (which is still covered but uses a longer poll window).
Quick usage
# Shortest path — OpenAI default voice
bun scripts/main.ts --text "Hello, world" --out ./hello.mp3
# Chinese, MiniMax
bun scripts/main.ts --provider minimax --text "大家好" --voice male-qn-qingse --out ./hello.mp3
# Long-form, Bailian (auto-splits by sentence)
bun scripts/main.ts --provider bailian --textfiles ./script.md --out ./narration.mp3
When to invoke this skill
- User asks to synthesize speech / TTS / read aloud / narrate / dub / make a voice-over.
- User asks to convert script / text / article into audio.
- User names a TTS voice or model.
Do not route here when the user wants to transcribe audio → text (that's STT, different domain), or edit / mix audio files (use a dedicated audio editor).
Step 0: Preflight (BLOCKING)
Locate EXTEND.md:
./.happy-skills/happy-audio-gen/EXTEND.md
$XDG_CONFIG_HOME/happy-skills/happy-audio-gen/EXTEND.md
~/.happy-skills/happy-audio-gen/EXTEND.md
If none found, run bun scripts/main.ts --setup and walk the user through references/config/first-time-setup.md.
Verify at least one provider has credentials (env var or 1Password reference).
Verify Bun is available. Fallback: npx -y bun.
Step 1: Choose provider
Preference order:
--provider <id>
- EXTEND.md
default_provider
- Auto-detect env vars:
openai > elevenlabs > bailian > minimax > siliconflow > playht
Pick by language / voice intent:
- English, natural + fast →
openai (gpt-4o-mini-tts / tts-1).
- Multilingual, voice cloning →
elevenlabs.
- Chinese, long-form →
bailian (qwen-tts auto-chunks long scripts) or minimax.
- Chinese dialect / voice design →
bailian (voice-design with qwen3-tts-vd) or siliconflow (CosyVoice2).
- Ultra-realistic, short-form →
playht (2.0).
Step 2: Fill parameters
--text or --textfiles: input. Always quote.
--out <path>: REQUIRED. Extension determines format (.mp3 / .wav / .ogg / .flac).
--voice <id>: provider-specific. See references/voices.md for the short list of well-known voices.
--rate 0.5..2.0: speaking rate.
--instruction "...": voice direction (only openai gpt-4o-mini-tts and siliconflow honor this).
--language <code>: en, zh, ja — only a few providers honor this explicitly.
Step 3: Run
bun scripts/main.ts \
--provider openai \
--model gpt-4o-mini-tts \
--voice alloy \
--text "..." \
--out ./out.mp3
JSON mode:
{ "success": true, "provider": "openai", "model": "gpt-4o-mini-tts", "voice": "alloy", "output": "/abs/out.mp3", "size_bytes": 76032, "format": "mp3" }
Step 4: Long text handling
happy-audio-gen automatically splits long input for providers that cap per-call length (Bailian ≤ 200 Chinese chars per call). Chunks are concatenated byte-for-byte on output.
- For best fidelity with concatenated MP3s, stitch the segments with ffmpeg afterward rather than relying on byte concat.
Step 5: Errors
[openai] OpenAI TTS 400 with invalid voice → the voice name is not supported by the model. Use one of alloy, ash, coral, echo, fable, onyx, nova, sage, shimmer.
[minimax] ... 2049 invalid api key → try MINIMAX_BASE_URL=https://api.minimaxi.com/v1 (different region).
[bailian] ... 400 DataInspectionFailed → Aliyun content filter. Surface to the user.
[elevenlabs] 401 → key invalid or subscription expired.
References
references/providers.md — per-provider env vars, default models, voice lists.
references/voices.md — curated voices for each provider.
references/error_codes.md — common errors and fixes.
references/config/first-time-setup.md
references/config/extend-schema.md
assets/EXTEND.template.md
1---2name: happy-audio-gen3description: Universal AI voice / text-to-speech skill supporting OpenAI TTS (gpt-4o-mini-tts, tts-1), ElevenLabs multilingual TTS with voice cloning, Bailian Qwen TTS (qwen-tts / qwen3-tts-vd with voice-design custom voices, long-text chunking built in), MiniMax speech-02-hd, SiliconFlow CosyVoice / SenseVoice, and PlayHT 2.0. Use this skill whenever the user asks to read text aloud, synthesize speech, generate narration, create voice-over, dub a script, or turn any text into audio (mp3 / wav / ogg / flac). Typical phrases include "read this aloud", "generate voice for ...", "create a narration of ...", "tts this", "把这段念出来", "做个配音", "合成语音", or mentions of voices / TTS model names like Alloy, Ash, Cherry, Rachel, CosyVoice, PlayHT. Always use this skill even if the user does not specify a provider — pick one from EXTEND.md defaults or available env keys.4---56# happy-audio-gen78Turns text into speech across 6 providers through one CLI. All providers are synchronous (TTS is fast — typically under 10 seconds) except Bailian's voice-design flow (which is still covered but uses a longer poll window).910## Quick usage1112```bash13# Shortest path — OpenAI default voice14bun scripts/main.ts --text "Hello, world" --out ./hello.mp31516# Chinese, MiniMax17bun scripts/main.ts --provider minimax --text "大家好" --voice male-qn-qingse --out ./hello.mp31819# Long-form, Bailian (auto-splits by sentence)20bun scripts/main.ts --provider bailian --textfiles ./script.md --out ./narration.mp321```2223## When to invoke this skill2425- User asks to synthesize speech / TTS / read aloud / narrate / dub / make a voice-over.26- User asks to convert script / text / article into audio.27- User names a TTS voice or model.2829Do **not** route here when the user wants to transcribe audio → text (that's STT, different domain), or edit / mix audio files (use a dedicated audio editor).3031## Step 0: Preflight (BLOCKING)32331. **Locate EXTEND.md**:34 - `./.happy-skills/happy-audio-gen/EXTEND.md`35 - `$XDG_CONFIG_HOME/happy-skills/happy-audio-gen/EXTEND.md`36 - `~/.happy-skills/happy-audio-gen/EXTEND.md`3738 If none found, run `bun scripts/main.ts --setup` and walk the user through `references/config/first-time-setup.md`.39402. **Verify at least one provider has credentials** (env var or 1Password reference).41423. **Verify Bun** is available. Fallback: `npx -y bun`.4344## Step 1: Choose provider4546Preference order:47481. `--provider <id>`492. EXTEND.md `default_provider`503. Auto-detect env vars: `openai > elevenlabs > bailian > minimax > siliconflow > playht`5152Pick by language / voice intent:5354- **English, natural + fast** → `openai` (gpt-4o-mini-tts / tts-1).55- **Multilingual, voice cloning** → `elevenlabs`.56- **Chinese, long-form** → `bailian` (qwen-tts auto-chunks long scripts) or `minimax`.57- **Chinese dialect / voice design** → `bailian` (voice-design with qwen3-tts-vd) or `siliconflow` (CosyVoice2).58- **Ultra-realistic, short-form** → `playht` (2.0).5960## Step 2: Fill parameters6162- **`--text`** or **`--textfiles`**: input. Always quote.63- **`--out <path>`**: REQUIRED. Extension determines format (`.mp3` / `.wav` / `.ogg` / `.flac`).64- **`--voice <id>`**: provider-specific. See `references/voices.md` for the short list of well-known voices.65- **`--rate 0.5..2.0`**: speaking rate.66- **`--instruction "..."`**: voice direction (only `openai` gpt-4o-mini-tts and `siliconflow` honor this).67- **`--language <code>`**: `en`, `zh`, `ja` — only a few providers honor this explicitly.6869## Step 3: Run7071```bash72bun scripts/main.ts \73 --provider openai \74 --model gpt-4o-mini-tts \75 --voice alloy \76 --text "..." \77 --out ./out.mp378```7980JSON mode:8182```json83{ "success": true, "provider": "openai", "model": "gpt-4o-mini-tts", "voice": "alloy", "output": "/abs/out.mp3", "size_bytes": 76032, "format": "mp3" }84```8586## Step 4: Long text handling8788- `happy-audio-gen` automatically splits long input for providers that cap per-call length (Bailian ≤ 200 Chinese chars per call). Chunks are concatenated byte-for-byte on output.89- For best fidelity with concatenated MP3s, stitch the segments with ffmpeg afterward rather than relying on byte concat.9091## Step 5: Errors9293- `[openai] OpenAI TTS 400` with `invalid voice` → the voice name is not supported by the model. Use one of `alloy`, `ash`, `coral`, `echo`, `fable`, `onyx`, `nova`, `sage`, `shimmer`.94- `[minimax] ... 2049 invalid api key` → try `MINIMAX_BASE_URL=https://api.minimaxi.com/v1` (different region).95- `[bailian] ... 400 DataInspectionFailed` → Aliyun content filter. Surface to the user.96- `[elevenlabs] 401` → key invalid or subscription expired.9798## References99100- `references/providers.md` — per-provider env vars, default models, voice lists.101- `references/voices.md` — curated voices for each provider.102- `references/error_codes.md` — common errors and fixes.103- `references/config/first-time-setup.md`104- `references/config/extend-schema.md`105- `assets/EXTEND.template.md`