Audio Theater

Turn a dialogue, a story with dialogue, or a one-line idea into multi-character audio using Gemini TTS, with word-level timecoded transcription (WhisperX or local faster-whisper), realistic ambient/foley SFX (ElevenLabs or the sound-effects skill), and instrumental score/music (MiniMax Music 2.6). Optional stereo spatialization (pan + distance + movement) and an ffmpeg mix with sidechain ducking (music ducks under voices+SFX). Three modes - theater (dramatized radio play), lipsync (clean per-line clips <=15s plus a manifest for the seedance-2 skill), and podcast (two hosts with intro/bed/outro music). With music it delivers three tracks (full mix, music-only, and a no-music stem for animated storyboards) plus a timecoded transcript and per-line voice + SFX/music files. Use for an audio drama, audio theater / audioteatro / radioteatro, a dramatized dialogue, a score for a story/podcast/storyboard, voiceover for lipsync, a TTS conversation, or a podcast episode from a script or idea.

puntorigen 57d824d 18 files · 193.9 KB Updated

File contents

puntorigen/avatar-skills/tree/main/audio-theater commit 57d824d722

Frequently asked questions

npx skillmds@latest add puntorigen/audio-theater