Podcast production with generative tools
This skill covers producing a finished, publishable podcast episode when some or all of the
audio is machine-generated. It is provider-neutral: it names TTS engines, music libraries, and
hosts only as illustrative options, never as the method. The craft — format design, writing for
the ear, voice direction, structure, rights, loudness, delivery, disclosure, and QA — is the
same regardless of which tool renders the audio.
An episode is done when it (1) serves a defined show and audience, (2) sounds intentional and
consistent, (3) uses only audio you have the right to distribute, (4) meets the loudness and
file specs of its target platforms, (5) ships correct RSS metadata plus a transcript, (6)
carries any legally or platform-required AI disclosure, and (7) has passed a listen-through QA.
Skipping any of these is a defect, not a shortcut.
1. Scope and activation
Use this skill when the task is a spoken-word audio program: a topic explainer show, an
interview series, a daily/weekly news brief, a narrative/documentary episode, a two-host
conversational show, an internal briefing feed, or a show bible/format design for any of those.
Do not use it for:
- Music production — songs, beds, or stings as the deliverable (that is a music/audio-generation
concern). This skill consumes music and treats it as a rights and loudness problem.
- A single TTS line — an IVR prompt, a notification, one voiceover clip with no show context.
- Video where the picture leads — if the deliverable is a video and audio is a track under it,
the video workflow owns pacing and delivery. This skill applies when audio is the product.
If a request is ambiguous ("make an audio version of this article"), confirm whether the output
is a one-off narration or an episode of a show. The two need different structure and metadata.
2. Format design (do this before writing a word)
The single most consequential early decision is format, because format dictates script
style, voice count, structure, and length. Common podcast formats and what each demands
(production heuristic, drawn from standard podcast-production practice — see Sources):
| Format |
Typical length |
Voices |
Script style |
Best for |
| Two-host conversational |
20–60 min |
2 (chemistry critical) |
Outlined, not fully written |
Ongoing relationship with audience; opinion/commentary |
| Interview |
25–75 min |
Host + guest(s) |
Host questions scripted; answers live |
Expertise, guests, evergreen back-catalog |
| Solo / monologue |
5–30 min |
1 |
Fully or mostly scripted |
Teaching, essays, focused explainers |
| Narrative / documentary |
20–45 min |
Narrator + tape/characters |
Fully scripted, structured in acts |
Storytelling, high production value |
| News brief / daily |
3–10 min |
1–2 |
Tightly scripted, dense |
Recurring, time-sensitive, habit-forming |
For a fully-synthetic show (every voice is TTS), the interview and two-host formats are the
hardest to make convincing because they depend on turn-taking realism and chemistry; solo,
narrative-narrator, and news-brief formats are the most forgiving. Weigh this when advising a
format for a synthetic show.
Design the show once, in a short show bible, before designing episodes:
- Premise and audience — one sentence: who it is for and the promise each episode keeps.
- Cadence and length — a length band the format supports and a realistic release schedule.
- Voice identity — which voice(s), their names/personas, and a locked voice configuration so
episode N sounds like episode 1 (see §4).
- Signature elements — the intro line, the outro/call-to-action, the music theme, and any
recurring segments. These are the show's "sonic branding" and must be reused verbatim.
- Disclosure stance — decided up front (see §6), because it affects the script and metadata.
3. Writing for the ear
Scripts for audio are not text documents read aloud. The ear has no scrollbar, cannot re-read,
and loses complex clauses. Rewrite prose into speech (production heuristics, widely taught in
podcast scriptwriting — see Sources):
- Short sentences, one idea each. Break compound sentences. The listener holds only the last
clause in memory.
- Contractions and spoken register. "It's / you'll / here's," not "it is / you will / here
is." Write how a competent host actually talks.
- Front-load the point. Say the conclusion, then support it. Never bury the payoff behind a
long subordinate clause.
- Signpost transitions aloud. "Three things. First… / But here's the catch… / So what does
that mean?" These are the audio equivalent of headings.
- Kill the unpronounceable. Spell out or rephrase acronyms, symbols, URLs, and numbers.
"twenty-sixteen L-U-F-S," not "-16 LUFS." "dot com slash join," not "/join."
- Read it out loud. Anything that trips a human reader will trip a synthetic voice too and
will bore a listener. For a synthetic show, read-aloud testing is doubly important because you
cannot ad-lib a save in post.
How much to script depends on format. Fully script the high-risk moments regardless of
format: the hook/cold open, segment transitions, sponsor reads, and the close. For conversational
and interview shows, outline the middle so the talk stays natural. For solo, narrative, and
synthetic shows, script fully — a synthetic voice reads exactly what you give it, so the script
is the performance.
When the script is the performance (fully-synthetic), also write performance direction into
the script: mark intended pauses, emphasis, and emotional tone per paragraph, because those become
your TTS instructions in §4.
4. Multi-voice casting and TTS direction
This is the craft that separates a convincing synthetic episode from an obviously robotic one.
The goal is voices that are distinct from each other, consistent across episodes, correctly
pronounced, and — for dialogue — plausibly reactive to each other.
4.1 Casting and consistency
- Lock voice configuration per persona. Save the exact voice identifier plus every generation
parameter (stability/consistency, similarity, style/expressiveness, speed) as the persona's
"voice lock." Reuse it for every episode. Voice drift — the same nominal voice sounding
different across episodes or even across a long file — is the top continuity failure in
synthetic shows (practitioner observation, corroborated across TTS vendor guidance — see Sources).
- Cast for contrast. In a two-voice show, pick voices that differ in pitch, pace, and timbre
so listeners can tell who is speaking without name tags. Two similar voices are worse than one.
- Segment long scripts by speaker/role and generate each speaker's lines in that speaker's
locked config, then assemble. This preserves consistency better than switching voices mid-request
on engines that support only one primary voice per call.
- Higher stability/consistency settings reduce drift and expressiveness together. For a
narrator that must sound identical for 30 minutes, bias toward stability. For a character that
needs range, accept more variation and regenerate takes until consistent (heuristic).
4.2 Pronunciation control
- Fix names, jargon, and foreign words explicitly. Do not hope the model guesses. Use the
engine's pronunciation mechanism: phoneme markup (IPA or CMU Arpabet via SSML
<phoneme>), a
custom pronunciation/lexicon dictionary, or, as a last resort, phonetic respelling in the text
("KAI-roh" for "Cairo"). Documented fact: SSML <phoneme> supports IPA and CMU Arpabet on major
engines; some newer expressive models honor a lexicon or respelling but not full SSML — verify
per engine (verified 2026-07-10 against ElevenLabs, Google Cloud TTS, and Azure Speech docs).
- Build a per-show pronunciation glossary in the show bible (recurring names, the show title,
sponsor names) and apply it every episode so "the host's name" never changes pronunciation.
- Numbers, dates, units, and symbols are read inconsistently. Normalize them in the script
("July tenth, twenty-twenty-six"), which also helps human clarity.
4.3 Emotional pacing and performance
- Punctuation is your primary pacing tool. Periods = full stop; commas = short breath;
ellipses = a hanging pause; question marks lift the final pitch. Sentence length sets rhythm —
short sentences accelerate, long ones slow down. Around 140–160 words per minute reads as
natural, engaged speech (production heuristic — see Sources).
- Apply emotion at the paragraph/section level, not word-by-word. Set a tone for a passage;
use per-word emphasis sparingly. Stacking many emotion directions or switching every sentence
produces unnatural tonal lurches (practitioner heuristic corroborated across TTS best-practice
docs — see Sources).
- Use the engine's expressive controls deliberately: SSML
<prosody> (rate/pitch/volume),
<break time="…"> for engineered pauses, <emphasis>, or an engine's audio/emotion tags. Where
the engine is prompt-driven rather than SSML-driven, put the direction in a delivery note the
engine reads. Verify which markup your chosen engine actually honors before authoring — support
is uneven (verified 2026-07-10).
- Generate multiple takes and choose. Synthetic delivery is stochastic; the first take is
rarely the best. Budget for regeneration.
4.4 Host chemistry in generated dialogue
Chemistry between synthetic hosts must be engineered, because the voices are generated
independently and do not actually react. Techniques (heuristics):
- Write the reactions in. Interruptions ("—wait, say that again"), back-channels ("mm-hm,"
"right,"), and callbacks to earlier lines. If it is not in the script, it is not in the audio.
- Vary turn length. Real conversation is uneven — a long point, a two-word reply. Scripting
even, equal turns is a tell.
- Control the seams in assembly. Turn-taking realism lives in the gaps. Tighten or overlap
the joins between speaker clips so the reply does not sound like it was recorded in a different
room a week later. Consistent room tone under both voices sells co-presence.
- Do not fake a live remote. If you add fake "over the phone" filtering to sell realness, that
edges toward deception; keep it clearly a produced show.
5. Episode structure and assembly
5.1 Structure
A conventional episode spine (production heuristic; adapt to format):
- Cold open (optional, 10–30 s) — the single strongest moment or hook, before any branding.
Earns the listen. Strong for narrative and social-clip-driven shows.
- Intro / theme (5–20 s) — show name, host, one-line promise, over the theme music. Keep it
short and identical every episode (sonic branding).
- Episode tease — what this episode delivers, in one or two sentences.
- Body / segments — the content, broken into clearly transitioned segments. Signpost each
segment change with a spoken transition and, optionally, a short music sting.
- Ad slots — see §5.3. Mark them structurally so they can be inserted/removed or
dynamically served.
- Outro (15–30 s) — recap, call to action (subscribe/share/link), sign-off, theme out.
Narrative shows use an act structure inside the body (setup → complication/turn → resolution or
open question) rather than flat segments.
5.2 Assembly and editing
- Level the dialogue first, then place music and SFX under it. Speech intelligibility wins over
music every time.
- Duck music under speech (sidechain or manual automation) so beds sit roughly 12–18 dB below
the voice during talk and come up in the gaps (heuristic; tune by ear).
- Trim synthetic artifacts — clipped word-onsets, unnatural breaths, and swallowed
syllables are common in TTS. Cut or regenerate the offending line.
- Consistent room tone / silence between clips. Dead-digital silence between assembled TTS
clips sounds unnatural; a low consistent floor reads as one continuous recording.
- Match levels across segments so the intro, body, and ads are not wildly different volumes
before you do the final loudness pass (§7).
5.3 Ad slots
- Baked-in vs. dynamic. Baked-in ads are part of the file forever; dynamically inserted ads are
stitched at request time by the host and can be updated or removed. If the show will run ads long
term, structure the timeline with clean, silent insertion points so ads can be dynamic.
- Disclosure carries into ads. A synthetic-voice ad read, or an AI-generated endorsement, may
trigger the same disclosure and consent duties as the show (see §6), plus advertising-law rules
about endorsements. Never synthesize a real person appearing to endorse something without rights.
6. Fully-synthetic vs. hybrid, and the disclosure duty
6.1 The production decision
- Fully-synthetic — every voice is TTS. Cheapest and fastest to iterate; best for solo,
narrative-narrator, and news-brief formats; weakest for spontaneity and true interview dynamics.
- Hybrid — recorded human voice(s) plus synthetic segments (e.g., a real host with a synthetic
co-host, synthetic narration around recorded interview tape, or synthetic reconstruction of
unavailable audio). Best of both, but multiplies the rights and disclosure surface: every
recorded human needs a release, and every synthetic voice needs a licensed/consented source.
Choose hybrid when the show needs a real human's authority, spontaneity, or an actual guest, and
fully-synthetic when scale, consistency, or cost dominate and the format tolerates it.
6.2 Disclosure and consent obligations
Treat these as requirements, not style choices. They are the highest-risk part of a synthetic
podcast.
- Voice cloning requires documented consent from the voice owner. Cloning or replicating a
real person's voice without permission exposes you to right-of-publicity and voice-rights
liability. As of 2026 multiple U.S. states protect voice as a distinct likeness right — Tennessee's
ELVIS Act (effective 2024) is the first to name AI voice replicas explicitly, and California,
Illinois, New York and others have related statutes; the federal TAKE IT DOWN Act (signed May
- targets nonconsensual synthetic depictions (documented fact, verified 2026-07-10 — see
Sources; not legal advice — confirm current law for your jurisdiction and use). Consent must be
specific, written, and documented; a verbal "yes" does not meet the bar. Never clone a public
figure, a guest, or a co-host's voice without a signed release scoped to the use.
- Platform AI-disclosure policies apply to synthetic voices (documented facts, verified
2026-07-10):
- YouTube requires creators to disclose realistic altered or synthetic content at upload,
explicitly including synthetic/cloned voices and AI voiceovers that could mislead a viewer into
thinking a real person spoke. A "How this content was made" label may be shown. Mass-produced,
low-effort AI content risks demonetization under 2025 monetization updates.
- Spotify does not down-rank AI-assisted content per se but bans unauthorized voice clones,
aggressively removes spam/low-quality mass-produced audio, and is adopting a DDEX-based AI
disclosure standard surfaced in-app. Its "Verified" program excludes profiles that primarily
represent AI personas.
- Apple Podcasts does not currently mandate a generic "this is AI" label for synthetic voices
in the base RSS spec, but its content policies still prohibit impersonation and require rights
to all audio. (Verify current policy at publish time — platform policy is volatile.)
- When in doubt, disclose. A brief, honest note ("voices in this episode are AI-generated" in
the show notes and/or a spoken line) is cheap insurance and increasingly expected. Regulatory
momentum (e.g., transparency laws emerging in 2025–2026) is toward more mandatory disclosure,
not less.
- Hybrid shows: label which parts are synthetic if a listener could otherwise be deceived about
a real person's words — for example, a synthetically reconstructed quote must be marked as a
recreation, never presented as authentic tape.
7. Loudness and delivery standards
This is a hard-numbers area. Get it wrong and platforms turn your show up or down, or reject it.
7.1 The target
Documented facts (verified 2026-07-10):
- Apple Podcasts recommends preconditioning so overall loudness is around -16 dB LKFS
(= LUFS) with ±1 dB tolerance, and true peak ≤ -1 dB FS, measured per ITU-R BS.1770.
(LKFS and LUFS are the same unit.)
- The commonly cited creator targets are -16 LUFS for a stereo file and -19 LUFS for a mono
file. The mono figure is lower on purpose: a mono file played through a stereo player is
perceived roughly 3 LU quieter, so mastering a mono file to ~-19 LUFS makes it sound like a
-16 LUFS stereo file. (Production heuristic reconciling the two numbers; the ~2–3 LU
perceptual offset is documented in AES TD1008 — see Sources.)
- Spotify normalizes podcasts to -14 LUFS at playback with a -1 dBTP ceiling; YouTube
normalizes toward roughly -14 LUFS as well. These platforms apply gain at playback rather than
re-encoding your file.
- AES TD1008 (2021, supersedes TD1004) — the streaming-loudness recommendation — sets a
distribution target of -18 LUFS for speech/"assorted" content and advises keeping
integrated loudness above -20 LUFS; it notes speech is perceived ~2–3 LU louder than music at
the same measured loudness. Note this is guidance for the distributor's normalization, not the
creator's master file.
Practical resolution (heuristic): Master one stereo file to -16 LUFS integrated, true peak
≤ -1.0 dBTP as the portable default that sits inside Apple's window and is only gently adjusted
by Spotify/YouTube. If delivering mono, target -19 LUFS for equivalent perceived loudness.
Never chase Spotify's -14 by squashing dynamics — platforms turn quiet content up; loud, over-
limited content cannot be turned back down cleanly and sounds fatiguing. Preserve dialogue dynamic
range; a spoken-word show does not need to be as loud as a mastered song.
7.2 File and encoding specs
Documented facts — Apple Podcasts (verified 2026-07-10):
- RSS-delivered episode audio: MP3 or AAC. Mono 64–128 kbps, stereo 128–256 kbps, at
44.1/48 kHz.
- Spoken-word content is often fine in mono — smaller files, and most listening is on a single
speaker or earbud. Reserve stereo for shows with meaningful music/SFX staging.
- Loudness is set before encoding: lossy compression does not change measured loudness, so
precondition first, then encode.
- (Apple's high-resolution WAV/FLAC specs apply to subscriber audio uploaded to Podcasts Connect,
not to the public RSS enclosure.)
Deliver a single file per episode via the <enclosure> in the RSS item (URL + byte length + MIME
type, all three required).
8. Metadata, chapters, transcripts, and RSS
A podcast is an RSS feed; the audio is just the enclosure. Getting the feed right is as much of
the deliverable as the audio (documented facts about the spec, verified 2026-07-10 — see Sources).
- Feed-level required elements: title, description, language, at least one
<itunes:category>,
<itunes:explicit>, artwork (Apple requires cover art typically 1400×1400 to 3000×3000 px, RGB
JPEG/PNG), and an owner/contact. Missing <itunes:explicit> or <itunes:category> gets a feed
rejected by Apple.
- Episode-level required elements: title, a
<enclosure> (URL, length in bytes, MIME type), a
globally unique <guid> that never changes, and a publish date. Changing a GUID duplicates
the episode for subscribers.
- Chapters — two mechanisms: (a) embedded in the file (ID3v2
CHAP frames in MP3, MP4 chpl
atoms in M4A), or (b) linked from the feed via the Podcasting 2.0 <podcast:chapters> tag
pointing to an external JSON chapters file. Chapters improve navigation and are surfaced by
several clients.
- Transcripts — the Podcasting 2.0
<podcast:transcript> element links a transcript file from
the RSS item. WebVTT (.vtt) is the preferred format because it supports speaker labels
(via cue identifiers), timing, and light styling, and has the widest client support; SRT is
accepted but simpler. Major clients (Apple, and Spotify as of late 2025) render feed-linked
transcripts. For a synthetic show you effectively have the transcript already — it is your
script — so shipping one is nearly free and removes any excuse not to.
- Namespaces: declare the
itunes and podcast (Podcasting 2.0) namespaces in the feed to use
those tags; validate the feed before publishing.
9. Accessibility
- Ship a transcript for every episode. It is the core accessibility feature of podcasting — it
serves Deaf and hard-of-hearing listeners, improves discoverability/SEO, and is trivially cheap
for a synthetic show (your script). Link it via
<podcast:transcript> (§8).
- Speaker-labeled transcripts (WebVTT cue identifiers) matter for multi-voice shows so a reader
can follow who said what.
- Clear show notes with a summary, key timestamps/chapters, and link text that makes sense out
of context.
- Avoid audio-only information that a transcript cannot carry — if a laugh, tone, or sound is
load-bearing for meaning, note it in the transcript ("[laughs]", "[phone rings]").
10. Rights and safety checklist
Every audio element must be one of: original, licensed for this use, or public-domain/appropriately
Creative-Commons-licensed with attribution honored.
- Music. You need rights to both the composition and the specific recording (master). Buying a
song, or a personal streaming subscription, does not grant podcast/sync rights. Use a
royalty-free / production-music library license that explicitly covers podcasts, a direct sync
license from the rights holders, or genuinely license-clear music. "Royalty-free" means one fee,
not free-of-charge, and not free-of-license-terms — read the license scope (episodes covered,
ad-supported use, term) (documented fact — see Sources).
- Sound effects — same principle; use SFX libraries whose license covers commercial podcast
distribution.
- Voices — every recorded human needs a release; every cloned/synthetic voice needs a licensed
or consented source and must respect the TTS provider's commercial-use terms (some voices/tiers
are non-commercial). Re-confirm the provider grants you rights to distribute generated audio
commercially.
- Third-party audio / clips — quoting tape, songs, or another show requires permission or a
valid fair-use/fair-dealing basis; do not assume short clips are automatically fine.
- Disclosure/consent for synthetic and cloned voices — see §6. This is the one that creates
legal exposure, not just a takedown.
11. Pre-publish QA (do not skip)
Run this before every release. A synthetic show especially needs a human (or careful agent)
listen-through, because generation errors are silent until you hear them.
- Full listen-through end to end, at listening volume, on earbuds. Catch mispronunciations,
voice drift, clipped words, robotic seams, awkward pacing, and level jumps.
- Pronunciation check against the show glossary — every name, the show title, the sponsor.
- Structure check — intro/outro present and identical to prior episodes; segments transition;
ad slots correct; cold open lands.
- Loudness/peak verification with a meter — integrated LUFS at target (~-16 LUFS stereo /
~-19 LUFS mono), true peak ≤ -1 dBTP, no clipping (§7).
- Encoding/format — correct codec, bitrate, channels, sample rate; file plays on a phone.
- Metadata — title, unique unchanged GUID, description, chapters, correct enclosure length in
bytes, artwork, categories, explicit flag (§8).
- Transcript attached, accurate, speaker-labeled (§9).
- Rights — every music/SFX/voice element cleared (§10).
- Disclosure — AI/synthetic disclosure present where required by law or platform, and consent
on file for any cloned voice (§6).
- Feed validation — run the feed through a validator; confirm it parses before it goes live.
12. Worked examples
The following are illustrative examples, not mandatory templates. They show how the decisions
above compose on a real brief. Adapt the specifics.
Example A — Fully-synthetic daily 5-minute news brief
Intent: A recurring weekday "3-minute-ish" brief summarizing one industry's news, single
synthetic host, published to Apple/Spotify/YouTube.
Format & bible: News-brief format (forgiving for synthetic). One host persona "Ava," a
mid-pitch, steady voice locked at high stability (news must sound identical daily). Glossary of
recurring company and person names with phoneme entries. Fixed 8-second music intro, 6-second
outro with a "subscribe" CTA. Disclosure stance: state "This is an AI-generated briefing" in the
show description and once in the outro.
Script (writing for the ear), abridged:
[COLD OPEN — no music]
Three things moved the market today, and the third one nobody saw coming.
[THEME 8s, duck under]
This is The Daily Brief for Tuesday, July eighth. I'm Ava. Here's what matters.
[BODY]
First. [company], read "ACK-mee", reported earnings after the bell...
(short sentences. One idea each. A pause... before the turn.)
[OUTRO, theme up]
That's your brief. Follow the show so tomorrow's finds you automatically.
This is an AI-generated briefing from [publisher]. See you tomorrow.
TTS direction: Generate the whole body in Ava's lock. Emotion at paragraph level: neutral-
authoritative for the body, a lift on the outro CTA. Numbers/dates spelled out in the script. Two
takes per day; pick the cleaner. Loudness: master mono to -19 LUFS, TP ≤ -1 dBTP; encode MP3
mono 96 kbps 44.1 kHz. Delivery: unique GUID per day; <podcast:transcript> VTT built from the
script; chapters unnecessary at this length. Disclosure: in description + spoken outro (covers
YouTube's synthetic-voice disclosure and general transparency).
Likely failure modes: drift if the voice lock is not reused; "ACK-mee" reverting to a spelling
pronunciation without the phoneme entry; over-limiting to chase Spotify's -14 and sounding harsh.
Example B — Hybrid two-host conversational tech show
Intent: Weekly 35-minute show. One real human host (recorded) plus one synthetic co-host
persona. This is the hard case: chemistry and disclosure both matter.
Decision: Hybrid, because the human brings spontaneity/authority and the synthetic co-host adds
a consistent "explainer" foil cheaply. Consent/rights: the synthetic co-host uses a voice the
producers licensed for commercial use (documented, provider terms confirmed); if it were modeled
on any real person, a signed release would be mandatory — it is not, it is a wholly synthetic
persona. Disclosure: the synthetic co-host is introduced honestly as an AI co-host in episode 1
and in every show description; a "How this content was made" disclosure is toggled on at YouTube
upload because a synthetic voice is present.
Chemistry engineering: The human's side is recorded live around a scripted skeleton. The
synthetic co-host's lines are written to react — interruptions, back-channels ("right, right"),
callbacks to what the human just said — then generated in the co-host's voice lock and edited into
the gaps with matched room tone so the exchange sounds co-present. Turn lengths are deliberately
uneven.
Assembly & loudness: Level both voices to match, duck the bed 12–15 dB under talk, master
stereo to -16 LUFS, TP ≤ -1 dBTP; encode AAC stereo 128 kbps. Delivery: chapters (JSON via
<podcast:chapters>) for the segment structure; speaker-labeled WebVTT transcript distinguishing
the human host and the AI co-host; clean silent ad-insertion points for dynamic ads.
Likely failure modes: the synthetic replies sounding recorded "elsewhere" (fix room tone and
tighten seams); even, equal turn lengths reading as fake; forgetting the platform AI-disclosure
toggle; the synthetic co-host reading a sponsor endorsement without the extra care ad reads require.
Example C — Narrative documentary episode with reconstructed audio
Intent: A 30-minute story where an unavailable historical quote is voiced synthetically.
The critical rule: A synthetically reconstructed voice or quote must be disclosed as a
recreation — never presented as authentic archival tape. Mark it in the narration ("what follows
is an AI re-creation of the words in the letter, not a recording") and in the transcript. Cloning a
specific real person's voice would additionally require rights/consent from their estate; a neutral
"reader" voice avoids that and is the safer choice unless rights are secured. Everything else
follows the narrative act structure (§5.1), full scripting (§3), and the standard rights, loudness
(-16 LUFS stereo), transcript, and QA passes.
Sources
Volatile facts (platform specs, policies, laws) verified 2026-07-10. Standards and craft
heuristics are labeled inline. Do not treat legal notes as legal advice; confirm current law and
policy at publish time.
Standards & loudness
RSS / metadata / transcripts / chapters
Platform AI-content policies
Voice rights & disclosure law (not legal advice)
TTS direction (best-practice docs, provider-neutral synthesis)
Music/SFX rights (practitioner references)
Format, structure & scriptwriting (practitioner references)
1---2name: podcast-production3description: Produce audio-first podcast episodes with generative tools — design the show and episode format (interview, narrative, news brief, two-host conversational), write scripts for the ear, decide between fully-synthetic and hybrid (recorded human + synthetic) production and meet the disclosure duty each triggers, cast and direct multi-voice TTS for consistency and chemistry, build episode structure (cold open, intro/outro, segments, ad slots), assemble and edit, clear music and SFX rights, hit podcast loudness and delivery standards, ship metadata/chapters/transcripts through RSS, and pass platform AI-content policy and QA before publishing. Use this skill whenever the deliverable is a podcast episode or a podcast show bible, or when an agent must make production, rights, disclosure, or delivery decisions for spoken-word audio. Not for music tracks, single-voice notification prompts, or video where picture leads (route audio-for-video to a video skill).4---56# Podcast production with generative tools78This skill covers producing a finished, publishable podcast episode when some or all of the9audio is machine-generated. It is provider-neutral: it names TTS engines, music libraries, and10hosts only as illustrative options, never as the method. The craft — format design, writing for11the ear, voice direction, structure, rights, loudness, delivery, disclosure, and QA — is the12same regardless of which tool renders the audio.1314An episode is *done* when it (1) serves a defined show and audience, (2) sounds intentional and15consistent, (3) uses only audio you have the right to distribute, (4) meets the loudness and16file specs of its target platforms, (5) ships correct RSS metadata plus a transcript, (6)17carries any legally or platform-required AI disclosure, and (7) has passed a listen-through QA.18Skipping any of these is a defect, not a shortcut.1920---2122## 1. Scope and activation2324Use this skill when the task is a spoken-word audio program: a topic explainer show, an25interview series, a daily/weekly news brief, a narrative/documentary episode, a two-host26conversational show, an internal briefing feed, or a show bible/format design for any of those.2728Do **not** use it for:2930- **Music production** — songs, beds, or stings as the deliverable (that is a music/audio-generation31 concern). This skill *consumes* music and treats it as a rights and loudness problem.32- **A single TTS line** — an IVR prompt, a notification, one voiceover clip with no show context.33- **Video where the picture leads** — if the deliverable is a video and audio is a track under it,34 the video workflow owns pacing and delivery. This skill applies when audio is the product.3536If a request is ambiguous ("make an audio version of this article"), confirm whether the output37is a one-off narration or an episode of a show. The two need different structure and metadata.3839---4041## 2. Format design (do this before writing a word)4243The single most consequential early decision is **format**, because format dictates script44style, voice count, structure, and length. Common podcast formats and what each demands45(production heuristic, drawn from standard podcast-production practice — see Sources):4647| Format | Typical length | Voices | Script style | Best for |48|---|---|---|---|---|49| **Two-host conversational** | 20–60 min | 2 (chemistry critical) | Outlined, not fully written | Ongoing relationship with audience; opinion/commentary |50| **Interview** | 25–75 min | Host + guest(s) | Host questions scripted; answers live | Expertise, guests, evergreen back-catalog |51| **Solo / monologue** | 5–30 min | 1 | Fully or mostly scripted | Teaching, essays, focused explainers |52| **Narrative / documentary** | 20–45 min | Narrator + tape/characters | Fully scripted, structured in acts | Storytelling, high production value |53| **News brief / daily** | 3–10 min | 1–2 | Tightly scripted, dense | Recurring, time-sensitive, habit-forming |5455For a **fully-synthetic** show (every voice is TTS), the interview and two-host formats are the56hardest to make convincing because they depend on turn-taking realism and chemistry; solo,57narrative-narrator, and news-brief formats are the most forgiving. Weigh this when advising a58format for a synthetic show.5960Design the **show** once, in a short show bible, before designing episodes:6162- **Premise and audience** — one sentence: who it is for and the promise each episode keeps.63- **Cadence and length** — a length band the format supports and a realistic release schedule.64- **Voice identity** — which voice(s), their names/personas, and a locked voice configuration so65 episode N sounds like episode 1 (see §4).66- **Signature elements** — the intro line, the outro/call-to-action, the music theme, and any67 recurring segments. These are the show's "sonic branding" and must be reused verbatim.68- **Disclosure stance** — decided up front (see §6), because it affects the script and metadata.6970---7172## 3. Writing for the ear7374Scripts for audio are not text documents read aloud. The ear has no scrollbar, cannot re-read,75and loses complex clauses. Rewrite prose into speech (production heuristics, widely taught in76podcast scriptwriting — see Sources):7778- **Short sentences, one idea each.** Break compound sentences. The listener holds only the last79 clause in memory.80- **Contractions and spoken register.** "It's / you'll / here's," not "it is / you will / here81 is." Write how a competent host actually talks.82- **Front-load the point.** Say the conclusion, then support it. Never bury the payoff behind a83 long subordinate clause.84- **Signpost transitions aloud.** "Three things. First… / But here's the catch… / So what does85 that mean?" These are the audio equivalent of headings.86- **Kill the unpronounceable.** Spell out or rephrase acronyms, symbols, URLs, and numbers.87 "twenty-sixteen L-U-F-S," not "-16 LUFS." "dot com slash join," not "/join."88- **Read it out loud.** Anything that trips a human reader will trip a synthetic voice too and89 will bore a listener. For a synthetic show, read-aloud testing is doubly important because you90 cannot ad-lib a save in post.9192**How much to script depends on format.** Fully script the high-risk moments regardless of93format: the hook/cold open, segment transitions, sponsor reads, and the close. For conversational94and interview shows, outline the middle so the talk stays natural. For solo, narrative, and95synthetic shows, script fully — a synthetic voice reads exactly what you give it, so the script96*is* the performance.9798When the script is the performance (fully-synthetic), also write **performance direction** into99the script: mark intended pauses, emphasis, and emotional tone per paragraph, because those become100your TTS instructions in §4.101102---103104## 4. Multi-voice casting and TTS direction105106This is the craft that separates a convincing synthetic episode from an obviously robotic one.107The goal is voices that are *distinct from each other*, *consistent across episodes*, correctly108*pronounced*, and — for dialogue — plausibly *reactive to each other*.109110### 4.1 Casting and consistency111112- **Lock voice configuration per persona.** Save the exact voice identifier plus every generation113 parameter (stability/consistency, similarity, style/expressiveness, speed) as the persona's114 "voice lock." Reuse it for every episode. Voice **drift** — the same nominal voice sounding115 different across episodes or even across a long file — is the top continuity failure in116 synthetic shows (practitioner observation, corroborated across TTS vendor guidance — see Sources).117- **Cast for contrast.** In a two-voice show, pick voices that differ in pitch, pace, and timbre118 so listeners can tell who is speaking without name tags. Two similar voices are worse than one.119- **Segment long scripts by speaker/role** and generate each speaker's lines in that speaker's120 locked config, then assemble. This preserves consistency better than switching voices mid-request121 on engines that support only one primary voice per call.122- **Higher stability/consistency settings reduce drift and expressiveness together.** For a123 narrator that must sound identical for 30 minutes, bias toward stability. For a character that124 needs range, accept more variation and regenerate takes until consistent (heuristic).125126### 4.2 Pronunciation control127128- **Fix names, jargon, and foreign words explicitly.** Do not hope the model guesses. Use the129 engine's pronunciation mechanism: phoneme markup (IPA or CMU Arpabet via SSML `<phoneme>`), a130 custom pronunciation/lexicon dictionary, or, as a last resort, phonetic respelling in the text131 ("KAI-roh" for "Cairo"). Documented fact: SSML `<phoneme>` supports IPA and CMU Arpabet on major132 engines; some newer expressive models honor a lexicon or respelling but not full SSML — verify133 per engine (verified 2026-07-10 against ElevenLabs, Google Cloud TTS, and Azure Speech docs).134- **Build a per-show pronunciation glossary** in the show bible (recurring names, the show title,135 sponsor names) and apply it every episode so "the host's name" never changes pronunciation.136- **Numbers, dates, units, and symbols** are read inconsistently. Normalize them in the script137 ("July tenth, twenty-twenty-six"), which also helps human clarity.138139### 4.3 Emotional pacing and performance140141- **Punctuation is your primary pacing tool.** Periods = full stop; commas = short breath;142 ellipses = a hanging pause; question marks lift the final pitch. Sentence length sets rhythm —143 short sentences accelerate, long ones slow down. Around **140–160 words per minute** reads as144 natural, engaged speech (production heuristic — see Sources).145- **Apply emotion at the paragraph/section level, not word-by-word.** Set a tone for a passage;146 use per-word emphasis sparingly. Stacking many emotion directions or switching every sentence147 produces unnatural tonal lurches (practitioner heuristic corroborated across TTS best-practice148 docs — see Sources).149- **Use the engine's expressive controls deliberately:** SSML `<prosody>` (rate/pitch/volume),150 `<break time="…">` for engineered pauses, `<emphasis>`, or an engine's audio/emotion tags. Where151 the engine is prompt-driven rather than SSML-driven, put the direction in a delivery note the152 engine reads. Verify which markup your chosen engine actually honors before authoring — support153 is uneven (verified 2026-07-10).154- **Generate multiple takes and choose.** Synthetic delivery is stochastic; the first take is155 rarely the best. Budget for regeneration.156157### 4.4 Host chemistry in generated dialogue158159Chemistry between synthetic hosts must be *engineered*, because the voices are generated160independently and do not actually react. Techniques (heuristics):161162- **Write the reactions in.** Interruptions ("—wait, say that again"), back-channels ("mm-hm,"163 "right,"), and callbacks to earlier lines. If it is not in the script, it is not in the audio.164- **Vary turn length.** Real conversation is uneven — a long point, a two-word reply. Scripting165 even, equal turns is a tell.166- **Control the seams in assembly.** Turn-taking realism lives in the *gaps*. Tighten or overlap167 the joins between speaker clips so the reply does not sound like it was recorded in a different168 room a week later. Consistent room tone under both voices sells co-presence.169- **Do not fake a live remote.** If you add fake "over the phone" filtering to sell realness, that170 edges toward deception; keep it clearly a produced show.171172---173174## 5. Episode structure and assembly175176### 5.1 Structure177178A conventional episode spine (production heuristic; adapt to format):1791801. **Cold open (optional, 10–30 s)** — the single strongest moment or hook, before any branding.181 Earns the listen. Strong for narrative and social-clip-driven shows.1822. **Intro / theme (5–20 s)** — show name, host, one-line promise, over the theme music. Keep it183 short and *identical* every episode (sonic branding).1843. **Episode tease** — what this episode delivers, in one or two sentences.1854. **Body / segments** — the content, broken into clearly transitioned segments. Signpost each186 segment change with a spoken transition and, optionally, a short music sting.1875. **Ad slots** — see §5.3. Mark them structurally so they can be inserted/removed or188 dynamically served.1896. **Outro (15–30 s)** — recap, call to action (subscribe/share/link), sign-off, theme out.190191Narrative shows use an act structure inside the body (setup → complication/turn → resolution or192open question) rather than flat segments.193194### 5.2 Assembly and editing195196- **Level the dialogue first**, then place music and SFX under it. Speech intelligibility wins over197 music every time.198- **Duck music under speech** (sidechain or manual automation) so beds sit roughly 12–18 dB below199 the voice during talk and come up in the gaps (heuristic; tune by ear).200- **Trim synthetic artifacts** — clipped word-onsets, unnatural breaths, and swallowed201 syllables are common in TTS. Cut or regenerate the offending line.202- **Consistent room tone / silence** between clips. Dead-digital silence between assembled TTS203 clips sounds unnatural; a low consistent floor reads as one continuous recording.204- **Match levels across segments** so the intro, body, and ads are not wildly different volumes205 before you do the final loudness pass (§7).206207### 5.3 Ad slots208209- **Baked-in vs. dynamic.** Baked-in ads are part of the file forever; dynamically inserted ads are210 stitched at request time by the host and can be updated or removed. If the show will run ads long211 term, structure the timeline with clean, silent insertion points so ads can be dynamic.212- **Disclosure carries into ads.** A synthetic-voice ad read, or an AI-generated endorsement, may213 trigger the same disclosure and consent duties as the show (see §6), plus advertising-law rules214 about endorsements. Never synthesize a real person appearing to endorse something without rights.215216---217218## 6. Fully-synthetic vs. hybrid, and the disclosure duty219220### 6.1 The production decision221222- **Fully-synthetic** — every voice is TTS. Cheapest and fastest to iterate; best for solo,223 narrative-narrator, and news-brief formats; weakest for spontaneity and true interview dynamics.224- **Hybrid** — recorded human voice(s) plus synthetic segments (e.g., a real host with a synthetic225 co-host, synthetic narration around recorded interview tape, or synthetic reconstruction of226 unavailable audio). Best of both, but multiplies the rights and disclosure surface: every227 recorded human needs a release, and every synthetic voice needs a licensed/consented source.228229Choose hybrid when the show needs a real human's authority, spontaneity, or an actual guest, and230fully-synthetic when scale, consistency, or cost dominate and the format tolerates it.231232### 6.2 Disclosure and consent obligations233234Treat these as **requirements**, not style choices. They are the highest-risk part of a synthetic235podcast.236237- **Voice cloning requires documented consent from the voice owner.** Cloning or replicating a238 real person's voice without permission exposes you to right-of-publicity and voice-rights239 liability. As of 2026 multiple U.S. states protect voice as a distinct likeness right — Tennessee's240 **ELVIS Act** (effective 2024) is the first to name AI voice replicas explicitly, and California,241 Illinois, New York and others have related statutes; the federal **TAKE IT DOWN Act** (signed May242 2025) targets nonconsensual synthetic depictions (documented fact, verified 2026-07-10 — see243 Sources; not legal advice — confirm current law for your jurisdiction and use). **Consent must be244 specific, written, and documented; a verbal "yes" does not meet the bar.** Never clone a public245 figure, a guest, or a co-host's voice without a signed release scoped to the use.246- **Platform AI-disclosure policies apply to synthetic voices** (documented facts, verified247 2026-07-10):248 - **YouTube** requires creators to disclose *realistic* altered or synthetic content at upload,249 explicitly including synthetic/cloned voices and AI voiceovers that could mislead a viewer into250 thinking a real person spoke. A "How this content was made" label may be shown. Mass-produced,251 low-effort AI content risks demonetization under 2025 monetization updates.252 - **Spotify** does not down-rank AI-assisted content per se but bans unauthorized voice clones,253 aggressively removes spam/low-quality mass-produced audio, and is adopting a DDEX-based AI254 disclosure standard surfaced in-app. Its "Verified" program excludes profiles that primarily255 represent AI personas.256 - **Apple Podcasts** does not currently mandate a generic "this is AI" label for synthetic voices257 in the base RSS spec, but its content policies still prohibit impersonation and require rights258 to all audio. (Verify current policy at publish time — platform policy is volatile.)259- **When in doubt, disclose.** A brief, honest note ("voices in this episode are AI-generated" in260 the show notes and/or a spoken line) is cheap insurance and increasingly expected. Regulatory261 momentum (e.g., transparency laws emerging in 2025–2026) is toward *more* mandatory disclosure,262 not less.263- **Hybrid shows: label which parts are synthetic** if a listener could otherwise be deceived about264 a real person's words — for example, a synthetically reconstructed quote must be marked as a265 recreation, never presented as authentic tape.266267---268269## 7. Loudness and delivery standards270271This is a hard-numbers area. Get it wrong and platforms turn your show up or down, or reject it.272273### 7.1 The target274275**Documented facts (verified 2026-07-10):**276277- **Apple Podcasts** recommends preconditioning so overall loudness is **around -16 dB LKFS278 (= LUFS) with ±1 dB tolerance**, and **true peak ≤ -1 dB FS**, measured per ITU-R BS.1770.279 (LKFS and LUFS are the same unit.)280- **The commonly cited creator targets are -16 LUFS for a stereo file and -19 LUFS for a mono281 file.** The mono figure is lower on purpose: a mono file played through a stereo player is282 perceived roughly 3 LU quieter, so mastering a mono file to ~-19 LUFS makes it *sound* like a283 -16 LUFS stereo file. (Production heuristic reconciling the two numbers; the ~2–3 LU284 perceptual offset is documented in AES TD1008 — see Sources.)285- **Spotify** normalizes podcasts to **-14 LUFS** at playback with a -1 dBTP ceiling; **YouTube**286 normalizes toward roughly -14 LUFS as well. These platforms apply gain at playback rather than287 re-encoding your file.288- **AES TD1008 (2021, supersedes TD1004)** — the streaming-loudness recommendation — sets a289 **distribution** target of **-18 LUFS for speech/"assorted" content** and advises keeping290 integrated loudness **above -20 LUFS**; it notes speech is perceived ~2–3 LU louder than music at291 the same measured loudness. Note this is guidance for the *distributor's normalization*, not the292 creator's master file.293294**Practical resolution (heuristic):** Master **one stereo file to -16 LUFS integrated, true peak295≤ -1.0 dBTP** as the portable default that sits inside Apple's window and is only gently adjusted296by Spotify/YouTube. If delivering **mono**, target **-19 LUFS** for equivalent perceived loudness.297Never chase Spotify's -14 by squashing dynamics — platforms turn quiet content *up*; loud, over-298limited content cannot be turned back down cleanly and sounds fatiguing. Preserve dialogue dynamic299range; a spoken-word show does not need to be as loud as a mastered song.300301### 7.2 File and encoding specs302303**Documented facts — Apple Podcasts (verified 2026-07-10):**304305- **RSS-delivered episode audio:** MP3 or AAC. Mono **64–128 kbps**, stereo **128–256 kbps**, at306 44.1/48 kHz.307- Spoken-word content is often fine in **mono** — smaller files, and most listening is on a single308 speaker or earbud. Reserve stereo for shows with meaningful music/SFX staging.309- Loudness is set *before* encoding: lossy compression does not change measured loudness, so310 precondition first, then encode.311- (Apple's high-resolution WAV/FLAC specs apply to subscriber audio uploaded to Podcasts Connect,312 not to the public RSS enclosure.)313314Deliver a single file per episode via the `<enclosure>` in the RSS item (URL + byte length + MIME315type, all three required).316317---318319## 8. Metadata, chapters, transcripts, and RSS320321A podcast *is* an RSS feed; the audio is just the enclosure. Getting the feed right is as much of322the deliverable as the audio (documented facts about the spec, verified 2026-07-10 — see Sources).323324- **Feed-level required elements:** title, description, language, at least one `<itunes:category>`,325 `<itunes:explicit>`, artwork (Apple requires cover art typically 1400×1400 to 3000×3000 px, RGB326 JPEG/PNG), and an owner/contact. Missing `<itunes:explicit>` or `<itunes:category>` gets a feed327 rejected by Apple.328- **Episode-level required elements:** title, a `<enclosure>` (URL, length in bytes, MIME type), a329 **globally unique `<guid>` that never changes**, and a publish date. Changing a GUID duplicates330 the episode for subscribers.331- **Chapters** — two mechanisms: (a) embedded in the file (ID3v2 `CHAP` frames in MP3, MP4 `chpl`332 atoms in M4A), or (b) linked from the feed via the Podcasting 2.0 `<podcast:chapters>` tag333 pointing to an external JSON chapters file. Chapters improve navigation and are surfaced by334 several clients.335- **Transcripts** — the Podcasting 2.0 `<podcast:transcript>` element links a transcript file from336 the RSS item. **WebVTT (`.vtt`) is the preferred format** because it supports speaker labels337 (via cue identifiers), timing, and light styling, and has the widest client support; SRT is338 accepted but simpler. Major clients (Apple, and Spotify as of late 2025) render feed-linked339 transcripts. For a **synthetic** show you effectively have the transcript already — it is your340 script — so shipping one is nearly free and removes any excuse not to.341- **Namespaces:** declare the `itunes` and `podcast` (Podcasting 2.0) namespaces in the feed to use342 those tags; validate the feed before publishing.343344---345346## 9. Accessibility347348- **Ship a transcript for every episode.** It is the core accessibility feature of podcasting — it349 serves Deaf and hard-of-hearing listeners, improves discoverability/SEO, and is trivially cheap350 for a synthetic show (your script). Link it via `<podcast:transcript>` (§8).351- **Speaker-labeled transcripts** (WebVTT cue identifiers) matter for multi-voice shows so a reader352 can follow who said what.353- **Clear show notes** with a summary, key timestamps/chapters, and link text that makes sense out354 of context.355- **Avoid audio-only information that a transcript cannot carry** — if a laugh, tone, or sound is356 load-bearing for meaning, note it in the transcript ("[laughs]", "[phone rings]").357358---359360## 10. Rights and safety checklist361362Every audio element must be one of: original, licensed for this use, or public-domain/appropriately363Creative-Commons-licensed with attribution honored.364365- **Music.** You need rights to *both* the composition and the specific recording (master). Buying a366 song, or a personal streaming subscription, does **not** grant podcast/sync rights. Use a367 royalty-free / production-music library license that explicitly covers podcasts, a direct sync368 license from the rights holders, or genuinely license-clear music. "Royalty-free" means one fee,369 not free-of-charge, and not free-of-license-terms — read the license scope (episodes covered,370 ad-supported use, term) (documented fact — see Sources).371- **Sound effects** — same principle; use SFX libraries whose license covers commercial podcast372 distribution.373- **Voices** — every recorded human needs a release; every cloned/synthetic voice needs a licensed374 or consented source and must respect the TTS provider's commercial-use terms (some voices/tiers375 are non-commercial). Re-confirm the provider grants you rights to *distribute* generated audio376 commercially.377- **Third-party audio / clips** — quoting tape, songs, or another show requires permission or a378 valid fair-use/fair-dealing basis; do not assume short clips are automatically fine.379- **Disclosure/consent** for synthetic and cloned voices — see §6. This is the one that creates380 legal exposure, not just a takedown.381382---383384## 11. Pre-publish QA (do not skip)385386Run this before every release. A synthetic show especially needs a human (or careful agent)387listen-through, because generation errors are silent until you hear them.3883891. **Full listen-through** end to end, at listening volume, on earbuds. Catch mispronunciations,390 voice drift, clipped words, robotic seams, awkward pacing, and level jumps.3912. **Pronunciation check** against the show glossary — every name, the show title, the sponsor.3923. **Structure check** — intro/outro present and identical to prior episodes; segments transition;393 ad slots correct; cold open lands.3944. **Loudness/peak verification** with a meter — integrated LUFS at target (~-16 LUFS stereo /395 ~-19 LUFS mono), true peak ≤ -1 dBTP, no clipping (§7).3965. **Encoding/format** — correct codec, bitrate, channels, sample rate; file plays on a phone.3976. **Metadata** — title, unique unchanged GUID, description, chapters, correct enclosure length in398 bytes, artwork, categories, explicit flag (§8).3997. **Transcript** attached, accurate, speaker-labeled (§9).4008. **Rights** — every music/SFX/voice element cleared (§10).4019. **Disclosure** — AI/synthetic disclosure present where required by law or platform, and consent402 on file for any cloned voice (§6).40310. **Feed validation** — run the feed through a validator; confirm it parses before it goes live.404405---406407## 12. Worked examples408409The following are **illustrative examples**, not mandatory templates. They show how the decisions410above compose on a real brief. Adapt the specifics.411412### Example A — Fully-synthetic daily 5-minute news brief413414**Intent:** A recurring weekday "3-minute-ish" brief summarizing one industry's news, single415synthetic host, published to Apple/Spotify/YouTube.416417**Format & bible:** News-brief format (forgiving for synthetic). One host persona "Ava," a418mid-pitch, steady voice locked at high stability (news must sound identical daily). Glossary of419recurring company and person names with phoneme entries. Fixed 8-second music intro, 6-second420outro with a "subscribe" CTA. Disclosure stance: state "This is an AI-generated briefing" in the421show description and once in the outro.422423**Script (writing for the ear), abridged:**424425```426[COLD OPEN — no music]427Three things moved the market today, and the third one nobody saw coming.428429[THEME 8s, duck under]430This is The Daily Brief for Tuesday, July eighth. I'm Ava. Here's what matters.431432[BODY]433First. [company], read "ACK-mee", reported earnings after the bell...434(short sentences. One idea each. A pause... before the turn.)435436[OUTRO, theme up]437That's your brief. Follow the show so tomorrow's finds you automatically.438This is an AI-generated briefing from [publisher]. See you tomorrow.439```440441**TTS direction:** Generate the whole body in Ava's lock. Emotion at paragraph level: neutral-442authoritative for the body, a lift on the outro CTA. Numbers/dates spelled out in the script. Two443takes per day; pick the cleaner. **Loudness:** master mono to -19 LUFS, TP ≤ -1 dBTP; encode MP3444mono 96 kbps 44.1 kHz. **Delivery:** unique GUID per day; `<podcast:transcript>` VTT built from the445script; chapters unnecessary at this length. **Disclosure:** in description + spoken outro (covers446YouTube's synthetic-voice disclosure and general transparency).447448**Likely failure modes:** drift if the voice lock is not reused; "ACK-mee" reverting to a spelling449pronunciation without the phoneme entry; over-limiting to chase Spotify's -14 and sounding harsh.450451### Example B — Hybrid two-host conversational tech show452453**Intent:** Weekly 35-minute show. One **real** human host (recorded) plus one **synthetic** co-host454persona. This is the hard case: chemistry and disclosure both matter.455456**Decision:** Hybrid, because the human brings spontaneity/authority and the synthetic co-host adds457a consistent "explainer" foil cheaply. **Consent/rights:** the synthetic co-host uses a voice the458producers **licensed for commercial use** (documented, provider terms confirmed); if it were modeled459on any real person, a signed release would be mandatory — it is not, it is a wholly synthetic460persona. **Disclosure:** the synthetic co-host is introduced honestly as an AI co-host in episode 1461and in every show description; a "How this content was made" disclosure is toggled on at YouTube462upload because a synthetic voice is present.463464**Chemistry engineering:** The human's side is recorded live around a scripted skeleton. The465synthetic co-host's lines are written to *react* — interruptions, back-channels ("right, right"),466callbacks to what the human just said — then generated in the co-host's voice lock and edited into467the gaps with matched room tone so the exchange sounds co-present. Turn lengths are deliberately468uneven.469470**Assembly & loudness:** Level both voices to match, duck the bed 12–15 dB under talk, master471**stereo to -16 LUFS**, TP ≤ -1 dBTP; encode AAC stereo 128 kbps. **Delivery:** chapters (JSON via472`<podcast:chapters>`) for the segment structure; speaker-labeled WebVTT transcript distinguishing473the human host and the AI co-host; clean silent ad-insertion points for dynamic ads.474475**Likely failure modes:** the synthetic replies sounding recorded "elsewhere" (fix room tone and476tighten seams); even, equal turn lengths reading as fake; forgetting the platform AI-disclosure477toggle; the synthetic co-host reading a sponsor endorsement without the extra care ad reads require.478479### Example C — Narrative documentary episode with reconstructed audio480481**Intent:** A 30-minute story where an unavailable historical quote is voiced synthetically.482483**The critical rule:** A synthetically reconstructed voice or quote **must be disclosed as a484recreation** — never presented as authentic archival tape. Mark it in the narration ("what follows485is an AI re-creation of the words in the letter, not a recording") and in the transcript. Cloning a486specific real person's voice would additionally require rights/consent from their estate; a neutral487"reader" voice avoids that and is the safer choice unless rights are secured. Everything else488follows the narrative act structure (§5.1), full scripting (§3), and the standard rights, loudness489(-16 LUFS stereo), transcript, and QA passes.490491---492493## Sources494495Volatile facts (platform specs, policies, laws) verified **2026-07-10**. Standards and craft496heuristics are labeled inline. Do not treat legal notes as legal advice; confirm current law and497policy at publish time.498499**Standards & loudness**500- Apple Podcasts — Audio requirements (formats, bitrates, -16 dB LKFS ±1, -1 dB FS true peak):501 https://podcasters.apple.com/support/893-audio-requirements502- AES TD1008 (2021), *Recommendations for Loudness of Internet Audio Streaming and On-Demand503 Distribution* (supersedes TD1004; -18 LUFS speech, keep above -20 LUFS, speech ~2–3 LU louder504 than music): https://aes2.org/wp-content/uploads/2024/01/20210924_TD1008_v3.13.pdf and505 https://aes.org/technical-council/technical-document-aestd1008/506- Production Advice, summary of AES TD1008: https://productionadvice.co.uk/td1008/507- Spotify — Loudness normalization (-14 LUFS, -1 dBTP):508 https://support.spotify.com/us/artists/article/loudness-normalization/509510**RSS / metadata / transcripts / chapters**511- Apple Podcasts — Podcast RSS feed requirements:512 https://podcasters.apple.com/support/823-podcast-requirements513- Podcasting 2.0 / Podcast Namespace spec (transcript, chapters, etc.):514 https://podcasting2.org/docs/podcast-namespace/1.0 and515 https://github.com/Podcast-Standards-Project/PSP-1-Podcast-RSS-Specification516517**Platform AI-content policies**518- YouTube — Disclosing altered or synthetic content (blog):519 https://blog.youtube/news-and-events/disclosing-ai-generated-content/ ; "How this content was520 made" help: https://support.google.com/youtube/answer/15447836521- Spotify — verification/AI-persona and spam enforcement (secondary, practitioner reporting):522 https://www.musicbusinessworldwide.com/spotify-extends-verified-by-spotify-badges-to-podcasts-further-cracking-down-on-ai-impersonators/523524**Voice rights & disclosure law (not legal advice)**525- Tennessee ELVIS Act overview (Art and Media Law): https://artandmedialaw.com/elvis-act/526- Right-of-publicity / voice-cloning risk map (Holon Law):527 https://holonlaw.com/entertainment-law/synthetic-media-voice-cloning-and-the-new-right-of-publicity-risk-map-for-2026/528- Deepfake & AI voice-cloning laws by state (Recording Law):529 https://www.recordinglaw.com/us-laws/deepfake-laws/530531**TTS direction (best-practice docs, provider-neutral synthesis)**532- ElevenLabs — Text-to-speech best practices:533 https://elevenlabs.io/docs/overview/capabilities/text-to-speech/best-practices534- Microsoft Azure — SSML voice/prosody/phoneme reference:535 https://learn.microsoft.com/en-us/azure/ai-services/speech-service/speech-synthesis-markup-voice536- Google Cloud Text-to-Speech — SSML reference:537 https://docs.cloud.google.com/text-to-speech/docs/ssml538539**Music/SFX rights (practitioner references)**540- The Podcast Haven — How to license music for a podcast:541 https://thepodcasthaven.com/how-to-license-music-for-a-podcast/542- Art and Media Law — Podcast music licensing: https://artandmedialaw.com/podcast-music-licensing/543544**Format, structure & scriptwriting (practitioner references)**545- Buzzsprout — How to write a podcast script: https://www.buzzsprout.com/blog/write-podcast-script-examples546- Riverside — Podcast structure: https://riverside.com/blog/podcast-structure547- Podnews — LUFS/LKFS for podcasters: https://podnews.net/article/lufs-lkfs-for-podcasters