Use get_run({ run_id, wait_seconds? }) to read a saved run before any retry; no idempotency key is accepted. Waiting defaults to 50 seconds and accepts 0–120. run_tool starts only and requires tool_id, input and one UUID per paid intent; retain that UUID for an exact retry after an uncertain start. Every state retains run_id.
Audio Edition
Turn written content into a concise episode written for the ear. The agent adapts rather than reads verbatim, the human approves the script before any audio spend, and scrollport supplies bounded narration and music primitives.
Use the six Scrollport control tools and only catalog tools that search_tools
currently returns as live.
Inputs and safe default
Accept pasted content, a post/newsletter URL or an RSS item. Ask for audience, tone and any words or names whose pronunciation matters.
- Default: one episode of five minutes or less.
- Trial-safe mode: at most 4,000 narration characters, one three-second instrumental sting, and no second render.
- Full-length mode: estimate and quote first; never silently expand the script.
- Feed hosting, distribution and publishing are out of scope.
Resolve the source before drafting the script. If the host environment can
safely retrieve a supplied public URL, it may do so. Otherwise plan the smallest
exa.search retrieval, obtain approval for that research spend, then run it
with the URL/title and page text enabled. Any optional fresh context must also
be gathered before the script checkpoint; label it as external and preserve the
author's claims separately. Never invent access to a paywalled or private
newsletter.
Catalog-tool selection
Discover and inspect web research, speech and music immediately before use. The expected launch tool ids are:
exa.searchfor optional source retrieval or enrichment;elevenlabs.eleven-v3for narration — one voice on Eleven v3;elevenlabs.eleven-music-v2for one bounded instrumental sting.
Narration default: a single narrator on Eleven v3. Use
elevenlabs.eleven-v3 with the same voice_id on every turn. The catalog
tool is named for its multi-speaker shape, but a single speaker is the adopted
route: v3 is the expressive model, and one narrator is what an episode wants.
Quote narration from the current inspected per-unit price and the complete input character count. Never discount the approved maximum using a historical billed-to-input ratio: a smaller settled charge is an observation, not a price guarantee. Record the final billed units and cost after each successful run.
Two voices remain available through the same capability by varying voice_id
per turn, but that is a deliberate departure from the default, not a fallback.
Dialogue requests are capped at 2,000 characters, so longer scripts must be
split at section boundaries and stitched agent-side.
Research, then script before audio spend
Before drafting, complete any required Exa source retrieval and optional context search with the smallest useful result count. Save the run, source text and citations. Research is useful only when the returned text supports the script; a non-empty search envelope is not enough.
Produce a broadcast script locally. For a newsletter, use short segments with spoken transitions; remove visual-only references, expand ambiguous acronyms, and attribute externally researched facts. Keep the script within the approved character budget. Do not paraphrase a source into claims it did not make.
Present this checkpoint:
- final script and character count;
- estimated duration;
- narration format and voice ids;
- exact narration chunks;
- music prompt and duration;
- inspected per-unit prices and maximum total USD;
- output choice: MP3 only or MP3 plus agent-side audiogram.
Wait for explicit approval of the script and plan. Any later script, voice, duration or music change invalidates that approval.
Save state:
{
"skill": "media-audio-edition",
"version": 1,
"status": "awaiting_script_approval",
"source": {"kind": "newsletter", "ref": "..."},
"script_sha256": "hex",
"script_characters": 0,
"narration": {"tool_id": "elevenlabs.eleven-v3", "voice_id": "...", "chunks": []},
"music": {"tool_id": "elevenlabs.eleven-music-v2", "music_length_ms": 3000},
"plan": [],
"completed": [],
"pending": "human script approval",
"spent_usd": "0.000000"
}
Generate
1. Narration
After approval, run elevenlabs.eleven-v3 chunk by chunk with one voice_id
throughout. Poll each run to terminal and download its MP3 artifact before
starting the next chunk. Save run id, final cost, artifact reference and script span. On an
uncertain result, poll the same run; never regenerate simply because a download
was slow.
Listen-check or inspect every artifact for non-empty audio, expected duration and obvious truncation. Pronunciation and performance quality remain a human judgement on any named-entity-heavy episode; the model choice is settled, the delivery on a given script is not.
2. Music
Generate one three-second instrumental sting by default. The prompt must state instrumental, mood and clean ending. Inspect the artifact before use; valid MP3 bytes are not proof that the sting fits the episode.
3. Assemble locally
Concatenate narration chunks and place the sting at the opening and/or closing using local audio tooling. This is agent-side work and creates no scrollport run. Do not loop the music underneath speech unless the human approved that mix.
For an audiogram, combine the final MP3 with a supplied or locally generated cover, waveform and timed captions using local tooling. If those tools are not available, return the MP3 and caption text; do not purchase a video capability as an undeclared fallback.
Resume and completion
On a terminal narration or music failure, preserve successful artifacts and the approved plan before correcting input or requesting a replacement run. If provider execution is uncertain, inspect the existing run before retrying. Assembly failures are local: retry assembly without regenerating narration or music.
On resume, verify the saved script hash still matches the approved script, then poll all non-terminal run ids. Reuse downloaded successful chunks. A completed episode reports:
- source and any added citations;
- approved script hash, character count and duration;
- narration and music tool ids, run ids and final costs;
- final MP3 path/artifact and optional audiogram path;
- total spend and skipped optional steps.
Completion requires a human-audible, non-truncated episode whose content matches the approved script. A provider success status alone is insufficient.