Voice & Sound Effects Generator
Generate voiceovers (TTS) and sound effects. For TTS, choose a concrete
provider and voice before calling submit_voice.
When to Use
- Generate voiceover/narration from text
- Create text-to-speech audio for videos
- Add, replace, or redo narration/voiceover for an existing video, timeline,
screen recording, slide animation, product demo, B-roll edit, MG explainer, or
other visual sequence
- Keep existing narration/voiceover aligned after trimming, speeding up, slowing
down, moving, reordering, or replacing the visuals it describes
- Offer and audition TTS voice choices when the user has not picked a concrete voice
- Generate custom sound effects from text descriptions only after checking the Sound Effects library first
TTS (Text-to-Speech)
If the current request has an existing visual target and the user wants
narration, voiceover, dubbing, or replacement speech for that target, read
references/video-sync.md before drafting new
narration, using existing narration text to generate TTS, or placing audio. Do
this even when the user did not explicitly say "sync" or "match the visuals";
the existence of a visual target means narration timing and meaning may need to
follow on-screen content. Use the normal standalone TTS path only when there is
no visual target or the user just wants an audio asset from text.
Also read references/video-sync.md when the timeline
already has narration/voiceover and the user asks to change the visuals while
keeping that voiceover aligned. This is a sync maintenance task even if no new
TTS is needed.
Use submit_voice to create a TTS audio asset. The current MCP tool contract is:
provider is required. Configured choices may be doubao, elevenlabs,
minimax, inworld, fishaudio, speechify, openai, gemini,
mistral, or cartesia. All providers are opt-in; use only providers shown
as configured in the capabilities prompt.
voiceId is required, concrete, and provider-specific. The only exception is
deliberate MiniMax timbreWeights mixing, where voiceId must be empty. Do
not mix catalogs.
- The curated catalog in references/voices.md covers
only Doubao, ElevenLabs, and MiniMax. Other providers have no bundled preset
or sample catalog in OpenChatCut. Require a concrete voice ID from the user or
their provider account; never invent a preset or
/voice-samples/... URL.
- AI SDK-backed fields are provider-specific: OpenAI supports
modelId,
speed, outputFormat, and instructions; Gemini supports modelId,
outputFormat, and instructions; Mistral supports modelId and
outputFormat; Cartesia supports modelId, speed, languageCode, and
outputFormat. Omit unsupported or unrequested fields.
- Inworld, Fish Audio, and Speechify accept only
voiceId plus optional
modelId. Do not pass expressive, speed, language, or output controls to
these providers.
submit_voice creates an audio asset only. Timeline placement, replacement,
trimming, and alignment happen later with timeline tools.
- For long narration, multiple
submit_voice calls can be useful: split at
natural pauses, sentence groups, or script beat boundaries when the workflow
benefits from separately timed or placed voice clips.
- Doubao supports
speedRatio, loudnessRatio, pitch, emotion,
emotionScale, performancePrompt, and explicitDialect, but not every
voice supports every expressive control. Check
references/voices.md before using them.
- ElevenLabs retains its official voice settings, language, seed, output,
normalization, pronunciation-dictionary, continuity, logging, and latency
controls. MiniMax retains its dedicated controls documented in
references/minimax-tts.md.
Doubao control support for current curated voices:
vivi, xiaohe, yunzhou, xiaotian, naiqimengwa, yingtaowanzi,
wenroumama, zhixingnv, dayi, jitangnv, liuchang, ruyayichen,
morgan, qingcang, huiben, popo, yuanboxiaoshu, baqiqingshu, and
tangseng support explicit emotion / emotionScale,
performancePrompt, and ASMR-style prompt directions.
shuanglangshaonian supports performancePrompt and COT/QA-style
instruction following, but does not support explicit emotion /
emotionScale or ASMR-style control.
explicitDialect is only supported by vivi and can be dongbei,
shaanxi, or sichuan.
ElevenLabs control support for current curated voices:
amelia, brittney, hope, jessica, arabella, jane, maria,
mark, frederick, peter, james, jon, sully, david, and alex
all support the same request-level controls; model-specific support is still
validated by ElevenLabs.
- These controls are not per-voice guarantees of a specific acting style.
Use the preset tags/samples to pick a naturally suitable voice, then use the
controls for moderate delivery changes.
- For ElevenLabs
eleven_v3, inline audio tags are available when the user
asks for expressive delivery such as emotion, tone, nonverbal cues, accent
hints, or local pacing. Official examples fit these useful TTS categories:
emotion/tone tags such as [happy], [sad], [angry], [excited],
[curious], [sarcastic], [crying], [annoyed], [appalled],
[thoughtful], [surprised], and [mischievously]; vocal delivery and
nonverbal cue tags such as [whispers], [laughs], [sighs], [exhales],
[inhales deeply], [clears throat], [snorts], [swallows],
[wheezing], and [coughs];
pacing/pause/local speed tags such as [slowly], [pause],
[short pause], [long pause], [rushed], and [drawn out]; and
accent/special-performance tags such as
[strong X accent], for example [strong French accent], plus [sings],
[singing], [woo], and [pirate voice]. Official examples are
non-exhaustive; similar auditory tags can be tried when the user explicitly
asks for that delivery and the tag describes how the voice should sound, not
a visual action. Write tags directly in text, close to the short phrase
they should affect. Treat tags as local guidance, not paragraph-wide controls.
- For pauses and pacing in
eleven_v3, use punctuation, text structure,
shorter generated segments, or local audio tags such as [short pause] and
[slowly] when needed.
// English / multilingual via ElevenLabs
submit_voice({
provider: "elevenlabs",
text: "Hello world",
voiceId: "peter",
});
// Chinese via Doubao
submit_voice({
provider: "doubao",
text: "你好世界",
voiceId: "liuchang",
});
// With speed adjustment (Doubao only)
submit_voice({
provider: "doubao",
text: "这是一段稍快的中文旁白。",
voiceId: "liuchang",
speedRatio: 1.5,
});
// With expressive Doubao controls
submit_voice({
provider: "doubao",
text: "这次事故提醒我们,安全永远不能侥幸。",
voiceId: "liuchang",
emotion: "sad",
emotionScale: 3,
performancePrompt: "痛心但克制,语速稍慢,像新闻专题旁白",
pitch: -1,
speedRatio: 0.92,
});
// With ElevenLabs delivery controls
submit_voice({
provider: "elevenlabs",
text: "The launch changed how teams plan their daily work.",
voiceId: "peter",
speed: 0.95,
stability: 0.4,
similarityBoost: 0.8,
outputFormat: "wav_44100",
});
// MiniMax TTS (when configured) — see references/minimax-tts.md
submit_voice({
provider: "minimax",
text: "欢迎使用视频编辑助手。",
voiceId: "female-yujie",
speed: 1,
name: "VO · welcome",
});
// Cartesia shape after the user confirms the exact account voice ID.
// confirmedCartesiaVoiceId represents that supplied value, not a preset.
submit_voice({
provider: "cartesia",
text: "A concise product introduction.",
voiceId: confirmedCartesiaVoiceId,
modelId: "sonic-3",
speed: 1,
languageCode: "en",
outputFormat: "mp3",
});
Voice Audition Before Generation
When the user needs TTS and has not already chosen a concrete voice, first
separate providers with curated OpenChatCut choices from providers that require
an account-specific voice ID.
For Doubao, ElevenLabs, or MiniMax, read
references/voices.md before recommending, rendering, or
submitting an option. Use it as the only source for curated preset IDs,
provider choice, display labels, tags, and bundled sample URLs. Do not create
voice options from memory, translated names, or broad user descriptions.
For Inworld, Fish Audio, Speechify, OpenAI, Gemini, Mistral, or Cartesia, do not
offer an invented audition list or sample URL. Ask the user for the concrete
voice ID from that configured provider. A broad description such as "warm
female" is not a valid voiceId.
First determine two separate languages:
- User conversation language: the language the user used to talk to you. Use
this for the surrounding reply,
form-visual label, visual-option name,
and summary.
- Target narration language: the language of the text being synthesized. Use
this only to choose provider and voice catalog.
The audition widget's submit button is fixed to the default label in this build
(submitLabel is accepted but not rendered); keep the question label and option
labels in the user conversation language, not the target narration language. For example:
English users see submit_label="Submit", Chinese users see
submit_label="提交", and Spanish users see submit_label="Enviar".
"help me generate ... voice over in Chinese" is an English conversation asking
for Chinese narration, so the audition widget copy stays in English while the
voice candidates come from Doubao.
For a curated provider:
- Filter
references/voices.md by target narration language / provider and
explicit requirements such as gender, age range, tone, and use case.
- If no preset matches all explicit requirements, say there is no exact match
and offer the closest supported presets with a clear caveat.
- Pick 2-4 matching curated presets.
- Load
widget-forms, then call ask_followup_questions with voice options
and real bundled audio samples.
- Wait for the user to choose.
- Call
submit_voice with the selected preset ID as voiceId.
For a provider without a curated OpenChatCut catalog, ask for a free-text,
concrete provider voice ID instead. Do not add media or synthesize a
/voice-samples/... path. Wait for the user to supply/confirm the exact ID
before calling submit_voice.
For each curated audition option, keep value, display label, media, and
summary tied to the same preset row from references/voices.md. Use only the
sample URLs recorded there. Keep value as the preset ID and media as its
matching sample URL. Write name and summary in the user's conversation
language. The target narration language only decides the provider/voice
catalog. After submission, map the display name back to the preset ID from the
same candidate list.
English request for Chinese narration:
<widget submit_label="Submit">
<form-visual
id="voiceId"
label="For Chinese voiceover, I recommend a few voices to try:"
required="true"
>
<visual-option
value="vivi"
name="Vivi"
media="/voice-samples/doubao-vivi.mp3"
aspect-ratio="16:5"
summary="Female / young / friendly, general"
/>
<visual-option
value="xiaohe"
name="Xiaohe"
media="/voice-samples/doubao-xiaohe.mp3"
aspect-ratio="16:5"
summary="Female / young / soft, clear"
/>
<visual-option
value="yunzhou"
name="Yunzhou"
media="/voice-samples/doubao-yunzhou.mp3"
aspect-ratio="16:5"
summary="Male / young / neutral, business"
/>
</form-visual>
</widget>
Chinese request for Chinese narration:
<widget submit_label="提交">
<form-visual
id="voiceId"
label="我推荐这几个中文旁白音色,先试听一下:"
required="true"
>
<visual-option
value="morgan"
name="Morgan"
media="/voice-samples/doubao-morgan.mp3"
aspect-ratio="16:5"
summary="男 / 中年 / 低沉知识解说"
/>
<visual-option
value="zhixingnv"
name="知性女声"
media="/voice-samples/doubao-zhixingnv.mp3"
aspect-ratio="16:5"
summary="女 / 中年 / 冷静知识讲解"
/>
<visual-option
value="vivi"
name="Vivi"
media="/voice-samples/doubao-vivi.mp3"
aspect-ratio="16:5"
summary="女 / 年轻 / 亲切通用口播"
/>
</form-visual>
</widget>
Sound Effects
For ordinary editing sound effects (SFX), do not generate first. Use the
built-in Sound Effects library before generating:
- Call
browse_library with category:"sound-effects" and a query such as
"whoosh", "camera shutter", "notification", "censor beep", or
"record scratch".
- Inspect the returned
library:sound:<id>.
- Place it with
edit_item, using fromFrame as the sound's
anchor/editorial moment frame:
browse_library({
category: "sound-effects",
query: "short whoosh transition",
});
edit_item({
adds: [
{
type: "audio",
assetId: "library:sound:whoosh-short",
fromFrame: 120,
trackId: "A1",
},
],
});
Only generate sound effects from text descriptions with submit_sound when:
- The user explicitly asks for a generated/original/custom sound.
- The requested sound is too specific for the existing Sound Effects library.
browse_library({ category:"sound-effects", query }) returns no suitable
match.
// Custom/generated sound effect after the library has no suitable match
submit_sound({ prompt: "A dog barking in the distance" });
// With custom duration (0.5-22 seconds)
submit_sound({
prompt: "Thunder and heavy rain",
durationSeconds: 15,
});
// High prompt adherence
submit_sound({
prompt: "Sci-fi laser gun firing",
promptInfluence: 0.8,
});
Tips for better results:
- Be specific: "A dog barking loudly" vs just "dog"
- Include context: "Footsteps on wooden floor in an empty room"
- Specify style: "Cinematic whoosh" or "8-bit game sound"
Parameters
TTS
| Field |
Description |
Notes |
provider |
doubao, elevenlabs, minimax, inworld, fishaudio, speechify, openai, gemini, mistral, or cartesia |
Required; configured choices only |
text |
Text to synthesize |
Required |
voiceId |
Concrete provider-specific voice ID |
Required except MiniMax timbre mix |
modelId |
Provider model override |
ElevenLabs, Inworld, Fish Audio, Speechify, OpenAI, Gemini, Mistral, Cartesia |
speed |
Speech speed |
ElevenLabs, MiniMax, OpenAI, Cartesia |
languageCode |
Language hint/code |
ElevenLabs, Cartesia |
outputFormat |
Provider-supported output format |
ElevenLabs, OpenAI, Gemini, Mistral, Cartesia |
instructions |
Natural-language delivery direction |
OpenAI, Gemini |
speedRatio |
Speech speed |
Doubao only |
name |
Media-pool asset name |
Optional |
Sound Effects
| Field |
Description |
Notes |
prompt |
Sound description |
Required |
durationSeconds |
Duration |
0.5-22 seconds |
promptInfluence |
Prompt adherence |
0-1 |
name |
Asset name |
Optional |
Voices
Use the submit_voice voiceId guide and
references/voices.md for the current curated preset
list, display labels, tags, and sample URLs.
Voice IDs are provider-specific — do NOT mix them
The curated catalog contains separate Doubao, ElevenLabs, and MiniMax IDs.
vivi / dayi are only Doubao; mark / amelia / james are only
ElevenLabs; female-yujie is only MiniMax. Inworld, Fish Audio, Speechify,
OpenAI, Gemini, Mistral, and Cartesia require a concrete provider-specific ID
confirmed by the user and have no bundled OpenChatCut samples.
Provider choice:
- Honor an explicit configured provider first.
- For Chinese narration, prefer a matching curated Doubao voice (or MiniMax
when configured/requested). Use another provider only after the user chooses
it and confirms its voice ID.
- For English / multilingual narration, prefer a curated ElevenLabs voice.
Use another provider only after the user chooses it and confirms its voice ID.
- Never offer a provider that is not shown as configured in capabilities.
Hard rules — what you must NOT do
- Never use a voice ID from a different provider.
- Never submit TTS while the voice is only described broadly; require a
concrete provider-specific ID confirmed by the user.
- Never recommend or render a curated TTS option before checking
references/voices.md.
- Never invent presets or sample URLs for Inworld, Fish Audio, Speechify,
OpenAI, Gemini, Mistral, or Cartesia.
- Never pass provider-specific fields to a provider that does not support them.
- Never claim stable age, regional accent, pronunciation dictionary, or exact
duration controls unless the selected provider exposes them.
- Never replace original recorded speech with TTS unless the user asks.
1---2name: voice-23description: Text-to-Speech (TTS), voiceover, narration placement/sync, and custom sound effects (SFX) generator. Use when the user wants generated speech from text, wants to add/replace/align narration or voiceover for an existing video/timeline, wants to keep existing voiceover synced after visual retiming edits, needs voice audition/selection, or explicitly wants a newly generated/custom sound effect that is not available in the Sound Effects library.4---56# Voice & Sound Effects Generator78Generate voiceovers (TTS) and sound effects. For TTS, choose a concrete9provider and voice before calling `submit_voice`.1011## When to Use1213- Generate voiceover/narration from text14- Create text-to-speech audio for videos15- Add, replace, or redo narration/voiceover for an existing video, timeline,16 screen recording, slide animation, product demo, B-roll edit, MG explainer, or17 other visual sequence18- Keep existing narration/voiceover aligned after trimming, speeding up, slowing19 down, moving, reordering, or replacing the visuals it describes20- Offer and audition TTS voice choices when the user has not picked a concrete voice21- Generate custom sound effects from text descriptions only after checking the Sound Effects library first2223## TTS (Text-to-Speech)2425If the current request has an existing visual target and the user wants26narration, voiceover, dubbing, or replacement speech for that target, read27[references/video-sync.md](references/video-sync.md) before drafting new28narration, using existing narration text to generate TTS, or placing audio. Do29this even when the user did not explicitly say "sync" or "match the visuals";30the existence of a visual target means narration timing and meaning may need to31follow on-screen content. Use the normal standalone TTS path only when there is32no visual target or the user just wants an audio asset from text.3334Also read [references/video-sync.md](references/video-sync.md) when the timeline35already has narration/voiceover and the user asks to change the visuals while36keeping that voiceover aligned. This is a sync maintenance task even if no new37TTS is needed.3839Use `submit_voice` to create a TTS audio asset. The current MCP tool contract is:4041- `provider` is required. Configured choices may be `doubao`, `elevenlabs`,42 `minimax`, `inworld`, `fishaudio`, `speechify`, `openai`, `gemini`,43 `mistral`, or `cartesia`. All providers are opt-in; use only providers shown44 as configured in the capabilities prompt.45- `voiceId` is required, concrete, and provider-specific. The only exception is46 deliberate MiniMax `timbreWeights` mixing, where `voiceId` must be empty. Do47 not mix catalogs.48- The curated catalog in [references/voices.md](references/voices.md) covers49 only Doubao, ElevenLabs, and MiniMax. Other providers have no bundled preset50 or sample catalog in OpenChatCut. Require a concrete voice ID from the user or51 their provider account; never invent a preset or `/voice-samples/...` URL.52- AI SDK-backed fields are provider-specific: OpenAI supports `modelId`,53 `speed`, `outputFormat`, and `instructions`; Gemini supports `modelId`,54 `outputFormat`, and `instructions`; Mistral supports `modelId` and55 `outputFormat`; Cartesia supports `modelId`, `speed`, `languageCode`, and56 `outputFormat`. Omit unsupported or unrequested fields.57- Inworld, Fish Audio, and Speechify accept only `voiceId` plus optional58 `modelId`. Do not pass expressive, speed, language, or output controls to59 these providers.60- `submit_voice` creates an audio asset only. Timeline placement, replacement,61 trimming, and alignment happen later with timeline tools.62- For long narration, multiple `submit_voice` calls can be useful: split at63 natural pauses, sentence groups, or script beat boundaries when the workflow64 benefits from separately timed or placed voice clips.65- Doubao supports `speedRatio`, `loudnessRatio`, `pitch`, `emotion`,66 `emotionScale`, `performancePrompt`, and `explicitDialect`, but not every67 voice supports every expressive control. Check68 [references/voices.md](references/voices.md) before using them.69- ElevenLabs retains its official voice settings, language, seed, output,70 normalization, pronunciation-dictionary, continuity, logging, and latency71 controls. MiniMax retains its dedicated controls documented in72 [references/minimax-tts.md](references/minimax-tts.md).7374Doubao control support for current curated voices:7576- `vivi`, `xiaohe`, `yunzhou`, `xiaotian`, `naiqimengwa`, `yingtaowanzi`,77 `wenroumama`, `zhixingnv`, `dayi`, `jitangnv`, `liuchang`, `ruyayichen`,78 `morgan`, `qingcang`, `huiben`, `popo`, `yuanboxiaoshu`, `baqiqingshu`, and79 `tangseng` support explicit `emotion` / `emotionScale`,80 `performancePrompt`, and ASMR-style prompt directions.81- `shuanglangshaonian` supports `performancePrompt` and COT/QA-style82 instruction following, but does not support explicit `emotion` /83 `emotionScale` or ASMR-style control.84- `explicitDialect` is only supported by `vivi` and can be `dongbei`,85 `shaanxi`, or `sichuan`.8687ElevenLabs control support for current curated voices:8889- `amelia`, `brittney`, `hope`, `jessica`, `arabella`, `jane`, `maria`,90 `mark`, `frederick`, `peter`, `james`, `jon`, `sully`, `david`, and `alex`91 all support the same request-level controls; model-specific support is still92 validated by ElevenLabs.93- These controls are not per-voice guarantees of a specific acting style.94 Use the preset tags/samples to pick a naturally suitable voice, then use the95 controls for moderate delivery changes.96- For ElevenLabs `eleven_v3`, inline audio tags are available when the user97 asks for expressive delivery such as emotion, tone, nonverbal cues, accent98 hints, or local pacing. Official examples fit these useful TTS categories:99 emotion/tone tags such as `[happy]`, `[sad]`, `[angry]`, `[excited]`,100 `[curious]`, `[sarcastic]`, `[crying]`, `[annoyed]`, `[appalled]`,101 `[thoughtful]`, `[surprised]`, and `[mischievously]`; vocal delivery and102 nonverbal cue tags such as `[whispers]`, `[laughs]`, `[sighs]`, `[exhales]`,103 `[inhales deeply]`, `[clears throat]`, `[snorts]`, `[swallows]`,104 `[wheezing]`, and `[coughs]`;105 pacing/pause/local speed tags such as `[slowly]`, `[pause]`,106 `[short pause]`, `[long pause]`, `[rushed]`, and `[drawn out]`; and107 accent/special-performance tags such as108 `[strong X accent]`, for example `[strong French accent]`, plus `[sings]`,109 `[singing]`, `[woo]`, and `[pirate voice]`. Official examples are110 non-exhaustive; similar auditory tags can be tried when the user explicitly111 asks for that delivery and the tag describes how the voice should sound, not112 a visual action. Write tags directly in `text`, close to the short phrase113 they should affect. Treat tags as local guidance, not paragraph-wide controls.114- For pauses and pacing in `eleven_v3`, use punctuation, text structure,115 shorter generated segments, or local audio tags such as `[short pause]` and116 `[slowly]` when needed.117118```ts119// English / multilingual via ElevenLabs120submit_voice({121 provider: "elevenlabs",122 text: "Hello world",123 voiceId: "peter",124});125126// Chinese via Doubao127submit_voice({128 provider: "doubao",129 text: "你好世界",130 voiceId: "liuchang",131});132133// With speed adjustment (Doubao only)134submit_voice({135 provider: "doubao",136 text: "这是一段稍快的中文旁白。",137 voiceId: "liuchang",138 speedRatio: 1.5,139});140141// With expressive Doubao controls142submit_voice({143 provider: "doubao",144 text: "这次事故提醒我们,安全永远不能侥幸。",145 voiceId: "liuchang",146 emotion: "sad",147 emotionScale: 3,148 performancePrompt: "痛心但克制,语速稍慢,像新闻专题旁白",149 pitch: -1,150 speedRatio: 0.92,151});152153// With ElevenLabs delivery controls154submit_voice({155 provider: "elevenlabs",156 text: "The launch changed how teams plan their daily work.",157 voiceId: "peter",158 speed: 0.95,159 stability: 0.4,160 similarityBoost: 0.8,161 outputFormat: "wav_44100",162});163164// MiniMax TTS (when configured) — see references/minimax-tts.md165submit_voice({166 provider: "minimax",167 text: "欢迎使用视频编辑助手。",168 voiceId: "female-yujie",169 speed: 1,170 name: "VO · welcome",171});172173// Cartesia shape after the user confirms the exact account voice ID.174// confirmedCartesiaVoiceId represents that supplied value, not a preset.175submit_voice({176 provider: "cartesia",177 text: "A concise product introduction.",178 voiceId: confirmedCartesiaVoiceId,179 modelId: "sonic-3",180 speed: 1,181 languageCode: "en",182 outputFormat: "mp3",183});184```185186## Voice Audition Before Generation187188When the user needs TTS and has not already chosen a concrete voice, first189separate providers with curated OpenChatCut choices from providers that require190an account-specific voice ID.191192For Doubao, ElevenLabs, or MiniMax, read193[references/voices.md](references/voices.md) before recommending, rendering, or194submitting an option. Use it as the only source for curated preset IDs,195provider choice, display labels, tags, and bundled sample URLs. Do not create196voice options from memory, translated names, or broad user descriptions.197198For Inworld, Fish Audio, Speechify, OpenAI, Gemini, Mistral, or Cartesia, do not199offer an invented audition list or sample URL. Ask the user for the concrete200voice ID from that configured provider. A broad description such as "warm201female" is not a valid `voiceId`.202203First determine two separate languages:204205- User conversation language: the language the user used to talk to you. Use206 this for the surrounding reply, `form-visual label`, `visual-option name`,207 and `summary`.208- Target narration language: the language of the text being synthesized. Use209 this only to choose provider and voice catalog.210211The audition widget's submit button is fixed to the default label in this build212(`submitLabel` is accepted but not rendered); keep the question label and option213labels in the user conversation language, not the target narration language. For example:214English users see `submit_label="Submit"`, Chinese users see215`submit_label="提交"`, and Spanish users see `submit_label="Enviar"`.216217"help me generate ... voice over in Chinese" is an English conversation asking218for Chinese narration, so the audition widget copy stays in English while the219voice candidates come from Doubao.220221For a curated provider:2222231. Filter `references/voices.md` by target narration language / provider and224 explicit requirements such as gender, age range, tone, and use case.2252. If no preset matches all explicit requirements, say there is no exact match226 and offer the closest supported presets with a clear caveat.2273. Pick 2-4 matching curated presets.2284. Load `widget-forms`, then call `ask_followup_questions` with voice options229 and real bundled audio samples.2305. Wait for the user to choose.2316. Call `submit_voice` with the selected preset ID as `voiceId`.232233For a provider without a curated OpenChatCut catalog, ask for a free-text,234concrete provider voice ID instead. Do not add `media` or synthesize a235`/voice-samples/...` path. Wait for the user to supply/confirm the exact ID236before calling `submit_voice`.237238For each curated audition option, keep `value`, display label, `media`, and239`summary` tied to the same preset row from `references/voices.md`. Use only the240sample URLs recorded there. Keep `value` as the preset ID and `media` as its241matching sample URL. Write `name` and `summary` in the user's conversation242language. The target narration language only decides the provider/voice243catalog. After submission, map the display name back to the preset ID from the244same candidate list.245246English request for Chinese narration:247248```html249<widget submit_label="Submit">250 <form-visual251 id="voiceId"252 label="For Chinese voiceover, I recommend a few voices to try:"253 required="true"254 >255 <visual-option256 value="vivi"257 name="Vivi"258 media="/voice-samples/doubao-vivi.mp3"259 aspect-ratio="16:5"260 summary="Female / young / friendly, general"261 />262 <visual-option263 value="xiaohe"264 name="Xiaohe"265 media="/voice-samples/doubao-xiaohe.mp3"266 aspect-ratio="16:5"267 summary="Female / young / soft, clear"268 />269 <visual-option270 value="yunzhou"271 name="Yunzhou"272 media="/voice-samples/doubao-yunzhou.mp3"273 aspect-ratio="16:5"274 summary="Male / young / neutral, business"275 />276 </form-visual>277</widget>278```279280Chinese request for Chinese narration:281282```html283<widget submit_label="提交">284 <form-visual285 id="voiceId"286 label="我推荐这几个中文旁白音色,先试听一下:"287 required="true"288 >289 <visual-option290 value="morgan"291 name="Morgan"292 media="/voice-samples/doubao-morgan.mp3"293 aspect-ratio="16:5"294 summary="男 / 中年 / 低沉知识解说"295 />296 <visual-option297 value="zhixingnv"298 name="知性女声"299 media="/voice-samples/doubao-zhixingnv.mp3"300 aspect-ratio="16:5"301 summary="女 / 中年 / 冷静知识讲解"302 />303 <visual-option304 value="vivi"305 name="Vivi"306 media="/voice-samples/doubao-vivi.mp3"307 aspect-ratio="16:5"308 summary="女 / 年轻 / 亲切通用口播"309 />310 </form-visual>311</widget>312```313314## Sound Effects315316For ordinary editing sound effects (SFX), do **not** generate first. Use the317built-in Sound Effects library before generating:3183191. Call `browse_library` with `category:"sound-effects"` and a query such as320 `"whoosh"`, `"camera shutter"`, `"notification"`, `"censor beep"`, or321 `"record scratch"`.3222. Inspect the returned `library:sound:<id>`.3233. Place it with `edit_item`, using `fromFrame` as the sound's324 anchor/editorial moment frame:325326```ts327browse_library({328 category: "sound-effects",329 query: "short whoosh transition",330});331332edit_item({333 adds: [334 {335 type: "audio",336 assetId: "library:sound:whoosh-short",337 fromFrame: 120,338 trackId: "A1",339 },340 ],341});342```343344Only generate sound effects from text descriptions with `submit_sound` when:345346- The user explicitly asks for a generated/original/custom sound.347- The requested sound is too specific for the existing Sound Effects library.348- `browse_library({ category:"sound-effects", query })` returns no suitable349 match.350351```ts352// Custom/generated sound effect after the library has no suitable match353submit_sound({ prompt: "A dog barking in the distance" });354355// With custom duration (0.5-22 seconds)356submit_sound({357 prompt: "Thunder and heavy rain",358 durationSeconds: 15,359});360361// High prompt adherence362submit_sound({363 prompt: "Sci-fi laser gun firing",364 promptInfluence: 0.8,365});366```367368**Tips for better results:**369370- Be specific: "A dog barking loudly" vs just "dog"371- Include context: "Footsteps on wooden floor in an empty room"372- Specify style: "Cinematic whoosh" or "8-bit game sound"373374## Parameters375376### TTS377378| Field | Description | Notes |379| --- | --- | --- |380| `provider` | `doubao`, `elevenlabs`, `minimax`, `inworld`, `fishaudio`, `speechify`, `openai`, `gemini`, `mistral`, or `cartesia` | Required; configured choices only |381| `text` | Text to synthesize | Required |382| `voiceId` | Concrete provider-specific voice ID | Required except MiniMax timbre mix |383| `modelId` | Provider model override | ElevenLabs, Inworld, Fish Audio, Speechify, OpenAI, Gemini, Mistral, Cartesia |384| `speed` | Speech speed | ElevenLabs, MiniMax, OpenAI, Cartesia |385| `languageCode` | Language hint/code | ElevenLabs, Cartesia |386| `outputFormat` | Provider-supported output format | ElevenLabs, OpenAI, Gemini, Mistral, Cartesia |387| `instructions` | Natural-language delivery direction | OpenAI, Gemini |388| `speedRatio` | Speech speed | Doubao only |389| `name` | Media-pool asset name | Optional |390391### Sound Effects392393| Field | Description | Notes |394| ----------------- | ----------------- | -------------- |395| `prompt` | Sound description | Required |396| `durationSeconds` | Duration | 0.5-22 seconds |397| `promptInfluence` | Prompt adherence | 0-1 |398| `name` | Asset name | Optional |399400## Voices401402Use the `submit_voice` `voiceId` guide and403[references/voices.md](references/voices.md) for the current curated preset404list, display labels, tags, and sample URLs.405406### Voice IDs are provider-specific — do NOT mix them407408The curated catalog contains separate Doubao, ElevenLabs, and MiniMax IDs.409`vivi` / `dayi` are only Doubao; `mark` / `amelia` / `james` are only410ElevenLabs; `female-yujie` is only MiniMax. Inworld, Fish Audio, Speechify,411OpenAI, Gemini, Mistral, and Cartesia require a concrete provider-specific ID412confirmed by the user and have no bundled OpenChatCut samples.413414Provider choice:415416- Honor an explicit configured provider first.417- For Chinese narration, prefer a matching curated Doubao voice (or MiniMax418 when configured/requested). Use another provider only after the user chooses419 it and confirms its voice ID.420- For English / multilingual narration, prefer a curated ElevenLabs voice.421 Use another provider only after the user chooses it and confirms its voice ID.422- Never offer a provider that is not shown as configured in capabilities.423424## Hard rules — what you must NOT do4254261. Never use a voice ID from a different provider.4272. Never submit TTS while the voice is only described broadly; require a428 concrete provider-specific ID confirmed by the user.4293. Never recommend or render a curated TTS option before checking430 [references/voices.md](references/voices.md).4314. Never invent presets or sample URLs for Inworld, Fish Audio, Speechify,432 OpenAI, Gemini, Mistral, or Cartesia.4335. Never pass provider-specific fields to a provider that does not support them.4346. Never claim stable age, regional accent, pronunciation dictionary, or exact435 duration controls unless the selected provider exposes them.4367. Never replace original recorded speech with TTS unless the user asks.