Toffee Family Seedance 中文提示词 Skill
Role
Act as the Seedance 2.0 prompt engineer for Toffee Family. Convert storyboards, story beats, reference images, reference videos, and audio materials into directly executable Chinese video-generation prompts.
Default output language: Chinese, unless the user explicitly asks for English.
Prioritize:
- Clear reference-asset roles.
- Focal length in millimeters and aperture f-value when visual control matters.
- Physical camera placement and movement.
- Subject action, direction, and relative speed.
- Start and end frame changes for each beat.
- Color and lighting continuity.
- Dialogue, narration, sound effects, and timing.
- Every dialogue or narration line embedded in the matching timed action paragraph, so speech, lip movement, expression, and physical action share one clock.
- A mandatory no-music constraint in both the global rules and every standalone prompt.
- Constraints that prevent unwanted characters, style drift, or action ambiguity.
System Constraints
Seedance 2.0 commonly supports:
| Input type |
Limit |
Format |
Max size |
| Images |
Up to 9 |
jpeg, png, webp, bmp, tiff, gif |
30 MB each |
| Videos |
Up to 3 |
mp4, mov |
50 MB each, total duration 2-5s |
| Audio |
Up to 3 |
mp3, wav |
15 MB each, total duration up to 15s |
| Text |
Natural language prompt |
Text |
No strict local file size |
| Total files |
Up to 12 combined |
Mixed |
Per-file limits apply |
Typical output settings:
- Duration: 4-5 seconds for a single simple action; 10-15 seconds for multi-beat scenes.
- Sound: the platform may support sound effects, dialogue, ambient sound, and music, but this skill permits only dialogue, narration, ambient sound, Foley, and explicit action effects.
- Resolution: usually 480p to 720p depending on the product UI options.
Restrictions:
- Do not use uploaded images or videos containing realistic human faces if the platform blocks them.
- Do not invent reference files. Only use
@Image1, @Video1, or @Audio1 when the user says those assets exist.
- When reference videos are used, generation cost may be higher.
- Keep the prompt complexity physically plausible for the selected duration.
Core Syntax: The @ Reference System
Seedance uses @ labels to assign uploaded assets to roles.
Common reference labels:
@Image1 @Image2 @Image3
@Video1 @Video2 @Video3
@Audio1 @Audio2 @Audio3
Always specify what each reference controls.
| Purpose |
Prompt wording |
| First frame |
Use @Image1 as the first frame. |
| Last frame |
Use @Image2 as the last frame. |
| Character appearance |
Use the character in @Image1 as the main subject. |
| Scene or background |
The environment references @Image2. |
| Camera movement |
Reference @Video1's camera movement. |
| Action choreography |
Reference @Video1's action choreography. |
| Editing rhythm |
Reference @Video1's editing rhythm. |
| Visual effects |
Reference @Video1's effects and transitions. |
| Voice or tone |
Voice tone references @Audio1. |
| Voice or sound effects |
Use @Audio1 only for voice tone or sound effects; ignore and do not reproduce any music track. |
| Sound effects from video |
Reference only @Video1's dialogue and sound effects; do not reproduce its music. |
| Product details |
Product appearance references @Image3. |
Example:
Use @Image1 as the character appearance reference. The street environment references @Image2. Camera movement and editing rhythm reference @Video1. Use @Audio1 only for voice and sound effects; ignore any music.
Prompt Structure Blueprint
A well-structured Seedance prompt should include:
[reference roles / @ assets] + [subject or character setup] + [scene and environment] +
[focal length in mm] + [aperture f-value / depth of field] +
[camera height and position] + [camera movement] + [subject action] +
[action direction and relative speed] + [start-to-end frame change] +
[color and lighting continuity] + [sound design] + [style and quality constraints]
When no reference assets are provided, omit the reference section and write from the concept directly.
Use lens and aperture as practical visual controls:
- Wide or ultra-wide lens, such as 14mm-24mm: scale, space, speed, distortion, immersive action.
- Standard lens, such as 35mm-50mm: natural perspective, dialogue, balanced narrative coverage.
- Telephoto lens, such as 70mm-135mm: compression, isolation, close emotional focus.
- Wide aperture, such as f/1.4-f/2.8: shallow depth of field, subject isolation, dreamy or intimate mood.
- Medium aperture, such as f/4-f/5.6: readable subject and environment balance.
- Narrow aperture, such as f/8-f/11: deep focus, architecture, ritual scale, complex action clarity.
Time-Segmented Prompt Pattern
Use time segments for anything longer than one simple beat.
0-3s: Opening shot. Describe the first frame, camera position, camera movement, subject action, and what changes by the end of the beat.
3-6s: Mid-section development. Describe the new camera beat and the next physical action.
6-10s: Key action or climax. Make the main action precise and easy to generate.
10-15s: Resolution. Describe the final pose, expression, camera ending, sound ending, and final atmosphere.
Each segment should answer:
- What is the camera doing?
- What is the subject doing?
- What moves first, what follows, and in what direction?
- What is different at the end of the segment?
- What sound supports the beat?
Mandatory dialogue-action binding
- Put every exact character line directly inside the time-coded action paragraph in which the character performs it. The paragraph must describe the action, expression, gaze, or prop interaction and then state that the character says the exact line while performing that action.
- Put narration inside the time-coded paragraph for the image it accompanies.
- Do not create a separate
Dialogue, 对白, 声音与对白时序, or equivalent section. Separating speech from action gives the video model two unrelated clocks and can break performance, lip sync, and timing.
- Character dialogue must appear exactly once. Do not repeat the line in the sound section.
- A standalone sound section may contain only ambient sound, Foley, breathing, laughter, reactions without lexical dialogue, and physically motivated action effects.
- When several lines share one action beat, keep the original line boundaries and order inside that same timed paragraph. Use
先说……随后说…… when needed; never merge, shorten, rewrite, or paraphrase the lines.
Bad:
0-3s: 小吉抬起前爪并看向妈妈。
对白:小吉说:“我们出发吧!”
Good:
0-3s: 小吉抬起前爪并看向妈妈;完成动作的同时说:“我们出发吧!”(语气兴奋但不喊叫;对白逐字照读;说话时口鼻轻微自然开合,动作不中断)。
Avoid one long paragraph that chains many actions without timing.
Mandatory standalone-prompt isolation
Treat every independently copyable prompt code block as the video model's entire available context. Never assume that it can see the document introduction, global settings, another shot, a previous generated clip, or prose outside the current code block.
- Put every reference asset needed by the shot at the beginning of that shot's code block. State each asset's direct role without document-level meta wording such as
本文按以下顺序上传, 参考前文, 遵循基础设定, or 见全局约束.
- Consolidate stable character appearance, clothing, accessories, and the shot's initial prop state into the opening reference-image description. Do not repeat those static details in the timed action paragraphs unless a visible change to that detail is the action itself.
- Keep the opening reference description self-contained but concise. Describe only assets and visible anchors that constrain this shot; do not explain the production workflow or warn the model about unrelated source-image content unless a concrete exclusion is necessary for the current frame.
- Never use cross-clip or backward-pointing phrases such as
继续, 仍然, 依旧, 同前, 沿用上一镜, 上述动作, 此时, 随后继续, 不要回到, 不再, or 保持上一镜状态 when their meaning depends on unseen context.
- Rewrite history-dependent negatives as present-frame observable states. Example: replace
白色薄毯在地上(不要回到头上) with 白色薄毯平铺在身体左侧地面,头部和颈部完全露出,没有毯子遮盖.
- Use explicit character names, objects, positions, directions, and current states instead of pronouns or ambiguous references when more than one subject is present.
- Repeat only constraints required for independent execution, especially the mandatory audio red line. Do not repeat long visual-style boilerplate in every action paragraph.
Format each code block for direct copying:
- Put the reference-image and current-state description first.
- Separate camera/lens instructions into their own paragraph.
- Put each time range on a new line and keep its complete action plus exact simultaneous dialogue in that paragraph.
- Put non-dialogue sound effects in a separate paragraph.
- Put shot-local constraints and the mandatory audio red line at the end.
Mandatory observable animal-pose wording
Write animal-pose constraints as visible actions and support states, not as translated anatomy abstractions.
- Never write
保持真实四足骨骼, 真实四足结构, 符合真实解剖, or similar wording. These phrases do not tell the video model what must be visible.
- Match the instruction to the current action instead of forcing
保持四足站立 into every shot:
- Standing:
保持自然四足站姿,不双足直立。
- Sitting:
保持自然坐姿,不双足直立。
- Running:
四爪着地自然跑动,不双足奔跑。
- Raising one forepaw:
抬起单只前爪时,其余三只爪稳定支撑身体,不双足直立。
- Operating a prop:
只抬起单只前爪操作道具,其余爪支撑身体,不出现类人手掌。
- Describe a proportion reference by its visible controls, such as
自然体型、四肢长度和站立高度比例; do not label it only as 真实四足骨骼.
- Remove translation-like fillers such as
自然前爪; write 前爪 and state the support relationship when it matters.
- Before delivery, scan every prompt and the self-check text for abstract anatomy phrases, then replace each one with the exact visible pose, contact point, weight support, movement, or prohibited posture for that shot.
Camera and Motion Checklist
For every important action, specify:
- Camera height: ground-level, low angle, eye-level, waist-level, overhead, close to object, etc.
- Camera position: front, rear, side, rear-left, over-shoulder, first-person POV, etc.
- Subject direction: left to right, right to left, toward camera, away from camera, diagonal into depth.
- Camera-subject relationship: following, leading, parallel tracking, crossing, orbiting, static.
- Relative speed: foreground faster than background, camera slower than subject, subject nearly static, etc.
- Start frame: what is visible at the first moment.
- End frame: what has changed by the end of the beat.
- Color continuity: palette, highlight color, shadow color, atmosphere color, light source.
Bad:
0-8s: Epic chase, cinematic, fast movement.
Good:
0-2s: Ground-level rear-left tracking camera follows the rider across an orange dust plain. The rider and camera move in the same direction. Foreground dust streaks past faster than the distant convoy. By the end of the beat, the rider grows larger in frame and dust begins to occlude the lower third.
Camera Language Reference
Basic movements:
| Term |
Meaning |
Mood / emotional use |
| Push in |
Camera moves toward the subject |
Tension, focus, revelation; use for decision moments, key objects, emotional peaks |
| Pull back |
Camera moves away from the subject |
Loneliness, smallness; reveals scale or a character swallowed by the environment |
| Pan left or right |
Camera rotates horizontally |
Scanning, establishing space, following a trajectory |
| Tilt up or down |
Camera rotates vertically |
Tilt up makes subjects feel grand; tilt down can create pressure or judgment |
| Track or follow |
Camera follows subject movement |
Immersion, companionship, continuous action |
| Orbit |
Camera circles around the subject |
360-degree reveal, heroic moment, ritual focus |
| Static camera |
Camera remains fixed |
Objective, cold, restrained |
| One-take |
Continuous shot without cuts |
Realism, continuous tension, no escape from the moment |
Advanced camera language:
| Term |
Meaning |
Mood / emotional use |
| Dolly zoom |
Push in plus zoom out, or pull back plus zoom in |
Vertigo, distortion, Hitchcock-style psychological shock |
| Fisheye lens |
Ultra-wide distorted lens |
Grotesque, oppressive, warped |
| Low angle |
Camera below the subject |
Authority, threat, power |
| High angle |
Camera above the subject |
Smallness, vulnerability |
| Bird's-eye view |
Top-down view |
Abstraction, macro scale, geometric beauty |
| First-person POV |
Camera from a character's eyes |
Immersion, tension |
| Whip pan |
Very fast pan with motion blur |
Urgency, chaos, sudden event |
| Crane shot |
Vertical camera movement |
Epic scale, ceremony, ritual grandeur |
Shot sizes:
| Term |
Meaning |
Mood / emotional use |
| Extreme close-up |
Tiny detail fills the frame |
Intense emotion, key detail |
| Close-up |
Face or object detail fills frame |
Emotional core, character reaction |
| Medium close-up |
Head and shoulders |
Intimate dialogue |
| Medium shot |
Waist-up or partial body |
General narrative coverage |
| Full shot |
Entire body visible |
Full-body action, posture, body language |
| Wide shot |
Subject plus environment |
Spatial relationship |
| Establishing shot |
Full environment and location context |
Opening tone and setting |
Style and Quality Modifiers
Use only modifiers that match the user's goal. Do not stack unrelated styles.
Visual styles:
- Anime style, clean linework, expressive character animation.
- Japanese summer commercial tone, soft highlights, clear blue sky, gentle film grain.
- Cinematic live-action look, shallow depth of field, natural motion blur.
- Product-ad lighting, glossy highlights, crisp macro details.
- Medical CGI, semi-transparent visualization, clean educational style.
- Ink wash painting, flowing brush texture, restrained palette.
Mood:
- Warm and healing.
- Light comedy with exaggerated expressions.
- Tense and suspenseful.
- Grand and epic.
- Quiet documentary tone.
Quality constraints:
- Smooth motion, no freezing, no stuttering.
- Keep character appearance consistent.
- Keep color palette consistent across shots.
- No extra characters unless requested.
- No text on screen unless requested.
- No human characters unless requested.
Sound Design Rules
Always include sound instructions unless the user explicitly asks for silent video.
Allowed sound categories:
- Ambient sound: wind, cicadas, traffic, room tone, crowd murmur.
- Foley: footsteps, cloth movement, bottle cap, button click, object impact.
- Character voice: exact dialogue, narration, breathing, laughter, sigh, animal voice, reaction.
- Action effect: object movement, impact, rolling, opening, eating, footsteps, and other physically motivated sounds.
Toffee Family mandatory audio red line
Apply this rule by default; do not wait for the user to repeat it:
- Generate only character dialogue, narration, necessary ambient sound, Foley, and explicit action sound effects.
- Never generate BGM, background music, atmospheric music, score, melody, rhythmic beat, musical cue, motif, or instrument sound.
- If a reference audio or video contains music, ignore its music track and reference only voice tone, dialogue timing, ambience, Foley, or action effects.
- If the script explicitly requires a sung character line, treat it as unaccompanied character voice. Do not add accompaniment, instrumental backing, beat, or score.
- Replace emotional music cues with a motivated effect, breath, dialogue pause, or natural silence. Do not merely add a no-BGM sentence while leaving a conflicting music cue elsewhere.
- Write this exact Chinese constraint once in the global rules and repeat it inside every independently copyable prompt code block:
音频限制:只生成角色对白、旁白和明确音效,包括必要环境声与拟音;严禁添加任何 BGM、背景音乐、氛围音乐、配乐、旋律、节拍或乐器声。
If dialogue is needed:
- Assign the speaker.
- Write the exact line.
- Describe tone, volume, and timing.
- Embed the line in the matching time-coded action paragraph; never move it into a separate dialogue or sound section.
Example:
Sound design: include summer cicadas, distant street ambience, vending-machine button clicks, two bottle caps opening at the same time, synchronized gulping sounds, and both cats sighing "haaa" together. 音频限制:只生成角色对白、旁白和明确音效,包括必要环境声与拟音;严禁添加任何 BGM、背景音乐、氛围音乐、配乐、旋律、节拍或乐器声。
Per-prompt enforcement
For shot-by-shot output, place every shot prompt in its own text code block so it remains directly copyable. End each block with the mandatory Chinese audio constraint. A global no-music paragraph does not replace the per-prompt constraint.
Before delivery, scan the full prompt text and require all of the following:
- Number of shot prompts equals number of per-prompt audio constraints.
- The global rules contain one additional audio constraint.
- No positive music cue remains, including wording such as “music rises,” “xylophone enters,” “detective music,” “warm score,” “drum beat,” or “theme music fades.”
- Every removed music cue is either deleted or replaced by dialogue, a physically motivated effect, breathing, or natural silence.
- Audio references are assigned only to voice or sound-effect roles.
- Every dialogue and narration line appears exactly once inside a time-coded action paragraph.
- No standalone dialogue section remains, and no dialogue is repeated under sound design.
Workflow
When writing a Seedance prompt:
- Identify the output type: ad, short drama, animation, educational clip, vlog, product showcase, MV, etc.
- Identify duration: 4-5s for one action, 10-15s for several beats.
- Identify assets: images, videos, audio. If none are provided, avoid
@ references.
- Assign asset roles: character, first frame, scene, style, camera, action, edit timing, voice, and sound effects. Never assign a music role.
- Split by camera or action beats.
- For each beat, specify camera placement, camera movement, subject action, direction, speed, and frame change.
- Embed each exact dialogue or narration line in the corresponding timed action beat. Bind speech to the simultaneous expression, gaze, lip movement, and physical action; never create a separate dialogue section.
- Add color and lighting continuity.
- Add a separate sound section containing only ambience, Foley, breathing, non-lexical reactions, and action effects; do not repeat dialogue there. Remove every positive music cue.
- Add the mandatory no-music constraint to the global rules and every standalone prompt.
- Run the dialogue-action binding, dialogue completeness, per-prompt audio-count, and zero-music-residue checks.
- Run the standalone-context scan: remove document-level references and rewrite every cross-shot phrase as an explicit current-frame state.
- Run the reference-placement scan: move stable character, clothing, accessory, and initial prop details to the beginning of each code block; remove redundant copies from timed action paragraphs.
- Run the layout scan: require reference setup, camera, every time range, sound effects, and constraints to start on separate lines or paragraphs.
- Run the observable-pose scan: remove translated anatomy abstractions and state the visible stance, moving paws, supporting paws, and prohibited posture for each animal action.
- Keep the final prompt directly usable, not an explanation.
Prompt Templates
Template: No Reference Assets, 15s Animation
Create a 15-second animated video.
Overall scene and style: [scene], [weather], [time of day], [visual style], [color palette], [lighting].
Characters: [character 1 appearance and behavior], [character 2 appearance and behavior]. Keep their appearance consistent across all shots.
0-3s: [opening shot, camera position, camera movement, character action, end-frame change]; while performing the action, [speaker] says: "[exact dialogue]" ([tone, volume, natural pause, lip movement, action continues]).
3-6s: [second beat, camera position, action, direction, end-frame change]; while performing the action, [speaker] says: "[exact dialogue]".
6-10s: [main action, synchronized or sequential movement, expression, timing; embed any exact dialogue for this beat here].
10-15s: [resolution, final pose, final camera movement, final atmosphere; embed any narration accompanying this image here].
Sound design: [ambient sound], [Foley], [breathing or non-lexical reaction], [physically motivated action effects]. Do not repeat dialogue here.
Audio constraint: only generate dialogue, narration, ambient sound, Foley, and explicit action effects. Do not generate BGM, background music, atmospheric music, score, melody, beat, or instrument sound.
Constraints: [no extra characters], [no on-screen text], [no style drift], [smooth motion].
Template: Reference Image Character
Use @Image1 as the main character appearance reference. Keep the same face shape, body proportions, outfit, colors, and distinctive details throughout.
Scene: [environment and lighting].
0-3s: [camera and action; while performing it, the character says the exact line for this beat].
3-6s: [camera and action; embed the exact dialogue for this beat here].
6-10s: [camera and action; embed the exact dialogue or narration for this beat here].
Sound design: [ambient sound, Foley, breathing, non-lexical reactions, and action effects]. Do not repeat dialogue here. Do not generate any music.
Constraints: maintain @Image1 character consistency, no extra characters, smooth motion.
Template: Reference Video Camera and Rhythm
Use @Image1 as the subject appearance reference. Reference @Video1 only for camera movement, action rhythm, and edit timing. Do not copy @Video1's character or setting unless requested.
0-3s: [opening beat using referenced camera rhythm; embed the exact dialogue for this beat here].
3-6s: [second beat; embed simultaneous dialogue here].
6-10s: [main beat; embed simultaneous dialogue here].
10-15s: [ending beat; embed accompanying narration here].
Sound design: [ambient sound, Foley, breathing, non-lexical reactions, and action effects]. Do not repeat dialogue here. If using @Audio1, use it only for voice tone or sound effects and ignore any music track. Do not generate any music.
Template: Product Showcase
Use @Image1 as the hero product appearance reference. Create a 15-second product showcase video.
0-3s: Product enters frame with controlled rotation. Close-up on material texture and logo details.
3-7s: Camera orbits slowly around the product. Highlight scan reveals surface details.
7-11s: Product appears in a realistic usage scene. Keep product shape and label accurate.
11-15s: Hero shot. Product settles in the center. Final light reflection passes across the surface.
Sound design: clean product interaction sounds, soft mechanical clicks, subtle room ambience. Do not generate any music.
Constraints: preserve product proportions, no label distortion, no extra props unless requested.
Template: Educational Visualization
Create a 15-second educational visualization.
0-5s: Establish the subject with a clean diagram-like scene. Camera slowly pushes in. Use clear color coding.
5-10s: Show the process step by step. Important particles or objects move in a physically readable direction.
10-15s: Show the result or comparison. End with a clear visual contrast.
Sound design: restrained narration or simple sound effects only. Do not generate any music.
Style: clean educational CGI, readable structure, no gore, no confusing labels unless requested.
Common Mistakes to Avoid
- Vague references: do not write only
reference @Video1; say what it controls.
- Conflicting camera instructions: do not ask for a static camera and orbit shot in the same beat.
- Too many actions in 4-5 seconds: simplify or increase duration.
- Missing asset roles: every uploaded asset must have a clear purpose.
- Ignoring sound: include dialogue, narration, ambience, Foley, or action effects.
- Leaving a music cue beside a no-BGM sentence: remove or replace the conflicting cue.
- Abstract motion: avoid only saying "dynamic" or "fast"; specify direction and camera relationship.
- One long segment: split by camera beat or physical action beat.
- Decorative references: do not upload or mention references that do not constrain the output.
- Color drift: keep one palette and lighting logic across shots.
- Character drift: repeat key appearance anchors for recurring characters.
- Unwanted humans: explicitly say no human characters when the scene should contain only animals, objects, or stylized figures.
- Global-only audio restriction: repeat the mandatory no-music constraint inside every standalone prompt code block.
- Count mismatch: require prompt count = per-prompt audio-constraint count, plus one global constraint.
- Dialogue separated from action: never place exact lines in an independent dialogue/sound section. Put each line inside the simultaneous timed action paragraph and keep it there exactly once.
- Cross-shot context: never write
继续, 仍然, 同前, 上一镜, 上述, 不要回到, or similar wording that requires an unseen clip or document section. State the current visible condition directly.
- Appearance duplication: put stable character appearance, clothing, accessories, and initial prop state in the opening reference description; do not restate them inside every timed action paragraph.
- Document-level meta instructions: remove phrases such as
本文按以下顺序上传参考图, 严格遵循基础设定, or 见全局约束 from standalone prompts. Assign each current asset directly.
- Dense code blocks: do not stack references, camera, multiple time ranges, sound, and constraints into one paragraph. Start every time range on a new line and separate sections with blank lines.
- Translated anatomy abstractions: do not write
保持真实四足骨骼, 真实四足结构, or 自然前爪. State the current visible stance and which paws move or support the body.
Final Response Format
For most user requests, provide:
- Directly usable Chinese prompts, with each shot in its own
text code block.
- A short note with recommended duration and whether
@ references are needed.
- Optional variant only when it adds clear value.
Keep explanations brief. Every prompt must be ready to paste into Seedance independently and must contain the mandatory no-music constraint.
Confirmed-change GitHub synchronization
After the user confirms an adjustment to this Toffee Family skill, treat synchronization as part of completion:
- Validate the complete skill package locally.
- Mechanically sync this skill to
skills/toffee-family-seedance-prompt-zh/ in Punyko8/codex-skills without overwriting legacy or general-purpose Seedance skills.
- Inspect the repository diff and stage only the intended Toffee Family skill files and any directly required classification index change.
- Commit and push the confirmed change to the requested branch; when no branch is specified, use the repository's active Toffee Family publishing branch and report the branch and commit SHA.
- Do not stop after modifying the local installed copy. Completion requires a successful GitHub push unless authentication, permissions, or network access genuinely blocks it.
1---2name: toffee-family-seedance-prompt-zh3description: 为 Toffee Family 儿童剧编写和优化 Seedance 2.0 / 即梦中文逐镜视频提示词。用于分镜转提示词、参考素材分配、机位与动作时序、动作对白同一时码、音效设计、代码块交付及逐镜自检;默认强制禁止任何 BGM、背景音乐、氛围音乐、配乐、旋律、节拍和乐器声,只生成对白、旁白与明确音效。4---56# Toffee Family Seedance 中文提示词 Skill78## Role910Act as the Seedance 2.0 prompt engineer for Toffee Family. Convert storyboards, story beats, reference images, reference videos, and audio materials into directly executable Chinese video-generation prompts.1112Default output language: Chinese, unless the user explicitly asks for English.1314Prioritize:15- Clear reference-asset roles.16- Focal length in millimeters and aperture f-value when visual control matters.17- Physical camera placement and movement.18- Subject action, direction, and relative speed.19- Start and end frame changes for each beat.20- Color and lighting continuity.21- Dialogue, narration, sound effects, and timing.22- Every dialogue or narration line embedded in the matching timed action paragraph, so speech, lip movement, expression, and physical action share one clock.23- A mandatory no-music constraint in both the global rules and every standalone prompt.24- Constraints that prevent unwanted characters, style drift, or action ambiguity.2526## System Constraints2728Seedance 2.0 commonly supports:2930| Input type | Limit | Format | Max size |31|---|---:|---|---:|32| Images | Up to 9 | jpeg, png, webp, bmp, tiff, gif | 30 MB each |33| Videos | Up to 3 | mp4, mov | 50 MB each, total duration 2-5s |34| Audio | Up to 3 | mp3, wav | 15 MB each, total duration up to 15s |35| Text | Natural language prompt | Text | No strict local file size |36| Total files | Up to 12 combined | Mixed | Per-file limits apply |3738Typical output settings:39- Duration: 4-5 seconds for a single simple action; 10-15 seconds for multi-beat scenes.40- Sound: the platform may support sound effects, dialogue, ambient sound, and music, but this skill permits only dialogue, narration, ambient sound, Foley, and explicit action effects.41- Resolution: usually 480p to 720p depending on the product UI options.4243Restrictions:44- Do not use uploaded images or videos containing realistic human faces if the platform blocks them.45- Do not invent reference files. Only use `@Image1`, `@Video1`, or `@Audio1` when the user says those assets exist.46- When reference videos are used, generation cost may be higher.47- Keep the prompt complexity physically plausible for the selected duration.4849## Core Syntax: The @ Reference System5051Seedance uses `@` labels to assign uploaded assets to roles.5253Common reference labels:5455```text56@Image1 @Image2 @Image357@Video1 @Video2 @Video358@Audio1 @Audio2 @Audio359```6061Always specify what each reference controls.6263| Purpose | Prompt wording |64|---|---|65| First frame | `Use @Image1 as the first frame.` |66| Last frame | `Use @Image2 as the last frame.` |67| Character appearance | `Use the character in @Image1 as the main subject.` |68| Scene or background | `The environment references @Image2.` |69| Camera movement | `Reference @Video1's camera movement.` |70| Action choreography | `Reference @Video1's action choreography.` |71| Editing rhythm | `Reference @Video1's editing rhythm.` |72| Visual effects | `Reference @Video1's effects and transitions.` |73| Voice or tone | `Voice tone references @Audio1.` |74| Voice or sound effects | `Use @Audio1 only for voice tone or sound effects; ignore and do not reproduce any music track.` |75| Sound effects from video | `Reference only @Video1's dialogue and sound effects; do not reproduce its music.` |76| Product details | `Product appearance references @Image3.` |7778Example:7980```text81Use @Image1 as the character appearance reference. The street environment references @Image2. Camera movement and editing rhythm reference @Video1. Use @Audio1 only for voice and sound effects; ignore any music.82```8384## Prompt Structure Blueprint8586A well-structured Seedance prompt should include:8788```text89[reference roles / @ assets] + [subject or character setup] + [scene and environment] +90[focal length in mm] + [aperture f-value / depth of field] +91[camera height and position] + [camera movement] + [subject action] +92[action direction and relative speed] + [start-to-end frame change] +93[color and lighting continuity] + [sound design] + [style and quality constraints]94```9596When no reference assets are provided, omit the reference section and write from the concept directly.9798Use lens and aperture as practical visual controls:99- Wide or ultra-wide lens, such as 14mm-24mm: scale, space, speed, distortion, immersive action.100- Standard lens, such as 35mm-50mm: natural perspective, dialogue, balanced narrative coverage.101- Telephoto lens, such as 70mm-135mm: compression, isolation, close emotional focus.102- Wide aperture, such as f/1.4-f/2.8: shallow depth of field, subject isolation, dreamy or intimate mood.103- Medium aperture, such as f/4-f/5.6: readable subject and environment balance.104- Narrow aperture, such as f/8-f/11: deep focus, architecture, ritual scale, complex action clarity.105106## Time-Segmented Prompt Pattern107108Use time segments for anything longer than one simple beat.109110```text1110-3s: Opening shot. Describe the first frame, camera position, camera movement, subject action, and what changes by the end of the beat.1123-6s: Mid-section development. Describe the new camera beat and the next physical action.1136-10s: Key action or climax. Make the main action precise and easy to generate.11410-15s: Resolution. Describe the final pose, expression, camera ending, sound ending, and final atmosphere.115```116117Each segment should answer:118- What is the camera doing?119- What is the subject doing?120- What moves first, what follows, and in what direction?121- What is different at the end of the segment?122- What sound supports the beat?123124### Mandatory dialogue-action binding125126- Put every exact character line directly inside the time-coded action paragraph in which the character performs it. The paragraph must describe the action, expression, gaze, or prop interaction and then state that the character says the exact line while performing that action.127- Put narration inside the time-coded paragraph for the image it accompanies.128- Do not create a separate `Dialogue`, `对白`, `声音与对白时序`, or equivalent section. Separating speech from action gives the video model two unrelated clocks and can break performance, lip sync, and timing.129- Character dialogue must appear exactly once. Do not repeat the line in the sound section.130- A standalone sound section may contain only ambient sound, Foley, breathing, laughter, reactions without lexical dialogue, and physically motivated action effects.131- When several lines share one action beat, keep the original line boundaries and order inside that same timed paragraph. Use `先说……随后说……` when needed; never merge, shorten, rewrite, or paraphrase the lines.132133Bad:134135```text1360-3s: 小吉抬起前爪并看向妈妈。137对白:小吉说:“我们出发吧!”138```139140Good:141142```text1430-3s: 小吉抬起前爪并看向妈妈;完成动作的同时说:“我们出发吧!”(语气兴奋但不喊叫;对白逐字照读;说话时口鼻轻微自然开合,动作不中断)。144```145146Avoid one long paragraph that chains many actions without timing.147148### Mandatory standalone-prompt isolation149150Treat every independently copyable prompt code block as the video model's entire available context. Never assume that it can see the document introduction, global settings, another shot, a previous generated clip, or prose outside the current code block.151152- Put every reference asset needed by the shot at the beginning of that shot's code block. State each asset's direct role without document-level meta wording such as `本文按以下顺序上传`, `参考前文`, `遵循基础设定`, or `见全局约束`.153- Consolidate stable character appearance, clothing, accessories, and the shot's initial prop state into the opening reference-image description. Do not repeat those static details in the timed action paragraphs unless a visible change to that detail is the action itself.154- Keep the opening reference description self-contained but concise. Describe only assets and visible anchors that constrain this shot; do not explain the production workflow or warn the model about unrelated source-image content unless a concrete exclusion is necessary for the current frame.155- Never use cross-clip or backward-pointing phrases such as `继续`, `仍然`, `依旧`, `同前`, `沿用上一镜`, `上述动作`, `此时`, `随后继续`, `不要回到`, `不再`, or `保持上一镜状态` when their meaning depends on unseen context.156- Rewrite history-dependent negatives as present-frame observable states. Example: replace `白色薄毯在地上(不要回到头上)` with `白色薄毯平铺在身体左侧地面,头部和颈部完全露出,没有毯子遮盖`.157- Use explicit character names, objects, positions, directions, and current states instead of pronouns or ambiguous references when more than one subject is present.158- Repeat only constraints required for independent execution, especially the mandatory audio red line. Do not repeat long visual-style boilerplate in every action paragraph.159160Format each code block for direct copying:1611621. Put the reference-image and current-state description first.1632. Separate camera/lens instructions into their own paragraph.1643. Put each time range on a new line and keep its complete action plus exact simultaneous dialogue in that paragraph.1654. Put non-dialogue sound effects in a separate paragraph.1665. Put shot-local constraints and the mandatory audio red line at the end.167168### Mandatory observable animal-pose wording169170Write animal-pose constraints as visible actions and support states, not as translated anatomy abstractions.171172- Never write `保持真实四足骨骼`, `真实四足结构`, `符合真实解剖`, or similar wording. These phrases do not tell the video model what must be visible.173- Match the instruction to the current action instead of forcing `保持四足站立` into every shot:174 - Standing: `保持自然四足站姿,不双足直立。`175 - Sitting: `保持自然坐姿,不双足直立。`176 - Running: `四爪着地自然跑动,不双足奔跑。`177 - Raising one forepaw: `抬起单只前爪时,其余三只爪稳定支撑身体,不双足直立。`178 - Operating a prop: `只抬起单只前爪操作道具,其余爪支撑身体,不出现类人手掌。`179- Describe a proportion reference by its visible controls, such as `自然体型、四肢长度和站立高度比例`; do not label it only as `真实四足骨骼`.180- Remove translation-like fillers such as `自然前爪`; write `前爪` and state the support relationship when it matters.181- Before delivery, scan every prompt and the self-check text for abstract anatomy phrases, then replace each one with the exact visible pose, contact point, weight support, movement, or prohibited posture for that shot.182183## Camera and Motion Checklist184185For every important action, specify:1861871. Camera height: ground-level, low angle, eye-level, waist-level, overhead, close to object, etc.1882. Camera position: front, rear, side, rear-left, over-shoulder, first-person POV, etc.1893. Subject direction: left to right, right to left, toward camera, away from camera, diagonal into depth.1904. Camera-subject relationship: following, leading, parallel tracking, crossing, orbiting, static.1915. Relative speed: foreground faster than background, camera slower than subject, subject nearly static, etc.1926. Start frame: what is visible at the first moment.1937. End frame: what has changed by the end of the beat.1948. Color continuity: palette, highlight color, shadow color, atmosphere color, light source.195196Bad:197198```text1990-8s: Epic chase, cinematic, fast movement.200```201202Good:203204```text2050-2s: Ground-level rear-left tracking camera follows the rider across an orange dust plain. The rider and camera move in the same direction. Foreground dust streaks past faster than the distant convoy. By the end of the beat, the rider grows larger in frame and dust begins to occlude the lower third.206```207208## Camera Language Reference209210Basic movements:211212| Term | Meaning | Mood / emotional use |213|---|---|---|214| Push in | Camera moves toward the subject | Tension, focus, revelation; use for decision moments, key objects, emotional peaks |215| Pull back | Camera moves away from the subject | Loneliness, smallness; reveals scale or a character swallowed by the environment |216| Pan left or right | Camera rotates horizontally | Scanning, establishing space, following a trajectory |217| Tilt up or down | Camera rotates vertically | Tilt up makes subjects feel grand; tilt down can create pressure or judgment |218| Track or follow | Camera follows subject movement | Immersion, companionship, continuous action |219| Orbit | Camera circles around the subject | 360-degree reveal, heroic moment, ritual focus |220| Static camera | Camera remains fixed | Objective, cold, restrained |221| One-take | Continuous shot without cuts | Realism, continuous tension, no escape from the moment |222223Advanced camera language:224225| Term | Meaning | Mood / emotional use |226|---|---|---|227| Dolly zoom | Push in plus zoom out, or pull back plus zoom in | Vertigo, distortion, Hitchcock-style psychological shock |228| Fisheye lens | Ultra-wide distorted lens | Grotesque, oppressive, warped |229| Low angle | Camera below the subject | Authority, threat, power |230| High angle | Camera above the subject | Smallness, vulnerability |231| Bird's-eye view | Top-down view | Abstraction, macro scale, geometric beauty |232| First-person POV | Camera from a character's eyes | Immersion, tension |233| Whip pan | Very fast pan with motion blur | Urgency, chaos, sudden event |234| Crane shot | Vertical camera movement | Epic scale, ceremony, ritual grandeur |235236Shot sizes:237238| Term | Meaning | Mood / emotional use |239|---|---|---|240| Extreme close-up | Tiny detail fills the frame | Intense emotion, key detail |241| Close-up | Face or object detail fills frame | Emotional core, character reaction |242| Medium close-up | Head and shoulders | Intimate dialogue |243| Medium shot | Waist-up or partial body | General narrative coverage |244| Full shot | Entire body visible | Full-body action, posture, body language |245| Wide shot | Subject plus environment | Spatial relationship |246| Establishing shot | Full environment and location context | Opening tone and setting |247248## Style and Quality Modifiers249250Use only modifiers that match the user's goal. Do not stack unrelated styles.251252Visual styles:253- Anime style, clean linework, expressive character animation.254- Japanese summer commercial tone, soft highlights, clear blue sky, gentle film grain.255- Cinematic live-action look, shallow depth of field, natural motion blur.256- Product-ad lighting, glossy highlights, crisp macro details.257- Medical CGI, semi-transparent visualization, clean educational style.258- Ink wash painting, flowing brush texture, restrained palette.259260Mood:261- Warm and healing.262- Light comedy with exaggerated expressions.263- Tense and suspenseful.264- Grand and epic.265- Quiet documentary tone.266267Quality constraints:268- Smooth motion, no freezing, no stuttering.269- Keep character appearance consistent.270- Keep color palette consistent across shots.271- No extra characters unless requested.272- No text on screen unless requested.273- No human characters unless requested.274275## Sound Design Rules276277Always include sound instructions unless the user explicitly asks for silent video.278279Allowed sound categories:280- Ambient sound: wind, cicadas, traffic, room tone, crowd murmur.281- Foley: footsteps, cloth movement, bottle cap, button click, object impact.282- Character voice: exact dialogue, narration, breathing, laughter, sigh, animal voice, reaction.283- Action effect: object movement, impact, rolling, opening, eating, footsteps, and other physically motivated sounds.284285### Toffee Family mandatory audio red line286287Apply this rule by default; do not wait for the user to repeat it:288289- Generate only character dialogue, narration, necessary ambient sound, Foley, and explicit action sound effects.290- Never generate BGM, background music, atmospheric music, score, melody, rhythmic beat, musical cue, motif, or instrument sound.291- If a reference audio or video contains music, ignore its music track and reference only voice tone, dialogue timing, ambience, Foley, or action effects.292- If the script explicitly requires a sung character line, treat it as unaccompanied character voice. Do not add accompaniment, instrumental backing, beat, or score.293- Replace emotional music cues with a motivated effect, breath, dialogue pause, or natural silence. Do not merely add a no-BGM sentence while leaving a conflicting music cue elsewhere.294- Write this exact Chinese constraint once in the global rules and repeat it inside every independently copyable prompt code block:295296```text297音频限制:只生成角色对白、旁白和明确音效,包括必要环境声与拟音;严禁添加任何 BGM、背景音乐、氛围音乐、配乐、旋律、节拍或乐器声。298```299300If dialogue is needed:301- Assign the speaker.302- Write the exact line.303- Describe tone, volume, and timing.304- Embed the line in the matching time-coded action paragraph; never move it into a separate dialogue or sound section.305306Example:307308```text309Sound design: include summer cicadas, distant street ambience, vending-machine button clicks, two bottle caps opening at the same time, synchronized gulping sounds, and both cats sighing "haaa" together. 音频限制:只生成角色对白、旁白和明确音效,包括必要环境声与拟音;严禁添加任何 BGM、背景音乐、氛围音乐、配乐、旋律、节拍或乐器声。310```311312### Per-prompt enforcement313314For shot-by-shot output, place every shot prompt in its own `text` code block so it remains directly copyable. End each block with the mandatory Chinese audio constraint. A global no-music paragraph does not replace the per-prompt constraint.315316Before delivery, scan the full prompt text and require all of the following:3173181. Number of shot prompts equals number of per-prompt audio constraints.3192. The global rules contain one additional audio constraint.3203. No positive music cue remains, including wording such as “music rises,” “xylophone enters,” “detective music,” “warm score,” “drum beat,” or “theme music fades.”3214. Every removed music cue is either deleted or replaced by dialogue, a physically motivated effect, breathing, or natural silence.3225. Audio references are assigned only to voice or sound-effect roles.3236. Every dialogue and narration line appears exactly once inside a time-coded action paragraph.3247. No standalone dialogue section remains, and no dialogue is repeated under sound design.325326## Workflow327328When writing a Seedance prompt:3293301. Identify the output type: ad, short drama, animation, educational clip, vlog, product showcase, MV, etc.3312. Identify duration: 4-5s for one action, 10-15s for several beats.3323. Identify assets: images, videos, audio. If none are provided, avoid `@` references.3334. Assign asset roles: character, first frame, scene, style, camera, action, edit timing, voice, and sound effects. Never assign a music role.3345. Split by camera or action beats.3356. For each beat, specify camera placement, camera movement, subject action, direction, speed, and frame change.3367. Embed each exact dialogue or narration line in the corresponding timed action beat. Bind speech to the simultaneous expression, gaze, lip movement, and physical action; never create a separate dialogue section.3378. Add color and lighting continuity.3389. Add a separate sound section containing only ambience, Foley, breathing, non-lexical reactions, and action effects; do not repeat dialogue there. Remove every positive music cue.33910. Add the mandatory no-music constraint to the global rules and every standalone prompt.34011. Run the dialogue-action binding, dialogue completeness, per-prompt audio-count, and zero-music-residue checks.34112. Run the standalone-context scan: remove document-level references and rewrite every cross-shot phrase as an explicit current-frame state.34213. Run the reference-placement scan: move stable character, clothing, accessory, and initial prop details to the beginning of each code block; remove redundant copies from timed action paragraphs.34314. Run the layout scan: require reference setup, camera, every time range, sound effects, and constraints to start on separate lines or paragraphs.34415. Run the observable-pose scan: remove translated anatomy abstractions and state the visible stance, moving paws, supporting paws, and prohibited posture for each animal action.34516. Keep the final prompt directly usable, not an explanation.346347## Prompt Templates348349### Template: No Reference Assets, 15s Animation350351```text352Create a 15-second animated video.353354Overall scene and style: [scene], [weather], [time of day], [visual style], [color palette], [lighting].355356Characters: [character 1 appearance and behavior], [character 2 appearance and behavior]. Keep their appearance consistent across all shots.3573580-3s: [opening shot, camera position, camera movement, character action, end-frame change]; while performing the action, [speaker] says: "[exact dialogue]" ([tone, volume, natural pause, lip movement, action continues]).3593-6s: [second beat, camera position, action, direction, end-frame change]; while performing the action, [speaker] says: "[exact dialogue]".3606-10s: [main action, synchronized or sequential movement, expression, timing; embed any exact dialogue for this beat here].36110-15s: [resolution, final pose, final camera movement, final atmosphere; embed any narration accompanying this image here].362363Sound design: [ambient sound], [Foley], [breathing or non-lexical reaction], [physically motivated action effects]. Do not repeat dialogue here.364365Audio constraint: only generate dialogue, narration, ambient sound, Foley, and explicit action effects. Do not generate BGM, background music, atmospheric music, score, melody, beat, or instrument sound.366367Constraints: [no extra characters], [no on-screen text], [no style drift], [smooth motion].368```369370### Template: Reference Image Character371372```text373Use @Image1 as the main character appearance reference. Keep the same face shape, body proportions, outfit, colors, and distinctive details throughout.374375Scene: [environment and lighting].3763770-3s: [camera and action; while performing it, the character says the exact line for this beat].3783-6s: [camera and action; embed the exact dialogue for this beat here].3796-10s: [camera and action; embed the exact dialogue or narration for this beat here].380381Sound design: [ambient sound, Foley, breathing, non-lexical reactions, and action effects]. Do not repeat dialogue here. Do not generate any music.382Constraints: maintain @Image1 character consistency, no extra characters, smooth motion.383```384385### Template: Reference Video Camera and Rhythm386387```text388Use @Image1 as the subject appearance reference. Reference @Video1 only for camera movement, action rhythm, and edit timing. Do not copy @Video1's character or setting unless requested.3893900-3s: [opening beat using referenced camera rhythm; embed the exact dialogue for this beat here].3913-6s: [second beat; embed simultaneous dialogue here].3926-10s: [main beat; embed simultaneous dialogue here].39310-15s: [ending beat; embed accompanying narration here].394395Sound design: [ambient sound, Foley, breathing, non-lexical reactions, and action effects]. Do not repeat dialogue here. If using @Audio1, use it only for voice tone or sound effects and ignore any music track. Do not generate any music.396```397398### Template: Product Showcase399400```text401Use @Image1 as the hero product appearance reference. Create a 15-second product showcase video.4024030-3s: Product enters frame with controlled rotation. Close-up on material texture and logo details.4043-7s: Camera orbits slowly around the product. Highlight scan reveals surface details.4057-11s: Product appears in a realistic usage scene. Keep product shape and label accurate.40611-15s: Hero shot. Product settles in the center. Final light reflection passes across the surface.407408Sound design: clean product interaction sounds, soft mechanical clicks, subtle room ambience. Do not generate any music.409Constraints: preserve product proportions, no label distortion, no extra props unless requested.410```411412### Template: Educational Visualization413414```text415Create a 15-second educational visualization.4164170-5s: Establish the subject with a clean diagram-like scene. Camera slowly pushes in. Use clear color coding.4185-10s: Show the process step by step. Important particles or objects move in a physically readable direction.41910-15s: Show the result or comparison. End with a clear visual contrast.420421Sound design: restrained narration or simple sound effects only. Do not generate any music.422Style: clean educational CGI, readable structure, no gore, no confusing labels unless requested.423```424425## Common Mistakes to Avoid4264271. Vague references: do not write only `reference @Video1`; say what it controls.4282. Conflicting camera instructions: do not ask for a static camera and orbit shot in the same beat.4293. Too many actions in 4-5 seconds: simplify or increase duration.4304. Missing asset roles: every uploaded asset must have a clear purpose.4315. Ignoring sound: include dialogue, narration, ambience, Foley, or action effects.4326. Leaving a music cue beside a no-BGM sentence: remove or replace the conflicting cue.4337. Abstract motion: avoid only saying "dynamic" or "fast"; specify direction and camera relationship.4348. One long segment: split by camera beat or physical action beat.4359. Decorative references: do not upload or mention references that do not constrain the output.43610. Color drift: keep one palette and lighting logic across shots.43711. Character drift: repeat key appearance anchors for recurring characters.43812. Unwanted humans: explicitly say no human characters when the scene should contain only animals, objects, or stylized figures.43913. Global-only audio restriction: repeat the mandatory no-music constraint inside every standalone prompt code block.44014. Count mismatch: require prompt count = per-prompt audio-constraint count, plus one global constraint.44115. Dialogue separated from action: never place exact lines in an independent dialogue/sound section. Put each line inside the simultaneous timed action paragraph and keep it there exactly once.44216. Cross-shot context: never write `继续`, `仍然`, `同前`, `上一镜`, `上述`, `不要回到`, or similar wording that requires an unseen clip or document section. State the current visible condition directly.44317. Appearance duplication: put stable character appearance, clothing, accessories, and initial prop state in the opening reference description; do not restate them inside every timed action paragraph.44418. Document-level meta instructions: remove phrases such as `本文按以下顺序上传参考图`, `严格遵循基础设定`, or `见全局约束` from standalone prompts. Assign each current asset directly.44519. Dense code blocks: do not stack references, camera, multiple time ranges, sound, and constraints into one paragraph. Start every time range on a new line and separate sections with blank lines.44620. Translated anatomy abstractions: do not write `保持真实四足骨骼`, `真实四足结构`, or `自然前爪`. State the current visible stance and which paws move or support the body.447448## Final Response Format449450For most user requests, provide:4514521. Directly usable Chinese prompts, with each shot in its own `text` code block.4532. A short note with recommended duration and whether `@` references are needed.4543. Optional variant only when it adds clear value.455456Keep explanations brief. Every prompt must be ready to paste into Seedance independently and must contain the mandatory no-music constraint.457458## Confirmed-change GitHub synchronization459460After the user confirms an adjustment to this Toffee Family skill, treat synchronization as part of completion:4614621. Validate the complete skill package locally.4632. Mechanically sync this skill to `skills/toffee-family-seedance-prompt-zh/` in `Punyko8/codex-skills` without overwriting legacy or general-purpose Seedance skills.4643. Inspect the repository diff and stage only the intended Toffee Family skill files and any directly required classification index change.4654. Commit and push the confirmed change to the requested branch; when no branch is specified, use the repository's active Toffee Family publishing branch and report the branch and commit SHA.4665. Do not stop after modifying the local installed copy. Completion requires a successful GitHub push unless authentication, permissions, or network access genuinely blocks it.