Video Prompt
The prompt layer of the video-prompt-craft toolkit: turn a shot (or a full shot list from video-storyboard) into
structured, platform-ready prompt text. This skill owns the platform syntax, the universal formula, and optimization —
it does NOT own narrative structure or shot design (that is video-storyboard's job).
三件套协作 / Three coordinated skills:
video-storyboard— idea → narrative shot list (story structure, shot language, visual intensity)video-prompt(this skill) — shot list → platform promptsvideo-review— multi-angle audit of storyboard & promptsIf a user hands you a rough idea with no shot plan and it needs multiple shots, suggest running
video-storyboardfirst, then return here to render each shot. For a single shot or a direct prompt request, proceed directly.
Core Principle
You are not describing the entire world — you are describing one shot through a camera lens. For Seedance 2.0: think like an AI director — spatial layer (what's in frame) + temporal layer (how it changes over time).
Every effective video prompt answers five questions:
| Question | Element |
|---|---|
| What am I looking at? | Subject + Environment |
| What is happening? | One primary action |
| How is it filmed? | Shot type + Camera movement |
| What does it feel like? | Style + Lighting + Color tone |
| What is the rhythm? | Pacing / Cut vs. long take |
Step 1: Classify the Task
Determine the generation mode before writing the prompt:
| Mode | Input | Prompt Focus |
|---|---|---|
| Text-to-Video (T2V) | Text only | Full description: subject + scene + motion + camera + style |
| Image-to-Video (I2V) | Image + text | Motion + camera only (subject/scene/style already in image) |
| First+Last Frame | 2 images + text | Bridge the two frames; describe the transition motion |
| Multi-shot / Storyboard | Text | Overall narrative + per-shot descriptions (镜头1/2/3) |
| Multimodal Reference | Images/videos/audio + text | Subject anchoring + @ reference syntax (see below) |
| Video Edit | Video + text | Describe what to add/remove/change; keep everything else implicit |
| Video Extension | Video + text | Describe what happens before/after; do NOT use "参考" prefix |
| Sound-synced | Text | Standard elements + explicit sound description |
Step 2: Construct the Prompt
Basic Formula (Universal)
[Subject description] + [Scene/Environment] + [Primary action] + [Shot type] + [Camera movement] + [Lighting/Style] + [Quality/Constraint tags]
Advanced Formula (Seedance 2.0)
精准主体 + 动作细节 + 场景环境 + 光影色调 + 镜头运镜 + 视觉风格 + 画质 + 约束条件
This maps to two internal model dimensions:
- Spatial layer: who/what, where, appearance, environment
- Temporal layer: action sequence, camera movement, timing, transitions
Rules
- One shot, one action — keep the primary motion singular and clear.
- Be concrete, not abstract — "a woman in a white sundress" beats "a beautiful person"; "golden hour warm sunlight" beats "nice lighting".
- Order matters — Subject first, then scene, then motion, then camera, then aesthetics.
- For I2V: skip what the image already shows — only describe motion and camera work.
- For multi-shot: use 镜头1/镜头2/镜头3 labels — do NOT force precise timestamps; let the model pace naturally.
- For sound: be explicit about silence — write "无台词" / "No dialogue." and "无背景音乐" / "No background music." if needed.
- Keep it under 300 words per shot — most models perform best with 50–200 words per shot.
- One camera move per shot — mixing push+orbit+pan in one shot increases instability.
Step 3: Seedance 2.0 — Subject Definition & Reference System
This section is critical for Seedance 2.0 multimodal reference tasks.
3a. Subject Definition
When referencing specific subjects in uploaded assets, always define them explicitly.
Syntax:
将<图片/视频N>中的[2-3个清晰静态特征]定义为<主体N>
Example:
将图片1中穿红色连衣裙、戴草帽的女人定义为主体1
Rules:
- Use 2–3 stable static features (clothing, hairstyle, appearance, type) for identification
- For multi-character scenes: name each character and link to their reference image
将视频1中的高个子男人定义为警察,矮个子男人定义为小偷 - For unnamed subjects: use
主体1@图片1inline every time the subject is referenced - Do NOT use Asset ID as a subject reference — always use
<图片/视频N> - Most important references go first in the prompt
3b. Reference Syntax (Multimodal Tasks)
| Reference type | Syntax | Purpose |
|---|---|---|
| Image subject | 参考<图片N>中的<主体> |
Character identity, visual style |
| Video action | 参考<视频N>中的动作/运镜/风格/音效 |
Motion, camera, effects, audio |
| Audio timbre | 参考<音频N>中的音色 |
Voice tone replication |
| First frame | 以<图片N>为首帧 |
Anchor opening frame |
| Last frame | 以<图片N>为尾帧 |
Anchor closing frame |
Input limits (Seedance 2.0):
- Images: 0–9
- Videos: 0–3
- Audio: 0–3
- Recommended configuration: 1–2 character photos + 1 scene photo + 1 camera reference video + 1 audio = 4–5 assets total
- Do not max out inputs — too many assets confuse the model's priority weighting
3c. Task Type Sentence Patterns
Multimodal Reference (生成新视频):
参考<图片/视频N>中的<动作/运镜/风格/音效>,生成...
Video Edit (修改已有视频):
严格编辑<视频N>,将其中的<原特征>修改为<新特征>
增加元素:清晰描述<元素特征> + <出现时机> + <出现位置>
删除元素:点明需要删除的元素,并在提示词中强调保持不变的部分
⚠️ For edit/extend tasks: use
<视频N>directly — do NOT write参考<视频N>or the model will treat it as reference-only.
Video Extension (延长视频):
向前/向后延长<视频N>,生成...
Track Completion (轨道补全/多片段串联):
<视频1> + <过渡画面描述> + 接 <视频2> + <过渡画面描述> + 接 <视频3>
Combined tasks:
参考<图片/视频N>的[参考维度],严格编辑<视频X>,[具体编辑内容]
Step 4: Multi-Shot Storyboard Structure
Break complex videos into numbered shots. The model's internal modeling decouples space and time, so explicit shot ordering helps.
Structure per shot:
- Camera move / transition type (全景缓慢推近 / 固定机位 / 镜头切至...)
- Subject action and expression (core character's key motion + emotional state)
- Location / spatial relationship
- Audio (SFX, voice, BGM for this shot)
Example:
镜头1:街巷侧拍,男人缓慢起跑,带有急促的呼吸感。
镜头2:男人撞翻水果摊,镜头快速摇动并给到男人惊恐的特写。
镜头3:男人翻过矮墙消失,镜头缓慢拉远定格在空荡的街道。
Note: Precise time segments (e.g. 0–3s) are not recommended for Seedance 2.0 — the model's timing support is unstable. Use shot labels instead and let the model pace naturally.
Step 5: Action Description
Motion Guidelines
- Specify body parts: hand, leg, head, shoulders, back
- Add intensity/speed: 缓慢抬手 / 快速转头 / 用力蹬地 / 微微低头
- Prefer smooth, continuous micro-motions over high-intensity actions (running full speed, big jumps, violent rolls)
- Describe motion continuity: 借着转身惯性顺势抬手 / 从停顿状态自然过渡到举手
Emotion → Physical Expression
Avoid abstract emotion words. Externalize emotion as body language:
| Abstract | Physical expression |
|---|---|
| 悲伤 | 低头、肩膀微微颤抖、眼眶泛红、手指无意识地攥紧衣角 |
| 喜悦 | 嘴角抑制不住地上扬、脚步变得轻快、下意识地哼起小曲 |
| 紧张/焦虑 | 频繁看手表、手指不停敲击桌面、呼吸急促、眼神闪躲 |
| 愤怒 | 双拳紧握、下颌线紧绷、胸口剧烈起伏、从牙缝里挤出话语 |
| 释然 | 长长舒了口气、紧绷肩膀完全放松、抬头望向远方 |
Step 6: Apply Aesthetic Controls
Consult references/aesthetics.md for the complete keyword dictionary. Key categories:
- Lighting: natural light, artificial light, special/effect light, light quality
- Shot size: extreme close-up → close-up → medium close-up → medium → medium full → full → extreme wide
- Camera angle: eye level, high angle, low angle, bird's eye, worm's eye, over-the-shoulder, Dutch angle
- Lens: telephoto, wide angle, fisheye, macro, anamorphic, tilt-shift
- Camera movement: push in, pull back, pan, tilt, dolly, orbit, crane, tracking, handheld, Steadicam
- Style: cinematic, photorealistic, anime, cyberpunk, vintage film, minimalist, stop-motion, clay, felt, pixel art
Step 7: Quality, Style & Constraint Tags
Always close your prompt with these three types of tags:
Quality Tags
高清,细节丰富,电影质感,色彩自然,光影柔和
Style Tags
赛博朋克冷蓝紫色调 / 复古胶片 / 日系清新 / 写实纪录片风
Constraint Tags (Critical — reduces artifacts)
| Issue to prevent | Tag |
|---|---|
| Unwanted subtitles | 保持无字幕 / 避免生成任何文字或字幕 |
| Logo / watermark | 不要生成Logo / 不要生成水印 |
| Duplicate characters | 视频全程禁止出现外形、着装、配饰完全一致的人物,禁止生成同款分身 |
| Style drift | 明确写出目标风格,如"2D日漫风格"、"3D国风漫画" |
Step 8: Special Features — Text Generation in Video
Seedance 2.0 can render text (slogans, subtitles, speech bubbles) in video.
Text generation formula:
「文字内容」+「出现时机」+「出现位置」+「出现方式」,「文字特征(颜色、风格)」
Symbol conventions for special content:
| Content type | Symbol | Example |
|---|---|---|
| 背景音乐 | () | (背景中播放着快节奏的摇滚乐) |
| 音效 | <> | <远处传来狗叫声> |
| 台词 | {} | {你好,世界} |
| 字幕 | 【】 | 【第一章:启程】 |
Text rendering tips:
- Use common characters; avoid rare characters (生僻字) and special symbols
- Supported: advertising slogans, subtitles (字幕), speech bubbles (气泡台词)
- For logo precision: upload the logo as a reference image
Step 9: Review Checklist
Before delivering the prompt, verify:
- Subject is concretely described (appearance, clothing, features)
- For multimodal: subject defined explicitly with
<图片N>binding - Scene/environment is specified
- There is exactly one primary action per shot
- Camera shot type is specified (close-up, medium, wide, etc.)
- Camera movement is specified (or explicitly "固定镜头 / static camera")
- Lighting and style are specified
- Constraint tags added (no subtitles, no watermark, etc.)
- For I2V: prompt does NOT redundantly describe what the image already shows
- For multi-shot: 镜头1/2/3 labels used (not forced timestamps)
- For edit/extend: no
参考prefix before the video reference - For sound-synced: sound explicitly described or silenced
- Prompt is under 300 words per shot
- Assets ≤ recommended count (4–5 total, not maxed out)
Output Format
When generating prompts, output in this structure:
## 模式 / Mode: [T2V / I2V / 首尾帧 / Multi-shot / Multimodal-Reference / Video-Edit / Video-Extension / Sound-synced]
## 平台 / Platform: [Any / Sora / Runway / Kling / Seedance / Wan / etc.]
## 宽高比 / Aspect ratio: [16:9 / 9:16 / 1:1 / 21:9 / 4:3]
## 时长 / Duration: [5s / 10s / 15s / etc.]
## 分辨率 / Resolution: [480p / 720p / 1080p / 4K]
### 提示词 / Prompt
[The actual prompt text here]
### Prompt (English)
[English translation — bilingual output for cross-platform portability]
If the user does not specify a platform, write in a generic style that works across all major platforms.
Tips for Common Problems
| Problem | Cause | Fix |
|---|---|---|
| Output deviates from intent | Missing elements | Add shot type + camera movement |
| Incoherent motion | No transition logic | Add "一镜到底" or explicit shot transitions |
| Character deformation | Prompt too complex | Split into multiple shorter clips |
| Physics violations | No constraints | Add "遵循自然重力" / "realistic fluid dynamics" |
| Inconsistent character | Vague subject | Add detailed features + "全程保持人物外形一致" |
| ID drift / face swap | Weak face reference | Add separate headshot (大头照) + full body; put face ref first |
| Duplicate characters | Vague multi-char definition | Name each character, add twin-prevention constraint tag |
| Style drift to realism | Style not specified | Add explicit style tag: "2D日漫风格" / "3D国风漫画" |
| Unwanted subtitles | Model auto-generates | Add "保持无字幕" constraint tag |
| End-of-video audio noise | Abrupt audio cutoff | Post-process with audio fade-out (剪映音量包络线) |
| Wrong Chinese pronunciation | Rare characters | Replace with phonetically identical common characters |
Language
- Write prompts in the same language the user uses.
- Always offer a bilingual version (Chinese + English) since different platforms prefer different languages.
- For Chinese: use
镜头推进notpush in; for English: usepush innot镜头推进. - For Chinese dialogue in prompts: keep all dialogue in one language — avoid mixing Chinese and English within a single character's lines (proper nouns excepted).
Platform Notes
| Platform | Model IDs | Key Tips |
|---|---|---|
| Seedance 2.0 | doubao-seedance-2-0-260128 |
Best quality; 4–15s; supports 4K, multimodal ref, edit, extend |
| Seedance 2.0 Fast | doubao-seedance-2-0-fast-260128 |
Faster + cheaper; 480p/720p only |
| Seedance 2.0 Mini | doubao-seedance-2-0-mini-260615 |
Lowest cost; 480p/720p |
| Sora | — | Narrative style, 100–300 words, cinematic / film grain |
| Runway | — | Detailed lighting, supports keyframe control |
| Kling / 可灵 | — | Strong with Chinese prompts and human close-ups |
| Wan | — | Sound-synced; write "Generate single shot." for single shot |
| Pika | — | Short, punchy, one clear action |
| Vidu | — | Balanced quality/speed, standard formula |
Video Extension Strategy
For videos longer than 15 seconds, chain multiple segments:
Option A — Video Extension (连续长镜头):
- Best for: dialogue scenes, emotion progression, single-location continuous takes
- Use
向后延长<视频N>for each extension - Fix seam artifacts: trim last 6 frames of prior clip + first 1 frame of next clip
Option B — Segment + Edit (分段拼接):
- Best for: action sequences, scene changes, montage, complex motion
- Generate each shot separately with consistent style/lighting tags
- Combine in editing software
Option C — Track Completion (轨道补全):
- Chain up to 3 video clips with transition descriptions:
<视频1> + <过渡画面描述> + 接 <视频2> + <过渡画面描述> + 接 <视频3>
Asset Configuration Strategy
Categorize your reference materials by function:
| Role | Asset type | Purpose |
|---|---|---|
| 角色锚定 | Character headshot (大头照) + full body | Lock character identity |
| 场景定调 | Scene image | Set environment and style |
| 运镜参考 | Reference video | Lock camera language and motion rhythm |
| 节奏氛围 | Audio clip | Control emotion, timbre, BGM |
Recommended: 4–5 assets total. Do not max out the input limit — too many assets cause the model to lose feature priority.
Examples
See references/examples.md for a library of examples across modes, styles, and platforms.
Optimization
See references/prompt-optimization.md for iterative refinement strategies and known failure modes.
Subagent 协作(可选)
若环境已注册 prompt-engineer 子 agent(插件 agents/ 目录自动加载),可委派逐镜提示词渲染:
Agent(subagent_type: "prompt-engineer", prompt: "把这份分镜表渲染成 Seedance 2.0 提示词:[shot list]")
降级规则:Agent 不可用时,主线程直接按本文件执行。
完成之后
- 提示词写好后,可用
video-review做多视角审查(结构/镜头语言/视觉/一致性),按报告返回本 skill 或video-storyboard修改。 - 若提示词是从
video-storyboard的分镜表来的,审查时对照分镜表的+/-、><、屏幕方向、视觉强度是否被如实翻译。
安装与整体架构见仓库根
README.md。