# Video Prompt

> Translate a shot / shot list into platform-ready prompts for AI video models (Sora, Runway, Kling, Seedance, Vidu, Pika, Wan). Use when the user has a scene, action, or storyboard and needs the actual prompt text — covering the universal prompt formula, Seedance 2.0 multimodal/edit/extend syntax, per-platform tips, aesthetic keywords, and iterative optimization. If the user only has a rough idea and needs it structured into shots first, use video-storyboard before this. Audit with video-review.

- Skill: `wanna-money/video-prompt` (Agent Skill, multi-file: 5 files)
- Install (CLI): `npx skillmds@latest add wanna-money/video-prompt`
- Raw SKILL.md: https://api.skillmd.com/api/skills/wanna-money/video-prompt/raw
- Safety review: pending
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: AI & ML
- License: MIT
- Author: wanna-money (https://skillmd.com/u/wanna-money)
- Updated: 2026-09-17
- Page: https://skillmd.com/skills/wanna-money/video-prompt

---


# Video Prompt

The **prompt layer** of the video-prompt-craft toolkit: turn a shot (or a full shot list from `video-storyboard`) into
structured, platform-ready prompt text. This skill owns the platform syntax, the universal formula, and optimization —
it does NOT own narrative structure or shot design (that is `video-storyboard`'s job).

> **三件套协作** / Three coordinated skills:
> - `video-storyboard` — idea → narrative shot list (story structure, shot language, visual intensity)
> - `video-prompt` (this skill) — shot list → platform prompts
> - `video-review` — multi-angle audit of storyboard & prompts
>
> If a user hands you a rough idea with no shot plan and it needs multiple shots, suggest running `video-storyboard` first,
> then return here to render each shot. For a single shot or a direct prompt request, proceed directly.

## Core Principle

> You are not describing the entire world — you are describing **one shot through a camera lens**.
> For Seedance 2.0: think like an **AI director** — spatial layer (what's in frame) + temporal layer (how it changes over time).

Every effective video prompt answers five questions:

| Question | Element |
|----------|---------|
| What am I looking at? | **Subject + Environment** |
| What is happening? | **One primary action** |
| How is it filmed? | **Shot type + Camera movement** |
| What does it feel like? | **Style + Lighting + Color tone** |
| What is the rhythm? | **Pacing / Cut vs. long take** |

---

## Step 1: Classify the Task

Determine the generation mode before writing the prompt:

| Mode | Input | Prompt Focus |
|------|-------|-------------|
| **Text-to-Video (T2V)** | Text only | Full description: subject + scene + motion + camera + style |
| **Image-to-Video (I2V)** | Image + text | Motion + camera only (subject/scene/style already in image) |
| **First+Last Frame** | 2 images + text | Bridge the two frames; describe the transition motion |
| **Multi-shot / Storyboard** | Text | Overall narrative + per-shot descriptions (镜头1/2/3) |
| **Multimodal Reference** | Images/videos/audio + text | Subject anchoring + @ reference syntax (see below) |
| **Video Edit** | Video + text | Describe what to add/remove/change; keep everything else implicit |
| **Video Extension** | Video + text | Describe what happens before/after; do NOT use "参考" prefix |
| **Sound-synced** | Text | Standard elements + explicit sound description |

---

## Step 2: Construct the Prompt

### Basic Formula (Universal)

```
[Subject description] + [Scene/Environment] + [Primary action] + [Shot type] + [Camera movement] + [Lighting/Style] + [Quality/Constraint tags]
```

### Advanced Formula (Seedance 2.0)

```
精准主体 + 动作细节 + 场景环境 + 光影色调 + 镜头运镜 + 视觉风格 + 画质 + 约束条件
```

This maps to two internal model dimensions:
- **Spatial layer**: who/what, where, appearance, environment
- **Temporal layer**: action sequence, camera movement, timing, transitions

### Rules

1. **One shot, one action** — keep the primary motion singular and clear.
2. **Be concrete, not abstract** — "a woman in a white sundress" beats "a beautiful person"; "golden hour warm sunlight" beats "nice lighting".
3. **Order matters** — Subject first, then scene, then motion, then camera, then aesthetics.
4. **For I2V: skip what the image already shows** — only describe motion and camera work.
5. **For multi-shot: use 镜头1/镜头2/镜头3 labels** — do NOT force precise timestamps; let the model pace naturally.
6. **For sound: be explicit about silence** — write "无台词" / "No dialogue." and "无背景音乐" / "No background music." if needed.
7. **Keep it under 300 words per shot** — most models perform best with 50–200 words per shot.
8. **One camera move per shot** — mixing push+orbit+pan in one shot increases instability.

---

## Step 3: Seedance 2.0 — Subject Definition & Reference System

This section is critical for Seedance 2.0 multimodal reference tasks.

### 3a. Subject Definition

When referencing specific subjects in uploaded assets, always define them explicitly.

**Syntax:**
```
将<图片/视频N>中的[2-3个清晰静态特征]定义为<主体N>
```

**Example:**
```
将图片1中穿红色连衣裙、戴草帽的女人定义为主体1
```

**Rules:**
- Use 2–3 stable static features (clothing, hairstyle, appearance, type) for identification
- For multi-character scenes: name each character and link to their reference image
  ```
  将视频1中的高个子男人定义为警察，矮个子男人定义为小偷
  ```
- For unnamed subjects: use `主体1@图片1` inline every time the subject is referenced
- Do NOT use Asset ID as a subject reference — always use `<图片/视频N>`
- Most important references go **first** in the prompt

### 3b. Reference Syntax (Multimodal Tasks)

| Reference type | Syntax | Purpose |
|---------------|--------|---------|
| Image subject | `参考<图片N>中的<主体>` | Character identity, visual style |
| Video action | `参考<视频N>中的动作/运镜/风格/音效` | Motion, camera, effects, audio |
| Audio timbre | `参考<音频N>中的音色` | Voice tone replication |
| First frame | `以<图片N>为首帧` | Anchor opening frame |
| Last frame | `以<图片N>为尾帧` | Anchor closing frame |

**Input limits (Seedance 2.0):**
- Images: 0–9
- Videos: 0–3
- Audio: 0–3
- Recommended configuration: 1–2 character photos + 1 scene photo + 1 camera reference video + 1 audio = 4–5 assets total
- **Do not max out inputs** — too many assets confuse the model's priority weighting

### 3c. Task Type Sentence Patterns

**Multimodal Reference (生成新视频):**
```
参考<图片/视频N>中的<动作/运镜/风格/音效>，生成...
```

**Video Edit (修改已有视频):**
```
严格编辑<视频N>，将其中的<原特征>修改为<新特征>
增加元素：清晰描述<元素特征> + <出现时机> + <出现位置>
删除元素：点明需要删除的元素，并在提示词中强调保持不变的部分
```
> ⚠️ For edit/extend tasks: use `<视频N>` directly — do NOT write `参考<视频N>` or the model will treat it as reference-only.

**Video Extension (延长视频):**
```
向前/向后延长<视频N>，生成...
```

**Track Completion (轨道补全/多片段串联):**
```
<视频1> + <过渡画面描述> + 接 <视频2> + <过渡画面描述> + 接 <视频3>
```

**Combined tasks:**
```
参考<图片/视频N>的[参考维度]，严格编辑<视频X>，[具体编辑内容]
```

---

## Step 4: Multi-Shot Storyboard Structure

Break complex videos into numbered shots. The model's internal modeling decouples space and time, so explicit shot ordering helps.

**Structure per shot:**
1. Camera move / transition type (全景缓慢推近 / 固定机位 / 镜头切至...)
2. Subject action and expression (core character's key motion + emotional state)
3. Location / spatial relationship
4. Audio (SFX, voice, BGM for this shot)

**Example:**
```
镜头1：街巷侧拍，男人缓慢起跑，带有急促的呼吸感。
镜头2：男人撞翻水果摊，镜头快速摇动并给到男人惊恐的特写。
镜头3：男人翻过矮墙消失，镜头缓慢拉远定格在空荡的街道。
```

> Note: Precise time segments (e.g. 0–3s) are **not recommended** for Seedance 2.0 — the model's timing support is unstable. Use shot labels instead and let the model pace naturally.

---

## Step 5: Action Description

### Motion Guidelines

- **Specify body parts**: hand, leg, head, shoulders, back
- **Add intensity/speed**: 缓慢抬手 / 快速转头 / 用力蹬地 / 微微低头
- **Prefer smooth, continuous micro-motions** over high-intensity actions (running full speed, big jumps, violent rolls)
- **Describe motion continuity**: 借着转身惯性顺势抬手 / 从停顿状态自然过渡到举手

### Emotion → Physical Expression

Avoid abstract emotion words. Externalize emotion as body language:

| Abstract | Physical expression |
|----------|---------------------|
| 悲伤 | 低头、肩膀微微颤抖、眼眶泛红、手指无意识地攥紧衣角 |
| 喜悦 | 嘴角抑制不住地上扬、脚步变得轻快、下意识地哼起小曲 |
| 紧张/焦虑 | 频繁看手表、手指不停敲击桌面、呼吸急促、眼神闪躲 |
| 愤怒 | 双拳紧握、下颌线紧绷、胸口剧烈起伏、从牙缝里挤出话语 |
| 释然 | 长长舒了口气、紧绷肩膀完全放松、抬头望向远方 |

---

## Step 6: Apply Aesthetic Controls

Consult [references/aesthetics.md](references/aesthetics.md) for the complete keyword dictionary. Key categories:

- **Lighting**: natural light, artificial light, special/effect light, light quality
- **Shot size**: extreme close-up → close-up → medium close-up → medium → medium full → full → extreme wide
- **Camera angle**: eye level, high angle, low angle, bird's eye, worm's eye, over-the-shoulder, Dutch angle
- **Lens**: telephoto, wide angle, fisheye, macro, anamorphic, tilt-shift
- **Camera movement**: push in, pull back, pan, tilt, dolly, orbit, crane, tracking, handheld, Steadicam
- **Style**: cinematic, photorealistic, anime, cyberpunk, vintage film, minimalist, stop-motion, clay, felt, pixel art

---

## Step 7: Quality, Style & Constraint Tags

Always close your prompt with these three types of tags:

### Quality Tags
```
高清，细节丰富，电影质感，色彩自然，光影柔和
```

### Style Tags
```
赛博朋克冷蓝紫色调 / 复古胶片 / 日系清新 / 写实纪录片风
```

### Constraint Tags (Critical — reduces artifacts)

| Issue to prevent | Tag |
|------------------|-----|
| Unwanted subtitles | 保持无字幕 / 避免生成任何文字或字幕 |
| Logo / watermark | 不要生成Logo / 不要生成水印 |
| Duplicate characters | 视频全程禁止出现外形、着装、配饰完全一致的人物，禁止生成同款分身 |
| Style drift | 明确写出目标风格，如"2D日漫风格"、"3D国风漫画" |

---

## Step 8: Special Features — Text Generation in Video

Seedance 2.0 can render text (slogans, subtitles, speech bubbles) in video.

**Text generation formula:**
```
「文字内容」+「出现时机」+「出现位置」+「出现方式」，「文字特征（颜色、风格）」
```

**Symbol conventions for special content:**

| Content type | Symbol | Example |
|-------------|--------|---------|
| 背景音乐 | （）| （背景中播放着快节奏的摇滚乐）|
| 音效 | <> | <远处传来狗叫声> |
| 台词 | {} | {你好，世界} |
| 字幕 | 【】 | 【第一章：启程】|

**Text rendering tips:**
- Use common characters; avoid rare characters (生僻字) and special symbols
- Supported: advertising slogans, subtitles (字幕), speech bubbles (气泡台词)
- For logo precision: upload the logo as a reference image

---

## Step 9: Review Checklist

Before delivering the prompt, verify:

- [ ] Subject is concretely described (appearance, clothing, features)
- [ ] For multimodal: subject defined explicitly with `<图片N>` binding
- [ ] Scene/environment is specified
- [ ] There is exactly one primary action per shot
- [ ] Camera shot type is specified (close-up, medium, wide, etc.)
- [ ] Camera movement is specified (or explicitly "固定镜头 / static camera")
- [ ] Lighting and style are specified
- [ ] Constraint tags added (no subtitles, no watermark, etc.)
- [ ] For I2V: prompt does NOT redundantly describe what the image already shows
- [ ] For multi-shot: 镜头1/2/3 labels used (not forced timestamps)
- [ ] For edit/extend: no `参考` prefix before the video reference
- [ ] For sound-synced: sound explicitly described or silenced
- [ ] Prompt is under 300 words per shot
- [ ] Assets ≤ recommended count (4–5 total, not maxed out)

---

## Output Format

When generating prompts, output in this structure:

```
## 模式 / Mode: [T2V / I2V / 首尾帧 / Multi-shot / Multimodal-Reference / Video-Edit / Video-Extension / Sound-synced]
## 平台 / Platform: [Any / Sora / Runway / Kling / Seedance / Wan / etc.]
## 宽高比 / Aspect ratio: [16:9 / 9:16 / 1:1 / 21:9 / 4:3]
## 时长 / Duration: [5s / 10s / 15s / etc.]
## 分辨率 / Resolution: [480p / 720p / 1080p / 4K]

### 提示词 / Prompt

[The actual prompt text here]

### Prompt (English)

[English translation — bilingual output for cross-platform portability]
```

If the user does not specify a platform, write in a generic style that works across all major platforms.

---

## Tips for Common Problems

| Problem | Cause | Fix |
|---------|-------|-----|
| Output deviates from intent | Missing elements | Add shot type + camera movement |
| Incoherent motion | No transition logic | Add "一镜到底" or explicit shot transitions |
| Character deformation | Prompt too complex | Split into multiple shorter clips |
| Physics violations | No constraints | Add "遵循自然重力" / "realistic fluid dynamics" |
| Inconsistent character | Vague subject | Add detailed features + "全程保持人物外形一致" |
| ID drift / face swap | Weak face reference | Add separate headshot (大头照) + full body; put face ref first |
| Duplicate characters | Vague multi-char definition | Name each character, add twin-prevention constraint tag |
| Style drift to realism | Style not specified | Add explicit style tag: "2D日漫风格" / "3D国风漫画" |
| Unwanted subtitles | Model auto-generates | Add "保持无字幕" constraint tag |
| End-of-video audio noise | Abrupt audio cutoff | Post-process with audio fade-out (剪映音量包络线) |
| Wrong Chinese pronunciation | Rare characters | Replace with phonetically identical common characters |

---

## Language

- Write prompts in the same language the user uses.
- Always offer a bilingual version (Chinese + English) since different platforms prefer different languages.
- For Chinese: use `镜头推进` not `push in`; for English: use `push in` not `镜头推进`.
- For Chinese dialogue in prompts: keep all dialogue in one language — avoid mixing Chinese and English within a single character's lines (proper nouns excepted).

---

## Platform Notes

| Platform | Model IDs | Key Tips |
|----------|-----------|---------|
| **Seedance 2.0** | `doubao-seedance-2-0-260128` | Best quality; 4–15s; supports 4K, multimodal ref, edit, extend |
| **Seedance 2.0 Fast** | `doubao-seedance-2-0-fast-260128` | Faster + cheaper; 480p/720p only |
| **Seedance 2.0 Mini** | `doubao-seedance-2-0-mini-260615` | Lowest cost; 480p/720p |
| **Sora** | — | Narrative style, 100–300 words, `cinematic` / `film grain` |
| **Runway** | — | Detailed lighting, supports keyframe control |
| **Kling / 可灵** | — | Strong with Chinese prompts and human close-ups |
| **Wan** | — | Sound-synced; write "Generate single shot." for single shot |
| **Pika** | — | Short, punchy, one clear action |
| **Vidu** | — | Balanced quality/speed, standard formula |

---

## Video Extension Strategy

For videos longer than 15 seconds, chain multiple segments:

**Option A — Video Extension (连续长镜头):**
- Best for: dialogue scenes, emotion progression, single-location continuous takes
- Use `向后延长<视频N>` for each extension
- Fix seam artifacts: trim last 6 frames of prior clip + first 1 frame of next clip

**Option B — Segment + Edit (分段拼接):**
- Best for: action sequences, scene changes, montage, complex motion
- Generate each shot separately with consistent style/lighting tags
- Combine in editing software

**Option C — Track Completion (轨道补全):**
- Chain up to 3 video clips with transition descriptions:
  ```
  <视频1> + <过渡画面描述> + 接 <视频2> + <过渡画面描述> + 接 <视频3>
  ```

---

## Asset Configuration Strategy

Categorize your reference materials by function:

| Role | Asset type | Purpose |
|------|-----------|---------|
| 角色锚定 | Character headshot (大头照) + full body | Lock character identity |
| 场景定调 | Scene image | Set environment and style |
| 运镜参考 | Reference video | Lock camera language and motion rhythm |
| 节奏氛围 | Audio clip | Control emotion, timbre, BGM |

**Recommended:** 4–5 assets total. Do not max out the input limit — too many assets cause the model to lose feature priority.

---

## Examples

See [references/examples.md](references/examples.md) for a library of examples across modes, styles, and platforms.

## Optimization

See [references/prompt-optimization.md](references/prompt-optimization.md) for iterative refinement strategies and known failure modes.

## Subagent 协作(可选)

若环境已注册 `prompt-engineer` 子 agent(插件 `agents/` 目录自动加载),可委派逐镜提示词渲染:

```
Agent(subagent_type: "prompt-engineer", prompt: "把这份分镜表渲染成 Seedance 2.0 提示词:[shot list]")
```

**降级规则**:Agent 不可用时,主线程直接按本文件执行。

---

## 完成之后

- 提示词写好后,可用 `video-review` 做多视角审查(结构/镜头语言/视觉/一致性),按报告返回本 skill 或 `video-storyboard` 修改。
- 若提示词是从 `video-storyboard` 的分镜表来的,审查时对照分镜表的 `+/-`、`><`、屏幕方向、视觉强度是否被如实翻译。

> 安装与整体架构见仓库根 `README.md`。

