Video Generation
Workflow Overview
- Phase 1: Initial → Gather requirements, STOP for user confirmation
- Phase 2: Global Definitions → Define style, characters, voices, BGM (text only, no images)
- Phase 3: Clip Planning → Segment into clips, plan each clip, determine reference image needs
- Phase 4: Reference Images → Generate reference images (MANDATORY before Phase 5)
- Phase 5: Execution → Generate keyframes, videos, audio
Critical Rules (MUST Follow)
Before starting, memorize these non-negotiable rules:
[PHASE 1 STOP] MUST ask questions to gather information. DO NOT assume or guess missing details—always ask the user. Never proceed without explicit user confirmation.
[DETAILED VIDEO PROMPT] Video prompts must include detailed transition_description (2-4 sentences). One-line prompts are insufficient.
[KEYFRAME DIFFERENCE] Last keyframe must show interpolatable change from first keyframe: subject position/pose, subject state (open/close, appear/disappear), or composition change. Subtle-only changes (lighting, background) while subject stays static cause unnatural video motion.
[PHASE 4 MANDATORY] MUST generate reference images before keyframes. Never skip Phase 4.
[ASPECT RATIO] ALL keyframes must use 16:9 or 9:16, and must be upright (not rotated). Never generate 1:1 or other ratios.
[NO TTS FOR ON-SCREEN] Never use TTS for on-screen dialogue or singing. Video model generates audio with lip sync.
[NARRATION CLIP BY CLIP] Generate off-screen narration separately for each clip, not all at once.
[AUDIO MIXING] When combining audio tracks (video audio, narration, BGM), preserve ALL tracks—overlay, never replace. Narration must be clearly audible and maintain consistent volume across all clips.
Image Generation Tools
| Tool |
Use When |
generate_image |
Create new images (with or without references) |
generate_image_variation |
Edit existing images |
Phase 1: Initial
Gather Information
| Field |
Description |
| Purpose |
Goal and target audience |
| Narrative arc |
Story structure and key points |
| Duration |
Total length in seconds |
| Aspect ratio |
16:9 or 9:16 only |
| Visual style |
Sub-genre aesthetic (e.g., "Makoto Shinkai anime", "Pixar 3D") |
| Reference materials |
Reference videos, images, brand guidelines |
| Language |
For dialogue and narration |
| Recurring elements |
Characters/objects with appearance descriptions |
| Dialogue/singing needs |
On-screen character audio |
| Narration needs |
Off-screen narrator (gender, tone, pace) |
Five-Dimension Expert Framework
Use these perspectives to guide your questions:
| Dimension |
Expert Role |
Key Questions |
| Strategy & Audience |
Creative Director |
Who is this for? What's the goal? What action should viewers take? |
| Narrative & Structure |
Screenwriter |
What's the story? Key moments? Emotional arc? |
| Visual Style |
Director + Art Director |
What look and feel? Reference videos/images? Color mood? |
| Shot Execution |
Cinematographer |
Any specific shots in mind? Product hero shots needed? |
| Sound Design |
Sound Designer |
Voiceover? Music mood? Dialogue? Sound effects? |
Ask questions across all dimensions. Prioritize based on user's initial description.
[MANDATORY STOP - DO NOT PROCEED WITHOUT USER CONFIRMATION]
Summarize gathered information and wait for user confirmation before Phase 2.
Phase 2: Global Definitions (Text Only)
Visual Style Specification
Define these 4 dimensions (applied to primary reference images in Phase 4):
| Dimension |
Example Values |
| Sub-genre |
Makoto Shinkai anime, Pixar 3D, cyberpunk noir |
| Rendering + Line |
2D hand-drawn with thick outlines, 3D cel-shading |
| Color + Lighting |
High saturation neon, soft diffused natural light |
| Detail density |
Minimalist, highly detailed backgrounds |
Example specification:
Sub-genre: Cyberpunk anime
Rendering + Line: 2D digital painting, thin glowing outlines
Color + Lighting: High saturation neon (pink, cyan, purple), dark backgrounds, rim lighting
Detail density: Highly detailed backgrounds, moderate character detail
Recurring Elements
For each character/object:
| Field |
Description |
| unique_identifier |
Name for reference |
| appearance |
Text description for prompts |
| outfit_description |
Clothing/accessories (characters) |
| language |
Spoken/sung language (if applicable) |
| mechanical_properties |
Physical behavior (if applicable) |
Voice Profiles
- On-screen: From character definitions (dialogue/singing)
- Off-screen narrator: name, gender, tone, pace, language
BGM Source Decision
| Scenario |
BGM Source |
| Music video / diegetic music (visible source) |
Embedded (in video prompt) |
| Background mood music |
Separate (Phase 5 BGM Preparation) |
| No music |
None |
If Separate, define: genre, instruments, tempo
Phase 3: Clip Planning
Segmentation Rules
- Clips: 4, 6, or 8 seconds only
- Each clip: one action, one scene
Per-Clip Specification
| Field |
Values |
| narrative_purpose |
establish / develop / climax / resolve / transition / supplementary (product shot, detail, reaction, insert, B-roll, POV) |
| pacing |
slow / moderate / fast |
| scene |
Environment description |
| content_action |
Subject + action + trajectory |
| transition_description |
[REQUIRED] Detailed transition process. Must include: subject appearance, movement trajectory, state changes, existence statements. 2-4 sentences minimum. |
| duration |
4 / 6 / 8 |
| camera_movement |
static / pan / tilt / dolly / zoom / crane / arc / handheld |
| first_keyframe_framing |
Shot size + angle + composition |
| first_keyframe_visible_content |
What's visible |
| last_keyframe_framing |
Shot size + angle + composition |
| last_keyframe_visible_content |
What's visible |
| last_keyframe_edit_from_first |
yes / no (see decision table below) |
| inter_clip_boundary |
continuous / scene_cut |
| first_keyframe_reuse |
yes / no |
| last_keyframe_required |
yes / no |
| on_screen_dialogue |
"Name: text" or "Name: [lyrics] (style)" or None |
| sound_effects |
Sources or None |
| bgm_source |
embedded / separate / none |
| bgm_cue |
If embedded: style, BPM, instruments. If separate: emotion, intensity |
| narration_cue |
Narrator text or None |
Field Dependencies
inter_clip_boundary = continuous → next clip's first_keyframe_reuse = yes
first_keyframe_reuse = yes → previous clip must have last_keyframe_required = yes
Keyframe Difference Requirement
When planning last_keyframe_visible_content, ensure interpolatable change from first_keyframe_visible_content:
- Subject position/pose change (movement, rotation, action)
- Subject state change (open/close, appear/disappear, expression)
- Composition change from camera movement (zoom, pan result)
[WARNING] Avoid last keyframes with only lighting or background changes while subject remains static—this causes unnatural video motion.
Decision: last_keyframe_edit_from_first
| Camera Movement |
First & Last Keyframe Overlap? |
Set to |
| static, small pan/tilt, zoom |
Yes (same scene area) |
yes |
| large pan, dolly, tracking, crane, arc |
No (different area) |
no |
transition_description Requirements
This field directly becomes part of the video prompt. The more detailed, the better.
Must include:
- Subject appearance: Key visual features that must remain consistent throughout
- Movement trajectory: How subject/camera moves through space and time
- State changes: How objects/environment change over the duration
- Existence statements: What is present throughout (prevents pop-in/pop-out)
Length guideline: 2-4 sentences minimum. One-line descriptions are insufficient.
transition_description Examples
| Insufficient |
Sufficient |
| "Open box revealing jar" |
"The frosted glass jar with gold lid is inside the box from the start, hidden by the closed cream-colored lid. Elegant hands with manicured nails lift the lid upward smoothly. As the lid rises, the jar gradually comes into view - first the gold cap edge, then the full jar nestled in champagne velvet." |
| "Person walks left to right" |
"Woman in white dress with brown hair starts at left edge of frame, walks steadily rightward at moderate pace, maintaining upright posture, reaches right edge by end of clip." |
| "Light turns on" |
"Room starts in complete darkness. Light gradually increases from the ceiling fixture at center, warm yellow glow spreading outward across the wooden furniture until fully illuminated." |
Physical Consistency Check
| Movement |
Constraint |
| Pan/Tilt/Zoom |
Camera fixed, content within rotational/zoom range |
| Dolly/Tracking/Crane |
Content physically traversable within duration |
| Arc |
Subject centered in both keyframes, environment allows orbit |
| Handheld |
Similar to Dolly but allows irregularity |
| Combined |
Must satisfy ALL involved movement constraints |
Common Mistakes:
| Mistake |
Correction |
| "Pan from corridor entrance to middle" |
Use "dolly forward" |
| First: room A, Last: room B |
Split into two clips |
| 6-second clip covering 100 meters |
Extend duration or reduce distance |
[MANDATORY] Reference Image Requirements
After all clips planned, list required reference images:
| Element |
Clips Using It |
Required Images |
| (name) |
Clip X (MS), Clip Y (CU) |
Full body, Face close-up |
[WARNING] Only generate what clips actually need. Do NOT generate all angles by default.
[MANDATORY] Phase 4: Reference Image Generation
MANDATORY. Do not skip to Phase 5.
Generation Order
Step 1: Primary reference (visual anchor)
- Tool:
generate_image (no references)
- Prompt MUST include: Full Visual Style Specification from Phase 2 + element description
- White background
- Ends with "no text, no watermarks, no logos, no labels, no annotations"
Step 2: Additional angles/shots
- Tool:
generate_image with primary reference as reference
- Prompt: New angle/shot only (style inherited from reference)
- White background
- Ends with "no text, no watermarks, no logos, no labels, no annotations"
[WARNING] Never generate additional refs without using primary ref as reference.
Phase 5: Execution
Global Rules
[CRITICAL] ALL keyframes: aspect ratio from Phase 1 (16:9 or 9:16). Never 1:1.
First Keyframe
first_keyframe_reuse = yes → Use previous clip's last keyframe (no generation)
first_keyframe_reuse = no → Generate new keyframe
If generating first keyframe:
Last Keyframe
last_keyframe_required = no → Skip
last_keyframe_required = yes:
last_keyframe_edit_from_first = yes → Edit mode
last_keyframe_edit_from_first = no → Generate mode
If EDIT mode:
If GENERATE mode:
Consistency Checklist (Easily Overlooked)
When generating last keyframe, verify:
Video Generation
Video prompt should be detailed. Even with keyframes, video models may drift during generation.
Prompt includes:
Audio in prompt:
| Type |
Include |
| On-screen dialogue |
"Name says: text" with tone, language |
| On-screen singing |
"Name sings: [lyrics]" with style, language |
| Sound effects |
Source + quality |
| Embedded BGM |
Style, BPM, instruments, mood |
Prompt ending by bgm_source:
- embedded → (no ending, music described in prompt body)
- separate/none → End with "No background music."
Example (music video with embedded BGM):
Hatsune Miku center stage, singing in Japanese with sweet electronic voice:
"ラララ、光の中で踊り出す", energetic J-pop at 140 BPM with synthesizer,
crowd cheering, concert atmosphere
[CRITICAL] Never use TTS for on-screen dialogue/singing. Video model generates audio with lip sync.
BGM Sourcing (if bgm_source = separate)
Method: Search and download from royalty-free music libraries (e.g., Pixabay, YouTube Audio Library).
[CRITICAL] Generating music with Python or any other tools is strictly prohibited. You must only use pre-existing, royalty-free tracks.
Match the downloaded music to the style defined in Phase 2.
Narration Generation (if narration exists)
[WARNING] Generate clip by clip, not all at once.
- TTS for off-screen narrator only
- Same voice profile across all clips
- Verify audio duration fits clip duration
Audio Summary
| Type |
Method |
Output |
| On-screen dialogue/singing |
Video model |
Embedded |
| Sound effects |
Video model |
Embedded |
| Embedded BGM |
Video model |
Embedded |
| Separate BGM |
Search only |
Separate track |
| Narration |
TTS (clip by clip) |
Separate track |
Audio Mixing (Final Assembly)
When combining multiple audio sources:
| Track |
Source |
| Video audio |
Embedded in video clips (dialogue, sound effects, embedded BGM) |
| Narration |
TTS generated (off-screen narrator) |
| Separate BGM |
Searched from royalty-free source |
[CRITICAL] Mixing rules:
- Preserve ALL audio tracks—overlay, never replace one with another
- Narration must be clearly audible—not drowned out by other tracks
- Narration volume must be consistent across all clips
1---2name: video-generator3description: Professional AI video production workflow. Use when creating videos, short films, commercials, or any video content using AI generation tools.4---56# Video Generation78## Workflow Overview9101. **Phase 1: Initial** → Gather requirements, STOP for user confirmation112. **Phase 2: Global Definitions** → Define style, characters, voices, BGM (text only, no images)123. **Phase 3: Clip Planning** → Segment into clips, plan each clip, determine reference image needs134. **Phase 4: Reference Images** → Generate reference images (MANDATORY before Phase 5)145. **Phase 5: Execution** → Generate keyframes, videos, audio1516---1718## Critical Rules (MUST Follow)1920Before starting, memorize these non-negotiable rules:21221. **[PHASE 1 STOP]** MUST ask questions to gather information. DO NOT assume or guess missing details—always ask the user. Never proceed without explicit user confirmation.23242. **[DETAILED VIDEO PROMPT]** Video prompts must include detailed transition_description (2-4 sentences). One-line prompts are insufficient.25263. **[KEYFRAME DIFFERENCE]** Last keyframe must show interpolatable change from first keyframe: subject position/pose, subject state (open/close, appear/disappear), or composition change. Subtle-only changes (lighting, background) while subject stays static cause unnatural video motion.27284. **[PHASE 4 MANDATORY]** MUST generate reference images before keyframes. Never skip Phase 4.29305. **[ASPECT RATIO]** ALL keyframes must use 16:9 or 9:16, and must be upright (not rotated). Never generate 1:1 or other ratios.31326. **[NO TTS FOR ON-SCREEN]** Never use TTS for on-screen dialogue or singing. Video model generates audio with lip sync.33347. **[NARRATION CLIP BY CLIP]** Generate off-screen narration separately for each clip, not all at once.35368. **[AUDIO MIXING]** When combining audio tracks (video audio, narration, BGM), preserve ALL tracks—overlay, never replace. Narration must be clearly audible and maintain consistent volume across all clips.3738---3940## Image Generation Tools4142| Tool | Use When |43|------|----------|44| `generate_image` | Create new images (with or without references) |45| `generate_image_variation` | Edit existing images |4647---4849## Phase 1: Initial5051### Gather Information5253| Field | Description |54|-------|-------------|55| Purpose | Goal and target audience |56| Narrative arc | Story structure and key points |57| Duration | Total length in seconds |58| Aspect ratio | 16:9 or 9:16 only |59| Visual style | Sub-genre aesthetic (e.g., "Makoto Shinkai anime", "Pixar 3D") |60| Reference materials | Reference videos, images, brand guidelines |61| Language | For dialogue and narration |62| Recurring elements | Characters/objects with appearance descriptions |63| Dialogue/singing needs | On-screen character audio |64| Narration needs | Off-screen narrator (gender, tone, pace) |656667### Five-Dimension Expert Framework6869Use these perspectives to guide your questions:7071| Dimension | Expert Role | Key Questions |72|-----------|-------------|---------------|73| **Strategy & Audience** | Creative Director | Who is this for? What's the goal? What action should viewers take? |74| **Narrative & Structure** | Screenwriter | What's the story? Key moments? Emotional arc? |75| **Visual Style** | Director + Art Director | What look and feel? Reference videos/images? Color mood? |76| **Shot Execution** | Cinematographer | Any specific shots in mind? Product hero shots needed? |77| **Sound Design** | Sound Designer | Voiceover? Music mood? Dialogue? Sound effects? |7879Ask questions across all dimensions. Prioritize based on user's initial description.8081> **[MANDATORY STOP - DO NOT PROCEED WITHOUT USER CONFIRMATION]**82> Summarize gathered information and wait for user confirmation before Phase 2.8384---8586## Phase 2: Global Definitions (Text Only)8788### Visual Style Specification8990Define these 4 dimensions (applied to primary reference images in Phase 4):9192| Dimension | Example Values |93|-----------|----------------|94| **Sub-genre** | Makoto Shinkai anime, Pixar 3D, cyberpunk noir |95| **Rendering + Line** | 2D hand-drawn with thick outlines, 3D cel-shading |96| **Color + Lighting** | High saturation neon, soft diffused natural light |97| **Detail density** | Minimalist, highly detailed backgrounds |9899**Example specification:**100101```102Sub-genre: Cyberpunk anime103Rendering + Line: 2D digital painting, thin glowing outlines104Color + Lighting: High saturation neon (pink, cyan, purple), dark backgrounds, rim lighting105Detail density: Highly detailed backgrounds, moderate character detail106```107108### Recurring Elements109110For each character/object:111112| Field | Description |113|-------|-------------|114| unique_identifier | Name for reference |115| appearance | Text description for prompts |116| outfit_description | Clothing/accessories (characters) |117| language | Spoken/sung language (if applicable) |118| mechanical_properties | Physical behavior (if applicable) |119120### Voice Profiles121122- **On-screen**: From character definitions (dialogue/singing)123- **Off-screen narrator**: name, gender, tone, pace, language124125### BGM Source Decision126127| Scenario | BGM Source |128|----------|------------|129| Music video / diegetic music (visible source) | **Embedded** (in video prompt) |130| Background mood music | **Separate** (Phase 5 BGM Preparation) |131| No music | **None** |132133**If Separate**, define: genre, instruments, tempo134135---136137## Phase 3: Clip Planning138139### Segmentation Rules140141- Clips: **4, 6, or 8 seconds only**142- Each clip: **one action, one scene**143144### Per-Clip Specification145146| Field | Values |147|-------|--------|148| **narrative_purpose** | establish / develop / climax / resolve / transition / supplementary (product shot, detail, reaction, insert, B-roll, POV) |149| **pacing** | slow / moderate / fast |150| **scene** | Environment description |151| **content_action** | Subject + action + trajectory |152| **transition_description** | **[REQUIRED]** Detailed transition process. Must include: subject appearance, movement trajectory, state changes, existence statements. 2-4 sentences minimum. |153| **duration** | 4 / 6 / 8 |154| **camera_movement** | static / pan / tilt / dolly / zoom / crane / arc / handheld |155| **first_keyframe_framing** | Shot size + angle + composition |156| **first_keyframe_visible_content** | What's visible |157| **last_keyframe_framing** | Shot size + angle + composition |158| **last_keyframe_visible_content** | What's visible |159| **last_keyframe_edit_from_first** | yes / no (see decision table below) |160| **inter_clip_boundary** | continuous / scene_cut |161| **first_keyframe_reuse** | yes / no |162| **last_keyframe_required** | yes / no |163| **on_screen_dialogue** | "Name: text" or "Name: [lyrics] (style)" or None |164| **sound_effects** | Sources or None |165| **bgm_source** | embedded / separate / none |166| **bgm_cue** | If embedded: style, BPM, instruments. If separate: emotion, intensity |167| **narration_cue** | Narrator text or None |168169### Field Dependencies170171- `inter_clip_boundary = continuous` → next clip's `first_keyframe_reuse = yes`172- `first_keyframe_reuse = yes` → previous clip must have `last_keyframe_required = yes`173174### Keyframe Difference Requirement175176When planning `last_keyframe_visible_content`, ensure interpolatable change from `first_keyframe_visible_content`:177- Subject position/pose change (movement, rotation, action)178- Subject state change (open/close, appear/disappear, expression)179- Composition change from camera movement (zoom, pan result)180181> **[WARNING]** Avoid last keyframes with only lighting or background changes while subject remains static—this causes unnatural video motion.182183### Decision: last_keyframe_edit_from_first184185| Camera Movement | First & Last Keyframe Overlap? | Set to |186|-----------------|-------------------------------|--------|187| static, small pan/tilt, zoom | Yes (same scene area) | `yes` |188| large pan, dolly, tracking, crane, arc | No (different area) | `no` |189190### transition_description Requirements191192This field directly becomes part of the video prompt. **The more detailed, the better.**193194**Must include:**1951. **Subject appearance**: Key visual features that must remain consistent throughout1962. **Movement trajectory**: How subject/camera moves through space and time1973. **State changes**: How objects/environment change over the duration1984. **Existence statements**: What is present throughout (prevents pop-in/pop-out)199200**Length guideline:** 2-4 sentences minimum. One-line descriptions are insufficient.201202### transition_description Examples203204| Insufficient | Sufficient |205|--------------|------------|206| "Open box revealing jar" | "The frosted glass jar with gold lid is inside the box from the start, hidden by the closed cream-colored lid. Elegant hands with manicured nails lift the lid upward smoothly. As the lid rises, the jar gradually comes into view - first the gold cap edge, then the full jar nestled in champagne velvet." |207| "Person walks left to right" | "Woman in white dress with brown hair starts at left edge of frame, walks steadily rightward at moderate pace, maintaining upright posture, reaches right edge by end of clip." |208| "Light turns on" | "Room starts in complete darkness. Light gradually increases from the ceiling fixture at center, warm yellow glow spreading outward across the wooden furniture until fully illuminated." |209210### Physical Consistency Check211212| Movement | Constraint |213|----------|------------|214| Pan/Tilt/Zoom | Camera fixed, content within rotational/zoom range |215| Dolly/Tracking/Crane | Content physically traversable within duration |216| Arc | Subject centered in both keyframes, environment allows orbit |217| Handheld | Similar to Dolly but allows irregularity |218| Combined | Must satisfy ALL involved movement constraints |219220**Common Mistakes:**221222| Mistake | Correction |223|---------|------------|224| "Pan from corridor entrance to middle" | Use "dolly forward" |225| First: room A, Last: room B | Split into two clips |226| 6-second clip covering 100 meters | Extend duration or reduce distance |227228### [MANDATORY] Reference Image Requirements229230After all clips planned, list required reference images:231232| Element | Clips Using It | Required Images |233|---------|----------------|-----------------|234| (name) | Clip X (MS), Clip Y (CU) | Full body, Face close-up |235236> **[WARNING]** Only generate what clips actually need. Do NOT generate all angles by default.237238---239240## [MANDATORY] Phase 4: Reference Image Generation241242**MANDATORY. Do not skip to Phase 5.**243244### Generation Order245246**Step 1: Primary reference (visual anchor)**247- Tool: `generate_image` (no references)248- Prompt MUST include: **Full Visual Style Specification** from Phase 2 + element description249- White background250- Ends with "no text, no watermarks, no logos, no labels, no annotations"251252**Step 2: Additional angles/shots**253- Tool: `generate_image` with **primary reference as reference**254- Prompt: New angle/shot only (style inherited from reference)255- White background256- Ends with "no text, no watermarks, no logos, no labels, no annotations"257258> **[WARNING]** Never generate additional refs without using primary ref as reference.259260---261262## Phase 5: Execution263264### Global Rules265266> **[CRITICAL]** ALL keyframes: aspect ratio from Phase 1 (16:9 or 9:16). Never 1:1.267268### First Keyframe269270```271first_keyframe_reuse = yes → Use previous clip's last keyframe (no generation)272first_keyframe_reuse = no → Generate new keyframe273```274275**If generating first keyframe:**276- [ ] Tool: `generate_image`277- [ ] References: Appropriate Phase 4 images278- [ ] Aspect ratio: 16:9 or 9:16279- [ ] Prompt includes:280 - [ ] Visual style (sub-genre + key characteristics, brief)281 - [ ] Scene environment282 - [ ] Framing (shot size + angle + lens)283 - [ ] Visible content284 - [ ] Subject appearance + outfit285- [ ] Prompt ends with: "no text, no watermarks, no logos, no annotations"286287### Last Keyframe288289```290last_keyframe_required = no → Skip291last_keyframe_required = yes:292 last_keyframe_edit_from_first = yes → Edit mode293 last_keyframe_edit_from_first = no → Generate mode294```295296**If EDIT mode:**297- [ ] Tool: `generate_image_variation`298- [ ] References: [first_keyframe, Phase 4 refs...]299- [ ] Prompt: "Edit this image: [changes only]"300- [ ] Do NOT repeat unchanged elements301302**If GENERATE mode:**303- [ ] Tool: `generate_image`304- [ ] References: [first_keyframe (scene ref), Phase 4 refs...]305- [ ] Aspect ratio: 16:9 or 9:16306- [ ] Prompt includes:307 - [ ] Visual style (brief)308 - [ ] Last keyframe framing + visible content309 - [ ] Subject appearance and end state310 - [ ] "Same location/environment as reference"311- [ ] Prompt ends with: "no text, no watermarks, no logos, no annotations"312313### Consistency Checklist (Easily Overlooked)314315When generating last keyframe, verify:316- [ ] **Interpolatable change**: Clear difference in subject position/pose, state, or composition (not just lighting/background)317- [ ] Same lighting direction and shadows as first keyframe318- [ ] Same color temperature (warm/cool)319- [ ] Same depth of field320- [ ] Same outfit, facial features, body proportions321- [ ] Environment details consistent322323### Video Generation324325**Video prompt should be detailed.** Even with keyframes, video models may drift during generation.326327**Prompt includes:**328- [ ] Visual style (brief)329- [ ] Pacing (slow / moderate / fast)330- [ ] **transition_description** from Phase 3 (detailed, 2-4 sentences)331- [ ] **Subject appearance** (key features for consistency)332- [ ] **Scene environment** (brief)333- [ ] Audio (see below)334335**Audio in prompt:**336337| Type | Include |338|------|---------|339| On-screen dialogue | "Name says: text" with tone, language |340| On-screen singing | "Name sings: [lyrics]" with style, language |341| Sound effects | Source + quality |342| Embedded BGM | Style, BPM, instruments, mood |343344**Prompt ending by bgm_source:**345- embedded → (no ending, music described in prompt body)346- separate/none → End with "No background music."347348**Example (music video with embedded BGM):**349```350Hatsune Miku center stage, singing in Japanese with sweet electronic voice: 351"ラララ、光の中で踊り出す", energetic J-pop at 140 BPM with synthesizer, 352crowd cheering, concert atmosphere353```354355> **[CRITICAL]** Never use TTS for on-screen dialogue/singing. Video model generates audio with lip sync.356357### BGM Sourcing (if bgm_source = separate)358359**Method:** Search and download from royalty-free music libraries (e.g., Pixabay, YouTube Audio Library).360361**[CRITICAL]** Generating music with Python or any other tools is strictly prohibited. You must only use pre-existing, royalty-free tracks.362363Match the downloaded music to the style defined in Phase 2.364365### Narration Generation (if narration exists)366367> **[WARNING]** Generate **clip by clip**, not all at once.368369- TTS for off-screen narrator only370- Same voice profile across all clips371- Verify audio duration fits clip duration372373### Audio Summary374375| Type | Method | Output |376|------|--------|--------|377| On-screen dialogue/singing | Video model | Embedded |378| Sound effects | Video model | Embedded |379| Embedded BGM | Video model | Embedded |380| Separate BGM | Search only | Separate track |381| Narration | TTS (clip by clip) | Separate track |382383### Audio Mixing (Final Assembly)384385When combining multiple audio sources:386387| Track | Source |388|-------|--------|389| Video audio | Embedded in video clips (dialogue, sound effects, embedded BGM) |390| Narration | TTS generated (off-screen narrator) |391| Separate BGM | Searched from royalty-free source |392393**[CRITICAL]** Mixing rules:394- Preserve ALL audio tracks—overlay, never replace one with another395- Narration must be clearly audible—not drowned out by other tracks396- Narration volume must be consistent across all clips