Directs 2026 T2V/I2V models (Veo 3, Sora, Kling, Runway Gen-4, Vidu, Hailuo, Luma, Pika) with shot recipes: subject/action/setting/camera/style channels, motion budget, negatives, and model-specific grammar. Use when writing or debugging generative video prompts rather than still-image copy. Not for ffmpeg assembly, screenwriting, Blender/Unreal, lip-sync reenactment, or calling the generation API; never spend the motion budget on unspecified camera wander.
Prompting a 2026 video model is directing, not describing. A still-image prompt answers what is in the frame; a video prompt must also answer what changes over time and how the camera observes it. The model allocates a finite budget of motion and coherence across the clip. If you do not direct that budget, it spends motion on hallucination — drifting textures, morphing anatomy, a wandering camera.
Governing discipline: spend the motion budget deliberately — lock what must stay still, specify what must move, name how the camera moves — then suppress everything else.
Self-contained package. This skill folder contains only SKILL.md. There are no scripts/, references/, or assets/ subfolders. Do not invent local helper paths. All procedures, templates, and checklists below are self-contained in this file.
When to Use
Activate this skill whenever the job-to-be-done is directing a generative video model with language — turning an intent ("a detective walks into a neon-lit alley, camera tracks behind her") into a prompt plus parameters that a 2026 text-to-video (T2V) or image-to-video (I2V) model will render as a deliberate, coherent shot across the clip's full duration.
This skill begins where motion begins and ends at a shippable shot recipe: full positive prompt, negative/suppression block, motion and camera settings, duration/aspect/fps, seed, and (when used) anchor frame paths. It does not write the story, does not call the generation API, and does not cut the timeline — it directs the frame over time so generation and assembly have a single clear contract per clip.
"Match the look of shot A in shot B" (style/lighting/lens continuity).
"Set motion intensity / camera-move strength / how do I use --motion or the motion slider?"
"Write a negative prompt to kill warping, flicker, and extra limbs."
"Which model for [photoreal dialogue / fast action / anime / long landscape], and how do I prompt it?"
"Safety filter rejected my prompt — rewrite while keeping intent."
Sitational conditions that raise priority
Output is a shot or sequence of shots with intended camera language, not a random clip.
Mode is I2V and motion must be choreographed from a fixed anchor frame.
Draft was rejected as "it looks AI" — usually motion, coherence, or camera-control failure.
Continuity across clips matters (same character, lighting, lens feel).
You must select among models and use model-specific grammar, not generic adjectives.
When NOT to use this skill
Out of scope
Route to
ffmpeg / post encode (mux, transcode, concat, color-space, subtitle burn-in)
Media pipeline — hand off after generation
Screenwriting (story, dialogue, arcs)
Writing skill — this skill consumes a beat and turns it into a shot prompt
Traditional 3D / offline renderers (Blender, Unreal, Maya)
Geometry is the source of truth; no diffusion prompt
Still-image only (character sheets, style locks, single frames)
Image-production skills; lock keyframe there → drive with I2V here
Lip-sync reenactment (MuseTalk, LivePortrait)
Dedicated driving stack, not T2V/I2V prompting
NLE editorial (cut timing, sound design, final grade of a timeline)
Post, not generation
Shot-list / continuity bible authoring for multi-scene pieces
cinematic-shot-listing-and-continuity produces the manifest; this skill phrases each row for a model
API submit → poll → download only
video-generation-api / provider skills — after the recipe exists
Prerequisites
Access to at least one 2026 T2V/I2V model (Veo 3, Sora, Kling, Runway Gen-4, Vidu, Hailuo, Luma, Pika) or its API/UI.
For I2V workflows: a clean anchor still (good subject/background separation, unoccluded face/hands, correct identity/style/lighting). Generate anchor stills via image-production skills if needed.
For multi-shot continuity: a shot list or continuity bible from cinematic-shot-listing-and-continuity before per-shot phrasing begins.
Windows host is primary (PowerShell). No local scripts are required — all procedures are self-contained.
Procedure
Step 1 — Lock one shot intent
Before writing any prompt text, define the shot boundary:
One camera, one continuous action, one duration. If the beat contains a cut, two locations, or three simultaneous primary actions, decompose into separate shots now.
Choose T2V vs I2V. Use I2V when identity, composition, or lighting must match a known frame. Use T2V for imaginative or greenfield scenes.
Choose target model. See the model temperament table below. Re-verify live capabilities (duration caps, audio, first/last-frame, Motion Brush) before production — product surfaces change.
Define duration. Default to 4–8 seconds per take (peak coherence window). Longer beats → multiple takes, stitch in post.
Step 2 — Fill the five channels
Every reliable 2026 video prompt is five ordered channels. Keep them mentally separate even when written as flowing prose; each owns a different failure mode.
Channel
Answers
Owns failure of…
Example tokens
Subject
Who/what is the focus
identity drift, subject morphing
"a silver-haired female astronaut," "a vintage red coupe"
Action
What happens over time
temporal incoherence, no/too-much motion
"slowly removes her helmet," "drifts around the corner"
Setting
Where / when / atmosphere
background melt, scene instability
"a rain-slicked Tokyo alley at night," "a dawn salt flat"
Cinematography
How the camera sees it (size, lens, move, fps feel)
camera wander, wrong scale, jitter
"low-angle medium close-up, 35mm, slow dolly-in"
Style
Rendered look (medium, grade, era, mood)
look inconsistency, off-tone
"cinematic, teal-orange grade, anamorphic, shot on film"
Canonical order: Subject → Action → Style / Setting → Cinematography. Swap Style and Setting when one dominates mood. Front-load the least-negotiable element (usually Subject + the single key Action). Models weight earlier tokens more heavily; a camera move buried at the end of a 90-word prompt is often ignored.
Channel anti-patterns:
Anti-pattern
Why it fails
Fix
Vague subject ("someone," "a person")
Model invents extras and identity
Count and specify: "a single woman in a red coat"
Three primary actions in one prompt
Motion budget fragments; mid-clip teleport
One primary action; demote the rest to secondary motion (hair, steam)
Atmosphere: volumetric light / god rays; haze; fog; lens flare; bloom; dust motes; rain; snow — production value and coherent secondary motion that masks minor instability.
Continuity rule: keep lighting/grade/lens tokens in the Style channel and freeze them verbatim across sibling shots. Vary content and camera; do not casually rewrite the look-tail.
Step 7 — For I2V, spend tokens on Action + Camera + invariants only
I2V is the highest-control production mode in 2026: the first frame locks identity, composition, lighting, and style. The model only animates outward.
Rule
Detail
Image = what; prompt = how it moves
Drop most Subject/Setting/Style prose; spend tokens on Action + Camera
Kling, Luma, Pika, others: author both endpoints; model fills the in-between. Use when destination state matters
Motion Brush / region mask
Runway and peers: paint where motion is allowed and direction; freeze the rest (open door, ripple water, drift clouds)
Clamp magnitude low
Over-driving I2V melts a perfect anchor
I2V contradiction trap: never prompt an action that fights the still (e.g. "sprints" when the image shows a seated person). That is a fast path to warping. Change the still, or choose a plausible micro-action from the pose.
Step 8 — Write negatives or in-prompt prohibitions
Suppress the failure modes the model is prone to. Cover at minimum:
Morphing / warping / melting anatomy
Extra limbs / fused fingers / duplicated subjects
Foot sliding / skating
Background melt / scene pop / mid-clip teleport
Flicker / jitter / texture shimmer
Watermark / on-screen text / subtitles / captions
Camera wander / unmotivated drift
Anchor invariants in positive form: "her face and outfit remain identical; the background architecture stays fixed."
Step 9 — Freeze the style/lighting tail across sibling shots
For multi-shot sequences, copy the Style channel tokens verbatim from shot to shot. Vary only content and camera. Use the same model/version and seed family. Share I2V reference images when a character recurs.
Step 10 — Record a reproducible recipe
Before shipping, record:
Model + version
Mode (T2V / I2V / first-last / Motion Brush)
Anchor frame path(s) if I2V
Full positive prompt
Full negative / suppression block
All parameters: motion value + confirmed scale, camera setting, fps, resolution, aspect ratio, duration, seed
Any UI panel settings (Motion Brush regions, camera sliders)
Step 11 — Run the verification gate
See the Verification section below. Do not ship until every applicable item passes.
Model-by-model temperament (2026)
Prompt each model for its bias. Re-verify live capabilities (duration caps, audio, first/last-frame, Motion Brush) before production recommendations — product surfaces change.
[Subject, counted and specific] [single primary Action, with speed adverb] in [concrete Setting with
architecture, time, weather]. [Shot size + angle], [one named camera move + speed], [lens/format].
[Style: medium, grade, era, mood, lighting direction/quality]. [Invariants: what stays fixed].
[Secondary motion: hair, steam, cloth, dust]. No morphing, no warping, no extra limbs, no foot
sliding, no background melt, no flicker, no watermark, no on-screen text.
I2V template (anchor-driven)
[Action + Camera only — what moves, how, how fast]. [Invariants: face, outfit, identity, background
architecture — remains identical and stable]. Minimal motion, no morphing, no warping.
Runway Motion Brush note (pair with short prompt)
Prompt: [short directive]
Motion Brush: paint [regions]; direction [vector]; strength [high on subject / low on bg]
Camera: [panel]
Worked example — Veo 3 T2V
A single woman in a red wool coat walks briskly through a rain-slicked Tokyo alley at night,
her reflection shimmering in puddles. Medium shot, eye-level, slow tracking shot moving right
to left, following her from behind, 35mm anamorphic, shallow depth of field. Cinematic
teal-and-orange grade, neon practicals reflecting on wet surfaces, volumetric haze, shot on
film with subtle grain. Her face and outfit remain identical throughout; the background
architecture stays fixed. Secondary motion: rain streaks, steam from a vent, her coat
billowing slightly. No morphing, no warping, no extra limbs, no foot sliding, no background
melt, no flicker, no watermark, no on-screen text.
Worked example — Kling I2V from anchor still
The woman slowly turns her head to the right, looking toward the neon sign. Camera pushes
in gently over 4 seconds. Her face, hair, outfit, and the alley architecture remain identical
and stable. Minimal motion, no morphing, no warping.
Pitfalls
Structural pitfalls
Failure mode
Typical cause
Fix
Mid-clip teleport / scene pop
Prompt implied sequence/cut
"single continuous take, one camera, no cuts"; decompose into separate shots; remove second location/action
Three primary actions in one prompt
Motion budget fragments
One primary action; demote rest to secondary motion
Camera move ignored / wanders
Buried or vague camera line; stacked moves
Front-load one named move + speed; use panel param; drop competing moves
Camera direction inverted
Ambiguous left/right POV
Disambiguate camera-POV + what enters frame; test one gen; flip term if model is consistent
Motion / coherence pitfalls
Failure mode
Typical cause
Fix
Morphing limbs / hands melting / fingers fusing
Motion too high; hands small/fast; complex articulation
Lower motion; "stable consistent anatomy, no morphing"; keep hands larger/slower; simplify action; I2V from clean hands-visible anchor
Sliding / gliding feet (foot skating)
Gait not grounded; budget on body not footfalls
"feet firmly planted with each step, weight on each footfall, no foot sliding"; slow subject; I2V or first/last to lock stride; lower motion
No unauthorized real-person likeness or brands; policy-safe intent; rights/consent confirmed where real people appear (no impersonation/deception).
Production / reproducibility
Recipe recorded — model+version, mode, anchors, full prompt + negative, all parameters (motion + scale, camera setting, fps, res, AR, duration, seed).
Handoff identified — stitch/encode/grade/caption skill receives the clip(s); generative prompting stops at the clip boundary.
Off-target results discarded — every delivered clip passes artifact checks (no morphing, no skating, no pops, intended camera move present).
If any box fails: identify the failure mode in the Pitfalls table, fix the input first, switch mechanism if three re-rolls fail, regenerate the affected shot, re-run the gate. Ship only fully passing shots or sequences.
Related skills
Need
Route
Multi-scene shot list + continuity bible
cinematic-shot-listing-and-continuity → then return here per shot
API submit / poll / download multi-vendor
video-generation-api / provider API skills
Full story → multi-clip film pipeline
story-to-video / video-ai-production (call this skill for per-clip craft)
Google Flow / Veo UI automation
google-flow-veo
Still character/style anchors
image-production / consistent-character skills
Lip-sync reenactment
dedicated lipsync / avatar pipeline
Concat, encode, filter existing MP4s
video-processing-pipeline / ffmpeg skills
Captions / kinetic type on finished footage
kinetic-typography-and-captions
This skill remains the per-shot directing and model-grammar craft layer those pipelines call into.
Mental model (one line)
Direct the motion budget: one shot, one camera, one action — lock invariants, name the move and its speed, clamp magnitude, suppress artifacts, freeze the look-tail across siblings, and record the recipe so the same intent regenerates the same directed clip.
1---2name: video-prompt-engineering-20263description: Directs 2026 T2V/I2V models (Veo 3, Sora, Kling, Runway Gen-4, Vidu, Hailuo, Luma, Pika) with shot recipes: subject/action/setting/camera/style channels, motion budget, negatives, and model-specific grammar. Use when writing or debugging generative video prompts rather than still-image copy. Not for ffmpeg assembly, screenwriting, Blender/Unreal, lip-sync reenactment, or calling the generation API; never spend the motion budget on unspecified camera wander.4---56# Video Prompt Engineering 2026
78## Overview
910Prompting a 2026 video model is **directing, not describing**. A still-image prompt answers *what is in the frame*; a video prompt must also answer *what changes over time* and *how the camera observes it*. The model allocates a finite budget of **motion** and **coherence** across the clip. If you do not direct that budget, it spends motion on hallucination — drifting textures, morphing anatomy, a wandering camera.
1112**Governing discipline:** spend the motion budget deliberately — lock what must stay still, specify what must move, name how the camera moves — then suppress everything else.
1314**Self-contained package.** This skill folder contains only `SKILL.md`. There are no `scripts/`, `references/`, or `assets/` subfolders. Do not invent local helper paths. All procedures, templates, and checklists below are self-contained in this file.
1516## When to Use
1718Activate this skill whenever the job-to-be-done is **directing a generative video model with language** — turning an intent ("a detective walks into a neon-lit alley, camera tracks behind her") into a **prompt plus parameters** that a 2026 text-to-video (T2V) or image-to-video (I2V) model will render as a deliberate, coherent shot across the clip's full duration.
1920This skill begins where **motion** begins and ends at a **shippable shot recipe**: full positive prompt, negative/suppression block, motion and camera settings, duration/aspect/fps, seed, and (when used) anchor frame paths. It does not write the story, does not call the generation API, and does not cut the timeline — it **directs the frame over time** so generation and assembly have a single clear contract per clip.
2122### Trigger keywords
2324text-to-video, T2V, image-to-video, I2V, Veo, Veo 3, Sora, Kling, Runway, Gen-4, Vidu, Hailuo, MiniMax, Luma, Dream Machine, Pika, video prompt, camera motion, dolly, tracking shot, crane, pan, tilt, orbit, motion scale, motion strength, negative prompt video, temporal coherence, morphing, sliding feet, keyframe motion, cinematic shot, establishing shot, Motion Brush, first last frame, push-in, pull-out, handheld, locked-off, whip pan, FPV.
2526### Concrete trigger examples
2728- "Write a Veo 3 / Sora / Kling / Runway Gen-4 / Vidu / Hailuo / Luma / Pika prompt for [scene]."
29- "Turn this image into a video — slow push-in, character turns her head."
30- "My generation has morphing hands / sliding feet / the background melts / mid-clip teleport — fix the prompt."
31- "Direct a cinematic sequence: establishing, then close-up, then tracking" (as *separate* one-shot prompts).
32- "Control the camera: crane-up reveal / whip-pan / slow dolly / orbit / locked-off."
33- "Match the look of shot A in shot B" (style/lighting/lens continuity).
34- "Set motion intensity / camera-move strength / how do I use `--motion` or the motion slider?"
35- "Write a negative prompt to kill warping, flicker, and extra limbs."
36- "Which model for [photoreal dialogue / fast action / anime / long landscape], and how do I prompt it?"
37- "Safety filter rejected my prompt — rewrite while keeping intent."
3839### Sitational conditions that raise priority
4041- Output is a **shot or sequence of shots** with intended camera language, not a random clip.
42- Mode is **I2V** and motion must be choreographed *from* a fixed anchor frame.
43- Draft was rejected as "it looks AI" — usually motion, coherence, or camera-control failure.
44- **Continuity across clips** matters (same character, lighting, lens feel).
45- You must **select among models** and use model-specific grammar, not generic adjectives.
4647### When NOT to use this skill
4849| Out of scope | Route to |
50|---|---|
51| **ffmpeg / post encode** (mux, transcode, concat, color-space, subtitle burn-in) | Media pipeline — hand off *after* generation |
52| **Screenwriting** (story, dialogue, arcs) | Writing skill — this skill consumes a beat and turns it into a *shot prompt* |
53| **Traditional 3D / offline renderers** (Blender, Unreal, Maya) | Geometry is the source of truth; no diffusion prompt |
54| **Still-image only** (character sheets, style locks, single frames) | Image-production skills; lock keyframe there → drive with I2V *here* |
55| **Lip-sync reenactment** (MuseTalk, LivePortrait) | Dedicated driving stack, not T2V/I2V prompting |
56| **NLE editorial** (cut timing, sound design, final grade of a timeline) | Post, not generation |
57| **Shot-list / continuity bible authoring** for multi-scene pieces | `cinematic-shot-listing-and-continuity` produces the manifest; this skill phrases each row for a model |
58| **API submit → poll → download only** | `video-generation-api` / provider skills — after the recipe exists |
5960## Prerequisites
6162- Access to at least one 2026 T2V/I2V model (Veo 3, Sora, Kling, Runway Gen-4, Vidu, Hailuo, Luma, Pika) or its API/UI.
63- For I2V workflows: a clean anchor still (good subject/background separation, unoccluded face/hands, correct identity/style/lighting). Generate anchor stills via image-production skills if needed.
64- For multi-shot continuity: a shot list or continuity bible from `cinematic-shot-listing-and-continuity` before per-shot phrasing begins.
65- Windows host is primary (PowerShell). No local scripts are required — all procedures are self-contained.
6667## Procedure
6869### Step 1 — Lock one shot intent
7071Before writing any prompt text, define the shot boundary:
72731. **One camera, one continuous action, one duration.** If the beat contains a cut, two locations, or three simultaneous primary actions, decompose into separate shots now.
742. **Choose T2V vs I2V.** Use I2V when identity, composition, or lighting must match a known frame. Use T2V for imaginative or greenfield scenes.
753. **Choose target model.** See the model temperament table below. Re-verify live capabilities (duration caps, audio, first/last-frame, Motion Brush) before production — product surfaces change.
764. **Define duration.** Default to 4–8 seconds per take (peak coherence window). Longer beats → multiple takes, stitch in post.
7778### Step 2 — Fill the five channels
7980Every reliable 2026 video prompt is five ordered **channels**. Keep them mentally separate even when written as flowing prose; each owns a different failure mode.
8182| Channel | Answers | Owns failure of… | Example tokens |
83|---|---|---|---|
84| **Subject** | Who/what is the focus | identity drift, subject morphing | "a silver-haired female astronaut," "a vintage red coupe" |
85| **Action** | What happens over time | temporal incoherence, no/too-much motion | "slowly removes her helmet," "drifts around the corner" |
86| **Setting** | Where / when / atmosphere | background melt, scene instability | "a rain-slicked Tokyo alley at night," "a dawn salt flat" |
87| **Cinematography** | How the camera sees it (size, lens, move, fps feel) | camera wander, wrong scale, jitter | "low-angle medium close-up, 35mm, slow dolly-in" |
88| **Style** | Rendered look (medium, grade, era, mood) | look inconsistency, off-tone | "cinematic, teal-orange grade, anamorphic, shot on film" |
8990**Canonical order:** Subject → Action → Style / Setting → Cinematography. Swap Style and Setting when one dominates mood. Front-load the least-negotiable element (usually Subject + the single key Action). Models weight earlier tokens more heavily; a camera move buried at the end of a 90-word prompt is often ignored.
9192**Channel anti-patterns:**
9394| Anti-pattern | Why it fails | Fix |
95|---|---|---|
96| Vague subject ("someone," "a person") | Model invents extras and identity | Count and specify: "a single woman in a red coat" |
97| Three primary actions in one prompt | Motion budget fragments; mid-clip teleport | One primary action; demote the rest to secondary motion (hair, steam) |
98| Setting only as "beautiful city" | Background melts | Concrete architecture, time of day, weather |
99| "Move the camera dramatically" | Unmotivated drift | One named move + speed |
100| Style keyword soup (`masterpiece, 8k, ultra detailed…`) | Dilutes real constraints | Prefer lighting/lens/grade terms that change the image |
101102### Step 3 — Set motion magnitude deliberately
103104Most platforms expose an explicit or implicit **motion intensity** control: how much the model may move per unit time.
105106| Control type | Models (typical) | How to drive |
107|---|---|---|
108| **Numeric / enum** | Kling, Runway, Pika, Luma (varies) | UI slider or param: 1–10, 1–100, or low/medium/high. **Always confirm the scale** — "5" on 1–10 is moderate; "5" on 1–100 is nearly static |
109| **Language-driven** | Veo, Sora | Verbs/adverbs *are* the slider: drifts/glides/slowly/gentle = low; races/whips/explosively/rapid = high |
110111**Magnitude heuristic:**
1121131. Start **low-to-moderate**.
1142. Climb only if the shot reads dead.
1153. Prefer reading action through **camera move** rather than subject thrash.
1164. Blowout (too high) is more common and uglier than deadness — warping, smear, subject "swimming."
1175. **I2V defaults lower than T2V** — the frame is already correct.
118119**Mental scale (map to platform after confirming units):**
120121| Intent | Language cues | Rough 1–10 | Rough 1–100 |
122|---|---|---|---|
123| Near-static portrait | "nearly static," "imperceptible," "minimal" | 1–2 | 5–15 |
124| Living still / gentle life | "slow," "gentle," "subtle drift" | 2–4 | 15–35 |
125| Standard cinematic | "smooth," "moderate," deliberate action | 4–6 | 35–55 |
126| Action / chase | "fast," "dynamic," "running speed" | 6–8 | 55–80 |
127| Extreme / whip / crash | "rapid," "explosive," "whip" | 8–10 | 80–100 |
128129### Step 4 — Name the camera move and speed
130131Models trained on captioned cinematography respond to **real camera vocabulary**. Vague "move the camera" → unmotivated drift.
132133| Move | Means | Prompt phrasing |
134|---|---|---|
135| **Pan** | Horizontal rotation, fixed pivot | "camera pans left to reveal…" |
136| **Tilt** | Vertical rotation, fixed pivot | "camera tilts up from her boots to her face" |
137| **Dolly / Push-in / Pull-out** | Camera *translates* toward/away | "slow dolly-in," "pull back to reveal the room" |
138| **Truck / Track** | Lateral translation, often following | "tracking shot moving left, following the runner" |
139| **Pedestal** | Vertical translation without tilt | "camera pedestals up over the desk" |
140| **Crane / Jib / Boom** | Large sweeping vertical + arc | "crane up and back to a high wide shot" |
141| **Zoom** | Optical focal-length change (flatter than dolly) | "slow zoom in" — choose deliberately vs dolly |
142| **Orbit / Arc** | Circles the subject | "camera orbits 180° around the statue" |
143| **Roll / Dutch** | Rotation around lens axis | "subtle dutch roll for unease" |
144| **Handheld** | Organic micro-shake | "handheld, subtle natural shake" |
145| **Static / Locked-off** | Tripod, no camera motion | "static locked-off shot" — *force* stillness |
146| **Whip pan / Crash zoom** | Very fast pan/zoom | "whip-pan right" — high budget; expect blur |
147| **FPV / Drone** | Flying first-person or aerial | "FPV drone shot diving through the canyon" |
148149**Two hard rules:**
1501511. **Name the speed** — "slow," "smooth," "rapid," "gentle." Velocity is a separate axis from move type.
1522. **One primary camera move per shot.** Stacking pan + tilt + zoom + orbit → mush. Complex move = multiple shots.
153154**Direction ambiguity:** "camera pans left" vs "subject moves left" confuses models; left/right may resolve from camera POV or subject POV. Disambiguate:
155156> "the camera pans to the left, revealing the door on the right side of the frame"
157158Verify direction on a first generation before mass-producing.
159160### Step 5 — State shot size and angle
161162| Size | Frames | Use |
163|---|---|---|
164| **EWS / extreme wide** | Vast landscape; subject tiny | Scale, world establish |
165| **WS / wide** | Full figure + environment | Geography of the scene |
166| **FS / full** | Head-to-toe | Body language, blocking |
167| **MS / medium** | Waist up | Conversation default |
168| **MCU / medium close-up** | Chest to head | Dialogue intimacy |
169| **CU / close-up** | Face fills frame | Emotion |
170| **ECU / extreme close-up** | Eyes, hands, object | Detail / insert |
171172| Angle | Connotation |
173|---|---|
174| Eye-level | Neutral |
175| Low angle | Power, threat, heroism |
176| High angle | Vulnerability, surveillance |
177| Overhead / bird's-eye | Abstraction, pattern |
178| Dutch / canted | Unease |
179| OTS / POV | Relationship / subjectivity |
180181State **size + angle + one move + speed** in every shot prompt (or explicit locked-off).
182183### Step 6 — Specify lighting and color vocabulary
184185Lighting is half of "cinematic." Prefer professional terms over mood adjectives alone.
186187- **Quality / direction:** soft / hard light; key, fill, rim/backlight; motivated lighting; golden-hour; blue-hour; overcast; high-key; low-key; chiaroscuro; practicals (neon, lamps, screens).
188- **Color / grade:** teal-and-orange; desaturated; monochrome; warm tungsten; cool moonlight; sodium-vapor orange; bleach-bypass; technicolor; "shot on 35mm film, subtle grain."
189- **Lens / format:** anamorphic (horizontal flares, oval bokeh); shallow DOF / bokeh; wide-angle; telephoto compression; macro; 24fps cinematic vs 60fps hyperreal; "shot on [camera/lens]" cues.
190- **Atmosphere:** volumetric light / god rays; haze; fog; lens flare; bloom; dust motes; rain; snow — production value *and* coherent secondary motion that masks minor instability.
191192**Continuity rule:** keep lighting/grade/lens tokens in the **Style** channel and freeze them *verbatim* across sibling shots. Vary content and camera; do not casually rewrite the look-tail.
193194### Step 7 — For I2V, spend tokens on Action + Camera + invariants only
195196I2V is the **highest-control** production mode in 2026: the first frame locks identity, composition, lighting, and style. The model only animates *outward*.
197198| Rule | Detail |
199|---|---|
200| **Image = what; prompt = how it moves** | Drop most Subject/Setting/Style prose; spend tokens on **Action + Camera** |
201| **Strong anchors constrain harder** | Clean composition, clear subject/background separation, unoccluded face/hands → stable motion |
202| **First/last-frame interpolation** | Kling, Luma, Pika, others: author both endpoints; model fills the in-between. Use when destination state matters |
203| **Motion Brush / region mask** | Runway and peers: paint *where* motion is allowed and *direction*; freeze the rest (open door, ripple water, drift clouds) |
204| **Clamp magnitude low** | Over-driving I2V melts a perfect anchor |
205206**I2V contradiction trap:** never prompt an action that fights the still (e.g. "sprints" when the image shows a seated person). That is a fast path to warping. Change the still, or choose a plausible micro-action from the pose.
207208### Step 8 — Write negatives or in-prompt prohibitions
209210Suppress the failure modes the model is prone to. Cover at minimum:
211212- Morphing / warping / melting anatomy
213- Extra limbs / fused fingers / duplicated subjects
214- Foot sliding / skating
215- Background melt / scene pop / mid-clip teleport
216- Flicker / jitter / texture shimmer
217- Watermark / on-screen text / subtitles / captions
218- Camera wander / unmotivated drift
219220Anchor invariants in **positive** form: "her face and outfit remain identical; the background architecture stays fixed."
221222### Step 9 — Freeze the style/lighting tail across sibling shots
223224For multi-shot sequences, copy the Style channel tokens **verbatim** from shot to shot. Vary only content and camera. Use the same model/version and seed family. Share I2V reference images when a character recurs.
225226### Step 10 — Record a reproducible recipe
227228Before shipping, record:
229230- Model + version
231- Mode (T2V / I2V / first-last / Motion Brush)
232- Anchor frame path(s) if I2V
233- Full positive prompt
234- Full negative / suppression block
235- All parameters: motion value + confirmed scale, camera setting, fps, resolution, aspect ratio, duration, seed
236- Any UI panel settings (Motion Brush regions, camera sliders)
237238### Step 11 — Run the verification gate
239240See the **Verification** section below. Do not ship until every applicable item passes.
241242## Model-by-model temperament (2026)
243244Prompt each model for its bias. Re-verify live capabilities (duration caps, audio, first/last-frame, Motion Brush) before production recommendations — product surfaces change.
245246| Model | Strengths | Prompt style | Best for |
247|---|---|---|---|
248| **Google Veo 3** | Prose comprehension; native **synced audio** (dialogue, SFX, ambience); physical realism | Rich **paragraph** natural language; explicit cinematography; optional dialogue in quotes + sound design | Photoreal, dialogue-driven, grounded shots |
249| **OpenAI Sora (2-class)** | Long descriptive narrative prompts; world coherence; imaginative scenes | Detailed cinematic paragraphs; language is the slider (few numeric knobs) | Stylized, surreal, or photoreal single takes |
250| **Kling (2.x)** | Motion realism; **camera controls**, motion settings, **start/end frame** | Concise structured fields + UI motion/camera params | Fast physical action; precise I2V |
251| **Runway Gen-4** | Camera sliders, **Motion Brush**, reference consistency | Short **directive** prompts + on-canvas controls | Controlled I2V, regional motion, production control |
252| **Vidu / Vidu 2** | Character/reference consistency; dynamic stylized/anime motion | Explicit camera + motion; style lock early | Anime/stylized character work |
253| **Hailuo (MiniMax)** | Aesthetic + director/camera following; often strong cost/quality | Clear beats + camera instructions | Bulk aesthetic clips, camera-led shots |
254| **Luma (Dream Machine)** | Fluid natural motion; keyframe first/last; solid I2V | Action + camera; lean magnitude | Naturalistic I2V, interpolations |
255| **Pika** | Accessible I2V; quick iterations | Short directive + motion params | Prototyping, quick I2V iterations |
256257## Examples
258259### T2V template (general)
260261```text
262[Subject, counted and specific] [single primary Action, with speed adverb] in [concrete Setting with
263architecture, time, weather]. [Shot size + angle], [one named camera move + speed], [lens/format].
264[Style: medium, grade, era, mood, lighting direction/quality]. [Invariants: what stays fixed].
265[Secondary motion: hair, steam, cloth, dust]. No morphing, no warping, no extra limbs, no foot
266sliding, no background melt, no flicker, no watermark, no on-screen text.
267```
268269### I2V template (anchor-driven)
270271```text
272[Action + Camera only — what moves, how, how fast]. [Invariants: face, outfit, identity, background
273architecture — remains identical and stable]. Minimal motion, no morphing, no warping.
274```
275276### Runway Motion Brush note (pair with short prompt)
277278```text
279Prompt: [short directive]
280Motion Brush: paint [regions]; direction [vector]; strength [high on subject / low on bg]
281Camera: [panel]
282```
283284### Worked example — Veo 3 T2V
285286```text
287A single woman in a red wool coat walks briskly through a rain-slicked Tokyo alley at night,
288her reflection shimmering in puddles. Medium shot, eye-level, slow tracking shot moving right
289to left, following her from behind, 35mm anamorphic, shallow depth of field. Cinematic
290teal-and-orange grade, neon practicals reflecting on wet surfaces, volumetric haze, shot on
291film with subtle grain. Her face and outfit remain identical throughout; the background
292architecture stays fixed. Secondary motion: rain streaks, steam from a vent, her coat
293billowing slightly. No morphing, no warping, no extra limbs, no foot sliding, no background
294melt, no flicker, no watermark, no on-screen text.
295```
296297### Worked example — Kling I2V from anchor still
298299```text
300The woman slowly turns her head to the right, looking toward the neon sign. Camera pushes
301in gently over 4 seconds. Her face, hair, outfit, and the alley architecture remain identical
302and stable. Minimal motion, no morphing, no warping.
303```
304305## Pitfalls
306307### Structural pitfalls
308309| Failure mode | Typical cause | Fix |
310|---|---|---|
311| **Mid-clip teleport / scene pop** | Prompt implied sequence/cut | "single continuous take, one camera, no cuts"; **decompose into separate shots**; remove second location/action |
312| **Three primary actions in one prompt** | Motion budget fragments | One primary action; demote rest to secondary motion |
313| **Camera move ignored / wanders** | Buried or vague camera line; stacked moves | Front-load one named move + speed; use panel param; drop competing moves |
314| **Camera direction inverted** | Ambiguous left/right POV | Disambiguate camera-POV + what enters frame; test one gen; flip term if model is consistent |
315316### Motion / coherence pitfalls
317318| Failure mode | Typical cause | Fix |
319|---|---|---|
320| **Morphing limbs / hands melting / fingers fusing** | Motion too high; hands small/fast; complex articulation | Lower motion; "stable consistent anatomy, no morphing"; keep hands larger/slower; simplify action; I2V from clean hands-visible anchor |
321| **Sliding / gliding feet (foot skating)** | Gait not grounded; budget on body not footfalls | "feet firmly planted with each step, weight on each footfall, no foot sliding"; slow subject; I2V or first/last to lock stride; lower motion |
322| **Temporal float / subject swims** | Over-motion; no invariants; weak scene | Clamp motion; "subject stays centered and stable, background fixed"; shorten clip; lower I2V strength |
323| **Background melt** | Aggressive camera; thin setting; high motion | Slow camera; concrete setting; "background architecture stays static"; Motion Brush freeze bg |
324| **Flicker / texture shimmer** | Fine detail (foliage, fabric, text) + motion; high fps | Reduce motion; avoid dense fine texture; "no flickering, stable textures"; try 24fps over 60 |
325| **Motion-scale blowout** | Wrong numeric scale (50 on 1–100 thinking 1–10) | Re-confirm scale; halve value; "smooth, controlled, minimal motion"; prefer camera-led action |
326| **Dead / static clip** | Magnitude too low; no actionable verb | Raise modestly; add micro-motion (blink, hair, steam); gentle camera life |
327328### Content / safety pitfalls
329330| Failure mode | Typical cause | Fix |
331|---|---|---|
332| **Extra people / wrong objects / invented text** | Under-constrained scene; vague nouns | Tighten subject + counts; negative "duplicate subjects, extra people, text, watermark"; shorten, front-load keys |
333| **Safety filter reject** | Violence, real public figures, brands, sensitive terms | Cinematic abstraction; generic roles not named people; drop brands; rights/consent for real likeness; no deception |
334| **Burned-in subtitles appear** | Training bias on captioned clips | "(no subtitles, no on-screen text, no captions)"; avoid quote-only dialogue without anti-caption line |
335336### I2V-specific pitfalls
337338| Failure mode | Typical cause | Fix |
339|---|---|---|
340| **I2V melts perfect anchor** | I2V motion too high | Drop to minimum; Motion Brush localized; first/last interpolation |
341| **I2V warps because action fights still** | Prompted motion impossible from pose | Change action to plausible micro-motion or regenerate anchor pose |
342343### Continuity pitfalls
344345| Failure mode | Typical cause | Fix |
346|---|---|---|
347| **Sequence continuity breaks** | Style/lighting/lens rewritten; different models/seeds | Freeze style tail verbatim; same model/version; shared I2V reference; reuse recipe |
348| **Lip-sync / dialogue mismatch** | Line too long; mouth under-constrained | Shorten dialogue; Veo quotes + delivery; precise lip-sync → reenactment pipeline |
349| **Multi-shot morph soup** | Hard cuts on single-shot model | Confirm multi-shot support; else one continuous action only; cut in post |
350351### Cross-cutting recovery order (always)
3523531. Fix **input** (anchor quality, action complexity, motion scale, one-shot constraint).
3542. Fix **mechanism** (T2V → I2V → first/last → Motion Brush → different model) after three failed re-rolls.
3553. Fix **wording** last (channel order, invariants, negatives).
3564. Re-run the verification gate; discard off-target clips — do not ship them.
357358## Verification
359360Do not ship a video prompt (or sequence package) until every applicable item passes. Hard gate, not a suggestion.
361362### Prompt structure & subject clarity
363364- [ ] **All five channels present and ordered** — Subject, Action, Setting, Cinematography, Style — most important front-loaded.
365- [ ] **Subject unambiguous and counted** — no vague nouns that invite extras.
366- [ ] **Exactly one shot, one camera, one continuous take** — no implied cut, second location, or sequence smuggled in.
367- [ ] **Single concrete primary action** (not three simultaneous primaries).
368369### Cinematography & motion control
370371- [ ] **Shot size and angle stated** (wide/medium/close-up; eye-level/low/high/dutch).
372- [ ] **Exactly one primary camera move named with a speed** — or explicit "static locked-off."
373- [ ] **Camera direction disambiguated** (camera-POV vs subject-POV) and verified on a first gen before mass production.
374- [ ] **Motion magnitude deliberate**; numeric value matches **confirmed scale** (1–10 vs 1–100 vs enum); default low, climb only if dead.
375- [ ] **fps / resolution / aspect / duration** set to model-native values and coherence budget (default 4–8 s per take).
376377### Look & continuity
378379- [ ] **Lighting and color/grade** specified with professional vocabulary.
380- [ ] **Style tail frozen verbatim** across sibling shots; identity via I2V/reference when a character recurs.
381382### I2V (when applicable)
383384- [ ] **Anchor clean** (separation, unoccluded face/hands, exact identity/style/lighting).
385- [ ] **Prompt spends on Action+Camera only**; translation/rotation named; **invariants stated**.
386- [ ] **First/last or Motion Brush** used when destination or regional motion matters; **I2V motion clamped low**.
387- [ ] Action is **physically plausible from the still**.
388389### Suppression & safety
390391- [ ] **Negative prompt or in-prompt prohibitions** cover morphing/warping, anatomy, flicker/jitter, foot-sliding, melt, duplicates, watermark/text, scene-cut.
392- [ ] **Invariants anchored** in positive form.
393- [ ] **No unauthorized real-person likeness or brands**; policy-safe intent; rights/consent confirmed where real people appear (no impersonation/deception).
394395### Production / reproducibility
396397- [ ] **Recipe recorded** — model+version, mode, anchors, full prompt + negative, all parameters (motion + scale, camera setting, fps, res, AR, duration, seed).
398- [ ] **Handoff identified** — stitch/encode/grade/caption skill receives the clip(s); generative prompting stops at the clip boundary.
399- [ ] **Off-target results discarded** — every delivered clip passes artifact checks (no morphing, no skating, no pops, intended camera move present).
400401If any box fails: identify the failure mode in the Pitfalls table, **fix the input first**, switch mechanism if three re-rolls fail, regenerate the affected shot, re-run the gate. Ship only fully passing shots or sequences.
402403## Related skills
404405| Need | Route |
406|---|---|
407| Multi-scene shot list + continuity bible | `cinematic-shot-listing-and-continuity` → then return here per shot |
408| API submit / poll / download multi-vendor | `video-generation-api` / provider API skills |
409| Full story → multi-clip film pipeline | `story-to-video` / `video-ai-production` (call this skill for per-clip craft) |
410| Google Flow / Veo UI automation | `google-flow-veo` |
411| Still character/style anchors | image-production / consistent-character skills |
412| Lip-sync reenactment | dedicated lipsync / avatar pipeline |
413| Concat, encode, filter existing MP4s | `video-processing-pipeline` / ffmpeg skills |
414| Captions / kinetic type on finished footage | `kinetic-typography-and-captions` |
415416This skill remains the **per-shot directing and model-grammar craft layer** those pipelines call into.
417418## Mental model (one line)
419420**Direct the motion budget:** one shot, one camera, one action — lock invariants, name the move and its speed, clamp magnitude, suppress artifacts, freeze the look-tail across siblings, and record the recipe so the same intent regenerates the same directed clip.
Run npx skillmds@latest add gabrielmoreira/video-prompt-engineering-2026 in your terminal (requires Node.js), paste this page's agent-chat prompt into Claude, Cursor, or any MCP-connected agent, or download the SKILL.md file and copy it into your agent's skills directory.
Directs 2026 T2V/I2V models (Veo 3, Sora, Kling, Runway Gen-4, Vidu, Hailuo, Luma, Pika) with shot recipes: subject/action/setting/camera/style channels, motion budget, negatives, and model-specific grammar. Use when writing or debugging generative video prompts rather than still-image copy. Not for ffmpeg assembly, screenwriting, Blender/Unreal, lip-sync reenactment, or calling the generation API; never spend the motion budget on unspecified camera wander. It is listed under AI & ML on SkillMD.
This skill has not completed SkillMD's automated safety review yet. Independent scanners report: SkillSpector: PASS, Skill Scanner: PASS. Capability flags: docs only. SkillMD never runs a skill's scripts for you; review the SKILL.md before installing.
This skill is tagged as working with Claude Code, Claude.ai, OpenAI Codex. SKILL.md is an open format, so most agents that read a skills directory can load it too.
Yes. Installing skills from SkillMD is free, and the skill stays under its author's original license.
gabrielmoreira (@gabrielmoreira) published this skill. Their other Agent Skills are listed on their SkillMD profile.