video-director
Turn a scene idea into the 11-block cinematic prompt the
hollywood-director persona ships. Output is provider-agnostic
prose; provider-specific tuning is handed off to
motion-choreographer.
When to use
- A scene idea, beat, or script line needs to become a cinematic
prompt ready to feed an image+video pipeline.
- A draft prompt reads as "AI video" — flat, centered, generic
golden-hour — and needs directorial choices on the page.
- Live-action / photoreal scenes. For animation acting beats, route
to
pixar-storyteller.
Do NOT use when:
- The deliverable is a still graphic or poster →
canvas-design.
- The work is provider-specific token tuning →
motion-choreographer after this skill has produced the blocks.
- A character identity must be re-used across scenes → run
character-consistency first
to lock identity tokens, then call this skill.
Procedure
Step 0: Inspect
- Confirm the input is a live-action / photoreal beat (not animation).
- If a
character.json exists under agents/reference/ai-video/<project>/characters/,
read the identity tokens — they are reused verbatim.
- Read the scene's intent in one sentence — what is the camera
witnessing, and why now?
Step 1: Draft the 11 blocks
Emit each block on its own labeled line. Blocks are mandatory and
in this order:
- SCENE — location, time-of-day, weather, era.
- CHARACTER — verbatim identity tokens from
character.json
when present; otherwise silhouette + wardrobe + signature prop.
- ACTION — verbs and beats. Anticipation → action → reaction
named separately. No "doing things" prose.
- CAMERA — position, height, distance, move (lock-off, dolly,
handheld, push, pull). Off-axis when on-axis is the AI default.
- LENS — focal length in mm (24 / 35 / 50 / 85 / 200) and
aperture intent (deep / shallow). "Cinematic" alone fails.
- LIGHTING — key, fill, back, practical sources named. "Golden
hour" requires an angle (low-east 15°, etc.).
- ENVIRONMENT MOTION — what the world does (wind, water,
crowd, traffic) and on which beat.
- SECONDARY MOTION — hair, cloth, dust, breath. Names the
reactive layer that sells the primary action.
- MOOD — one emotional read; no compound moods.
- DURATION — seconds. Integer or one decimal.
- NEGATIVE CONSTRAINTS — clichés to reject, in load-bearing
order (top survives truncation). Always names: centered framing,
symmetric composition, generic "cinematic", soap-opera contrast.
Step 2: Self-review
- Lens length present? "Cinematic" without mm → fail.
- Lighting direction present? "Golden hour" without angle → fail.
- ACTION names beats, not adjectives?
- NEGATIVE block has at least 4 entries, load-bearing on top?
- CHARACTER block reuses identity tokens verbatim when a lock exists?
Any "no" → revise that block before handing off.
Step 3: Validate
- Output is plain text, one labeled block per line, ready for
scripts/ai-video/lib/parse-blueprint.sh (Phase 3 Step 5).
- No provider tokens (no
--aspect, no --model). That is
motion-choreographer's job.
Output format
scenes/<id>/prompt.txt — 11 labeled blocks, one per line,
ready for the blueprint parser.
scenes/<id>/review.md — one-paragraph rationale per
non-obvious directorial choice (lens, light angle, camera move).
Gotcha
- The model defaults to centered, on-axis, symmetric — name an
off-axis or rule-of-thirds camera or it will silently center.
- "Golden hour" alone reads as a sunset GIF; require a sun angle.
- ACTION written as a paragraph of adjectives ("dynamically", "powerfully")
fails — adapters need verbs with beat counts.
- When
character.json exists, paraphrasing identity tokens breaks
Character Lock — copy them verbatim.
- Negative constraints in the truncated tail get dropped — load-
bearing ones go first.
Do NOT
- Do NOT emit provider-specific tokens (aspect ratio, model id,
duration flags) — that is
motion-choreographer's scope.
- Do NOT collapse anticipation / action / reaction into one verb.
- Do NOT use "cinematic" without lens + lighting + camera move.
- Do NOT invent character details when a
character.json exists.
Policies
Paths, enforcement model, and the full set: the
media policy preamble.
The 11-block cinematic prompt is live-action shape — real-person and brand-impersonation risks are the highest in the cluster. Before emitting:
likeness — when the prompt names or visually identifies a real person on camera.
public-figures — when the subject is a recognised public figure.
brand-impersonation — when the prompt copies a journalism / broadcaster / regulated-industry visual identity.
style — when LIGHT / LENS choices are anchored to a named living cinematographer's signature.
disclosure — every distributed live-action AI clip carries the non-removable AI-generation disclosure.
Refuse-and-surface at the directorial layer; live-action realism amplifies every policy gap downstream.
1---2name: video-director3description: Use when a live-action beat becomes the 11-block cinematic prompt — lens, lighting, negatives. Triggers 'cinematic prompt', 'film-grade scene'. Animated → pixar-storyteller; 12-block → scene-expander.4---56# video-director78> Turn a scene idea into the **11-block cinematic prompt** the9> `hollywood-director` persona ships. Output is provider-agnostic10> prose; provider-specific tuning is handed off to11> [`motion-choreographer`](../motion-choreographer/SKILL.md).1213## When to use1415- A scene idea, beat, or script line needs to become a cinematic16 prompt ready to feed an image+video pipeline.17- A draft prompt reads as "AI video" — flat, centered, generic18 golden-hour — and needs directorial choices on the page.19- Live-action / photoreal scenes. For animation acting beats, route20 to [`pixar-storyteller`](../pixar-storyteller/SKILL.md).2122Do NOT use when:2324- The deliverable is a still graphic or poster → `canvas-design`.25- The work is provider-specific token tuning →26 `motion-choreographer` after this skill has produced the blocks.27- A character identity must be re-used across scenes → run28 [`character-consistency`](../character-consistency/SKILL.md) first29 to lock identity tokens, then call this skill.3031## Procedure3233### Step 0: Inspect34351. Confirm the input is a live-action / photoreal beat (not animation).362. If a `character.json` exists under `agents/reference/ai-video/<project>/characters/`,37 read the identity tokens — they are reused verbatim.383. Read the scene's intent in one sentence — what is the camera39 witnessing, and why now?4041### Step 1: Draft the 11 blocks4243Emit each block on its own labeled line. Blocks are mandatory and44in this order:45461. **SCENE** — location, time-of-day, weather, era.472. **CHARACTER** — verbatim identity tokens from `character.json`48 when present; otherwise silhouette + wardrobe + signature prop.493. **ACTION** — verbs and beats. Anticipation → action → reaction50 named separately. No "doing things" prose.514. **CAMERA** — position, height, distance, move (lock-off, dolly,52 handheld, push, pull). Off-axis when on-axis is the AI default.535. **LENS** — focal length in mm (24 / 35 / 50 / 85 / 200) and54 aperture intent (deep / shallow). "Cinematic" alone fails.556. **LIGHTING** — key, fill, back, practical sources named. "Golden56 hour" requires an angle (low-east 15°, etc.).577. **ENVIRONMENT MOTION** — what the world does (wind, water,58 crowd, traffic) and on which beat.598. **SECONDARY MOTION** — hair, cloth, dust, breath. Names the60 reactive layer that sells the primary action.619. **MOOD** — one emotional read; no compound moods.6210. **DURATION** — seconds. Integer or one decimal.6311. **NEGATIVE CONSTRAINTS** — clichés to reject, in load-bearing64 order (top survives truncation). Always names: centered framing,65 symmetric composition, generic "cinematic", soap-opera contrast.6667### Step 2: Self-review68691. Lens length present? "Cinematic" without mm → fail.702. Lighting direction present? "Golden hour" without angle → fail.713. ACTION names beats, not adjectives?724. NEGATIVE block has at least 4 entries, load-bearing on top?735. CHARACTER block reuses identity tokens verbatim when a lock exists?7475Any "no" → revise that block before handing off.7677### Step 3: Validate78791. Output is plain text, one labeled block per line, ready for80 `scripts/ai-video/lib/parse-blueprint.sh` (Phase 3 Step 5).812. No provider tokens (no `--aspect`, no `--model`). That is82 `motion-choreographer`'s job.8384## Output format85861. **`scenes/<id>/prompt.txt`** — 11 labeled blocks, one per line,87 ready for the blueprint parser.882. **`scenes/<id>/review.md`** — one-paragraph rationale per89 non-obvious directorial choice (lens, light angle, camera move).9091## Gotcha9293- The model defaults to centered, on-axis, symmetric — name an94 off-axis or rule-of-thirds camera or it will silently center.95- "Golden hour" alone reads as a sunset GIF; require a sun angle.96- ACTION written as a paragraph of adjectives ("dynamically", "powerfully")97 fails — adapters need verbs with beat counts.98- When `character.json` exists, paraphrasing identity tokens breaks99 Character Lock — copy them verbatim.100- Negative constraints in the truncated tail get dropped — load-101 bearing ones go first.102103## Do NOT104105- Do NOT emit provider-specific tokens (aspect ratio, model id,106 duration flags) — that is `motion-choreographer`'s scope.107- Do NOT collapse anticipation / action / reaction into one verb.108- Do NOT use "cinematic" without lens + lighting + camera move.109- Do NOT invent character details when a `character.json` exists.110111## Policies112113Paths, enforcement model, and the full set: the114[media policy preamble](../../../agents/settings/policies/media/README.md).115116The 11-block cinematic prompt is live-action shape — real-person and brand-impersonation risks are the highest in the cluster. Before emitting:117118- **`likeness`** — when the prompt names or visually identifies a real person on camera.119- **`public-figures`** — when the subject is a recognised public figure.120- **`brand-impersonation`** — when the prompt copies a journalism / broadcaster / regulated-industry visual identity.121- **`style`** — when LIGHT / LENS choices are anchored to a named living cinematographer's signature.122- **`disclosure`** — every distributed live-action AI clip carries the non-removable AI-generation disclosure.123124Refuse-and-surface at the directorial layer; live-action realism amplifies every policy gap downstream.125