Prompt expander
Turn a rough idea into a complete, generation-ready cinematic prompt — by deliberately
deciding every dimension a clip can express, then writing only the relevant choices
as natural prose.
Written against Gemini Omni Flash in Google Flow (~10s clips with synced audio;
accepts text + reference image/audio/video; supports multi-turn conversational editing).
The dimension model transfers to any video generator; the mechanics in §B of the audio
reference are Omni-specific.
When to use: someone hands you a logline and wants it "expanded" · you are writing
or polishing a Flow prompt · you want to edit an existing clip (use the Editing path,
not full expansion).
The dimensions — think through ALL of them
For every prompt, consciously decide each axis, then keep only what serves the idea:
| # |
Dimension |
Reference |
| 1 |
Subject — who/what, 5–8 specific descriptors |
subject-action-continuity.md |
| 2 |
Action / physics — one beat, strong verbs, secondary motion |
same |
| 3 |
Setting — location, era, weather, depth layers |
same |
| 4 |
Camera — shot size, angle, movement, lens |
visual-language.md |
| 5 |
Light & colour — setup, time of day, palette, film stock |
same |
| 6 |
Style & mood — visual style, tone word, atmosphere |
same |
| 7 |
Audio — dialogue / SFX / ambient / music, or explicit silence |
audio-and-mechanics.md |
| 8 |
Beat — the single emotional moment the clip delivers |
subject-action-continuity.md |
Read the reference for any axis you are unsure how to fill. Do not paste the menus
into the prompt — choose a few options and write prose.
Workflow
- Parse the idea: pull out subject, action, setting, and any stated mood, genre or
audio.
- Classify: a fresh generation, or an edit of an existing clip? An edit goes
to the Editing path below.
- Fill the gaps: for each of the 8 dimensions, choose options that fit the mood.
Use the defaults table for anything unspecified.
- Respect the constraints (audio reference §B5): one beat per ~10s · always specify
audio · phrase exclusions positively · do not rely on rendered on-screen text · keep
references to a few.
- Write the prompt as flowing prose in this order:
camera → subject → action → setting → light/style → audio.
Add a timestamp breakdown only if there are genuinely 2–3 beats.
- Lint the result —
veo-prompt-lint — and fix any
ERRORs. If that is unavailable, self-check against the Common mistakes list below.
- Present three things: the prompt block, a short Levers list of the assumptions
you made so they can be tweaked, and 1–2 suggested conversational edits for
iterating.
Sensible defaults (when the user is silent)
| Axis |
Default |
| Style |
photorealistic, cinematic |
| Camera |
slow push-in, eye-level, 50mm, shallow depth of field |
| Light |
motivated natural light matched to the setting's time of day |
| Motion |
24fps real-time, natural real-world physics |
| Audio |
matching ambient bed + one tied SFX; no dialogue unless implied |
| Aspect |
16:9 (9:16 if "shorts / vertical / phone" is implied) |
| Beat |
one clear reveal or reaction |
State the defaults you used in the Levers list. Never silently bury a consequential
choice.
Ask vs. assume
Default to assuming and proceeding, noting it in Levers. Ask only when a choice is
both consequential and genuinely ambiguous — almost always one of: realistic vs.
animated/stylised, dialogue vs. silent, vertical vs. widescreen. Bundle those
questions; do not drip them one at a time.
Editing path (an existing clip)
Use the surgical rule from the audio reference §B1:
- One change per turn. Isolate a single dimension; verify before the next edit.
- Freeze-list: name explicitly what stays.
Change <X>. Keep the camera, subject, location, timing, and audio identical.
- Batch changes as a numbered list, ending with the same freeze line.
- Remember: editing the spoken audio of a clip is currently restricted.
Worked example
Idea: "a lonely astronaut finds a flower on a dead planet"
Expanded prompt:
Slow push-in, low-angle medium shot on a weathered astronaut in a scuffed white EVA
suit, cracked gold visor reflecting a dim red sun, kneeling on cracked ochre desert
rock. He slowly reaches a trembling gloved hand toward a single impossible blue flower
pushing up through the dust; grit drifts in the low gravity. Vast empty rust-coloured
plain stretches to a hazy horizon behind him, distant dust devils. Bleak desaturated
palette with one vivid blue accent, harsh low sunlight, 35mm film grain, melancholic
and awestruck. Ambient: faint wind and the hiss of suit air. SFX: soft servo whir as
his hand moves. No dialogue, no music.
Levers: photorealistic (could go painterly) · 16:9 (could be 9:16) · silent (could
add a breath or VO) · single push-in (could add a [00:06-00:10] insert close-up on the
flower).
Suggested edit: Change the flower from blue to glowing bioluminescent teal. Keep the camera, subject, location, timing, and audio identical.
Common mistakes
- Pasting whole menu tables into the prompt — pick a few options, write prose.
- Forgetting audio → the model invents unwanted speech.
- Cramming a multi-scene story into 10 seconds — one beat.
- "no people / no text" phrasing → flip it positive ("a deserted street", "a clear sky").
- "locked tripod" → use "static shot".
The reference menus
| File |
Covers |
visual-language.md |
Shot sizes, angles, camera movements, lenses and optics, composition, motion feel, lighting setups, time of day, colour temperature, palettes and grading, film stock, visual styles, mood, texture |
subject-action-continuity.md |
Subject descriptor slots (people, creatures, objects), action verbs and pacing, physics cues, setting layers, continuity and multi-shot sequencing, emotional beats |
audio-and-mechanics.md |
Dialogue, SFX, ambient, music, sync, the silence gotcha — plus Omni Flash mechanics: conversational editing, reference inputs, exclusions, timestamps, constraints |
1---2name: prompt-expander3description: Use when someone hands you a rough video idea, logline or one-liner and wants a full cinematic prompt, when writing or polishing a Gemini Omni Flash / Google Flow prompt, or when planning a conversational edit of an existing clip. Turns "a lonely astronaut finds a flower" into a complete, generation-ready prompt by deciding every dimension a clip can express.4---56# Prompt expander78Turn a rough idea into a complete, generation-ready cinematic prompt — by deliberately9deciding **every dimension a clip can express**, then writing only the relevant choices10as natural prose.1112Written against **Gemini Omni Flash** in Google Flow (~10s clips with synced audio;13accepts text + reference image/audio/video; supports multi-turn conversational editing).14The dimension model transfers to any video generator; the mechanics in §B of the audio15reference are Omni-specific.1617**When to use:** someone hands you a logline and wants it "expanded" · you are writing18or polishing a Flow prompt · you want to edit an existing clip (use the Editing path,19not full expansion).2021---2223## The dimensions — think through ALL of them2425For every prompt, consciously decide each axis, then keep only what serves the idea:2627| # | Dimension | Reference |28|---|---|---|29| 1 | **Subject** — who/what, 5–8 specific descriptors | [`subject-action-continuity.md`](references/subject-action-continuity.md) |30| 2 | **Action / physics** — one beat, strong verbs, secondary motion | same |31| 3 | **Setting** — location, era, weather, depth layers | same |32| 4 | **Camera** — shot size, angle, movement, lens | [`visual-language.md`](references/visual-language.md) |33| 5 | **Light & colour** — setup, time of day, palette, film stock | same |34| 6 | **Style & mood** — visual style, tone word, atmosphere | same |35| 7 | **Audio** — dialogue / SFX / ambient / music, or explicit silence | [`audio-and-mechanics.md`](references/audio-and-mechanics.md) |36| 8 | **Beat** — the single emotional moment the clip delivers | [`subject-action-continuity.md`](references/subject-action-continuity.md) |3738Read the reference for any axis you are unsure how to fill. **Do not paste the menus39into the prompt** — choose a few options and write prose.4041---4243## Workflow44451. **Parse** the idea: pull out subject, action, setting, and any stated mood, genre or46 audio.472. **Classify:** a fresh **generation**, or an **edit** of an existing clip? An edit goes48 to the Editing path below.493. **Fill the gaps:** for each of the 8 dimensions, choose options that fit the mood.50 Use the defaults table for anything unspecified.514. **Respect the constraints** (audio reference §B5): one beat per ~10s · always specify52 audio · phrase exclusions positively · do not rely on rendered on-screen text · keep53 references to a few.545. **Write** the prompt as flowing prose in this order:55 *camera → subject → action → setting → light/style → audio*.56 Add a timestamp breakdown only if there are genuinely 2–3 beats.576. **Lint** the result — [`veo-prompt-lint`](../veo-prompt-lint/SKILL.md) — and fix any58 ERRORs. If that is unavailable, self-check against the Common mistakes list below.597. **Present** three things: the prompt block, a short **Levers** list of the assumptions60 you made so they can be tweaked, and 1–2 suggested **conversational edits** for61 iterating.6263---6465## Sensible defaults (when the user is silent)6667| Axis | Default |68|---|---|69| Style | photorealistic, cinematic |70| Camera | slow push-in, eye-level, 50mm, shallow depth of field |71| Light | motivated natural light matched to the setting's time of day |72| Motion | 24fps real-time, natural real-world physics |73| Audio | matching ambient bed + one tied SFX; **no dialogue** unless implied |74| Aspect | 16:9 (9:16 if "shorts / vertical / phone" is implied) |75| Beat | one clear reveal or reaction |7677**State the defaults you used in the Levers list.** Never silently bury a consequential78choice.7980## Ask vs. assume8182Default to assuming and proceeding, noting it in Levers. Ask **only** when a choice is83both consequential and genuinely ambiguous — almost always one of: **realistic vs.84animated/stylised**, **dialogue vs. silent**, **vertical vs. widescreen**. Bundle those85questions; do not drip them one at a time.8687---8889## Editing path (an existing clip)9091Use the surgical rule from the audio reference §B1:9293- **One change per turn.** Isolate a single dimension; verify before the next edit.94- **Freeze-list:** name explicitly what stays.95 `Change <X>. Keep the camera, subject, location, timing, and audio identical.`96- **Batch changes as a numbered list**, ending with the same freeze line.97- Remember: editing the *spoken audio* of a clip is currently restricted.9899---100101## Worked example102103**Idea:** "a lonely astronaut finds a flower on a dead planet"104105**Expanded prompt:**106107> Slow push-in, low-angle medium shot on a weathered astronaut in a scuffed white EVA108> suit, cracked gold visor reflecting a dim red sun, kneeling on cracked ochre desert109> rock. He slowly reaches a trembling gloved hand toward a single impossible blue flower110> pushing up through the dust; grit drifts in the low gravity. Vast empty rust-coloured111> plain stretches to a hazy horizon behind him, distant dust devils. Bleak desaturated112> palette with one vivid blue accent, harsh low sunlight, 35mm film grain, melancholic113> and awestruck. Ambient: faint wind and the hiss of suit air. SFX: soft servo whir as114> his hand moves. No dialogue, no music.115116**Levers:** photorealistic (could go painterly) · 16:9 (could be 9:16) · silent (could117add a breath or VO) · single push-in (could add a `[00:06-00:10]` insert close-up on the118flower).119120**Suggested edit:** `Change the flower from blue to glowing bioluminescent teal. Keep121the camera, subject, location, timing, and audio identical.`122123---124125## Common mistakes126127- Pasting whole menu tables into the prompt — pick a few options, write prose.128- Forgetting audio → the model invents unwanted speech.129- Cramming a multi-scene story into 10 seconds — one beat.130- "no people / no text" phrasing → flip it positive ("a deserted street", "a clear sky").131- "locked tripod" → use "static shot".132133---134135## The reference menus136137| File | Covers |138|---|---|139| [`visual-language.md`](references/visual-language.md) | Shot sizes, angles, camera movements, lenses and optics, composition, motion feel, lighting setups, time of day, colour temperature, palettes and grading, film stock, visual styles, mood, texture |140| [`subject-action-continuity.md`](references/subject-action-continuity.md) | Subject descriptor slots (people, creatures, objects), action verbs and pacing, physics cues, setting layers, continuity and multi-shot sequencing, emotional beats |141| [`audio-and-mechanics.md`](references/audio-and-mechanics.md) | Dialogue, SFX, ambient, music, sync, the silence gotcha — plus Omni Flash mechanics: conversational editing, reference inputs, exclusions, timestamps, constraints |