Veo (video prompt craft)
Write prompts that get cinematic clips out of Veo — Google's video model (Veo 3.1; Fast/Lite
tiers for cheap iteration). This is the prompt-craft layer of a three-layer setup, and the video
counterpart to nano-banana:
- Connection/API (model IDs, async generate-and-poll, the generate → upload to WoopSocial Media →
attach flow) →
tools/integrations/veo.md.
- Prompt craft (this skill) → how to direct the right clip well.
- In-skill application →
reels-script's veo-prompt-pack.
Fast-moving area — re-verify model names/specs quarterly.
Reach for Veo when… (match the job)
Its real strengths: native synchronized audio (dialogue + lip-sync, SFX, ambient, music — the
differentiator), cinematic realism/physics, image-to-video, reference-image consistency,
first/last-frame control, and native 9:16 vertical / 4K. Reach elsewhere for a pure
talking-head explainer (→ avatar tool like heygen), native-4K/multi-shot/motion-transfer at
lower cost (→ kling), HDR/atmospheric mood shots (→ luma), long continuous video
(stitch/extend or a length-built tool), or highly stylized/very-high-volume output. Don't
default it to every job. (Details: references/when-and-how-to-prompt.md.)
Step 0 — Read the brand + the job
Load brand-profile.md (visual style, palette, tone). Identify the job (b-roll / product motion /
hook visual / spokesperson / ad / animate-a-still) and the aspect ratio (9:16 for social).
Step 1 — Direct the shot (natural language, not quality spam)
Describe a shot like a director: subject · action · scene · camera (type/movement/angle/lens) ·
lighting/mood · style · audio · timing. One clear motion per clip. Set 9:16. Ground in the brand.
See references/when-and-how-to-prompt.md.
Step 2 — Use the superpower: audio
Veo generates native synchronized audio in one pass — describe it explicitly: dialogue in
quotes (+ who/tone), SFX, ambient, music mood. Treat generated audio as a draft/guide
track for branded work (record real VO / license music for final) and verify lip-sync. Direct
the camera/motion for the cinematic feel. See references/audio-and-camera.md.
Step 3 — Inputs, length, recipes
- Image-to-video: animate a still (e.g. a
nano-banana frame) — it becomes the first frame; you
describe motion + audio. Use reference images / first-and-last frame for consistency/control.
- Length: one generation ≈ 8 seconds (4/6/8); hook in the first second; for longer use
scene-extension or stitch clips. Plan short beats.
- Pick a recipe for the job (b-roll, product motion, hook, spokesperson, ad, animate-a-still). See
references/inputs-length-and-recipes.md.
Step 4 — Iterate cheaply, then verify, disclose, ship
- Iterate at low-res / Fast to lock the prompt; finalize the keeper at high-res/4K (4K costs
~40–60% more time/$). Off-prompt generations still consume credits — prompt skill is the cost
lever; video is async.
- Verify every clip (artifacts, lip-sync, physics) before publishing.
- Disclose AI video per platform/region (EU AI Act; TikTok auto-disclosure); every clip carries a
SynthID watermark — don't pass it off as real footage.
- Ship: generate per the integration guide → upload to WoopSocial Media → attach via
scheduling-and-queue. WoopSocial doesn't generate video.
Quality bar — self-check
- Did I match the tool to the job (and route talking-heads/long-form elsewhere)?
- Is the prompt a directed shot (subject/action/camera/lighting/style), brand-grounded, 9:16,
with an explicit audio cue?
- Did I plan for 8-second clips (hook first) and iterate cheap → finalize?
- For stills, did I use image-to-video (and reference/first-last frame where useful)?
- Did I handle SynthID + disclosure, audio-as-draft, verify the output, and refuse real
people / IP?
- Did I point to
tools/integrations/veo.md for the API + WoopSocial flow (no claim WoopSocial
generates video)?
Edge cases & pushback
- Talking-head explainer / lots of dialogue → suggest an avatar tool (
heygen); don't force Veo.
- "Make a 40-second video" → ~8s per gen; scene-extend/stitch; plan short beats.
- "Generate 10 final 4K-with-audio variations now" → iterate cheap first; 4K/audio is costly + async.
- "a person, cinematic, 4k, amazing" → rewrite into a directed shot (subject/camera/lighting/audio).
- Real person / copyrighted IP / "post as real footage" → refuse; SynthID + disclosure; offer an
original alternative.
- "Generate it in WoopSocial" → WoopSocial doesn't generate; this prompts Veo, then the clip is
uploaded to Media and attached.
Related
tools/integrations/veo.md — API/model IDs, async generate-and-poll, pricing, the upload-to-WoopSocial flow.
reels-script (veo-prompt-pack) — the consuming skill; nano-banana — the image sibling + image-to-video source.
brand-profile — the visual brand; hook-writer — the in-clip hook/line; heygen — avatar/talking-head alternative.
ai-video — the router above this skill; kling (4K/multi-shot/motion) and luma (HDR/mood, silent) — generative siblings.
scheduling-and-queue — attach the video to a post and publish.
References
references/when-and-how-to-prompt.md — when to reach for Veo vs other tools, and the shot/prompt anatomy.
references/audio-and-camera.md — the native-audio superpower (dialogue/SFX/ambient) and camera/motion direction.
references/inputs-length-and-recipes.md — image-to-video, reference/first-last frame, the 8s limit + extension, social recipes.
references/examples.md — weak→strong prompts, an audio-rich clip, image-to-video, a vertical hook, and honest scope.
1---2name: veo-33description: Use to write great prompts for Google Veo (Veo 3.1) to generate or animate video for social media — the video-prompt-craft mini-skill, the video counterpart to nano-banana. Run when the user wants a Veo / AI video prompt, a short video clip for a post (b-roll, product-in-motion, hook visual, spokesperson/UGC clip, ad), to animate a still image into video, or a vertical Reel/TikTok/Short clip. Reads brand-profile for brand style. Veo's standout is native synchronized audio in one pass, plus image-to-video and native 9:16 vertical. Teaches the prompt anatomy, audio prompting, the image-to-video pipeline, and the 8-second constraint + extension/stitching. Honest: iterate cheap then finalize, disclose AI video (SynthID watermark), never generate real identifiable people or copyrighted IP. This is the prompt-craft layer — the API/connection and the generate -> upload to WoopSocial Media -> attach flow live in tools/integrations/veo.md; the consuming pack is reels-script's veo-prompt-pack.4license: MIT5---6
7# Veo (video prompt craft)
8
9Write prompts that get cinematic clips out of **Veo** — Google's video model (**Veo 3.1**; Fast/Lite
10tiers for cheap iteration). This is the **prompt-craft** layer of a three-layer setup, and the **video
11counterpart to `nano-banana`**:
12
13- **Connection/API** (model IDs, async generate-and-poll, the *generate → upload to WoopSocial Media →
14 attach* flow) → `tools/integrations/veo.md`.
15- **Prompt craft** (this skill) → how to direct the right clip well.
16- **In-skill application** → `reels-script`'s veo-prompt-pack.
17
18> Fast-moving area — re-verify model names/specs quarterly.
19
20## Reach for Veo when… (match the job)
21
22Its real strengths: **native synchronized audio** (dialogue + lip-sync, SFX, ambient, music — the
23differentiator), **cinematic realism/physics**, **image-to-video**, **reference-image consistency**,
24**first/last-frame** control, and **native 9:16 vertical / 4K**. Reach **elsewhere** for a **pure
25talking-head explainer** (→ avatar tool like `heygen`), **native-4K/multi-shot/motion-transfer at
26lower cost** (→ `kling`), **HDR/atmospheric mood shots** (→ `luma`), **long continuous video**
27(stitch/extend or a length-built tool), or **highly stylized/very-high-volume** output. Don't
28default it to every job. (Details: `references/when-and-how-to-prompt.md`.)
29
30## Step 0 — Read the brand + the job
31
32Load `brand-profile.md` (visual style, palette, tone). Identify the **job** (b-roll / product motion /
33hook visual / spokesperson / ad / animate-a-still) and the **aspect ratio** (9:16 for social).
34
35## Step 1 — Direct the shot (natural language, not quality spam)
36
37Describe a shot like a director: **subject · action · scene · camera (type/movement/angle/lens) ·
38lighting/mood · style · audio · timing**. One clear motion per clip. Set **9:16**. Ground in the brand.
39See `references/when-and-how-to-prompt.md`.
40
41## Step 2 — Use the superpower: audio
42
43Veo generates **native synchronized audio** in one pass — describe it explicitly: **dialogue in
44quotes** (+ who/tone), **SFX**, **ambient**, **music** mood. Treat generated audio as a **draft/guide
45track** for branded work (record real VO / license music for final) and **verify** lip-sync. Direct
46the **camera/motion** for the cinematic feel. See `references/audio-and-camera.md`.
47
48## Step 3 — Inputs, length, recipes
49
50- **Image-to-video:** animate a still (e.g. a `nano-banana` frame) — it becomes the first frame; you
51 describe motion + audio. Use **reference images** / **first-and-last frame** for consistency/control.
52- **Length:** one generation ≈ **8 seconds** (4/6/8); **hook in the first second**; for longer use
53 **scene-extension or stitch** clips. Plan short beats.
54- Pick a **recipe** for the job (b-roll, product motion, hook, spokesperson, ad, animate-a-still). See
55 `references/inputs-length-and-recipes.md`.
56
57## Step 4 — Iterate cheaply, then verify, disclose, ship
58
59- **Iterate at low-res / Fast** to lock the prompt; **finalize the keeper at high-res/4K** (4K costs
60 ~40–60% more time/$). Off-prompt generations still consume credits — **prompt skill is the cost
61 lever**; video is **async**.
62- **Verify** every clip (artifacts, lip-sync, physics) before publishing.
63- **Disclose** AI video per platform/region (EU AI Act; TikTok auto-disclosure); every clip carries a
64 **SynthID** watermark — don't pass it off as real footage.
65- **Ship:** generate per the integration guide → **upload to WoopSocial Media → attach** via
66 `scheduling-and-queue`. WoopSocial doesn't generate video.
67
68## Quality bar — self-check
69
70- Did I **match the tool to the job** (and route talking-heads/long-form elsewhere)?
71- Is the prompt a **directed shot** (subject/action/camera/lighting/style), **brand-grounded**, **9:16**,
72 with an **explicit audio cue**?
73- Did I plan for **8-second** clips (hook first) and **iterate cheap → finalize**?
74- For stills, did I use **image-to-video** (and reference/first-last frame where useful)?
75- Did I handle **SynthID + disclosure**, **audio-as-draft**, **verify the output**, and **refuse real
76 people / IP**?
77- Did I point to **`tools/integrations/veo.md`** for the API + WoopSocial flow (no claim WoopSocial
78 generates video)?
79
80## Edge cases & pushback
81
82- **Talking-head explainer / lots of dialogue** → suggest an avatar tool (`heygen`); don't force Veo.
83- **"Make a 40-second video"** → ~8s per gen; scene-extend/stitch; plan short beats.
84- **"Generate 10 final 4K-with-audio variations now"** → iterate cheap first; 4K/audio is costly + async.
85- **"a person, cinematic, 4k, amazing"** → rewrite into a directed shot (subject/camera/lighting/audio).
86- **Real person / copyrighted IP / "post as real footage"** → refuse; SynthID + disclosure; offer an
87 original alternative.
88- **"Generate it in WoopSocial"** → WoopSocial doesn't generate; this prompts Veo, then the clip is
89 uploaded to Media and attached.
90
91## Related
92
93- `tools/integrations/veo.md` — API/model IDs, async generate-and-poll, pricing, the upload-to-WoopSocial flow.
94- `reels-script` (veo-prompt-pack) — the consuming skill; `nano-banana` — the image sibling + image-to-video source.
95- `brand-profile` — the visual brand; `hook-writer` — the in-clip hook/line; `heygen` — avatar/talking-head alternative.
96- `ai-video` — the router above this skill; `kling` (4K/multi-shot/motion) and `luma` (HDR/mood, silent) — generative siblings.
97- `scheduling-and-queue` — attach the video to a post and publish.
98
99## References
100
101- `references/when-and-how-to-prompt.md` — when to reach for Veo vs other tools, and the shot/prompt anatomy.
102- `references/audio-and-camera.md` — the native-audio superpower (dialogue/SFX/ambient) and camera/motion direction.
103- `references/inputs-length-and-recipes.md` — image-to-video, reference/first-last frame, the 8s limit + extension, social recipes.
104- `references/examples.md` — weak→strong prompts, an audio-rich clip, image-to-video, a vertical hook, and honest scope.