scripting-and-storyboarding
The pre-production system — start from the spine, write both columns, envision every beat, number the
shots, and edit on paper first. The agent writes the plan; the human shoots and judges; WoopSocial
publishes the finished video. (Craft skill — no tool file.)
The POV: videos are won in pre-production — the cheapest edit is on paper
Chaotic shoots and endless edits are a pre-production deficit, not a talent deficit: the plan is "where you
catch problems before they become expensive fixes on set or in post." Four truths carry this skill. (1) A
words-only script plans half the video — the visual half gets improvised, and improvisation defaults to a
static talking face; the two-column AV script (audio | visual) writes what's seen, beat by beat, so the
change-every-few-seconds rhythm is planned, not hoped for. (2) The storyboard is a decision document, not
art — its only question is "what's on screen at this beat?", answered by a shot table, stick figures, or AI
previz frames (2026 pipelines make character-consistent boards in minutes — the tested benchmark: 12 frames
from ~3 hours to 40–60 minutes, with explicit camera intent and a character bible as the craft rules).
(3) The shot list regrouped BY SETUP is the batching trick — two weeks of content in one afternoon comes
from sorting shots by location/framing/outfit, never by video. (4) Runtime math doesn't negotiate —
~130–150 wpm means a 300-word script is a 2-minute video; the stopwatch pass and the paper cut happen before
the shoot. And the 2026 twist: AI video sequences need more pre-production, not less — every panel becomes
a generation's shot brief, and previz frames cost cents where generations cost credits.
Read these first
- idea-generation-and-ideation — the idea being productionized.
- short-form-video-script (or the long-form skill) — the retention spine this turns into production.
- brand-profile + design-and-templates — visual language, end cards.
The framework: SCENE
(Depth: references/the-scene-framework.md.)
- S — Start from the spine: one job, audience, format target, beat outline — before any words.
- C — Columns: audio + visual: the two-column AV script with [VO]/[SFX]/[Action] tags; if the visual isn't
written, it will be improvised; no "me talking" twice in a row.
- E — Envision every beat: the cheapest board that answers "what's on screen" — shot table → stick figures
→ AI previz frames (explicit camera intent, character bible + ref image, micro-beats, directional not
final); skip boards honestly where a shot list suffices; for AI video, each panel = the shot brief.
- N — Number the shots: shot list with setups → regroup by setup for batching (sequence setups, buffer
takes on hooks, label footage
vid#-shot#); runtime math + feasibility pass.
- E — Edit on paper first: table read with a stopwatch, cut the middle not the spine, the honesty pass
(no staged-as-candid, no lifted scripts) — then hand off to the shoot, the edit, and WoopSocial.
The reality (verify-quarterly)
Stable craft: the two-column AV script (broadcast standard), ~130–150 wpm runtime math, the
storyboard-as-decision-document, the by-setup batch regroup. The 2026 AI layer (attribute): script→board
pipelines matured (Boords — free tier, sign-off layer; Storyboarder.ai — 250K+ creators, animatics; LTX,
mStudio, Studiovity; free Wonderunit); character consistency is now baseline; the practitioner benchmark cut
12-frame boards from ~3 hrs to 40–60 min using LLM-beats → visual directives → seeded frames — with tags
([VO]/[SFX]/[Action]), explicit camera intent ("the AI will guess — don't make it"), locked palette, and a
character bible; boards stay directional; traditional illustration runs ~$50–300/frame (the economics behind
the shift). AI-video sequences: panel = shot brief (luma's DREAM); board first, generate second. Attribute
all; verify-quarterly. Full detail: references/scripting-and-storyboarding-2026-reality.md; the templates
and two worked examples: references/templates-and-examples.md.
Honest scope (never violate)
- The agent writes every planning artifact (beats, AV script, board/briefs, shot list, batch plan, timing
pass, footage map); the human shoots/generates, judges every take, and approves — the agent cannot see
footage or operate a camera and never fabricates "that take works." AI previz follows the image rules
(no unpermitted likeness; original/consented characters; disclosure if frames publish).
- Production integrity: no staged-as-candid content (actors posing as unaffiliated strangers = deceptive
endorsement; skits are fine disclosed); no lifted scripts (structure study yes, verbatim no); honest
runtime/feasibility. WoopSocial publishes the finished video; it does not script, storyboard, shoot, or
edit. (Full scope:
references/scope-and-connections.md.)
Distinct from its siblings (route correctly)
scripting-and-storyboarding (this) = the pre-production system · short-form-video-script = the
retention-words craft (WATCH writes the spine; SCENE productionizes it) · talking-head-and-piece-to-camera
= the on-camera delivery · capcut / descript = the edit executing the AV script's visual plan · luma /
ai-video = the generations whose shot briefs the panels become · flux / image-prompt = previz frames +
reference stills · storytelling-and-narrative = the narrative angle/WHAT this schedules into shots ·
idea-generation-and-ideation = supplies the idea.
Where this connects
Reads first: idea-generation-and-ideation + short-form-video-script/long-form + brand-profile +
design-and-templates. Feeds: the shoot day (the human), talking-head-and-piece-to-camera, luma
(shot briefs), capcut/descript (footage map + AV script), content-calendar (the batch). Publishes via:
the finished video → scheduling-and-queue → WoopSocial. Measure with: shoot efficiency + edit speed +
published retention via analytics-and-reporting — never fabricated.
Definition of done
A shootable, editable plan: the spine confirmed (one job, beat outline from the script skill), the two-column
AV script written with every visual beat specified ([VO]/[SFX]/[Action] tags; no "me talking" twice in a row),
the storyboard produced at the cheapest sufficient fidelity (shot table, stick figures, or AI previz frames
with explicit camera intent + a character bible — directional, never final art; skipped honestly where a shot
list suffices), a numbered shot list regrouped by setup with the batch plan, buffer takes, and a labeled
footage map, the runtime verified by stopwatch math (~130–150 wpm) and cut on paper before the shoot, and the
honesty pass held (no staged-as-candid, no lifted scripts, feasibility stated straight); AI-video panels
written as generation shot briefs with previz-before-credits economics; the human shooting and judging,
the edit receiving a clean handoff, and the finished video publishing via WoopSocial; nothing staged as
real, nothing plagiarized, no fabricated production claims; and correctly distinguished from
short-form-video-script, talking-head-and-piece-to-camera, capcut/descript, and luma.
1---2name: scripting-and-storyboarding3description: The pre-production system — turn a video idea into a shootable, editable plan: the two-column AV script, the storyboard as a decision document, the numbered shot list, the batch-shoot plan, and the paper edit. Use when shoots are chaotic, edits drag, videos feel static ("just me talking"), someone wants to batch-film a week of content in one session, needs a storyboard but can't draw, is planning a multi-shot AI video sequence, or scripts keep running long. Uses the SCENE framework. Reads idea-generation-and-ideation + short-form-video-script or the long-form skill + brand-profile first. A words-only script plans half the video; the storyboard is a decision document, not art; the shot list grouped by setup enables batch shooting. The agent writes the plan; the HUMAN shoots/generates and judges; WoopSocial publishes. Never stages candid-as-real moments or lifts another creator's script. Distinct from short-form-video-script, talking-head-and-piece-to-camera, capcut/descript, luma.4---5
6# scripting-and-storyboarding
7
8The **pre-production system** — start from the spine, write both columns, envision every beat, number the
9shots, and edit on paper first. The agent writes the plan; the **human shoots and judges**; **WoopSocial
10publishes** the finished video. (Craft skill — no tool file.)
11
12## The POV: videos are won in pre-production — the cheapest edit is on paper
13Chaotic shoots and endless edits are a pre-production deficit, not a talent deficit: the plan is "where you
14catch problems before they become expensive fixes on set or in post." Four truths carry this skill. **(1) A
15words-only script plans half the video** — the visual half gets improvised, and improvisation defaults to a
16static talking face; the **two-column AV script** (audio | visual) writes what's *seen*, beat by beat, so the
17change-every-few-seconds rhythm is planned, not hoped for. **(2) The storyboard is a decision document, not
18art** — its only question is "what's on screen at this beat?", answered by a shot table, stick figures, or AI
19previz frames (2026 pipelines make character-consistent boards in minutes — the tested benchmark: 12 frames
20from ~3 hours to 40–60 minutes, with explicit camera intent and a character bible as the craft rules).
21**(3) The shot list regrouped BY SETUP is the batching trick** — two weeks of content in one afternoon comes
22from sorting shots by location/framing/outfit, never by video. **(4) Runtime math doesn't negotiate** —
23~130–150 wpm means a 300-word script is a 2-minute video; the stopwatch pass and the paper cut happen *before*
24the shoot. And the 2026 twist: **AI video sequences need more pre-production, not less** — every panel becomes
25a generation's shot brief, and previz frames cost cents where generations cost credits.
26
27## Read these first
281. **idea-generation-and-ideation** — the idea being productionized.
292. **short-form-video-script** (or the long-form skill) — the retention spine this turns into production.
303. **brand-profile** + **design-and-templates** — visual language, end cards.
31
32## The framework: SCENE
33(Depth: `references/the-scene-framework.md`.)
34- **S — Start from the spine:** one job, audience, format target, beat outline — before any words.
35- **C — Columns: audio + visual:** the two-column AV script with [VO]/[SFX]/[Action] tags; if the visual isn't
36 written, it will be improvised; no "me talking" twice in a row.
37- **E — Envision every beat:** the cheapest board that answers "what's on screen" — shot table → stick figures
38 → AI previz frames (explicit camera intent, character bible + ref image, micro-beats, directional not
39 final); skip boards honestly where a shot list suffices; for AI video, each panel = the shot brief.
40- **N — Number the shots:** shot list with setups → **regroup by setup** for batching (sequence setups, buffer
41 takes on hooks, label footage `vid#-shot#`); runtime math + feasibility pass.
42- **E — Edit on paper first:** table read with a stopwatch, cut the middle not the spine, the honesty pass
43 (no staged-as-candid, no lifted scripts) — then hand off to the shoot, the edit, and WoopSocial.
44
45## The reality (verify-quarterly)
46Stable craft: the two-column AV script (broadcast standard), ~130–150 wpm runtime math, the
47storyboard-as-decision-document, the by-setup batch regroup. The 2026 AI layer (attribute): script→board
48pipelines matured (Boords — free tier, sign-off layer; Storyboarder.ai — 250K+ creators, animatics; LTX,
49mStudio, Studiovity; free Wonderunit); character consistency is now baseline; the practitioner benchmark cut
5012-frame boards from ~3 hrs to 40–60 min using LLM-beats → visual directives → seeded frames — with tags
51([VO]/[SFX]/[Action]), explicit camera intent (*"the AI will guess — don't make it"*), locked palette, and a
52character bible; boards stay directional; traditional illustration runs ~$50–300/frame (the economics behind
53the shift). AI-video sequences: panel = shot brief (luma's DREAM); board first, generate second. **Attribute
54all; verify-quarterly.** Full detail: `references/scripting-and-storyboarding-2026-reality.md`; the templates
55and two worked examples: `references/templates-and-examples.md`.
56
57## Honest scope (never violate)
58- **The agent** writes every planning artifact (beats, AV script, board/briefs, shot list, batch plan, timing
59 pass, footage map); the **human** shoots/generates, judges every take, and approves — the agent cannot see
60 footage or operate a camera and never fabricates "that take works." **AI previz** follows the image rules
61 (no unpermitted likeness; original/consented characters; disclosure if frames publish).
62- **Production integrity:** no staged-as-candid content (actors posing as unaffiliated strangers = deceptive
63 endorsement; skits are fine **disclosed**); no lifted scripts (structure study yes, verbatim no); honest
64 runtime/feasibility. **WoopSocial publishes** the finished video; it does not script, storyboard, shoot, or
65 edit. (Full scope: `references/scope-and-connections.md`.)
66
67## Distinct from its siblings (route correctly)
68**scripting-and-storyboarding (this)** = the pre-production system · **short-form-video-script** = the
69retention-words craft (WATCH writes the spine; SCENE productionizes it) · **talking-head-and-piece-to-camera**
70= the on-camera delivery · **capcut / descript** = the edit executing the AV script's visual plan · **luma /
71ai-video** = the generations whose shot briefs the panels become · **flux / image-prompt** = previz frames +
72reference stills · **storytelling-and-narrative** = the narrative angle/WHAT this schedules into shots ·
73**idea-generation-and-ideation** = supplies the idea.
74
75## Where this connects
76Reads first: **idea-generation-and-ideation** + **short-form-video-script**/long-form + **brand-profile** +
77**design-and-templates.** Feeds: the shoot day (the human), **talking-head-and-piece-to-camera**, **luma**
78(shot briefs), **capcut/descript** (footage map + AV script), **content-calendar** (the batch). Publishes via:
79the finished video → **scheduling-and-queue → WoopSocial.** Measure with: shoot efficiency + edit speed +
80published retention via **analytics-and-reporting** — never fabricated.
81
82## Definition of done
83A shootable, editable plan: the spine confirmed (one job, beat outline from the script skill), the two-column
84AV script written with every visual beat specified ([VO]/[SFX]/[Action] tags; no "me talking" twice in a row),
85the storyboard produced at the cheapest sufficient fidelity (shot table, stick figures, or AI previz frames
86with explicit camera intent + a character bible — directional, never final art; skipped honestly where a shot
87list suffices), a numbered shot list regrouped by setup with the batch plan, buffer takes, and a labeled
88footage map, the runtime verified by stopwatch math (~130–150 wpm) and cut on paper before the shoot, and the
89honesty pass held (no staged-as-candid, no lifted scripts, feasibility stated straight); AI-video panels
90written as generation shot briefs with previz-before-credits economics; the human shooting and judging,
91the edit receiving a clean handoff, and the finished video publishing via WoopSocial; **nothing staged as
92real, nothing plagiarized, no fabricated production claims**; and correctly distinguished from
93short-form-video-script, talking-head-and-piece-to-camera, capcut/descript, and luma.