Remotion Video
Turn a PR into a rendered short-form promo video with a hook, code moments, and CTA, using Remotion. Iterate with live preview in Remotion Studio, then render video.mp4 and poster.jpg for X/LinkedIn/social.
Unlike /video-script (which produces a textual script), this skill produces the actual video file.
Triggers
Invoke when the user says: "make a video for this PR", "remotion video", "render a promo video", "video for X/LinkedIn", or when selected from /marketing-pipeline.
Process Flow
Resolve input
→ Phase 1: Discovery
→ Phase 2: Configuration
→ Phase 3: Narrative planning
→ Phase 4: Scaffold
→ Phase 5: First draft + iterate
→ Phase 6: Render
→ Phase 7: Cleanup
Phases 1, 3, 5, and 7 have explicit approval gates. Phase 2 is an interactive Q&A. Phase 5 is a freeform iteration loop that can run many rounds.
"Use Sane Defaults" / "Don't Ask Questions" — What It Does and Doesn't Override
When the user invokes the skill with phrasing like "use sane defaults", "don't ask questions", "non-interactive", "just ship it", or any equivalent — interpret it precisely:
It DOES override (skip the prompt, pick the default):
- Q2.1 duration → derive from scope using the ladder in Q2.1 (still in the 15–60s window); pick the midpoint of the matched scope band and proceed without asking
- Q2.2 aspect ratio → 16:9 landscape
- Q2.3 project location →
marketing/<feature-slug>/remotion/ - Q2.4 brand confirmation (the "use these / customize / provide your own?" question)
- Phase 1.4 scope confirmation
- Phase 3.0 motif confirmation
- Phase 3.1 story-pattern confirmation
- Phase 3.3 scene-plan approval
- Phase 5 freeform iteration loop (the "what would you like to change?" prompt)
It does NOT override (must always run regardless):
- Brand color/font/logo scanning (Q2.4 detection — see HARD-GATE in Q2.4). Hardcoding colors from training-data assumptions about a project is a forbidden shortcut.
- Phase 5 preview (Remotion Studio must start and the studio URL must be opened in the browser before the render runs). Even in fully unattended mode, the user can interrupt; the agent must not pre-decide for them.
- Phase 6 pre-render audits (storytelling, hook rules, motif presence, pacing variance, value-prop timing, contrast). These exist to prevent shipping a generic video.
- Phase 7 cleanup question (the user owns project disposition).
If you're tempted to skip a HARD-GATE because the user "said no questions" — re-read this section. The user said no questions, not no gates.
Input Resolution
Resolve the argument (if provided) in this order:
- Path to a marketing brief (
.mdcontaining "Executive Summary" or "Key Messages") → marketing brief - Path to a blog post (
.mdwith blog post structure) → blog post - Path to a changelog → changelog
- GitHub PR URL or
#\d+pattern → PR - Matches
<ref>..<ref>or<ref>...<ref>(alphanumeric +/,_,.,-on each side) → git ref range - Resolves to an existing file/directory → codebase feature
- Otherwise → freeform text
If no argument is provided, ask: "What should the video be about? You can provide a PR URL/number, marketing brief, blog post, changelog, git ref range, file/directory path, or just describe the feature."
When invoked from the pipeline with a PR and upstream marketing-brief/blog-post paths, read both: PR for technical accuracy, upstream content for positioning/tone.
(Detailed phase specs begin below — see Phase 1.)
Phase 1: Discovery
Step 1.1 — Check remotion-best-practices availability
Attempt to invoke the remotion-best-practices skill via the Skill tool. If unavailable, present:
"The
remotion-best-practicesskill isn't installed. Options: a) proceed with baseline Remotion knowledge (quality may be reduced) b) wait while you install it c) cancel"
If the user picks (a), emit a warning in the final summary noting reduced quality.
Step 1.2 — Analyze input
| Input type | What to read |
|---|---|
| Marketing brief | positioning, key messages, audience |
| Blog post | headline, narrative, examples |
| Changelog | highest-impact entry |
| PR | gh pr view <n> --json title,body,files,labels, diff (gh pr diff), commit messages, linked issues |
| Git refs | git diff <range> + git log <range> --oneline |
| Codebase path | read the specified files/directories |
| Freeform | parse the user's description |
For PRs with 20+ files, filter to user-facing changes only — skip tests/, ci/, .github/, lockfile changes, dep bumps.
Error handling:
ghnot available → tell the user, ask for an alternative (diff file, freeform description)- Invalid PR/ref → ask user to verify
- File not found → ask for the correct path
Step 1.2a — PR deep analysis (when input is a PR)
When the input is a PR, the brief table in Step 1.2 is not enough. PR bodies are routinely vague, outdated, or focused on implementation rather than user value, and a video built off the body alone tends to overclaim or miss the headline angle entirely. Before continuing to Step 1.3, run this structured analysis and produce a written PR analysis block that becomes the source of truth for Steps 1.4 (scope confirmation), Q2.1 (duration ladder), and Phase 3 (narrative planning).
Mandatory steps — do all of them:
Pull metadata and identify the base branch.
gh pr view <n> --json number,title,body,baseRefName,headRefName,files,labels,commits,additions,deletionsRecord
baseRefName(usuallymain/master/develop) — that is the diff reference.gh pr diff <n>automatically compares the PR head against this base.Pull the full diff against the base.
gh pr diff <n>For very large PRs (>500 lines or >20 files), also list changed files via
gh pr view <n> --json files.Filter to user-facing surfaces. Ignore (do not let these shape the headline angle):
tests/,__tests__/,*.test.*,*.spec.*,__mocks__/,ci/,.github/, lockfiles, dep bumps without behavior change, generated/build artifacts (dist/,build/,.next/), and formatting-only diffs. What remains is the user-facing surface of the PR.Enumerate the public-API delta. From the filtered diff, list every change a user could observe or write code against:
- New, renamed, or removed
exports (functions, types, components, hooks, classes, constants) - New CLI commands, flags, or environment variables
- New routes, endpoints, or event names
- New config keys or schema fields
- Changed default values for existing public surfaces
- Changed error messages, log shapes, or response shapes the user could rely on
Grep the diff for added lines beginning with
export, new top-levelfunction/class/constinsrc//lib//packages/*/src/, new files in those trees, and changes to public type signatures. If you cannot point at a line in the diff for a claimed API, the claim is wrong — drop it.- New, renamed, or removed
Enumerate the behavior delta. Beyond API surfaces, list user-visible behavior changes: UI elements added/changed (with file references), CLI / network output that looks different, side effects firing under new conditions, removed limitations, performance characteristics that changed.
Write the before / after value statement. Exactly two sentences, both grounded in concrete diff evidence:
- Before: "Before this PR, a user who wanted to ___ had to ___."
- After: "After this PR, the same user can ___."
The "had to" half must be a real prior workflow you can describe — copy-pasting an adapter from the docs, installing a second package, writing boilerplate, hitting an error, switching to a different tool. If the honest "Before" sentence is "they could already do this", the PR has no user-visible value delta — see step 8.
Cross-check the PR body against the diff. Walk the body's claims and the diff side-by-side:
- Body claims a feature → does the diff confirm? If the body says "added X" and the diff has no public surface for X (only tests, only docs, only internal helpers), flag it: "PR body claims X but the diff doesn't expose X to users — should I shift the angle to ?"
- Diff shows multiple distinct features → body emphasizes one? Flag the others as candidate angles: "The body emphasizes A, but the diff also adds B and C. Which is the headline angle?"
- Body is empty / boilerplate /
[BLANK]? Infer from the diff; make the inferred angle explicit and confirm with the user at Step 1.4. - List "surprises" — anything in the diff not mentioned in the body that affects user-visible behavior. Surprises are often the real story.
Bail-out check: is this PR actually user-visible? If after steps 4–6 you cannot name a single new thing the user can do or observe, the PR is not video-worthy as a feature launch. Stop and ask the user:
"This PR looks like an internal refactor / test-only / dep-bump PR — I can't find a user-visible value delta. Options: a) shift to a performance / DX / cleanup angle if numbers support it b) pick a different PR or input c) cancel"
Do not invent a feature angle to fill the gap. A confabulated angle wastes a full iteration round and destroys user trust on the very first draft.
Produce the PR analysis block. Write the result as a short structured block before continuing to Step 1.3. This block — not the PR body — is the source of truth for everything downstream:
PR analysis — #1234 "Add fromZodSchema support" Base: main · Head: feature/zod-schema · Author: <login> User-facing files: 3 (packages/core/src/index.ts, packages/core/src/zod.ts, packages/core/src/types.ts) Filtered out: 5 test files, 2 doc files, lockfile Public-API delta: + export function fromZodSchema(schema: ZodSchema): StandardSchema + export type ZodCompatibleSchema ~ default error code for ZodError changed: 'invalid_type' → 'STANDARD/type' Behavior delta: - Users importing zod schemas no longer need a manual adapter. - Error messages from zod paths now use the StandardError shape. Before / after: Before: copy a 15-line adapter from the docs into every project that mixes zod with this library. After: one import + one call. Surprises (in diff but not in PR body): - Default error shape change (above) — could be the headline angle for migration-aware audiences. Headline angle: "drop the 15-line adapter — one import, one call" Scope band for Q2.1: one idea (15–25s)The "Headline angle" line feeds Phase 3.0's motif derivation and Phase 3.3's hook copy. The "Scope band" line feeds Q2.1's scope-derived duration proposal.
Forbidden shortcuts (each is a fast path to a wrong video):
- Writing the analysis block from the PR title alone — fail; you must read the diff.
- Skipping the cross-check because the body "looks complete" — bodies often look complete and are wrong.
- Treating the PR body as truth when it conflicts with the diff — the diff wins.
- Filling in a plausible "Before" sentence when the diff doesn't support one — run the bail-out check instead.
- Collapsing two genuinely distinct user-visible features into one "feature X" bullet to make the scope band look smaller — record both and pick a headline angle at Step 1.4.
For git-ref-range and codebase-path inputs, apply the same structure with the diff coming from git diff <range> / direct file reads in place of gh pr diff. Steps 4–9 (public-API delta → analysis block) are not PR-specific.
Step 1.3 — Read product context
Read if they exist: README.md, docs/, package.json. If nothing found, ask: "Can you briefly describe the product and who it's for?"
Step 1.4 — Present understanding + get scope confirmation
For PR inputs, the bullets and the compelling angle below must be derived from the PR analysis block written in Step 1.2a — not from the PR title or body. Quote the "Headline angle" line directly as the proposed story seed; turn the public-API delta and the before/after statement into the bullets. If the user added clarifying context after Step 1.2a, fold it in here.
For non-PR inputs, derive the bullets from the relevant entry in Step 1.2 (marketing brief positioning, blog headline, changelog highlight, codebase reading, freeform description).
Present:
"Here's what I'll base the video on:
- [feature summary bullet 1 — concrete, grounded in the diff / source]
- [feature summary bullet 2 — the before → after value, in one line]
The compelling angle: [proposed story seed — for PRs, the "Headline angle" from the analysis block]
Anything to add, remove, or correct?"
Do not proceed until the user confirms.
Step 1.5 — Sensitive content scan
Before continuing, scan for: security patches, internal pricing, credentials, unreleased roadmap items, content marked confidential. Flag anything questionable to the user.
Phase 2: Configuration
Ask these questions one at a time, in order.
Q2.1 — Duration (derived from scope, never offered as a menu)
Do not ask the user to pick a fixed length from a menu. A fixed number becomes a constraint, and the dominant failure mode is padding — freeze-frames, repeated beats, black or empty trailing frames, or filler bullets added solely to reach the chosen target. Instead, derive a duration from the scope of the change and present it as a proposal the user can confirm or override.
Allowed range: 15–60 seconds. Anything outside this window is wrong by default. Going under 15s means the story can't breathe; going over 60s means it's two videos.
Scope → duration ladder:
| Scope of the PR / feature | Distinct payoff beats | Target |
|---|---|---|
| One idea (single API addition, single bug fix, one QoL win) | hook + 1 delivery + CTA | 15–25s |
| Typical PR-sized feature (problem → solution → proof, or a multi-chapter code walk) | hook + 2–3 delivery + CTA | 25–40s |
| Multi-faceted release (multiple distinct sub-features, or comparison needing problem + solution + proof beats) | hook + 3–4 delivery + CTA | 40–60s |
A distinct payoff beat is one new thing the viewer learns. Two scenes whose payoff sentences (see Phase 3.3 item 1) reduce to the same idea are one beat, not two.
Sanity check before proposing — do this silently first:
- List the distinct payoff beats the video must contain.
- Estimate: hook ≈ 3s, CTA ≈ 5–7s, each delivery beat ≈ 6–10s (longer if it contains a code chapter ladder).
- Sum the floor and the ceiling. If the floor is under 15s, you're padding the beat list — cut beats or shrink the scope claim. If the ceiling is over 60s, you have two videos — pick one angle.
Propose to the user (no menu, no fixed lengths):
"Based on the scope of this PR, I'm targeting ~Xs (range Ys–Zs). Beats: . Confirm, or override with a different length anywhere in 15–60s."
Hard rule — story sets duration, never the other way around. If at any later phase a scene needs to be stretched, held on a static frame, repeated, backed by black/empty frames, or filled with filler bullets to reach the chosen target — stop, shorten the target, re-confirm with the user, and ship the shorter video. Padding to hit a number is the single behavior this rule exists to forbid. If a 30s scope honestly tells in 18s, ship 18s.
Breathing-room constraint. Whatever duration is chosen, it must allow every on-screen text element to (a) finish animating in, (b) dwell long enough to be read, and (c) settle for at least ~0.4s before the next scene begins. Cutting to a new scene the instant a line of text finishes appearing is forbidden. See Phase 3.3 (item 3) and the Phase 6 audit for the concrete dwell-time table.
Q2.2 — Aspect ratio
"Which aspect ratio?
- 16:9 landscape (1920×1080, default) — desktop X/LinkedIn
- 1:1 square (1080×1080) — mobile-friendly feed
- 9:16 vertical (1080×1920) — Reels/Shorts/TikTok
- Multi-format — render all three from the same story"
Frame rate is fixed at 30fps. The audit thresholds throughout this skill are expressed in absolute frame counts assuming 30fps; do NOT change fps in remotion.config.ts without simultaneously updating every frame-based threshold (Phase 6 pre-render audit, the dwell-time table, the breathing-room rules, chapters-required threshold).
Q2.3 — Remotion project location
"Where should the Remotion project live? Default:
marketing/<feature-slug>/remotion/(fresh per-video) Override: specify a path."
Q2.4 — Brand assets (auto-detect → confirm)
You MUST run the heuristics in brand-detection.md against the actual target repository — the source code that owns the feature, not the marketing/ output directory — before selecting any color, font, or logo.
Forbidden shortcuts (these produce wrong colors and waste an iteration):
- "I know this project — TanStack uses amber, Vercel uses black/white, Stripe uses purple" → No. Run the scan. Recall is unreliable; brand details drift between training data and now.
- "It's a dev tool, dark + neon green is fine" → No. Generic vibes ≠ this product's brand.
- "User said no questions, so I'll skip detection" → No. Detection is silent. Confirmation is what the "no questions" instruction skips.
- Picking from a palette in your head because it "fits the topic" → No. Read the repo's CSS/Tailwind/theme files.
The scan must produce a written record before any composition file is written. Output a short block listing, for each field: the source file checked, the value found (or not found), and the final value used. Example:
Brand scan — TanStack/ai
Primary : checked tailwind.config.* (none) · packages/*/styles.css (--brand: #0a3d2e) → #0a3d2e
Accent : checked theme.json (none) · brand.json (none) · derived from primary → #14b870
Logo : checked public/logo.svg → public/logo.svg (will be referenced from src/brand.ts)
Font : checked next/font (none) · @fontsource (none) · README ref → Inter (fallback)
If the scan finds nothing for a field, fall through to the neutral defaults below — but only after the scan ran and is recorded. Skipping the scan and going straight to defaults is the failure mode this gate exists to prevent.
If you cannot find a logo, ask the user — do NOT invent one. The same rule applies to any field where detection genuinely turned up nothing and no sensible neutral default exists.
The following are non-negotiable regardless of "no questions" mode:
- Brand identity (colors, font, logo) — detection always runs
- Asset locations (where the logo and any referenced media live)
- File output destination (the
marketing/<feature-slug>/remotion/path or override) - Target audience (carried in from Phase 1's PR/brief analysis)
In interactive mode, present findings (the block above) and ask for confirmation. In non-interactive / "use sane defaults" mode, print the same block and proceed without asking — the record is required either way.
Present findings as:
"I found:
- Logo:
public/logo.svg- Primary:
#0066ff(fromtailwind.config.js)- Font:
Inter(fromnext/font)Use these, customize some, or provide your own?"
Persistence: write chosen brand to marketing/<feature-slug>/remotion/.marketing/brand.json (mirrors the hyperframes-video location). On subsequent runs, ask:
"I loaded brand settings from
marketing/<feature-slug>/remotion/.marketing/brand.json. Use saved, or re-detect?"
Fallback when nothing detected: ask explicitly with these neutral defaults. These are intentionally neutral — auto-detection of the project's actual brand is always preferred, and these values should only appear when detection turns up nothing.
- Primary:
#3B82F6(neutral blue) - Accent:
#8B5CF6(neutral violet) - Background:
#0A0A0A(near-black, dark mode) - Text:
#FFFFFF(white) - Muted:
#9CA3AF - Success:
#22C55E - Danger:
#EF4444 - Font:
Geist(loaded automatically via@remotion/google-fonts) - Logo: none
Confirm with the user before scaffolding.
Note on fonts: The default scaffold loads Geist via @remotion/google-fonts and wires it into brand.font.family. To use a different Google Font, edit src/fonts.ts to import from a different @remotion/google-fonts/<Name> module. brand.font.googleFont is the display name used by the skill to decide which module to load. Non-Google fonts require manual wiring.
Phase 3: Narrative Planning
Step 3.0 — Derive the signature motif from context
Before picking a story pattern, identify the signature visual motif that will carry the narrative across scenes. This is the single most important call for not looking generic.
- Finish this sentence in one verb: "This feature lets developers ___ something."
- Examples: compose music, sync state, validate inputs, secure tokens, deploy functions, route requests.
- Translate the verb into a physical / spatial metaphor: a waveform for audio composition, packets traveling edges for sync, squiggle underlines for type validation, a padlock for auth, a rocket/checkpoint line for deploy, a switchboard for routing.
- Pick the state-change axis: the motif must visibly change between Problem and Solution scenes (broken ↔ unified, empty ↔ full, disconnected ↔ connected). Pick ONE axis and hold it across the video.
- Look up the motif in
references/visual-motifs.md— the catalog maps 20+ common verbs to motifs, state axes, and example custom elements. If the verb isn't listed, apply the heuristic in that file.
Confirm with user:
"The core verb is compose music, so I'll use an animated waveform as the signature motif — smooth/clean in the hook and solution, jagged/red in the problem. This thread will appear in scenes 1, 2, and 4. Approve, pick a different motif, or let me propose alternatives?"
Do NOT skip this step. A video without a derived motif defaults to generic bullet-list storytelling and fails the generic test (Storytelling Rule 10).
Step 3.1 — Detect story pattern
Scan the PR/input for signals and pick one of 5 patterns. See the detection signals table in patterns/README.md — that file is the single source of truth.
Load the matching pattern spec from patterns/<pattern>.md.
Confirm with user:
"This looks like an [API/Library feature] PR. I'll use that story template. Override? Options: api-library / ui / performance / bugfix / generic / describe a custom pattern"
Step 3.2 — Decide code sourcing (hybrid)
Per pattern:
- api-library-feature → synthesize realistic usage examples that show how developers will actually use the feature
- ui-feature → screenshots / mock components (no code block scenes)
- performance-win → metric cards + optional code
- bug-fix → before (broken) + after (working) snippets, synthesized if raw diff is noisy
- generic-fallback → bullet benefits, no code
Offer user override:
"For code snippets, I'll synthesize realistic usage examples rather than paste raw diff. Override: use-diff / synthesize / mix"
Step 3.2b — Ground synthesized code in the real library
Before synthesizing usage code for a PR, verify the library's actual public API. Do not invent method names, argument shapes, or import paths.
For each snippet the skill plans to include:
- Locate the real library code locally (e.g.,
packages/<name>/src/index.ts, docs examples indocs/, test fixtures intests/). - Confirm every imported name exists as exported.
- Confirm every method/function signature matches (argument names, shape, async vs sync).
- Prefer patterns from the library's own docs over inferred shapes.
If the library isn't available locally and the skill can't verify, ask the user before synthesizing. A wrong API in the first draft destroys user trust and wastes a full iteration round.
Step 3.3 — Self-improve the draft, then present for approval
This step has two parts: a silent self-improvement loop that you run before the user sees anything, and the user-facing approval gate that follows.
Do not show the user the first thing you wrote. The first version of a scene plan is almost always weaker than the second. The agent's job here is to hand the user the strongest plan it can build given the rules in this skill, not the first draft.
Self-improvement loop (silent — run before presenting)
After drafting an initial scene plan, run a deliberate improvement pass against the rule sections listed below. Iterate at least twice. Stop only when one full pass produces zero further changes — the draft has stabilized.
On every pass, for each scene and for the plan as a whole, ask: "Does this satisfy this rule? If not, can I rewrite, merge, drop, split, or reorder a scene to fix it?" Apply the fix in-place, then continue the pass.
The five Core checks below must be satisfied — every plan, no exceptions. On top of those, scan against the broader rule sections at the end of this step.
Core checks (every plan must satisfy all five):
Per-scene payoff: for each scene, write one sentence of the form "The new thing a viewer knows at the end of this scene is ___." If two scenes produce the same sentence, one is redundant — merge or cut. If a scene's sentence is vague (e.g., "the product is good"), the scene is filler — redesign.
Pacing variance: scene durations must reflect cognitive load, not a uniform slice. Reference shape for a 30s target — scale proportionally for shorter (15–25s) or longer (40–60s) videos:
- Hook: ~10–12% of total (≈3s @ 30s; ≈1.8s @ 15s; ≈6s @ 60s) — a single punch
- Problem / setup: ~15–20% (≈5s @ 30s) — enough to land one concrete claim
- Delivery (code / swap / comparison): ~45–55%, with internal chapters if the beat exceeds ~8s (>240 frames at 30fps)
- CTA: ~18–25%, never below 4s (120 frames at 30fps) regardless of total — the CTA always breathes and never rushes
Reject plans where the shortest and longest scene differ by less than ~2×. Equal-slice plans are the single strongest "AI-generated" tell. Reject plans where the CTA is under 4s — a rushed CTA destroys the conversion the rest of the video bought.
Breathing room (no rush-cuts on text): a scene must not transition out while a text element is still being read. Minimum on-screen dwell time, measured from the moment the element finishes animating in to the moment the scene begins transitioning out:
Element Minimum dwell Short headline / one phrase (≤6 words) ≥ 1.5s (≥45 frames at 30fps) Long headline / single sentence (7–14 words) ≥ 2.5s (≥75 frames) Two-line text / short paragraph (15–30 words) ≥ 3.5s (≥105 frames) Code chapter (per chapter, after focus lands) ≥ 3s (≥90 frames) CTA URL / handle (must be clearly readable) ≥ 3s (≥90 frames) Every scene must also include a ~0.4s settle hold (≥12 frames at 30fps) between the last animation completing and the scene transitioning out. Cutting on the same frame an animation finishes is forbidden — the eye needs a beat to confirm what it saw, and transitions that arrive on the resolve-frame feel cluttered and amateur. If the proposed
durationFramescannot accommodate these minima, shorten the beat list, do not shrink the dwell times.Value prop by ~t=8s (or by ~25–30% of total duration, whichever is earlier): by the end of scene 2, the viewer must know what the feature does, who it's for, and why it matters. If that's not true with the current plan, restructure before scaffolding. Do NOT bury the value in the delivery scene.
Motif presence and state-change: the signature motif chosen in Phase 3.0 must appear in at least 2 scenes (typically 3: hook + problem + CTA) and visibly change state between at least one adjacent pair (e.g., clean → glitchy → clean again).
Broader rule sections to scan on every improvement pass. Search this file (or the linked file) for each section heading and walk it against the current draft:
- Hook enforcement (
hooks/hook-rules.md) — applied to the HookTitle scene. If it fails any check, rewrite the headline; don't just record the failure. - Storytelling & Visual Uniqueness Rules (the numbered "Rule N" sections later in this file) — generic-test, no plain bullet lists, signature motif as load-bearing element, anti-clickbait, side-by-side contrasts, insight tagline, counter-expectation beat, pattern-interrupt vs information, first-10s value prop, et al.
- Code Scene Rules — chapters mandatory for ≥5s code beats (>150 frames at 30fps), synchronized per-chapter narration, per-chapter dwell (~3s, ~90 frames), line-length per scene type, diagnostic-comment color, elide unimportant config with
/*…*/, pre-break long imports. - Layout Rules — single alignment per scene, foreground readability over decoration, no accidental overlap, hero-text size on aspect changes.
- Visual Cognition Rules — ≤4 visual chunks per frame, one pre-attentive cue per focal element, reading-flow matched to scene type, reading-saccade limits.
- Anti-padding rule (Q2.1) — if any scene exists solely to fill time, cut it and shorten the target. Do not carry it forward.
- Pattern spec loaded in Step 3.1 (
patterns/<pattern>.md) — the scene sequence and payoffs should be coherent with the template; deviations must have a clear justification. - Step 3.0 motif & Step 3.2b grounded code — verify motif state-change still applies and any synthesized code still maps to real exports/signatures after the rewrites.
Improvements you may apply silently during the loop (no user question needed):
- Rewrite headlines, payoff sentences, captions, and CTA copy for stronger hook discipline and clearer payoff
- Merge two scenes whose payoff sentences collapse to the same idea
- Drop a scene whose only function is to fill time, and shorten the target accordingly
- Split a delivery scene into chaptered sub-beats when a single block exceeds ~8s (>240 frames at 30fps)
- Re-assign motif state per scene to produce a visible state-change between at least one adjacent pair
- Reorder scenes to move the value prop earlier
- Adjust per-scene
durationFramesto satisfy pacing variance and breathing-room minima — only by trimming, never by stretching or padding - Replace generic bullets with concrete artifacts (real API shapes, real error messages, real metrics) drawn from the input
Changes that must wait for the user — flag them, do not apply silently:
- Moving the target duration outside the ±20% band confirmed at Q2.1
- Changing the signature motif chosen at Step 3.0
- Changing the story pattern chosen at Step 3.1
- Adding or removing major scenes that alter the headline narrative claim
Stop condition. End the loop when one full pass over all rule sections (Core checks + broader sections) produces zero changes. If you hit five passes without stabilizing, you have a structural problem the loop can't fix — stop and ask the user for guidance instead of churning.
Present the (already-improved) scene plan for approval
When presenting to the user, include a brief "Self-review notes" line near the top of the response so the user can see the improvement work was actually done. List the 2–5 most material changes the loop applied. Example:
Self-review notes: tightened the hook line from "Add validation easily" to "Swap validation libs with one line"; merged a redundant "why it matters" scene into the problem setup; trimmed delivery 14s → 11s so the CTA dwells ≥4s.
If the loop produced no material changes (rare — usually means the first draft was already strong, or the agent isn't pushing hard enough), say so explicitly: "Self-review notes: draft was stable on first pass; no rewrites needed."
Example output:
"Here's the plan (30s target):
- HookTitle (0–3s, 90f) —
"Swap validation libs with one line"— payoff: there's one line that replaces N SDKs- ProblemSetup (3–8s, 150f) — three evidence cards with concrete conflicting API shapes — payoff: viewer sees the real API-shape conflict they live with today
- LibrarySwap (8–22s, 420f) — shared Standard Schema code, import line cycles zod → valibot → arktype — payoff: viewer sees the "one line change" literally happen on screen
- CTAEndScreen (22–30s, 240f) —
"Ship it"+ link to standardschema.dev — payoff: viewer knows exactly where to go nextPacing: hook 3s / problem 5s / delivery 14s / CTA 8s (ratio ~4.7×) — passes variance check. Motif: schema-interop glyph (interlocking rings) appears in scenes 1, 2, 4; rings are disconnected in scene 2, unified in 1 and 4.
Approve or adjust any section?"
Do not scaffold until the user approves the scene plan.
See patterns/README.md for how patterns map to scene plans.
Step 3.4 — Plan scene transitions (match-cuts)
Hard cuts between scenes are acceptable; match-cuts are what make a video feel crafted. When the signature motif (Phase 3.0) appears in adjacent scenes, the motif must carry over as a match-cut, not restart from zero.
Rules:
- When the same motif component appears in scene N and scene N+1, render it in the last ~8 frames of scene N with the entering state of scene N+1 already beginning. The motif's position, scale, and core geometry must be continuous across the cut — only its state (color, amplitude, opacity, shape) changes.
- When adjacent scenes use different motifs, a text/background element may bridge them: e.g., the last word of scene N's tagline becomes the first word of scene N+1's caption, kept at the same position during the transition.
- When scenes have NO common element, use a directional motion cue — a brand-primary bar sweeping left-to-right that covers the cut — never a generic fade-to-black.
SceneBackgroundvariant must change between adjacent scenes (no two adjacent scenes share a variant). This is enforced at pre-render.
Write the planned transitions into the scene plan before scaffolding:
Transition 1→2: waveform carries over at the same position; color fades from brand.primary → brand.danger; amplitude jitters from smooth to glitchy over 8 frames. Transition 2→3: glitch waveform contracts into a flat line → that line becomes the top border of the code card in scene 3. Transition 3→4: active-provider pill in scene 3 slides down and morphs into the URL pill of the CTA.
Step 3.5 — UI moment (when a UI surface exists)
If the feature has a visible UI surface — a generated image, a rendered audio player, a dashboard, a settings toggle, a diff view — the video must include at least one beat that shows that surface. Research on product-launch video is consistent: "show the product in action" is the strongest single predictor of viewer recall.
Options (in order of preference):
- Real screen capture: a short 2–3s clip of the feature running in the example app. Embedded via
<Video src={...}/>in a Custom scene. - Static screenshot with kinetic overlay: a high-res shot of the UI with brand-primary-tinted call-out boxes or arrows animating in. Easiest to author; works for any UI.
- Mock UI rendered in React: a stylized recreation of the UI inside the scene — buttons, progress bars, output previews — using only the scene's brand palette. Preferred when no screen capture is available and the UI is simple enough to fake convincingly.
Place the UI moment in the delivery scene (Scene 3 in the default structure), ideally at its midpoint so the viewer gets code-plus-result. A code-only delivery feels like a reference doc; code-plus-UI feels like a demo.
Skip this step only if the feature is purely API-level with no user-visible surface (a parser, a compiler pass, a type-level utility).
Step 3.6 — Per-aspect narrative adjustment (multi-format)
When the user picked option 4 (Multi-format) in Q2.2, the skill must plan per-aspect narrative variants — a 9:16 vertical video is not just a cropped 16:9.
Rule 0 (the load-bearing one): Fill the canvas. Don't just shrink content.
The most common failure mode when porting a 16:9 layout to 1:1 or 9:16: the agent keeps the original element sizes, only changes Composition width and height, and ships a video where content occupies ~50% of the new canvas with huge dead margins on top/bottom (vertical) or sides (square).
The right move is the opposite of intuition: when the canvas gets smaller in one axis, the content's per-element size needs to get bigger, not smaller. Less competing content = each element earns more visual space.
Concrete defaults when porting from 1920×1080 (landscape) to other aspects:
- Hero text (hook/CTA headlines): same px size or +10–25%. A 132px landscape headline becomes ~116–144px on 1:1 and ~140–160px on 9:16. Going down to 88px is wrong.
- Body text & captions: +20–40% for 9:16 (a 22px caption → 28–32px). On 1:1, hold or grow slightly.
- Padding & margins: increase scene padding 1.5–2× to consume edge space. A 32px landscape scene padding becomes ~70px on 1:1 and ~140–200px on 9:16 (especially top/bottom on 9:16).
- Inter-element gaps: the
Stack-style flexgapof 56px on landscape becomes ~80px on 9:16. Whitespace is doing brand work — keep it generous. - Element heights / min-heights: code-card mount-card heights, browser-mock chat min-height — all should grow on 9:16 to use the tall canvas.
- Code font: code that was 18px on landscape often goes to ~17–18px on 1:1 (less width to spare) but bigger on 9:16 where there's less content competing — 16–20px is fine.
Implementation pattern in Remotion: read the current aspect from useVideoConfig() (or a context) inside each scene/Panel and switch sizing values via a small byAspect({ landscape, square, vertical }) helper, instead of hardcoding pixel sizes. The same Story data drives all three compositions; only the per-aspect sizing layer differs.
Run the fill check mentally before rendering and as a hard pre-render audit:
At the hero frame (peak content) of every scene at every aspect, the bounding box of the foreground content reaches within ~80–100px of every canvas edge. If any scene at any aspect has more than ~120px of dead margin on either axis, the per-aspect sizing is wrong — the content is undersized.
If you're tempted to keep all the landscape sizes "since they look fine on landscape" — stop. They look fine on landscape because they fill landscape. They will not fill a different aspect at the same sizes.
Layout-shape adjustments
- Horizontal-heavy scenes: side-by-side
BeforeAfter, multi-pill pill rows, and wideLibrarySwaplayouts must switch to stacked (top/bottom) or single-item-at-a-time layouts on 9:16. Mark these scenes with anaspectOverrideshint during scene planning, and branch the JSX in the scene component onuseVideoConfig().width / heightratio. - Multi-column grids: a 6-col framework grid becomes 4-col on 1:1, 3-col on 9:16. A 5-card mount row becomes a 3-col grid (with 2 wrapping to a second row) on 1:1, or a 2-col grid (with the 5th wrapping) on 9:16.
- Motif sizing: waveform/thread widths should scale with composition width, n
…(truncated)