Video Editing — a decision system, not a tutorial
You make editing decisions that maximize per-second retention, rhythm, clarity, narrative impact, perceived quality, and shareability. You are software-agnostic: you reason in techniques and intent; a tool (Premiere/DaVinci/CapCut/After Effects/Final Cut, or an executor skill like /video-use) performs the clicks. When this skill says "punch in" or "matte the text," translate to whatever the available tool calls it.
Companion references ship with this skill: knowledge-base.md (the full consolidated theory by area — hook, retention, pacing, structure, motion, text, sound, transitions, B-roll, CTA, emotion, clarity, density, payoff, loopability, color, composition, depth, distribution, character/avatar & drawn overlays) and rules.md (the 20 highest-leverage rules of the whole corpus).
0. North star (read first, overrides everything below when in tension)
- Retention is the lever; length follows. There is no magic duration. Optimal length = the longest you can hold attention. Never pad. Every second must earn the next.
- Match the experience the audience came for. Editing intensity lives on an authenticity ↔ stimulation axis. A calm "hanging-out" vlog and a rapid MrBeast-style edit both work — for different audiences. Disrupting what the viewer came for is the #1 way to make them leave. Decide the target experience before choosing techniques.
- The edit serves the idea/script; it can't save a bad one. Order: idea → script → edit. Vet the idea ("would someone watch if they didn't know you?"). Fix upstream first.
- Less is more. Simple, well-executed, well-framed, well-graded beats complex over-creativity. When unsure, remove a layer, not add one. Effects must complement what's said; "can use ≠ should use."
- Show, don't tell. Visualize every important spoken line (B-roll, icon, stock, motion graphic). Visual delivery beats verbal.
1. First: classify the job (drives every later choice)
Before editing, answer four questions:
| Question | Options | Why it matters |
|---|---|---|
| Length | short-form (≤~60s) / long-form | short → HPC structure, payoff at end; long → 4-pillar, A/B-roll variety |
| Style family | talking-head · motion-graphics · cinematic/aesthetic | sets pace, text density, motion vocabulary, finishing |
| Goal | retention/views · conversion (sell/CTA) · brand/identity | conversion → funnel (proof B-roll); brand → strong fixed signature |
| Audience experience | authentic/calm ↔ stimulating/fast | sets where on the intensity axis to edit (north star #2) |
Style-family quick configs:
- Talking-head: transcript/continuous-speech cut, J-cuts, eyeline-matched punch-ins, centered captions near face, show-don't-tell B-roll, whoosh/click sound, micro-zoom to avoid static.
- Motion-graphics: design system (fixed fonts/palette/shapes), premium-motion formula, staggered hierarchical reveals, expression-linked/continuous-camera moves, motion-sound coupling, finishing FX (chromatic aberration/grain) masked off text.
- Cinematic/aesthetic: rule-of-thirds + negative-space framing, repurposed B-roll, shot sequences, music-first + cut-on-dips, borders/grain, shoot-for-the-grade, sparse text.
- Cross-cutting registers (drop into any family): animated avatar/character for personality, and hand-drawn/organic overlays for warmth — see §6b.
2. Structure frameworks (pick by length)
Short-form — HPC (payoff at the END):
- H — Hook (first ~3–5s, the most important seconds). State the topic AND open loops (raise unanswered questions). Fast, no pauses, music, captions, peak energy. Don't reveal the answer.
- P — Progression. Deliver the journey toward the promised payoff; keep the hook's promise (no bait-and-switch, no anticlimax).
- C — Climax. Pay off at the end. The #1 short-form mistake is paying off too early — anticipation collapses, viewer scrolls.
Long-form — 4-pillar "addictive" spine: 0. Match the audience's experience (north star #2).
- Visual variety — alternate A-roll / B-roll / motion-graphics to sustain attention; capped by clarity (over-cutting = "visual mush"). Shot length = how long it stays interesting, not minimized.
- Visual continuity — graphics enter motivated (move in, or justify a pop with a sound); keep the focal point consistent across cuts (eye-trace); invisible cuts. Break continuity only as a deliberate pattern interrupt.
- Immersive audio — sound emotion toolkit + music as an emotional engine.
Conversion / service funnel reel: problem→rescue hook → process B-roll (= social proof, makes competence visible) → benefit statement → soft CTA. The CTA is not the lever; the proof is.
3. Hook system (decides whether they stay)
A hook is multi-channel: first frame + ambient/music + vocal delivery + on-screen text — not just the line. Build them together.
- What to say (claim taxonomy): big DESIRE · SOCIAL PROOF · CONTROVERSY · TOO-GOOD-TO-BE-TRUE · PROBLEM→RESCUE. Pick the angle that fits the idea.
- How to hold: open loops (questions) + anticipation rope (withhold the payoff; one short bridge, not stacked teases).
- Framing: prefer raw/personal over tired listicle formulas — audiences have a "hook guard" against formulaic openers.
- Delivery: peak energy; hype up before recording and dump it into the hook.
- Visual hook styles: rapid flash-frame burst (many 2-frame clips + ramp-clicks + riser, reel starts on the peak) for montage; clean shot sequence for cinematic; strong claim + face-fast for talking-head.
- Payoff placement rule: borrowed/known context → open hot (cut to climax); original/unknown context → build then pay off late.
4. Pacing & retention rules
- Kill dead air. Remove every silence; for talking-head, tighten so speech is unbroken — continuous talking holds viewers.
- Cut waste, not story. First pass adds (memes/animation/rewrites), second pass cuts everything not serving the goal. A deliberate pause is fine if you make the stop worth waiting for ("higher standards, not shorter attention").
- Beat-map the music. Mark beats first; land text, cuts, and reveals on beats; cut scene changes on the song's natural dips.
- Motivated movement only. Zoom to emphasize a statement; track to follow movement; micro-zoom/drift to keep static shots alive. Never motion for motion's sake.
- Loopability (short-form): a seamless loop and loopable audio drive rewatches and watch-time.
5. Text system
- Three roles: caption (only on key lines — never text every second), center/emphasis text (styled, glow/color), diegetic/integrated text (behind subject, on a surface, inside an object, as a matte).
- Placement (eye-travel rule): anchor text to the subject's face; never bury captions at the far screen bottom, never over the eyes, always clear of platform UI. Face/eyes upper-third, subtitles lower-third by default — but anchor to the face, not a fixed zone.
- Style: brand font; avoid the hard default drop shadow (a subtle soft shadow/outline for legibility over busy footage is fine); ≤~3 words (or one word/beat for kinetic captions); convert auto-captions to editable graphics; trim captions where not needed.
- Emphasis: recolor/animate only the single payload word — or, in the size-hierarchy caption style, scale the payload word large (
2–2.5×) with supporting words small (70%) stacked above/below it (uniform supporting size throughout). Emphasis by size+position, not only color. - Hero text depth: long shadow with alpha-falloff (smooth, not blocky), only on highlighted words; or a bottom-gradient scrim (masked black, high feather) behind captions for legibility.
- Entrance: overshoot-and-settle (e.g. 70%→110%→100%); or the caption fade-in recipe = transform(position)+blur+opacity keyframes (blur 50–100→0, offset→center, 0→100% over ~15–20f), eased + smoothed graph; vary entrance direction (bottom-up vs left-right) across words. Save as a preset; vary it to avoid monotony.
6. Motion system
- Premium-motion formula = overshoot + motion-blur + easing + stagger + secondary motion. Motion blur on every moving element and eased/smoothed curves are non-negotiable. Stagger entrances (never all at once); delay child elements until their parent lands; add secondary motion (arcs, shake, bumps) so paths aren't robotic.
- Focus continuity is exact: match the subject's eyeline to the same screen coordinates across cuts regardless of zoom. Keep the moving subject centered (camera-follow).
- Reveals: hide/blur → approach/zoom → reveal · path-length draw-on · overshoot entrance · box-open + cursor · matte wipe.
- Depth: cast-shadow primitive (duplicate → black → heavy blur → low opacity → behind); stacking order + drop shadow; behind-subject elements (duplicate + background-remove); blurred foreground/background for fake depth-of-field.
- High-end / 3D: continuous never-stopping camera journey (chain moves, 3 keyframes); cut at peak velocity (motion blur hides the cut); shape/object matte transitions; expression-link related motions (spin↔move speed, shadow↔−rotation) for automatic consistency.
- Build efficiently: build one element fully, duplicate, re-skin (color/text/icon). Keep a personal asset pack (hooks, text in/out, overlays, SFX, zoom presets) — in short-form this is a required pipeline step, not optional.
6b. Character/avatar animation & hand-drawn overlays
Avatar/mascot (adds personality; faceless or not): build from a stock PNG body + your circle-cropped face/logo head; keep every character the same size + position so it's swappable; keep a library.
- Easing → intent: ease-in = entrance, ease-out = exit, ease-in-out = traverse. The velocity graph IS a speed curve: graph height = speed; smooth/wave it for smooth motion (linear = harsh).
- Liveliness: overshoot-and-settle entrances; oscillating wiggle (value → +40 → −50 → +20 → original) for idle; motion blur on every move (shutter-angle dials amount, 360 for spins).
- Compound motion: stack several overlapping transforms (each a simple motion), not one overloaded transform.
- Reuse: save animations as presets (anchor to in/out point so timing doesn't stretch); subject-swap — drop a different character into the saved animation for zero-rework variants.
- Depth on flat avatars: opposing inner-glow + inner-shadow (erase opposite corners) for fake-3D light; drop shadow; nested breathing idle. Track an element to a moving head (manual keyframes or AI-track).
Hand-drawn/organic overlays (warm, scrapbook register; can convey a FEELING the footage can't): frame-by-frame onion-skin (draw → 25–50% opacity → new layer → repeat; start from the main object). Key drawings over footage with blend modes — Add/Screen (key black), Color Dodge (interact with footage color), Lighten (key black for light art), Darken (key white). Boil/handwritten text (trace ~3 frames, or screen-record handwriting on black + key + pencil SFX). Custom-drawn shape transitions (expanding star/heart keyed over footage; duplicate+reverse for symmetric in/out).
B-roll : A-roll ≈ 3–4 : 1 for faceless/gaming/explainer. Transitions only at topic/music/chapter boundaries.
7. Sound system (≈ half of perceived quality; most viewers also watch muted → captions mandatory)
- Phase it: do a dedicated sound pass after picture lock. 3 steps: (1) whooshes on every movement — "edit as if everything moves through thick air"; (2) textured SFX matched to each element (click/mechanical/UI/glitch); (3) risers + hits for anticipation at transitions.
- Motion-sound coupling: every visual movement gets a matched sound; peak-align whooshes to motion peaks. A pop/appear must be justified by a sound.
- Emotion toolkit: risers (build anticipation — ONLY before a real payoff, or they lose credibility), hits (release/emphasize; reverse a hit to build tension), drones (mystery/suspense).
- Isolate the pass: mute VO + music while placing SFX; then duck SFX/music well under VO.
- Music as emotion engine: choose early/pre-production (shoot to its rhythm); map a mood per segment; cut on beats/dips; pause music to jolt/spotlight; fade to signal an ending; sync a swell to a topic shift; use stems for control. Favor loopable, clean-cut, trending-but-not-overused audio.
8. Color, composition, finishing
- Grade for the look. Grade early when the look drives creative choices or you reuse one LUT across clips; grade late for per-clip precision. Shoot-for-the-grade (enough light + color in frame). A strong grade is itself a scroll-stopper.
- Composition retains even when slow: rule-of-thirds, subject/horizon on grid lines, foreground for depth, text in negative space, color contrast (e.g. orange/teal). One focal element per scene (object/character/text-centered); everything else supports it; decorative detail only if subtle and subordinate.
- Finishing stack for cohesion: the same vignette + grain + subtle zoom-blur (and borders, if that's the signature) on every scene. Mask finishing FX off the text to keep it legible.
- One signature, not more effects. A fixed identity (theme + locked text palette + framing/border) builds memory. The clutter anti-pattern (many fonts/colors/effects, no unique filter) is the #1 amateur mistake.
9. Transitions
- Default to a clean hard cut. Fancy transitions usually look worse, especially between similar shots.
- Smooth dialogue cuts with J/L-cuts (audio leads/trails; place the lead at a phrase boundary).
- Hide cuts with full-screen transitions or by cutting at peak motion.
- Motivated transitions only: overlay (additive/screen blend), shape-matte wipe, or an object carrying you into the next scene.
10. Distribution awareness
- Tool-neutral ranking: platforms rank the video, not the editing app. Pick the tool that lets you produce quality fastest.
- Platform-fit: the same edit can do 8M on one platform and 2k on another. Ask what THIS platform rewards (IG: trends/audio/watch-time/saves/shares; YT Shorts: search/subs/niche) and adapt.
- Virality = emotion + timing + shareability: a topic people already care about, posted when it's on their minds, with a clear reason to share. Editing amplifies an already-resonant moment.
11. Intervention priority (what to fix FIRST when an edit underperforms)
Work top-down; don't polish low items while high ones are broken:
- Idea/topic + audience-experience match — is this for the right audience, at the right time, with a real reason to care/share?
- Hook (first 3–5s) — claim + open loop + multi-channel + energy; payoff NOT given away.
- Retention/pacing — dead air killed; no "mush"; payoff placed at the end (short) or variety+continuity sustained (long).
- Clarity / show-don't-tell — is each point delivered visually; is text legible and eye-travel minimal?
- Sound — motion-sound coupling, emotion toolkit, music mood/sync.
- Motion/text polish — eased curves + motion blur, eyeline match, overshoot, focal composition.
- Color/finish/signature — grade, finishing stack, one consistent identity.
12. Anti-patterns (errors to avoid)
- Forcing length / padding; cutting for cutting's sake (visual mush); paying off too early.
- Flashy/formulaic hook on an authenticity audience; revealing the payoff in the hook.
- Text every second; captions at the screen bottom far from the face; many fonts/colors/effects; no consistent identity.
- Effects because they exist (not complementing the line); identical animation on every element.
- Motion with no matching sound (feels empty) or sound with no motion (noise); risers with no payoff (cry wolf).
- Linear/robotic keyframes; missing motion blur; mismatched eyelines; graphics that pop in unmotivated.
- Designing SFX with everything playing; music that doesn't fit or can't loop.
- Over-editing an inherently authentic/calm piece; ignoring platform-fit.
13. Checklists
Build (in order — phase separation: cut → structure → visual → text → sound → color/finish):
- Idea vetted; audience experience + style family + goal chosen
- Cut: silences gone / speech continuous; J-cuts on phrase boundaries; add-then-subtract done
- Structure: HPC (short) or 4-pillar (long); payoff placed correctly
- Hook: claim + open loop + multi-channel + energy; face/visual fast
- Visual: show-don't-tell B-roll on key lines; focal composition; eye-travel minimal
- Motion: eased + motion-blur + overshoot + stagger; eyeline matched; motivated only
- Text: 3-role system; legible; ≤3 words / 1-per-beat; emphasis on payload word
- Sound: isolated pass; motion-sound coupling; emotion toolkit; music mood/sync; ducked
- Color/finish: grade; finishing stack; one signature; FX masked off text
- Distribution: platform-fit; loopable; timely
Final review:
- First 3s stop the scroll on their own?
- Any second that doesn't earn the next? (cut it)
- Is the payoff at the end (short) and does the body keep the hook's promise?
- Does every important line have a visual?
- Could a viewer follow it muted (captions) and would sound add for those who hear it?
- Is there a simpler version that's just as strong? (less-is-more)
Losing-rhythm signals: dead air; a slow stretch with no escalating reason; a delayed payoff with nothing building; eye forced to travel between scattered elements; a transition that distracts from the point.
Over-edited signals: effect with no narrative reason; text on screen constantly; every line animated identically; many fonts/colors; flashy transition worse than a cut; finishing FX over the text; motion with no sound or sound with no motion.
14. Using with an executor (e.g. /video-use)
- This skill decides WHAT and WHY (structure, hook, pacing, which technique, intervention priority). Hand the executor concrete, tool-neutral instructions: "eyeline-matched punch-in on the claim word; overshoot-pop the caption; whoosh peak-aligned to the zoom; cut at the music dip; grade warm; vignette+grain finish."
- Production-correctness (sync, safe margins, export at source resolution/frame rate, duck levels, no copyright music) is hard and owned by the executor; everything above is artistic direction.
- If used alone, apply the same instructions manually in any editor.