Cut List and Subtitles
Everything produced here is already in the script. This instruction translates it into the order used when working in the editing program.
To have the cut executed instead of described, use /remotion-cut in Claude Code — it consumes this cut list directly.
Procedure
- Read in the script. If only the video file is available, ask for the matching script – without a script this instruction doesn't apply.
- Ask what hasn't been decided – before cutting, not after. Color, size, and position of subtitles are brand decisions, not cutting decisions. If a value is in neither the prompt nor the visual-style file, ask for it instead of guessing: a wrongly guessed color costs a whole render pass and a round of review.
- Put the beats in cutting order. That's usually the script order, with one important exception: if a later beat has a stronger image than the hook, suggest moving it up as a cold open – as a suggestion, not a unilateral move.
- For each beat, mark the cut point in the spoken text: the last word before the cut. That's what you actually search for while editing.
- Give the overlays from the TEXT fields timestamps.
- Generate the subtitle correction template (see below).
Selecting Takes
In the raw footage, every sentence appears multiple times. The selection is the actual cut — after that, the ordering is just craft.
With multiple attempts at the same sentence, the rule is: the last complete take. People repeat a sentence because they misspoke or got stuck; the last attempt is the one they meant to keep. Exception only if the last one is audibly worse — then say explicitly why you're using an earlier one.
Cut outright:
- Slips of the tongue and the run-ups before them — everything up to the last clean run-through
- Laughter, unless it's at the end of the sentence and carries the point. If it starts on the last syllable, end the clip right after the word — the jump cut hides the rest. Check the region list first for a second take of that sentence
- Aborted sentences, even if the beginning sounds good
- Stage directions and self-talk ("again", "wait", "what was that")
- Filler words at the start of a sentence — "So", "Um", "Yeah, okay", "Right" — the cut sits after the filler word, not before it
- Throat-clearing, swallowing, deep inhale before the first word
- Repeated half-sentences within a take ("that's — that's important")
Don't cut out: the visible microphone, small hand movements — and, in the standard cut, normal-length breathing pauses between sentences. Whether the breathing stays or goes is the account's decision, not the editor's: if the visual-style or format file says "cut without silences", the section below applies.
Set the margins (standard cut): the cut sits in the speech pause, not on the word. About 0.15 s before the first sound at the front, 0.25 s after the last at the back — otherwise it sounds choppy. Measure the pauses in the audio, don't estimate.
Different for a stalling pause mid-sentence: this isn't framed, it's removed. What remains is about 0.05 s on each side, roughly 0.1 s total — a normal word gap. With nothing left at all, the following word sounds abrupt instead of emphasized.
Cut Without Silences
When the brief says to remove every spot "where the waveform is flat" — breaths, stalls, pauses — these values apply (measured on a professionally edited reel and on our own third video):
Silence RMS below −40 dB in 10 ms windows, at least 0.10 s. Room tone sits at
−60…−80 dB, speech at −15…−25 dB — that is where the waveform is
actually flat. The −36 dB of the pause detection above is too high
for this: it swallows quiet word endings.
Cut only silences of 0.30 s or more. A jump cut that gains four frames is
more visible than the pause it removes. A professional edit keeps
nearly everything under 0.3 s and removes everything above.
Residual gap 0.08 to 0.13 s (0.05–0.07 s after the last sound, 0.03–0.06 s before
the next). The professional edit goes down to 0.07 s (0.03 + 0.02).
Clip margins 0.03–0.06 s at the front, 0.06–0.08 s at the back.
Audio fade 1 frame instead of 2 as soon as a margin is under 0.07 s — otherwise
the fade lands on the first sound.
Jump cut every removed stretch is a cut and needs the zoom change from the
visual-style file, or it reads as a stutter.
Result on our third video: 92 s → 83 s, 24 clips → 37 parts. The professional edit of comparable material sits at 19 % silence, ours at 24 %.
Word endings are not breaths. A quiet, voiced block right after a word is almost always the end of the word, not a breath or a filler — four out of four candidates were speech: an unstressed final syllable (twice), the aspirated "‑t" of a word ending in "‑ct", the word "the". Pattern to know: a plosive closure ("k", "t") before the last syllable shows as a 0.1 s dip to −45…−55 dB, followed by a flat block — acoustically indistinguishable from an "uhh". Cut only real silence. Soft sentence endings ("‑ful", "‑ly") ring out up to 0.4 s beyond the −40 dB point — at the end of a sentence, err on the side of too much tail; it costs nothing.
Don't hunt fillers ("uhh", "um") acoustically. Level, zero-crossing rate and pitch find word endings, not fillers; three renders were rolled back in the end. If someone hears a filler, pin the spot to a raw-footage time together. The zero-crossing rate (per 10 ms at 16 kHz) is still useful for one thing: above 80 is a fricative or aspiration ("sh…", "‑ct"), below 25 a voiced sound.
Speech Recognition Swallows Filler Words — and Can't Hear Them Even When Prompted
Whisper does not output "uhh" or "um", not even with a prompt full of fillers, and it places its word timings right across such sounds. Fragments under about one second are worthless as a test: either empty ("is not a mistake" on 0.8 s) or completed from the language model (a full "‑ing" from 0.03 s of audio). Whether a word is clipped only shows in a transcription of the whole render in context — and even that does not hear what a human hears. Token timings are too coarse for gap analysis (±0.3 s).
A take that looks clean when transcribed as a whole is not clean. Recognition smooths things over: a "So, right." before the actual sentence simply doesn't show up in the output. Anyone who only transcribes the whole clip carries the false start into the finished video.
So, per clip:
- Get speech regions via pause detection (about −36 dB, minimum pause 0.14 s).
- Transcribe each region individually, not the clip as a whole.
- Suspicious pattern for a false start: the first region is shorter than 1.2 s and is followed by a pause over 0.28 s.
The pattern also catches harmless cases — a direct address ("And to the guys out there,") or the first item in a list. So listen to or transcribe the region and only then decide — don't cut blind.
Cut List Output Format
# CUT · <Title>
SOURCE: <Script name> · LENGTH: <Seconds>s · CUTS: <n>
## 1 · HOOK · 00:00.0
Clip: <Video code> B1, Take <n>
In at: "<first word>"
Out at: "<last word before the cut>" ← hard cut
Overlay: "COLD BREW ≠ COLD COFFEE" in 00:00.3 out 00:02.0
Sound: Music from 00:00.0, beat on the cut
Note: <only if there's something to note here>
After that:
Overlay overview – all overlays in chronological order with fade-in and fade-out time, so they can be set in one pass.
Sound overview – music entries, changes, silence, effects, each with a timestamp.
Rules for Overlays
- Fade in no earlier than 0.3 seconds after the cut, otherwise it flickers.
- Leave it on screen for at least 1.2 seconds, otherwise it isn't readable. Rule of thumb: 0.4 seconds per word, minimum 1.2.
- Never two overlays at the same time.
- The last overlay ends at the latest 0.5 seconds before the video ends.
Subtitles
What actually helps here — and what doesn't. Editing programs and Instagram generate subtitles automatically from the real audio, accurate to the second. A file calculated from planned timecodes can't compete with that: after shooting and editing, the times are never right.
The weakness of auto-subtitles lies elsewhere: they mangle technical terms, proper names, and numbers and break lines in the middle of phrases. That's exactly where this output earns its keep.
So deliver a correction template, not a timing file:
SUBTITLES · Correction Template
Order as spoken, one line per subtitle.
1 Does your cold brew taste bitter?
2 Then you're probably
3 doing this.
...
CAUTION with automatic recognition:
AeroPress · often becomes "Air Press"
V60 · often becomes "V 60" or "v sixty"
1:8 · often becomes "one to eight"
Convert point sizes, don't copy them as-is. When a subtitle size is given in pt, it usually means the point size on the phone: a 1080 px wide portrait frame corresponds to 360 pt, so pt × 3 = px. 11 pt is therefore 33 px. Have the number confirmed before rendering – the difference between 33 px and 58 px is substantial, design-wise.
Rules for the lines:
- One line is one unit of meaning, maximum seven words.
- At most 30 characters per line – calculated for portrait format, not desktop.
- Break at a natural speech pause, never in the middle of a phrase.
- Numbers, units, and proper names exactly as they're spoken.
The list under CAUTION contains every technical term, every number, and every proper name from the script where automatic recognition, from experience, tends to get it wrong. That's the actual value of this output.
SRT File
Only on explicit request – for instance when cutting without sound or producing a foreign-language version. Then derive the times from the beat timecodes, distribute them within the beat proportionally to word count, cue duration 0.8 to 3.0 seconds, and state clearly that the times will need to be adjusted afterward.
Remotion
If FORMATS.md or a visual-style file describes a Remotion project, the video is
rendered there instead of in an editing program. Then the cut list must be
machine-actionable, not just readable.
Also output a timeline table — times in frames at 30 fps, not seconds, because Remotion works in frames:
CLIP SOURCE-IN SOURCE-OUT DURATION ZOOM OVERLAY
01 00:01:53.9 00:02:00.7 204f 1.00 title
02 00:02:40.8 00:02:49.1 249f 1.14 –
03 00:03:18.7 00:03:22.0 100f 1.04 quote
Rules that follow from this structure:
- Zoom has two jobs. The start value per clip carries the cut: two consecutive clips never get the same one, otherwise the cut reads as a mistake. The movement within the clip carries the take: roughly one percent per second of clip length, 2 to 7 percent in total. A one-percent move across an entire take is invisible — the image just sits still and the cut lives on the edges alone.
- The end value is the limit, not the start value. It decides how tight the frame gets, and must be checked against the safe zone (see below).
- As you zoom, the head drifts out of frame unless the face sits dead center.
Compensate with
translateY = (scale − 1) × offset, whereoffset ≈ (0.5 − faceY) × 1920and faceY is the face's vertical position as a fraction of frame height — measure it once per setup, don't guess. - No transitions. Hard cuts, a 2-frame audio fade on every edge (1 frame when a margin is under 0.07 s, see "Cut Without Silences").
- Zoom can also have two states instead of a ladder of steps: default 100 % and zoomed 130 %, alternating on every cut, plus the drift. Which variant applies is in the account's format file; at 130 %, measure the zone.
- Always cut sub-clips from the raw footage (raw time = clip start + offset), never as a second encode from finished clips. A tail may reach past the old clip end into the raw footage.
- Generate the timeline and the documentation table, don't maintain them by hand. Keep manual decisions (a longer tail, a forced cut) as per-clip overrides in the plan so the rest stays reproducible and only affected parts get re-cut.
- Remotion's bundled ffmpeg lacks
astats,fps,showspectrumpic,tileand scene detection;cropneeds an even width, frame rate goes through-r. Level, zero-crossing rate and pitch are computed in plain Python overwave/array. Cuts in someone else's video are found via the frame difference of tiny grayscale frames (head movements produce false hits). - Overlays are post-production, not props — title and quote cards are created during rendering. None of it is printed or held in hand.
- Only use the documented overlay types. A new type is a design decision, not a cutting decision — propose it, don't introduce it unilaterally.
- Overlays sit on the word, not on the cut. A number fades in when it's spoken, not when the clip begins.
- Never attach overlays to clip indices, always to clip identifiers. If a clip is split later — say, to remove a stalling pause — every index after it shifts, and the overlay silently ends up on the wrong take. The error produces no warning.
Safe Zone
Every element stays in the safe zone — no exceptions. The platform lays its controls over the video; what's underneath is invisible while watching but visible during rendering. So the mistake only shows up once the video is live.
For Instagram Reels at 1080×1920:
clear: x 60 … 1021 y 296 … 1533
from y ≥ 1179 the right side is also occupied → clear only up to x 886
At the top, the header area covers 296 px; at the bottom, the caption and control bar
take up 386 px. The clear area isn't a rectangle: from y = 1179 the action
column (like, comment, share, menu) intrudes on the right.
Subtitles usually sit at exactly this height. So for anything below 1179 that's centered on the middle of the frame, a maximum text width of 692 px applies — otherwise the last word slides under the share button.
If the account's visual-style file defines its own values, those apply instead. State the position for every overlay and check it against these limits instead of estimating.
The Speaking Person Also Belongs in the Zone
The zone applies to every element in the frame, not just text. Too tight a crop pushes the head into the area the platform covers. This doesn't show up when watching the cut on its own — only in comparison with the raw footage, or once the video is already live.
Measure after every render, don't estimate:
Cut out the center column of the frame as a narrow strip, and from the top find the
first row after which a good dozen rows stay dark — that's the hairline.
Measure at 75% of the clip duration, not the middle: that's where the zoom move is
almost finished and the frame is tightest.
Target: a clear margin below the upper zone boundary, at least about 150 px.
Calibrate the threshold to the wall's brightness, don't guess: a bright white wall can easily sit at luma 150, in which case detection triggers on the very first row.
Animation
Motion serves readability, not effect.
Subtitles track along with the spoken text, highlighted word by word: the word currently being spoken is emphasized in color, the rest stays put. The block itself doesn't move — no flying in or out, no position changes — and stays on screen across cut points, because it belongs to the audio, not the clip. In practice that means: the first subtitle page of a clip begins with the clip, not with the first word, otherwise a gap flashes at every cut point.
Highlight whole words, not syllables. Speech recognition returns sub-word pieces ("what" + "ever"). Colored individually, the emphasis runs through the middle of a word. Merge before display: anything that doesn't start with a space belongs to the previous piece.
Set subtitles in regular weight, not bold. The outline carries the readability. Bold reads as loud — use it only where the visual style explicitly calls for it.
Scale the outline with the font size. At small sizes, a letter stem is only a few pixels wide; an outline that was correct for large text then eats into the core and the text color disappears inside it — the type looks gray even though the color is set correctly. Rule of thumb: at most one third of the stem width, with a soft shadow taking over the rest. When in doubt, measure on the rendered image what share of the bright pixels actually carries the intended tint.
Cards and overlays appear in stages, not all at once. Fading everything in simultaneously reads as a freeze frame, no matter how clean the curve is:
Frame 0 Card 0.90 → 1.00 over 10 frames, soft spring, no overshoot
Frame 5 Text 0.96 → 1.00, multi-line rows offset by ~4 frames each
Frame 13 Sticker Pop, see below
Exit: simply fade out the whole card. No rotating, no wobbling.
Stickers placed on top are allowed to pop. An element that visibly sits on the card — a paperclip, a badge — gets a real spring with overshoot: it snaps past 1.00 and settles back. The scaling anchor point sits at the edge where the element is attached; scaled from the center, it drifts out of place as it pops open.
The sticker is the only element allowed to overshoot. The card and text ease in without overshoot — the pop belongs to the element sitting on top, not to the surface.
Pop sound, if the visual style calls for it. Effect libraries have clicks and memes, but rarely a pop; it can be synthesized: a good 100 ms, pitch falling exponentially from about 600 to 150 Hz, a very fast attack, short decay, a few milliseconds of noise for the attack of the "P".
The sound must land in a speech gap and stay well under the voice level. Cross-check before delivery: measure the effect's level against the voice level, don't estimate by ear.
Numbers don't count up. A counter draws attention to itself instead of the number.
Evidence stays still. A screenshot shown as proof is meant to be read — any motion on it is a distraction.
For every overlay, state when it appears and how long it stays, in frames. The overlay sits on the word, not on the cut.
Final Check on the Finished Video
Watch it again after rendering and confirm or flag each point individually. These are exactly the errors that survive the first pass:
Sound
- No slip of the tongue, no doubled sentence start, no leftover "um" – checked per speech region, not just on the finished piece: recognition swallows filler words and then falsely reports a take as clean
- No stalling pause left in the middle of a sentence
- Any effect sound, if present, falls in a speech gap and stays under the voice
- No laughter and no cutoff missed during editing
- No stage direction left in the audio
- No clicks or pops at the cut points (measure the sample jump at every edge)
- No clipped sound at the beginning or end
- No clipped word ending — especially quiet endings ("‑ing", "‑ful", "‑ct"): transcribe the rendered audio in full and read it against the script
- With a cut without silences: list the remaining pauses with timestamps in the docs
Picture
- No two consecutive clips with the same start zoom
- The zoom move within the clips is visible, not just present on paper
- At 75% of every clip's duration, the head sits with margin below the upper zone boundary – measured, not estimated
- No visible jump in brightness or color between clips
Text
- Subtitles match the spoken words verbatim — especially for technical terms
- No subtitle stays on screen longer than the sentence, none appears too early
- Title and quote cards spelled correctly, including special characters
- Nothing sits in the control zone: 296 px top, 386 px bottom, 60 px sides
- Below y 1179, nothing extends past x 886 — that's where the action column sits
- Subtitles in regular weight, not bold
- The highlighting runs in sync with the spoken words, not ahead or behind
- The highlighting jumps to whole words, not syllables
- The text color is actually visible in the rendered image and hasn't been swallowed by the outline
Content
- The hook is fully readable within the first two seconds
- An evidence screenshot, if present, is sharp enough that its source line is legible
- The payoff promised in the script actually happens
- The video ends on a sentence, not in a pause
Report issues with a timestamp, so they can be jumped to directly. If everything checks out, say so in one sentence — no list of passed items.
Wrap-up
In two to three lines, name what specifically needs attention when cutting this particular video – the one spot where it can go wrong. No general editing tips.