# Cut

> Derives the cut list from a script and the raw footage, selects the usable takes, and describes the cut so it can be rendered in Remotion – cut order, cut points, text overlays with fade-in and fade-out times, sound cues – and delivers a subtitle correction template with the terms automatic recognition fails on. Use when a video needs to be cut, subtitles are needed, or overlays need to be set.

- Skill: `bendemartin97/cut` (Agent Skill)
- Install (CLI): `npx skillmds@latest add bendemartin97/cut`
- Raw SKILL.md: https://api.skillmd.com/api/skills/bendemartin97/cut/raw
- Safety review: pending
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: bendemartin97 (https://skillmd.com/u/bendemartin97)
- Updated: 2026-09-17
- Page: https://skillmd.com/skills/bendemartin97/cut

---


# Cut List and Subtitles

Everything produced here is already in the script. This instruction translates it into the order used when working in the editing program.

To have the cut executed instead of described, use /remotion-cut in Claude Code — it consumes this cut list directly.

## Procedure

1. Read in the script. If only the video file is available, ask for the matching script – without a script this instruction doesn't apply.
2. **Ask what hasn't been decided – before cutting, not after.** Color, size,
   and position of subtitles are brand decisions, not cutting decisions.
   If a value is in neither the prompt nor the visual-style file, ask for it
   instead of guessing: a wrongly guessed color costs a whole render pass and
   a round of review.
3. Put the beats in cutting order. That's usually the script order, with one important exception: if a later beat has a stronger image than the hook, suggest moving it up as a cold open – as a suggestion, not a unilateral move.
4. For each beat, mark the **cut point in the spoken text**: the last word before the cut. That's what you actually search for while editing.
5. Give the overlays from the TEXT fields timestamps.
6. Generate the subtitle correction template (see below).

## Selecting Takes

In the raw footage, every sentence appears multiple times. The selection is the actual
cut — after that, the ordering is just craft.

**With multiple attempts at the same sentence, the rule is: the last complete take.**
People repeat a sentence because they misspoke or got stuck; the last attempt is
the one they meant to keep. Exception only if the last one is audibly worse — then say explicitly
why you're using an earlier one.

**Cut outright:**

- **Slips of the tongue and the run-ups before them** — everything up to the last clean run-through
- **Laughter**, unless it's at the end of the sentence and carries the point. If it
  starts on the last syllable, end the clip right after the word — the jump cut hides
  the rest. Check the region list first for a second take of that sentence
- **Aborted sentences**, even if the beginning sounds good
- **Stage directions and self-talk** ("again", "wait", "what was that")
- **Filler words at the start of a sentence** — "So", "Um", "Yeah, okay", "Right" — the cut sits after the filler word, not before it
- **Throat-clearing, swallowing, deep inhale** before the first word
- **Repeated half-sentences** within a take ("that's — that's important")

**Don't cut out:** the visible microphone, small hand movements — and, in the
standard cut, normal-length breathing pauses between sentences. Whether the breathing
stays or goes is the account's decision, not the editor's: if the visual-style or
format file says "cut without silences", the section below applies.

**Set the margins (standard cut):** the cut sits in the speech pause, not on the
word. About 0.15 s before the first sound at the front, 0.25 s after the last at the
back — otherwise it sounds choppy. Measure the pauses in the audio, don't estimate.

Different for a **stalling pause mid-sentence**: this isn't framed, it's removed. What
remains is about 0.05 s on each side, roughly 0.1 s total — a normal word gap. With
nothing left at all, the following word sounds abrupt instead of emphasized.

### Cut Without Silences

When the brief says to remove every spot "where the waveform is flat" — breaths,
stalls, pauses — these values apply (measured on a professionally edited reel and on
our own third video):

```
Silence       RMS below −40 dB in 10 ms windows, at least 0.10 s. Room tone sits at
              −60…−80 dB, speech at −15…−25 dB — that is where the waveform is
              actually flat. The −36 dB of the pause detection above is too high
              for this: it swallows quiet word endings.
Cut           only silences of 0.30 s or more. A jump cut that gains four frames is
              more visible than the pause it removes. A professional edit keeps
              nearly everything under 0.3 s and removes everything above.
Residual gap  0.08 to 0.13 s (0.05–0.07 s after the last sound, 0.03–0.06 s before
              the next). The professional edit goes down to 0.07 s (0.03 + 0.02).
Clip margins  0.03–0.06 s at the front, 0.06–0.08 s at the back.
Audio fade    1 frame instead of 2 as soon as a margin is under 0.07 s — otherwise
              the fade lands on the first sound.
Jump cut      every removed stretch is a cut and needs the zoom change from the
              visual-style file, or it reads as a stutter.
```

Result on our third video: 92 s → 83 s, 24 clips → 37 parts. The professional edit
of comparable material sits at 19 % silence, ours at 24 %.

**Word endings are not breaths.** A quiet, voiced block right after a word is almost
always the end of the word, not a breath or a filler — four out of four candidates
were speech: an unstressed final syllable (twice), the aspirated "‑t" of a word ending
in "‑ct", the word "the". Pattern to know: a plosive closure ("k", "t") before the last
syllable shows as a 0.1 s dip to −45…−55 dB, followed by a flat block — acoustically
indistinguishable from an "uhh". **Cut only real silence.** Soft sentence endings
("‑ful", "‑ly") ring out up to 0.4 s beyond the −40 dB point — at the end of a
sentence, err on the side of too much tail; it costs nothing.

**Don't hunt fillers ("uhh", "um") acoustically.** Level, zero-crossing rate and
pitch find word endings, not fillers; three renders were rolled back in the end. If
someone hears a filler, pin the spot to a raw-footage time together. The
zero-crossing rate (per 10 ms at 16 kHz) is still useful for one thing: above 80 is a
fricative or aspiration ("sh…", "‑ct"), below 25 a voiced sound.

### Speech Recognition Swallows Filler Words — and Can't Hear Them Even When Prompted

Whisper does not output "uhh" or "um", not even with a prompt full of fillers, and it
places its word timings right across such sounds. **Fragments under about one second
are worthless as a test:** either empty ("is not a mistake" on 0.8 s) or completed
from the language model (a full "‑ing" from 0.03 s of audio). Whether a word is
clipped only shows in a transcription of the whole render in context — and even that
does not hear what a human hears. Token timings are too coarse for gap analysis
(±0.3 s).

**A take that looks clean when transcribed as a whole is not clean.** Recognition
smooths things over: a "So, right." before the actual sentence simply doesn't show up
in the output. Anyone who only transcribes the whole clip carries the false start into
the finished video.

So, per clip:

1. Get speech regions via pause detection (about −36 dB, minimum pause 0.14 s).
2. **Transcribe each region individually**, not the clip as a whole.
3. Suspicious pattern for a false start: the first region is shorter than 1.2 s and
   is followed by a pause over 0.28 s.

The pattern also catches harmless cases — a direct address ("And to the guys out
there,") or the first item in a list. So listen to or transcribe the region and only
then decide — don't cut blind.

## Cut List Output Format

```
# CUT · <Title>
SOURCE: <Script name> · LENGTH: <Seconds>s · CUTS: <n>

## 1 · HOOK · 00:00.0
Clip:     <Video code> B1, Take <n>
In at:    "<first word>"
Out at:   "<last word before the cut>"  ← hard cut
Overlay:  "COLD BREW ≠ COLD COFFEE"  in 00:00.3  out 00:02.0
Sound:    Music from 00:00.0, beat on the cut
Note:     <only if there's something to note here>
```

After that:

**Overlay overview** – all overlays in chronological order with fade-in and fade-out time, so they can be set in one pass.

**Sound overview** – music entries, changes, silence, effects, each with a timestamp.

## Rules for Overlays

- Fade in no earlier than 0.3 seconds after the cut, otherwise it flickers.
- Leave it on screen for at least 1.2 seconds, otherwise it isn't readable. Rule of thumb: 0.4 seconds per word, minimum 1.2.
- Never two overlays at the same time.
- The last overlay ends at the latest 0.5 seconds before the video ends.

## Subtitles

**What actually helps here — and what doesn't.** Editing programs and Instagram
generate subtitles automatically from the real audio, accurate to the second. A file
calculated from planned timecodes can't compete with that: after shooting and editing,
the times are never right.

The weakness of auto-subtitles lies elsewhere: **they mangle technical terms, proper
names, and numbers** and break lines in the middle of phrases. That's exactly where
this output earns its keep.

So deliver **a correction template**, not a timing file:

```
SUBTITLES · Correction Template
Order as spoken, one line per subtitle.

1  Does your cold brew taste bitter?
2  Then you're probably
3  doing this.
...

CAUTION with automatic recognition:
  AeroPress · often becomes "Air Press"
  V60 · often becomes "V 60" or "v sixty"
  1:8 · often becomes "one to eight"
```

**Convert point sizes, don't copy them as-is.** When a subtitle size is given in pt,
it usually means the point size on the phone: a 1080 px wide portrait frame
corresponds to 360 pt, so **pt × 3 = px**. 11 pt is therefore 33 px. Have the number
confirmed before rendering – the difference between 33 px and 58 px is substantial,
design-wise.

Rules for the lines:

- One line is one unit of meaning, **maximum seven words**.
- At most 30 characters per line – calculated for portrait format, not desktop.
- Break at a natural speech pause, never in the middle of a phrase.
- Numbers, units, and proper names exactly as they're spoken.

The list under CAUTION contains every technical term, every number, and every proper
name from the script where automatic recognition, from experience, tends to get it
wrong. That's the actual value of this output.

### SRT File

Only on explicit request – for instance when cutting without sound or producing a
foreign-language version. Then derive the times from the beat timecodes, distribute
them within the beat proportionally to word count, cue duration 0.8 to 3.0 seconds,
and **state clearly that the times will need to be adjusted afterward**.

## Remotion

If `FORMATS.md` or a visual-style file describes a Remotion project, the video is
rendered there instead of in an editing program. Then the cut list must be
**machine-actionable**, not just readable.

Also output a timeline table — times in **frames at 30 fps**, not seconds, because
Remotion works in frames:

```
CLIP  SOURCE-IN   SOURCE-OUT   DURATION   ZOOM   OVERLAY
01    00:01:53.9  00:02:00.7   204f       1.00   title
02    00:02:40.8  00:02:49.1   249f       1.14   –
03    00:03:18.7  00:03:22.0   100f       1.04   quote
```

Rules that follow from this structure:

- **Zoom has two jobs.** The *start value* per clip carries the cut: two
  consecutive clips never get the same one, otherwise the cut reads as a mistake. The
  *movement within* the clip carries the take: roughly one percent per second of clip
  length, 2 to 7 percent in total. A one-percent move across an entire take is
  invisible — the image just sits still and the cut lives on the edges alone.
- **The end value is the limit, not the start value.** It decides how tight
  the frame gets, and must be checked against the safe zone (see below).
- As you zoom, the head drifts out of frame unless the face sits dead center.
  Compensate with `translateY = (scale − 1) × offset`, where
  `offset ≈ (0.5 − faceY) × 1920` and faceY is the face's vertical position as a
  fraction of frame height — measure it once per setup, don't guess.
- **No transitions.** Hard cuts, a 2-frame audio fade on every edge (1 frame when
  a margin is under 0.07 s, see "Cut Without Silences").
- **Zoom can also have two states** instead of a ladder of steps: default 100 % and
  zoomed 130 %, alternating on every cut, plus the drift. Which variant applies is
  in the account's format file; at 130 %, measure the zone.
- **Always cut sub-clips from the raw footage** (raw time = clip start + offset),
  never as a second encode from finished clips. A tail may reach past the old clip
  end into the raw footage.
- **Generate the timeline and the documentation table, don't maintain them by
  hand.** Keep manual decisions (a longer tail, a forced cut) as per-clip overrides in
  the plan so the rest stays reproducible and only affected parts get re-cut.
- **Remotion's bundled ffmpeg lacks `astats`, `fps`, `showspectrumpic`, `tile` and
  scene detection;** `crop` needs an even width, frame rate goes through `-r`. Level,
  zero-crossing rate and pitch are computed in plain Python over `wave`/`array`. Cuts
  in someone else's video are found via the frame difference of tiny grayscale
  frames (head movements produce false hits).
- Overlays are **post-production, not props** — title and quote cards are
  created during rendering. None of it is printed or held in hand.
- Only use the documented overlay types. A new type is a design decision, not
  a cutting decision — propose it, don't introduce it unilaterally.
- **Overlays sit on the word, not on the cut.** A number fades in when it's
  spoken, not when the clip begins.
- **Never attach overlays to clip indices, always to clip identifiers.** If a
  clip is split later — say, to remove a stalling pause — every index after it
  shifts, and the overlay silently ends up on the wrong take. The error produces no
  warning.

## Safe Zone

**Every element stays in the safe zone — no exceptions.** The platform lays its
controls over the video; what's underneath is invisible while watching but visible
during rendering. So the mistake only shows up once the video is live.

For Instagram Reels at 1080×1920:

```
clear:  x 60 … 1021        y 296 … 1533
from y ≥ 1179 the right side is also occupied → clear only up to x 886
```

At the top, the header area covers 296 px; at the bottom, the caption and control bar
take up 386 px. **The clear area isn't a rectangle:** from `y = 1179` the action
column (like, comment, share, menu) intrudes on the right.

Subtitles usually sit at exactly this height. So for anything below 1179 that's
centered on the middle of the frame, a maximum text width of **692 px** applies —
otherwise the last word slides under the share button.

If the account's visual-style file defines its own values, those apply instead. State
the position for every overlay and check it against these limits instead of
estimating.

### The Speaking Person Also Belongs in the Zone

**The zone applies to every element in the frame, not just text.** Too tight a crop
pushes the head into the area the platform covers. This doesn't show up when watching
the cut on its own — only in comparison with the raw footage, or once the video is
already live.

Measure after every render, don't estimate:

```
Cut out the center column of the frame as a narrow strip, and from the top find the
first row after which a good dozen rows stay dark — that's the hairline.
Measure at 75% of the clip duration, not the middle: that's where the zoom move is
almost finished and the frame is tightest.

Target: a clear margin below the upper zone boundary, at least about 150 px.
```

Calibrate the threshold to the wall's brightness, don't guess: a bright white
wall can easily sit at luma 150, in which case detection triggers on the very
first row.

## Animation

Motion serves readability, not effect.

**Subtitles track along with the spoken text**, highlighted word by word: the word
currently being spoken is emphasized in color, the rest stays put. The block itself
doesn't move — no flying in or out, no position changes — and stays on screen across
cut points, because it belongs to the audio, not the clip. In practice that means: the
first subtitle page of a clip begins with the clip, not with the first word, otherwise
a gap flashes at every cut point.

**Highlight whole words, not syllables.** Speech recognition returns sub-word pieces
("what" + "ever"). Colored individually, the emphasis runs through the middle of a
word. Merge before display: anything that doesn't start with a space belongs to the
previous piece.

**Set subtitles in regular weight, not bold.** The outline carries the readability.
Bold reads as loud — use it only where the visual style explicitly calls for it.

**Scale the outline with the font size.** At small sizes, a letter stem is only a
few pixels wide; an outline that was correct for large text then eats into the core
and the text color disappears inside it — the type looks gray even though the color is
set correctly. Rule of thumb: at most one third of the stem width, with a soft shadow
taking over the rest. When in doubt, measure on the rendered image what share of the
bright pixels actually carries the intended tint.

**Cards and overlays appear in stages, not all at once.** Fading everything in
simultaneously reads as a freeze frame, no matter how clean the curve is:

```
Frame  0   Card         0.90 → 1.00 over 10 frames, soft spring, no overshoot
Frame  5   Text         0.96 → 1.00, multi-line rows offset by ~4 frames each
Frame 13   Sticker      Pop, see below
```

Exit: simply fade out the whole card. No rotating, no wobbling.

**Stickers placed on top are allowed to pop.** An element that visibly sits on the
card — a paperclip, a badge — gets a real spring with overshoot: it snaps past 1.00
and settles back. The scaling anchor point sits at the **edge where the element is
attached**; scaled from the center, it drifts out of place as it pops open.

The sticker is the **only** element allowed to overshoot. The card and text ease in
without overshoot — the pop belongs to the element sitting on top, not to the surface.

**Pop sound, if the visual style calls for it.** Effect libraries have clicks and
memes, but rarely a pop; it can be synthesized: a good 100 ms, pitch falling
exponentially from about 600 to 150 Hz, a very fast attack, short decay, a few
milliseconds of noise for the attack of the "P".

The sound must land in a **speech gap** and stay well under the voice level.
Cross-check before delivery: measure the effect's level against the voice level,
don't estimate by ear.

**Numbers don't count up.** A counter draws attention to itself instead of the number.

**Evidence stays still.** A screenshot shown as proof is meant to be read — any
motion on it is a distraction.

For every overlay, state **when it appears and how long it stays**, in frames. The
overlay sits on the word, not on the cut.

## Final Check on the Finished Video

Watch it again after rendering and confirm or flag **each point individually**. These
are exactly the errors that survive the first pass:

**Sound**
- No slip of the tongue, no doubled sentence start, no leftover "um" – **checked
  per speech region, not just on the finished piece**: recognition swallows filler
  words and then falsely reports a take as clean
- No stalling pause left in the middle of a sentence
- Any effect sound, if present, falls in a speech gap and stays under the voice
- No laughter and no cutoff missed during editing
- No stage direction left in the audio
- No clicks or pops at the cut points (measure the sample jump at every edge)
- No clipped sound at the beginning or end
- **No clipped word ending** — especially quiet endings ("‑ing", "‑ful", "‑ct"):
  transcribe the rendered audio in full and read it against the script
- With a cut without silences: list the remaining pauses with timestamps in the docs

**Picture**
- No two consecutive clips with the same start zoom
- The zoom move within the clips is visible, not just present on paper
- **At 75% of every clip's duration, the head sits with margin below the upper
  zone boundary** – measured, not estimated
- No visible jump in brightness or color between clips

**Text**
- Subtitles match the spoken words verbatim — especially for technical terms
- No subtitle stays on screen longer than the sentence, none appears too early
- Title and quote cards spelled correctly, including special characters
- Nothing sits in the control zone: 296 px top, 386 px bottom, 60 px sides
- Below y 1179, nothing extends past x 886 — that's where the action column sits
- Subtitles in regular weight, not bold
- The highlighting runs in sync with the spoken words, not ahead or behind
- The highlighting jumps to whole words, not syllables
- The text color is actually visible in the rendered image and hasn't been swallowed
  by the outline

**Content**
- The hook is fully readable within the first two seconds
- An evidence screenshot, if present, is sharp enough that its source line is legible
- The payoff promised in the script actually happens
- The video ends on a sentence, not in a pause

Report issues with a **timestamp**, so they can be jumped to directly. If everything
checks out, say so in one sentence — no list of passed items.

## Wrap-up

In two to three lines, name what specifically needs attention when cutting this particular video – the one spot where it can go wrong. No general editing tips.

