Reel Studio — SkynetLabs Master Reel Skill
End-to-end production for every professional short-form reel. Default style: saddamh1. Decoded forensic-level.
Repo
<repo>/interview-clip-engine/
Master playbook (READ FIRST every session)
references/SADDAMH1-MASTER-PLAYBOOK.md — 12 sections, single source of truth.
- Visual ID card (Montserrat Bold 700,
#FFFFFF/#F7E043/#000000stroke 2px) - Hook architecture (5 tactics + 10 verbatim hooks)
- Body architecture (25-30 cuts/min, push-in 1.0→1.06)
- Caption rules (lower-third, color rotation green=money/red=loss/yellow=CTA/gold=default/cyan=tech)
- B-roll cards (light/dark/red/purple/green + 10 verbatim patterns)
- SFX timing (exact dB + ms per type)
- Audio mix (one canonical ffmpeg line)
- Recording playbook (Pocket 3, framing, wardrobe, energy)
- 5 script templates (problem-solution / list-of-N / contrarian / story-payoff / comment-bait)
- Decision tree (topic → template → BGM → card → color)
- 12 anti-patterns
- 10-item pre-ship checklist
LOCKED defaults (2026-05-19 — user approved clip_03)
Reference clip: clips/DJI_20260518_0023.silcut_clip_03_.mp4 (visa for digital nomad, 74s).
The user approved this version ("always do such work") — these rules apply to EVERY new reel.
Mandatory thumbnail title card (first card, every reel)
start_sec = 0.2,duration_sec = 3.5— covers the IG/YT/TT thumbnail window.text= full episode title (10-14 words OK, render auto-shrinks via font ladder 120→60).style = "light"by default (white bg + black text + red accent on 1-2 nouns).illustration_prompt= 1 scenic nature anchor — see below.
B-roll illustration relevance (MANDATORY)
- Every
illustration_promptMUST reference a subject noun from the clip's hook/topic — NOT random food/sand/abstract. - For nomad/visa/freelance/travel topics → AT LEAST 1 scenic nature anchor per clip: rice terrace · beach drone · palm sunset · scooter street · jungle · cliff · roadside cafe · infinity pool.
- Prompt template:
minimalist editorial photo of <topic-subject>, <scenic location>, golden hour, muted dark palette, soft warm glow, vertical 9:16, cinematic depth of field, no text no words no captions
Card cadence
- 4-6 cards per ~75s clip (was 3 — too sparse).
- Card 1 = thumbnail title.
- Cards 2-5 = key beat punctuation, ~5-10s gap.
- Last card = payoff/CTA in final 4s.
Contrast + legibility (MANDATORY — 2026-05-19 fix)
Card text MUST be readable at 1.5s pause-frame inspection. Original render failed: black text on dark-scrim scenic image = invisible.
- Light style + bg_image → use WHITE scrim
(255,255,255,140)to LIFT image, NOT dark scrim. Black text + red accent then pops. - Dark style + bg_image → keep dark scrim
(0,0,0,160). White text + yellow accent. - Every word renders with a 4px stroke outline (white on light cards, black on dark cards) + soft drop shadow — survives any busy bg.
- Pre-ship QA: scrub to each card start+1s. If text is mostly the bg color, regen the card.
- B-roll image scenes must be BRIGHT enough — avoid all-dark cliff/night/silhouette shots. Aim for golden hour / midday / turquoise / sunlit subjects. Dark cliff silhouette REJECTED 2026-05-19.
Fade timing (slow OUT for readability — 2026-05-19 fix)
- Fade-in:
0.35s(was 0.20 — slightly slower entrance feels editorial) - Fade-out:
0.80s(was 0.25 — viewer needs time to finish reading) - Card duration: 2.8s minimum for body cards, 4.5s for title + payoff/CTA cards (was 2.0/3.5 — too fast).
- Both fade durations encoded in
pipeline/stage_08c_broll_cards.py::build_overlay_filter.
Design + color (do NOT drift)
- Caption: Montserrat Bold, lower-third,
#FFFFFFbase,#F7E043exact yellow accent (NOT#FFFF00), 2px stroke, 4px shadow, pop-karaoke, sentence case (NOT all-caps). - B-roll text card colors: dark = black/white/
#F7E043· light = white/black/#FF3C3C. - Push-in zoom: 1.00→1.06 every shot.
- Voice EQ chain:
highpass=94, eq 200/-2, eq 3500/+2, eq 8000/+4.4, acompressor -18/3/8/180/+3, deesser. - BGM duck: vol 0.22 + sidechain compressor (threshold=0.05 ratio=8 attack=20 release=400) — sidechain ducks dynamically under voice peaks, more transparent than fixed vol cut.
- Loudnorm: I=-14 LRA=11 TP=-1.5 for IG/TT Reels (was I=-16 — IG Reels spec is -14 LUFS, not podcast -16). YT Shorts also -14. Use I=-16 only for podcast/long-form.
- End-card slate: always on.
Speed defaults (LOCKED 2026-05-25, REVISED 2026-05-25 eve)
- Multi-clip interview merge: 1.0x NATIVE (REVISED — was 1.15x). Mike Chang reel pack test showed 1.15-1.25x produces perceived lip-sync drift even when video/audio technically aligned. atempo<0.9 audio time-stretch introduces phase artifacts the eye reads as out-of-sync.
- Single saddamh1 clip: 1.0x (raw pace, no speedup — pause beats already cut).
- Kinetic-stoic: 1.0x (per-beat pacing is the format).
- NEVER atempo source talking-head audio. If clip too long, TRIM (in/out points) instead of speeding. Lip-sync is sacred.
- Reasoning: 1.25x kills retention on talking-head IG reels. Research (May 2026): IG Reels sweet spot retention = source-pace + light BGM, NOT speed-ramp.
Brightness lift for IG re-compression (MANDATORY)
- IG aggressively re-compresses, crushes dark frames + shifts tone-map. Pre-lift required.
- Apply at finalize stage:
eq=brightness=0.05:contrast=1.08:saturation=1.10 - Verify: scrub to 5/30/50/70% timestamps, sample frames. If a frame is still <40% avg luma, the B-roll source itself is too dark — SWAP it, do not just lift more (lifting kills contrast).
- BANNED B-roll sources for bright-feel reels: night cityscapes, cliff silhouette, dark cave, black tarmac, low-key portrait. Use golden hour / midday / turquoise water / sunlit subjects instead.
Background music (MANDATORY for engagement, May 2026 research)
- Always layer BGM under voice — silent reels lose 30-40% completion rate on IG.
- Default track pool:
assets/bgm/mixkit-cinematic-*.mp3(7 tracks, 100-226s, Mixkit Free license commercial-OK). - BPM target: 100-130 for talking-head (matches natural speech rhythm). Tested pick:
mixkit-cinematic-871.mp3(light, uplifting). - Mix: BGM at vol 0.22, sidechain-ducked under voice (see Design + color above). Fade out last 2s.
- IG 2026 algorithm: rewards original audio + voiceover combos — your voice IS the original audio, BGM under it is fine.
Name strap rotation (LOCKED 2026-05-25 — Mike Chang pack)
For interview-style or "guest + me" reels, rotate 4 straps (top center) w/ 0.4s alpha fade-in/out:
- Subject strap (0-T1): guest name + credibility (
MIKE CHANG / 7 MILLION+ FOLLOWERS) - Host strap (T1-T2): your name + positioning (
YOUR NAME / FOUNDER SKYNETLABS · CLAUDE CODE EXPERT) - Mid-roll CTA (T2-T3): build-tease (
EDITED BY CLAUDE CODE / GUYS - DM FOR FULL GUIDE) - Outro CTA (T3-end): connect ask (
LETS CONNECT / DM FOR INTERVIEW)
Per-clip duration timing table:
| Clip ≈ | T1 | T2 | T3 |
|---|---|---|---|
| 20-25s | 8s | 14s | 20s |
| 25-30s | 8s | 16s | 23s |
| 30-35s | 10s | 18s | 25-28s |
| 35-40s | 10s | 20s | 30-32s |
Adjust to natural sentence breaks — strap should never change mid-sentence.
Color palette (LOCKED 2026-05-25)
| Use | Hex | Text color | Notes |
|---|---|---|---|
| Brand primary pill | #F7E043 gold |
black | Subject strap default |
| Premium accent | #00897B teal |
white | outro strap (REPLACES ugly red) |
| Premium dark | #1A1A2E charcoal |
gold/white | Card body bg |
| Emphasis only | #E53935 red |
white | Word-level color pop, NEVER full strap bg |
| Host strap top | #4FC3F7 cyan |
black | Your name |
| Host strap sub | #0D47A1 deep blue |
white | Your credentials |
| Mid CTA pill | #E57345 orange |
black | Claude Code message |
| Sub text default | black @0.85 |
white | Always under top pill |
BANNED: #FF3C3C bright red as strap bg (cheap/alarmist). Use teal #00897B instead for any "urgent" CTA.
Intro/outro animated cards (LOCKED 2026-05-25)
For multi-clip packs, prepend 3s animated intro + (optionally) append 6s animated outro to give narrative arc.
Intro card recipe (3s):
- Bg: bright scenic B-roll image w/ slow Ken Burns zoom
zoompan=z='min(zoom+0.0008,1.06)':d=90:s=1080x1920:fps=30 - White scrim
drawbox=color=white@0.55:t=fillfor text legibility - Staggered text reveals via alpha (4-5 lines, each 0.4s gap):
- Line 1 at 0.2-0.5s
- Line 2 at 0.6-0.9s (bigger, color pop)
- Line 3 at 1.0-1.3s (smaller, context)
- Line 4 at 1.4-1.7s (BIGGEST payoff word)
- Line 5 at 1.8-2.2s (punctuation/question mark)
- BGM faded in 0.3s, faded out 0.4s before end
Outro card recipe (6s):
- Same bg + Ken Burns + scrim
- 6-8 cascading text reveals (1s apart) building to CTA
- Final strap: teal pill
FOLLOW @SKYNETLABS(the handle, always w/@) - All text stays on screen until last 0.5s (then global fade)
Concat: ffmpeg -i intro.mp4 -i main.mp4 -i outro.mp4 -filter_complex "[0:v][0:a][1:v][1:a][2:v][2:a]concat=n=3:v=1:a=1[v][a]" re-encodes ALL → seamless audio/video sync, same codec params end-to-end.
Dark B-roll mask-overlay technique (LOCKED 2026-05-25)
If source clip has dark B-roll cards baked in from older saddamh1 pipeline runs (night cityscapes, dark forests):
- Scan with
ffmpeg signalstatsat 0.3s intervals, identify windows with YAVG<60 - Generate bright scenic B-roll images via Pollinations or use
work/*/broll_bg/*.png - Overlay during dark window:
[scene]overlay=0:0:enable='between(t,X,Y)' - Source:
-loop 1 -i scene.png(NO-t— image stream must outlast overlay window) - Crop to vertical:
scale=1080:1920:force_original_aspect_ratio=increase,crop=1080:1920,setsar=1
drawtext gotchas (caught 2026-05-25)
| Gotcha | Symptom | Fix |
|---|---|---|
% literal in text |
drawtext silently fails to render | Strip % or write "percent" |
\@ in single-quoted string |
bash escape leaks | Use double quotes OR keep @ as-is (drawtext doesn't interpret it) |
-t N on -loop 1 input |
overlay enable past N → nothing renders | Drop -t, use -shortest on output |
Newlines in -filter_complex |
parse error | Flatten to single line OR use here-doc carefully |
| atempo<0.9 on speech | lip-sync looks off even when timing aligned | Don't speed-shift voice. Trim instead. |
Reusable templates
Copy & adapt from templates/merge_pack/:
brand_main.sh— 4-strap rotation render for N clipsbuild_intros_outro.sh— animated intro/outro card generatorconcat_final.sh— intro+main+outro stitcherREADME.md— full quick-start + locked defaults
Trigger: when user has 2+ interview clips and wants branded social-pack output.
3 modes
Mode 1: interview
Long-form raw video (interview, podcast, lecture) → 5-10 viral 9:16 shorts.
cd <repo>/interview-clip-engine
python run.py --input raw/episode.mp4
python run.py --url "https://youtube.com/watch?v=..."
Mode 2: saddamh1 (default for new content)
User records solo talking-head per generated script → I edit indistinguishable saddamh1 output.
# Step 1: generate script
python script_gen.py --topic "your topic" --duration 45 --tone tough-love
# → outputs/scripts/<slug>.md with HOOK + BEATS + PAYOFF + CTA + camera brief + B-roll cues + SFX + BGM + post caption
# Step 2: user records per brief, saves MP4 in raw/
# Step 3: edit
python run.py --input raw/my-take.mp4 --single-speaker
Mode 3: tts (voiceover, no recording)
Script → ElevenLabs TTS → animated captions + B-roll cards + BGM.
- Adapt from
_VIDEO-INVENTORY/PENDING/voiceover-batch-2026-05-16/pipeline - Status: PENDING migration into pipeline/10_tts_voiceover.py
Mode 5: mograph (motion-graphics explainer, no face cam) ⭐ NEW 2026-05-17
Decoded forensically from @beingmayy (After Effects MORPHING tutorial, 33s 16:9) + @mister.usb (Mac mini replaces streaming-stack, 67s 9:16). Flat-bg explainer w/ bold typography + UI mockup chips + 3D Apple emojis + product photos w/ soft blue halos. Voiceover-driven, NO face cam.
Playbook: references/mograph/MOGRAPH-MASTER-PLAYBOOK.md (12 sections —
visual ID, slide grammar, 11-chip library, glow specs, hook bank, anti-patterns,
pre-ship checklist).
11 slide types (consumed from script JSON):
typography— bold mixed-weight word reveal (withweight_mixfor multi-size rows)chip-timer-pill·chip-search·chip-imessage·chip-button·chip-youtube·chip-track-order·chip-iphone·chip-phone-screen·chip-timelineproduct-photo— centered w/ soft blue halo + headline aboveicon-halo-cluster— 3+ icons w/ halos in triangle layoutend-cta— avatar + handle + socials row + animated FOLLOW + cursor
Run:
# Manual — author script JSON yourself, render:
python mograph_reel.py --script examples/mograph-sap-n8n.json
# Auto — 1-line topic → Gemini drafts full slide JSON (mograph schema) → render:
python mograph_script_gen.py --topic "SAP just bought into n8n at $5.2B" --render
python mograph_script_gen.py --batch outputs/topics-week-21.txt --render
# → clips/mograph_*.mp4 (~25-30s, 9:16, 10-13 slides, hard cuts)
4 ready samples in examples/:
mograph-sap-n8n.json(13 slides, 27.4s) — SAP buys n8n at $5.2Bmograph-apple-claude-wwdc.json(12 slides, 24.8s) — Apple opens Siri to Claudemograph-aeo-geo-killed.json(11 slides, 25.2s) — Google's AEO/GEO is still SEOmograph-chip-showcase.json(13 slides, 22.8s) — exercises all 11 chip primitives
Topic fit: "X tool hit $Y valuation" / "Big company surprising move" / "Old vs new way" / "X pays for itself" / "You don't need to be Y to do Z". NOT good for personal stories (no face = no emotion anchor).
Latest test (2026-05-17 PM): SAP buys n8n at $5.2B sample reel rendered clean — 13 slides distinct, FOLLOW CTA fires w/ cursor + 4 socials (IG/YT/TT/LI).
Reference assets: references/mograph/refs/ (2 source videos staged).
Decode scripts: references/mograph/analyze_mograph.py + deep_decode_mograph.py
(pending Gemini re-quota — playbook synthesized from manual frame-by-frame decode).
Format: kinetic-stoic (text-only, no face)
Inspired by @naval / @ryanholiday / @dailystoic. Premium thought-leader reels.
- Cream #F4F1EA bg + charcoal text + gold #8B6F47 accent
- Fraunces Bold serif (downloaded to
assets/fonts/) - Word-by-word reveal w/ cross-dissolve between beats
- No source video — pure synthesis from quote text
- Ambient piano BGM ducked
- 5-10s per reel typically
python kinetic_reel.py --quote "Do what's needed. Not what you want." --emphasis "needed,want"
python kinetic_reel.py --batch outputs/kinetic-quotes-pack.txt --emphasis "obstacle,path,consistency"
python kinetic_reel.py --quote "..." --voice tts.wav # add VO
Use when: B2B / agency / luxury client targeting. Batchable from existing story posts. Zero recording required.
Mode 6: aeo-daily (skynet-aeo-engine bridge) ⭐ NEW 2026-05-19
Wires daily AEO content into reel-studio. 3 variants per AEO daily output → 5 channel slots.
Bridge: <repo>/skynet-aeo-engine/scripts/build_videos.py
| Variant | Duration | Style | Source script (extended schema) | Fallback (legacy) | Channels |
|---|---|---|---|---|---|
aeo-daily-biz-pro |
60-75s (target 67s) | saddamh1 talking-head, business voice, F7E043 yellow + green-on-money | copy.business.linkedin.post |
copy.li_post |
ig-pro, yt, tt-pro |
aeo-daily-travel-narrative |
45-60s (target 52s) | kinetic-stoic text reel + scenic DJI Ken-Burns (NO face, NO VO) | copy.travel.ig_travel_1.caption |
synth from copy.anchor |
ig-travel-1 |
aeo-daily-travel-tiktok |
30-45s (target 38s) | saddamh1-lite Hinglish, warm gold/coral/turquoise palette | copy.travel.tt_travel_1.script |
synth from copy.anchor |
tt-travel-1 |
End-card handles (mandatory):
- biz reels →
example.com(agency) - travel reels →
@yourhandle(personal)
Voice-lint guard: travel variants HARD-FAIL if script contains aeo / agency / client / skynetlabs / linkedin / ghl / n8n / saas / mrr / fiverr / upwork. Scrub copy.json before re-run.
Source-MP4 lookup (real pipeline):
- Talking-head expected at
interview-clip-engine/raw/aeo-daily-YYYY-MM-DD.mp4 - If absent → biz-pro + travel-tiktok fall back to
tts_reel.py(auto ElevenLabs > Edge > pyttsx3) travel-narrativeNEVER needs source MP4 (kinetic_reel.py + Ken-Burns layer)
Run:
# Smoke test (ffmpeg colorbars, no deps) — proves orchestration
cd <repo>/skynet-aeo-engine
python scripts/build_videos.py --smoke
# Real pipeline (today)
python scripts/build_videos.py
# Specific date
python scripts/build_videos.py --date 2026-05-19
# One variant only
python scripts/build_videos.py --only travel-tiktok
Output (predictable for schedulers):
skynet-aeo-engine/outputs/<date>/business/ig-pro/reel.mp4
skynet-aeo-engine/outputs/<date>/business/yt/reel.mp4
skynet-aeo-engine/outputs/<date>/business/tt-pro/reel.mp4
skynet-aeo-engine/outputs/<date>/travel/ig-travel-1/reel.mp4
skynet-aeo-engine/outputs/<date>/travel/tt-travel-1/reel.mp4
skynet-aeo-engine/outputs/<date>/video_build_report.json
New run.py flags (added 2026-05-19):
--script-text "..."— persist script text alongside the work dir--slug ig-pro— predictable output filename override--synth-colorbars— emit ffmpeg colorbars MP4 at variant's target duration + AR (smoke-test gate)
Smoke test results 2026-05-19 (3 colorbar mp4s):
- biz-pro → 1080×1920 × 67.0s ✓ (fans to 3 channel slots)
- travel-narrative → 1080×1920 × 52.0s ✓
- travel-tiktok → 1080×1920 × 38.0s ✓
3 caption variants (style mixing for fatigue prevention)
Per Agent C competitor research — rotate variants every 4 reels:
| Variant | When | Inspired by | Look |
|---|---|---|---|
saddamh1-default |
75% of reels | saddamh1 | Lower-third, Sentence case, #F7E043 yellow + cyan/green accents, dense color-pop |
iman-premium |
every 4th reel | @imangadzhi | lowercase Inter Bold 52px, minimal color (white + 1 accent), ambient pad BGM -22dB, gentle push-in 1.0→1.03, teal/orange or warm-muted grade, 24fps cinematic, AR toggle 9:16/16:9 |
bartlett-podcast |
interview cutdowns | Steven Bartlett (Diary of a CEO) | stacked 2-cam 1080×960+1080×960 (top:speaker close / bottom:wide both), diarization-driven cam switch (120ms lead + 800ms min hold), dual-color caps (host #F7E043 / guest #FFFFFF), 3s hook ribbon w/ name + EP#, podcast BGM fade-out at 2s, "Watch full episode" end card |
| substance-caps | personal-brand / founder talking-head | Submagic "Hormozi 2" + UK coach reels (decoded 2026-06-01) | MID-SCREEN 2-line stack, lead words WHITE + punch word GOLD #E8C87E bigger, Montserrat Black caps word-pop, full-frame espresso #3C2422 break-cards (lowercase gold word) as pattern-interrupts, occasional Playfair-italic soft phrase, warm-clean grade |
Run via --variant <name>. Default = saddamh1-default.
substance-caps (NEW 2026-06-01) — proven clone of the "stop polishing, start substance" reference reel.
Preset: config/presets/substance-caps.yaml. Standalone renderer: tools/substance_caps_render.py
(midcaps-twotone ASS + espresso break-cards + serif soft-phrases in ONE ffmpeg pass).
Proof: work/substance-caps-PROOF.html (ref-vs-clone side-by-side). Fonts shipped in assets/fonts/
(Montserrat-Black, Anton, PlayfairDisplay-Italic). To run on a clip: whisper word-stamps → substance_caps_render.py.
TODO: fold the midcaps-twotone branch into pipeline/stage_08_burn_caps.py keyed by captions.style
so run.py --variant substance-caps routes the full multi-platform pipeline.
v0.4.0 2026-05-19 — iman-premium + bartlett-podcast upgraded from caption-only to full-spec modes:
- Config presets:
config/presets/iman-premium.yaml(99 lines) +config/presets/bartlett-podcast.yaml(132 lines) - Reference playbooks:
references/iman-premium/IMAN-PREMIUM-PLAYBOOK.md(12 sections) +references/bartlett-podcast/BARTLETT-PODCAST-PLAYBOOK.md(12 sections) - Wired today: captions, voice EQ, BGM, push-in, name strap (bartlett), B-roll cards (bartlett)
- Wired v0.4.0 (2026-05-19 PM):
stage_11_end_card.py(shared, 2 flavors — clean-fade-handle for iman + watch-full-episode for bartlett, Pillow slate + ffmpeg concat w/ audio fade-out) ·stage_06b_multicam_stack.py(bartlett signature, vstack top:cam_b 1080×960 + bottom:cam_a 1080×960, single-cam pass-through fallback if--cam-babsent, diarization-driven swap = v2 TODO) ·stage_07c_color_grade.py(iman LUT apply vialut3d=, ffmpegeq+colorbalance+curvesfallback w/ preset-driven dict from YAMLfallback_filter) - Wired CLI flags:
--cam-b,--handle,--hook-name,--hook-episode,--ar 9:16|16:9,--skip-grade,--skip-endcard. Pipeline routing inrun.py: iman → 07c + 11, bartlett → 06b (if--cam-b) + 11. Drop a.cubeLUT atassets/luts/teal-orange-cinematic.cubeto swap fallback for cinematic grade. - Still stubbed:
stage_08dual-color speaker routing (host yellow / guest white via diarization tags per word),stage_08e_hook_ribbon(3s top-third overlay w/ speaker name + EP#), BGM fade-out-at-mark in stage_09 finalize (afade=t=out:st=0:d=2on BGM track only) - Smoke v0.4.0 (3-sec ffmpeg colorbars source): stage_07c teal-orange fallback grade renders 1080×1920 → 1080×1920 ✓. stage_06b 2-cam vstack renders 2× landscape → 1080×1920 ✓. stage_06b single-cam fallback (no
--cam-b) → pass-through copy ✓. stage_11 iman flavor: 3s clip + 2.0s slate → 5.03s output ✓. stage_11 bartlett flavor: 3s clip + 2.5s slate → 5.54s output ✓. - Hand-validate: real talking-head clip (e.g.
raw/DJI_iman.MP4) end-card text legibility at iPhone preview size (handle@yourhandleInter Bold 64px → may want bigger), then drop a real.cubeLUT for the cinematic grade pass.
Stack ($0 forever)
| Component | Tool | Purpose |
|---|---|---|
| Silence kill | unsilence (replacing auto-editor) |
30-50% runtime save |
| Transcribe | faster-whisper large-v3 GPU | Word timestamps |
| Diarize | pyannote-audio 3.1 | Speaker turns (interview mode) |
| Hook detect | Gemini 2.5 Flash native video | 1hr ctx free tier |
| Hook timestamp snap | custom (fuzzy match Whisper) | Fix Gemini ±15s drift |
| Cut | ffmpeg | Frame-accurate |
| Reframe | MediaPipe face-track | Horizontal → 9:16 |
| Zoom | ffmpeg zoompan (1.0→1.06 push-in) | saddamh1 signature |
| B-roll cards | Pillow + ffmpeg overlay | Gemini picks card text per clip |
| Captions | ASS karaoke (upgrade to pycaps planned) | Word-by-word selective highlight |
| SFX layer | ffmpeg amix | impact + pop on emphasis + ding on numbers |
| Voice EQ | ffmpeg afilter chain | highpass + presence + de-ess + compressor |
| BGM | Mixkit cinematic ducked-low | Sidechain compress |
| Loudness | alimiter + loudnorm -16 LUFS | Platform-spec |
All free, all local-first, all Windows-tested on RTX 4060.
API keys (in interview-clip-engine/.env)
GEMINI_API_KEY=... # https://aistudio.google.com/app/apikey (free)
HF_TOKEN=... # https://huggingface.co/settings/tokens + accept pyannote license
GROQ_API_KEY= # optional Whisper fallback
Decision tree (topic → settings)
| Topic family | Template | BGM mood | Card style | Highlight color | Variant |
|---|---|---|---|---|---|
| Money / income | T1 problem-solution | motivational-uplift | dark + green accent | green (#00FF00) |
saddamh1-default |
| Skill / future-threat | T3 contrarian | cinematic-tense | red bg | red (#FF3B30) |
saddamh1-default |
| Discipline / mindset | T4 story-payoff | dramatic-cello | dark | yellow (#F7E043) |
iman-premium |
| List of N | T2 list-of-N | upbeat | light + numbered | cyan (#00FFFF) |
saddamh1-default |
| Comment-bait reveal | T5 comment-bait | trap-lite | dark + yellow CTA | yellow (#F7E043) |
saddamh1-default |
| Interview cutdown | n/a (interview mode) | ambient | lower-third name strap | white + 1 accent | bartlett-podcast |
Anti-patterns (NEVER do)
- Crossfade transitions — saddamh1 uses 97% hard cuts
- Captions covering faces — always lower-third
- ALL CAPS everywhere — only KEYWORDS uppercase
- Zoom-OUT on talking head (only on B-roll card reveals)
- BGM louder than -20 dB under voice
- Rainbow highlighting (only Gemini-tagged emphasis words colored)
- Mid-word phrase cuts in captions
- Skip alimiter before loudnorm (= peak clipping)
- Double loudnorm (= 2-pass instability)
- < 4K source (Pocket 3 native is fine)
- Card duration > 30% of clip total
- Sentence-case caption MarginV padding wrong (= covers face)
Pre-ship checklist (10 items)
- ✅ Hook timestamp snapped to real speech (not Gemini's ±15s guess)
- ✅ Captions in lower-third, not covering faces
- ✅ Only Gemini emphasis words colored (no rainbow)
- ✅ Peak ≤ -1.5 dBTP (alimiter active)
- ✅ Loudness target -16 LUFS (±2 OK for IG/TT)
- ✅ Push-in zoom 1.0→1.06 active
- ✅ B-roll cards 1-3 per clip, max 30% screen time
- ✅ SFX impact on hook, pops on emphasis words
- ✅ Voice EQ 6-stage chain applied
- ✅ BGM ducked ≤ -20 dB under voice peaks
Quick reference — common invocations
# Interview → 5-10 shorts
python run.py --input raw/long_interview.mp4
# YouTube interview URL
python run.py --url "https://youtube.com/watch?v=ABC"
# My talking-head recording, single speaker
python run.py --input raw/my_take.mp4 --single-speaker
# Skip silence cut (short clips)
python run.py --input raw/short.mp4 --skip-silence
# Generate script for new topic
python script_gen.py --topic "why most freelancers stay broke"
# Generate 10 scripts batch
python script_gen.py --batch outputs/topics-week-21.txt
# Re-render with different variant
python run.py --input raw/my_take.mp4 --variant iman-premium
Pipeline stages
01_silence_cut auto-editor (→ unsilence upgrade)
02_transcribe faster-whisper word-stamps
03_diarize pyannote (skipped for single-speaker)
04_hook_detect Gemini Flash native video
05_edl_build hook timestamp snap + word merge
06_cut_clips ffmpeg frame-accurate
07_reframe MediaPipe face-track 9:16
07b_zoomout push-in 1.0→1.06 (saddamh1 signature)
08c_broll_cards Gemini picks + Pillow renders + ffmpeg overlays
08_burn_caps ASS karaoke selective highlight
09_finalize voice EQ + BGM duck + alimiter + loudnorm
09b_sfx impact + pop + ding overlays
Per-stage skip flags: --skip-silence, --skip-reframe, --skip-zoom, --skip-broll, --skip-caps, --skip-bgm, --skip-sfx.
Roadmap (next sprint)
- Absorb
unsilencelib → replace auto-editor (1-day, top OSS win) - Absorb
pycaps→ upgrade ASS karaoke to CSS-styled animated captions (2-3 days) - Absorb
opensource-clippingB-roll fetch (Pexels API) + auto-thumbnail (1-2 days) - Absorb
bilingualsub→ stacked EN+UR subs for Pakistan reels (1 day) - Migrate voiceover-batch →
pipeline/10_tts_voiceover.py(mode 3 unlock) - Build
iman-premium+bartlett-podcastcaption variants - Multi-version A/B render (3 hook variants per clip)
- Auto-thumbnail generator for IG/YT
- Beat-synced cuts (librosa beat_track + ffmpeg concat)
Deprecated (use this skill instead)
→ merged heresaddamh1-replicator→ merged hereinterview-clipper→ superseded/video-editcommand