YouTube Shorts Generator
End-to-end pipeline that turns one long video into N viral-ready vertical clips. Each clip ships with a viral score (0–100), an opening hook line, and a one-sentence reason it should perform.
Reference implementation: https://github.com/SamurAIGPT/AI-Youtube-Shorts-Generator
When to use this skill
- "Generate shorts from this YouTube video"
- "Find the most viral 60-second clips in this podcast"
- "Auto-crop this interview to 9:16"
- "Give me TikTok clips from this lecture"
If the user only wants transcription, summarization, or thumbnails — this is the wrong skill.
Inputs to collect before running
Ask once, then proceed:
- Source — YouTube URL (preferred) or path/URL to an mp4
num_clips — default 3
aspect_ratio — default 9:16 (also: 1:1, 4:5)
language — default auto-detect (forwarded to MuAPI Whisper as ISO-639-1)
- Output JSON path — optional; if set, dump full result there
If the user gave a URL and nothing else, use defaults and don't block on questions.
Prerequisites (verify before first run)
- Python 3.10+
- A MuAPI key — set
MUAPI_API_KEY in .env. Powers download, transcription, highlight ranking, and clipping. If missing, stop and ask the user for it; do not invent one.
pip install -r requirements.txt inside a venv
If the repo isn't cloned yet, clone https://github.com/SamurAIGPT/AI-Youtube-Shorts-Generator.git into the working directory.
Pipeline (what to execute)
Run the eight stages in order. Each maps to a module in shorts_generator/.
- Download (
downloader.py) — pull the source video at the requested resolution (360/480/720/1080, default 720).
- Transcribe (
transcriber.py) — MuAPI /openai-whisper runs Whisper server-side and returns timestamped verbose_json segments. Billed per minute of audio.
- Classify content type — LLM tags the video (podcast / interview / tutorial / vlog / lecture / monologue) and density. Tune the highlight prompt per type.
- Chunk if long (
highlights.py) — videos > LONG_VIDEO_THRESHOLD (1800s default) are split into CHUNK_SIZE_SECONDS (1200s default) windows with CHUNK_OVERLAP_SECONDS (60s default) overlap so cross-boundary highlights aren't missed.
- Rank highlights — LLM scans each chunk through
VIRALITY_CRITERIA:
- Hook moments — strong opening line that stops the scroll
- Emotional peaks — laughter, anger, vulnerability, awe
- Opinion bombs — spicy, contrarian, debate-bait takes
- Revelation moments — "wait, what?" reframes
- Conflict — disagreement, tension, callouts
- Quotable lines — tight, screenshot-worthy phrasing
- Story peaks — climax of a narrative arc
- Practical value — actionable insight a viewer will save
Each candidate gets
start_time, end_time, score 0–100, title, hook_sentence, virality_reason. Aim for 30–75s clips unless content dictates otherwise.
- Dedupe — collapse overlaps. Rule: if two candidates overlap > 50%, keep the higher score, drop the other.
- Top-N selection — sort surviving candidates by score, take
num_clips.
- Vertical auto-crop (
clipper.py) — render each highlight at aspect_ratio. Auto-handles face tracking and screen recordings; no Haar cascades.
Invocation
CLI (the standard path):
python main.py "<YOUTUBE_URL>" \
--num-clips 5 \
--aspect-ratio 9:16 \
--output-json result.json
Python API (when embedding in another pipeline):
from shorts_generator import generate_shorts
result = generate_shorts(
"<URL>",
num_clips=5,
aspect_ratio="9:16",
)
for short in result["shorts"]:
print(short["score"], short["title"], short["clip_url"])
Batch mode — urls.txt with one URL per line:
xargs -a urls.txt -I{} python main.py "{}"
CLI flags reference
| Flag |
Default |
Notes |
--num-clips |
3 |
How many shorts to render |
--aspect-ratio |
9:16 |
9:16 for TikTok/Reels, 1:1 square, anything else by flag |
--format |
720 |
Source download resolution |
--language |
auto |
Whisper language code (e.g. en) |
--output-json |
— |
Dump full result (transcript + all candidates + clip URLs) |
Output schema
{
"source_video_url": "...",
"transcript": { "duration": 1873.4, "segments": [...] },
"highlights": [ /* every candidate, before top-N cut */ ],
"shorts": [
{
"title": "The one mistake that cost me $50K",
"start_time": 124.3,
"end_time": 187.6,
"score": 92,
"hook_sentence": "Nobody talks about this, but it killed my first startup...",
"virality_reason": "Opens with a number + regret, peaks on a contrarian lesson",
"clip_url": "https://.../short_1.mp4"
}
]
}
When reporting back to the user, surface for each clip: rank, score, time range, title, hook, and clip URL. Skip the raw transcript unless asked.
Tunable knobs
shorts_generator/highlights.py
VIRALITY_CRITERIA — reorder or extend signals
HIGHLIGHT_SYSTEM_PROMPT — duration sweet spot, hook rules, JSON schema
CHUNK_SIZE_SECONDS — 1200s default
LONG_VIDEO_THRESHOLD — 1800s default
CHUNK_OVERLAP_SECONDS — 60s default
shorts_generator/config.py (or env vars)
MUAPI_POLL_INTERVAL — 5s
MUAPI_POLL_TIMEOUT — 1800s
Whisper transcription
Audio is transcribed by MuAPI's /openai-whisper endpoint (server-side whisper-1, billed per minute). The CLI passes --language straight through; leave it empty for auto-detection, or pass an ISO-639-1 code (e.g. en) to lock it.
Failure modes — handle, don't paper over
- Whisper produced no segments — likely no detectable speech or a hard language. Retry with
--language <code> (correct ISO-639-1) before declaring failure.
- API key missing or rejected — surface the exact error; never fabricate a key.
- Job timed out — bump
MUAPI_POLL_TIMEOUT and retry; don't silently truncate.
- Highlight ranker returned <
num_clips — return what survived dedupe with a note; don't pad with low-score filler.
Done criteria
The skill is done when:
result["shorts"] has up to num_clips entries, each with a working clip_url.
- The user has been shown the ranked list (score, time range, title, hook, URL).
- If
--output-json was set, the file exists and parses.
If any clip URL 404s on a HEAD check, re-run just the crop stage for that highlight rather than re-running the whole pipeline.
1---2name: youtube-shorts-generator3description: Generate viral 9:16 YouTube Shorts (or TikTok/Reels clips) from a long-form YouTube URL or local video. Triggers on requests like "make shorts from this video", "extract viral clips from this YouTube link", "auto-clip this podcast", "find the best moments and crop vertical". Pipeline downloads the source, transcribes via MuAPI /openai-whisper, ranks highlights through a virality framework (hook / emotional peak / opinion bomb / revelation / conflict / quotable / story peak / practical value), dedupes overlapping candidates, and vertically auto-crops the top N as mp4s.4---5
6# YouTube Shorts Generator
7
8End-to-end pipeline that turns one long video into N viral-ready vertical clips. Each clip ships with a viral score (0–100), an opening hook line, and a one-sentence reason it should perform.
9
10Reference implementation: https://github.com/SamurAIGPT/AI-Youtube-Shorts-Generator
11
12## When to use this skill
13
14- "Generate shorts from this YouTube video"
15- "Find the most viral 60-second clips in this podcast"
16- "Auto-crop this interview to 9:16"
17- "Give me TikTok clips from this lecture"
18
19If the user only wants transcription, summarization, or thumbnails — this is the wrong skill.
20
21## Inputs to collect before running
22
23Ask once, then proceed:
241. **Source** — YouTube URL (preferred) or path/URL to an mp4
252. **`num_clips`** — default 3
263. **`aspect_ratio`** — default `9:16` (also: `1:1`, `4:5`)
274. **`language`** — default auto-detect (forwarded to MuAPI Whisper as ISO-639-1)
285. **Output JSON path** — optional; if set, dump full result there
29
30If the user gave a URL and nothing else, use defaults and don't block on questions.
31
32## Prerequisites (verify before first run)
33
34- Python 3.10+
35- A MuAPI key — set `MUAPI_API_KEY` in `.env`. Powers download, transcription, highlight ranking, and clipping. If missing, stop and ask the user for it; do not invent one.
36- `pip install -r requirements.txt` inside a venv
37
38If the repo isn't cloned yet, clone `https://github.com/SamurAIGPT/AI-Youtube-Shorts-Generator.git` into the working directory.
39
40## Pipeline (what to execute)
41
42Run the eight stages in order. Each maps to a module in `shorts_generator/`.
43
441. **Download** (`downloader.py`) — pull the source video at the requested resolution (`360`/`480`/`720`/`1080`, default `720`).
452. **Transcribe** (`transcriber.py`) — MuAPI `/openai-whisper` runs Whisper server-side and returns timestamped `verbose_json` segments. Billed per minute of audio.
463. **Classify content type** — LLM tags the video (podcast / interview / tutorial / vlog / lecture / monologue) and density. Tune the highlight prompt per type.
474. **Chunk if long** (`highlights.py`) — videos > `LONG_VIDEO_THRESHOLD` (1800s default) are split into `CHUNK_SIZE_SECONDS` (1200s default) windows with `CHUNK_OVERLAP_SECONDS` (60s default) overlap so cross-boundary highlights aren't missed.
485. **Rank highlights** — LLM scans each chunk through `VIRALITY_CRITERIA`:
49 - **Hook moments** — strong opening line that stops the scroll
50 - **Emotional peaks** — laughter, anger, vulnerability, awe
51 - **Opinion bombs** — spicy, contrarian, debate-bait takes
52 - **Revelation moments** — "wait, what?" reframes
53 - **Conflict** — disagreement, tension, callouts
54 - **Quotable lines** — tight, screenshot-worthy phrasing
55 - **Story peaks** — climax of a narrative arc
56 - **Practical value** — actionable insight a viewer will save
57 Each candidate gets `start_time`, `end_time`, `score` 0–100, `title`, `hook_sentence`, `virality_reason`. Aim for 30–75s clips unless content dictates otherwise.
586. **Dedupe** — collapse overlaps. Rule: if two candidates overlap > 50%, keep the higher score, drop the other.
597. **Top-N selection** — sort surviving candidates by score, take `num_clips`.
608. **Vertical auto-crop** (`clipper.py`) — render each highlight at `aspect_ratio`. Auto-handles face tracking and screen recordings; no Haar cascades.
61
62## Invocation
63
64CLI (the standard path):
65
66```bash
67python main.py "<YOUTUBE_URL>" \
68 --num-clips 5 \
69 --aspect-ratio 9:16 \
70 --output-json result.json
71```
72
73Python API (when embedding in another pipeline):
74
75```python
76from shorts_generator import generate_shorts
77
78result = generate_shorts(
79 "<URL>",
80 num_clips=5,
81 aspect_ratio="9:16",
82)
83for short in result["shorts"]:
84 print(short["score"], short["title"], short["clip_url"])
85```
86
87Batch mode — `urls.txt` with one URL per line:
88
89```bash
90xargs -a urls.txt -I{} python main.py "{}"
91```
92
93## CLI flags reference
94
95| Flag | Default | Notes |
96|------|---------|-------|
97| `--num-clips` | `3` | How many shorts to render |
98| `--aspect-ratio` | `9:16` | `9:16` for TikTok/Reels, `1:1` square, anything else by flag |
99| `--format` | `720` | Source download resolution |
100| `--language` | auto | Whisper language code (e.g. `en`) |
101| `--output-json` | — | Dump full result (transcript + all candidates + clip URLs) |
102
103## Output schema
104
105```json
106{
107 "source_video_url": "...",
108 "transcript": { "duration": 1873.4, "segments": [...] },
109 "highlights": [ /* every candidate, before top-N cut */ ],
110 "shorts": [
111 {
112 "title": "The one mistake that cost me $50K",
113 "start_time": 124.3,
114 "end_time": 187.6,
115 "score": 92,
116 "hook_sentence": "Nobody talks about this, but it killed my first startup...",
117 "virality_reason": "Opens with a number + regret, peaks on a contrarian lesson",
118 "clip_url": "https://.../short_1.mp4"
119 }
120 ]
121}
122```
123
124When reporting back to the user, surface for each clip: rank, score, time range, title, hook, and clip URL. Skip the raw transcript unless asked.
125
126## Tunable knobs
127
128- `shorts_generator/highlights.py`
129 - `VIRALITY_CRITERIA` — reorder or extend signals
130 - `HIGHLIGHT_SYSTEM_PROMPT` — duration sweet spot, hook rules, JSON schema
131 - `CHUNK_SIZE_SECONDS` — 1200s default
132 - `LONG_VIDEO_THRESHOLD` — 1800s default
133 - `CHUNK_OVERLAP_SECONDS` — 60s default
134- `shorts_generator/config.py` (or env vars)
135 - `MUAPI_POLL_INTERVAL` — 5s
136 - `MUAPI_POLL_TIMEOUT` — 1800s
137
138## Whisper transcription
139
140Audio is transcribed by MuAPI's `/openai-whisper` endpoint (server-side `whisper-1`, billed per minute). The CLI passes `--language` straight through; leave it empty for auto-detection, or pass an ISO-639-1 code (e.g. `en`) to lock it.
141
142## Failure modes — handle, don't paper over
143
144- **Whisper produced no segments** — likely no detectable speech or a hard language. Retry with `--language <code>` (correct ISO-639-1) before declaring failure.
145- **API key missing or rejected** — surface the exact error; never fabricate a key.
146- **Job timed out** — bump `MUAPI_POLL_TIMEOUT` and retry; don't silently truncate.
147- **Highlight ranker returned <`num_clips`** — return what survived dedupe with a note; don't pad with low-score filler.
148
149## Done criteria
150
151The skill is done when:
1521. `result["shorts"]` has up to `num_clips` entries, each with a working `clip_url`.
1532. The user has been shown the ranked list (score, time range, title, hook, URL).
1543. If `--output-json` was set, the file exists and parses.
155
156If any clip URL 404s on a HEAD check, re-run just the crop stage for that highlight rather than re-running the whole pipeline.