YouTube Transcribe
Smart YouTube video transcription with automatic fallback:
- Captions first — extracts existing subtitles (manual or auto-generated) via yt-dlp. Fast, free, no compute.
- Whisper fallback — when no captions exist, downloads audio and transcribes locally with the best available Whisper backend.
When to Use
Use this skill when the user wants to:
- Get a transcript or text version of a YouTube video
- Understand what a YouTube video says without watching it
- Summarize, analyze, or take notes from a YouTube video
- Extract subtitles or captions from a video
Triggers
- "transcribe this YouTube video"
- "what does this video say"
- "get the transcript of [YouTube URL]"
- "summarize this YouTube video" (transcribe first, then process)
- Any YouTube URL shared with a request to understand its content
Requirements
Required:
yt-dlp — for caption extraction and audio download
python3
For Whisper fallback (when no captions available):
ffmpeg — for audio processing
- One of these Whisper backends (auto-detected in priority order):
mlx-whisper — Apple Silicon native, fastest on Mac (pip install mlx-whisper)
faster-whisper — CTranslate2 backend, fast on CUDA/CPU (pip install faster-whisper)
openai-whisper — Original Whisper, universal fallback (pip install openai-whisper)
Usage
Basic — transcribe a video
python3 {baseDir}/scripts/transcribe.py "https://www.youtube.com/watch?v=VIDEO_ID"
Specify language for captions
python3 {baseDir}/scripts/transcribe.py "URL" --language zh
Force Whisper (skip caption check)
python3 {baseDir}/scripts/transcribe.py "URL" --force-whisper
JSON output
python3 {baseDir}/scripts/transcribe.py "URL" --format json
Save to file
python3 {baseDir}/scripts/transcribe.py "URL" --output transcript.txt
Options
| Flag |
Default |
Description |
--language |
auto |
Preferred subtitle/transcription language (e.g. zh, en, ja) |
--format |
text |
Output format: text, json, srt, vtt |
--output |
stdout |
Save transcript to file |
--force-whisper |
false |
Skip caption extraction, go straight to Whisper |
--backend |
auto |
Whisper backend: auto, mlx, faster-whisper, whisper |
--model |
auto |
Whisper model size: auto, large-v3, medium, small, base, tiny |
Environment Variables
| Variable |
Description |
YT_WHISPER_BACKEND |
Override Whisper backend selection |
YT_WHISPER_MODEL |
Override Whisper model size |
Auto-Detection
Whisper Backend (priority order)
- MLX Whisper — detected via
import mlx_whisper. Best for Apple Silicon.
- faster-whisper — detected via
import faster_whisper. Best for CUDA GPU, good on CPU.
- OpenAI Whisper — detected via
import whisper. Universal fallback.
Model Size (based on available RAM)
| RAM |
Model |
VRAM/RAM Usage |
| ≥16GB |
large-v3 |
~6-10GB |
| ≥8GB |
medium |
~5GB |
| ≥4GB |
small |
~2.5GB |
| <4GB |
base |
~1.5GB |
Caption Language Priority
When --language is not specified, captions are searched in this order:
- Video's original language
- Chinese variants:
zh-Hant, zh-Hans, zh-TW, zh-CN, zh
- English:
en
- Any available language
Output Formats
text (default)
Plain text transcript, one continuous block.
json
{
"video_id": "ZSnYlbIYpjs",
"title": "Video Title",
"channel": "Channel Name",
"duration": 708,
"language": "zh",
"method": "captions",
"transcript": [
{"start": 0.0, "end": 5.2, "text": "..."},
...
],
"full_text": "Complete transcript as single string"
}
srt / vtt
Standard subtitle formats with timestamps.
1---2name: youtube-transcribe3description: Transcribe YouTube videos with smart fallback: extracts captions first (fast, free), falls back to local Whisper transcription when no captions available. Auto-detects best Whisper backend (MLX/faster-whisper/openai-whisper) and model size based on hardware. Use when the user shares a YouTube link and wants to know what it says, get a transcript, summarize, or analyze video content. Keywords: YouTube, transcribe, transcript, subtitles, captions, speech-to-text, whisper, mlx, video to text.4---56# YouTube Transcribe78Smart YouTube video transcription with automatic fallback:91. **Captions first** — extracts existing subtitles (manual or auto-generated) via yt-dlp. Fast, free, no compute.102. **Whisper fallback** — when no captions exist, downloads audio and transcribes locally with the best available Whisper backend.1112## When to Use1314Use this skill when the user wants to:15- Get a transcript or text version of a YouTube video16- Understand what a YouTube video says without watching it17- Summarize, analyze, or take notes from a YouTube video18- Extract subtitles or captions from a video1920## Triggers2122- "transcribe this YouTube video"23- "what does this video say"24- "get the transcript of [YouTube URL]"25- "summarize this YouTube video" *(transcribe first, then process)*26- Any YouTube URL shared with a request to understand its content2728## Requirements2930**Required:**31- `yt-dlp` — for caption extraction and audio download32- `python3`3334**For Whisper fallback (when no captions available):**35- `ffmpeg` — for audio processing36- One of these Whisper backends (auto-detected in priority order):37 1. `mlx-whisper` — Apple Silicon native, fastest on Mac (pip install mlx-whisper)38 2. `faster-whisper` — CTranslate2 backend, fast on CUDA/CPU (pip install faster-whisper)39 3. `openai-whisper` — Original Whisper, universal fallback (pip install openai-whisper)4041## Usage4243### Basic — transcribe a video4445```bash46python3 {baseDir}/scripts/transcribe.py "https://www.youtube.com/watch?v=VIDEO_ID"47```4849### Specify language for captions5051```bash52python3 {baseDir}/scripts/transcribe.py "URL" --language zh53```5455### Force Whisper (skip caption check)5657```bash58python3 {baseDir}/scripts/transcribe.py "URL" --force-whisper59```6061### JSON output6263```bash64python3 {baseDir}/scripts/transcribe.py "URL" --format json65```6667### Save to file6869```bash70python3 {baseDir}/scripts/transcribe.py "URL" --output transcript.txt71```7273## Options7475| Flag | Default | Description |76|------|---------|-------------|77| `--language` | auto | Preferred subtitle/transcription language (e.g. `zh`, `en`, `ja`) |78| `--format` | `text` | Output format: `text`, `json`, `srt`, `vtt` |79| `--output` | stdout | Save transcript to file |80| `--force-whisper` | false | Skip caption extraction, go straight to Whisper |81| `--backend` | auto | Whisper backend: `auto`, `mlx`, `faster-whisper`, `whisper` |82| `--model` | auto | Whisper model size: `auto`, `large-v3`, `medium`, `small`, `base`, `tiny` |8384## Environment Variables8586| Variable | Description |87|----------|-------------|88| `YT_WHISPER_BACKEND` | Override Whisper backend selection |89| `YT_WHISPER_MODEL` | Override Whisper model size |9091## Auto-Detection9293### Whisper Backend (priority order)941. **MLX Whisper** — detected via `import mlx_whisper`. Best for Apple Silicon.952. **faster-whisper** — detected via `import faster_whisper`. Best for CUDA GPU, good on CPU.963. **OpenAI Whisper** — detected via `import whisper`. Universal fallback.9798### Model Size (based on available RAM)99| RAM | Model | VRAM/RAM Usage |100|-----|-------|----------------|101| ≥16GB | `large-v3` | ~6-10GB |102| ≥8GB | `medium` | ~5GB |103| ≥4GB | `small` | ~2.5GB |104| <4GB | `base` | ~1.5GB |105106## Caption Language Priority107108When `--language` is not specified, captions are searched in this order:1091. Video's original language1102. Chinese variants: `zh-Hant`, `zh-Hans`, `zh-TW`, `zh-CN`, `zh`1113. English: `en`1124. Any available language113114## Output Formats115116### text (default)117Plain text transcript, one continuous block.118119### json120```json121{122 "video_id": "ZSnYlbIYpjs",123 "title": "Video Title",124 "channel": "Channel Name",125 "duration": 708,126 "language": "zh",127 "method": "captions",128 "transcript": [129 {"start": 0.0, "end": 5.2, "text": "..."},130 ...131 ],132 "full_text": "Complete transcript as single string"133}134```135136### srt / vtt137Standard subtitle formats with timestamps.