Transcript Fetcher
Purpose: Given a video URL, produce a clean plain-text transcript
ready for downstream summarization or analysis. Single responsibility:
fetch + clean. Composes well with summarizing-meetings (run this
first, then feed the resulting .txt to that skill).
Dependencies: yt-dlp is vendored in scripts/.venv (installed
by scripts/install.sh) — it is NOT on $PATH and never will be. Do
NOT which yt-dlp, pip install yt-dlp, or brew install yt-dlp
globally — a missing $PATH entry does not mean yt-dlp is absent, it
means the check was wrong; this exact false-negative once caused a
manual out-of-skill workaround. The two canonical readiness probes are
scripts/.venv/bin/python -m yt_dlp --version and
scripts/.venv/bin/python scripts/fetch.py doctor (§4 Script Contract;
the latter also reports ffmpeg + ASR backends — see the ASR portability
note under §5 Safety Boundaries). Callers that shell in (e.g. a
downstream integrator) MUST invoke the venv interpreter directly
(scripts/.venv/bin/python), never a $PATH python.
1. Red Flags (Anti-Rationalization)
STOP and READ THIS if you are thinking:
- "I'll just paste the URL into the model and ask it to transcribe" -> WRONG. Models do not have audio access and will hallucinate. Always run
fetch.pyand read back the resulting.txt. - "Manual
rufailed, I'll just useenauto-translation, no warning needed" -> WRONG. Auto-translated English subtitles destroy idioms, names, and technical terms. Thequality_flag: english_auto_translationin the stat MUST be surfaced to the user. - "I'll skip writing the JSON stat sidecar, the .txt is enough" -> WRONG. The sidecar records WHICH track was picked. Without it, downstream cannot tell whether the transcript is high-quality manual subs or low-quality auto-translation.
- "The transcript has weird
>>markers, I'll strip them" -> WRONG. Those are speaker-turn boundaries. Removing them collapses multi-speaker meetings into a single voice and ruins downstream attribution. - "Auto-generated Russian (
ru-orig) is garbage, I'll prefereninstead" -> WRONG.ru-origis the actual Russian audio transcribed;ru(without-orig) is often an English auto-translation back to Russian.ru-orig>ru>en. - "This talk is only on a proprietary player, so I'll ASR it" -> CHECK FOR A MIRROR FIRST. Conference and webinar recordings are routinely re-published by the same organiser on YouTube with auto-captions, even when the marketing page embeds its own player. Search the organiser's channel for the exact talk title before spending hours of ASR. Real case: a 9-talk Yandex AI Studio Series page embedded a caption-less Yandex VH/Strm player, but all 5 underlying streams were on the official Yandex Cloud YouTube channel with full-coverage
ru-origcaptions — ~7.6 h of ASR avoided. Prefer the mirror for cost, not quality: a spot-check on the same talk found the local MacWhisper output was actually cleaner than the YouTube captions (no rolling-caption duplication), it just costs hours. ASR is the fallback, not the opening move. - "I'll add
ffmpegto the pip deps just in case" -> WRONG. The caption path (WebVTT parsing) is pure Python and needs nothing extra;ffmpegis a soft-optional external system tool, never a pip dependency. It IS genuinely required for the X ASR path on HLS sources (Broadcasts/Spaces) — yt-dlp uses it to extract a clean audio-onlym4a, and the skill fails fast (exit 7) when it is absent there — but it is detected at runtime (install_components.py), not bundled. Do not pull heavy packages intorequirements.txt.
2. Capabilities
- Fetch YouTube and Vimeo captions via
yt-dlp(no audio download). - Fetch X.com / Twitter — native status video AND Broadcasts/Spaces.
The X provider is captions-first: it reuses embedded
subtitles/automatic_captionswhen present, and only when none exist does it download the smallest media and transcribe via ASR. Fully automatic — no mode switch. ASR runs through a pluggable backend chain: MacWhisper (mw) → Whisper CLI → whisper.cpp → opt-in OpenAI/compatible cloud.ffmpegis required for the X ASR path on HLS sources (Broadcasts/Spaces): with it the smallest media is extracted to a clean audio-onlym4a; without it the skill fails fast (exit 7) because yt-dlp's no-ffmpeg HLS output is not a valid container the ASR engine can open. (For non-HLS progressive media, or when embedded captions exist, ffmpeg is not needed.) - Fetch Yandex VH/Strm recordings (
runtime.strm.yandex.ru/player/episode/<ID>,frontend.vh.yandex.ru/player/<ID>) — the player behind Yandex Cloud webinars and streams. ASR-only by design: that player carries no caption track at all (verified five ways — no subtitle key in the config JSON, zero#EXT-X-MEDIAtags in the HLS master, only video+lang="rus"audio AdaptationSets in DASH,yt-dlp --list-subsreports none, and a live player reportsvideo.textTracks: []), so a caption ladder would be dead code. yt-dlp has no extractor for this host, but its generic extractor takes the signed manifest directly — the adapter is therefore only a resolver (episode id -> config JSON -> signed DASH/HLS URL) and reuses the shared download/ASR pipeline unchanged. No auth: the config endpoint answers with zero headers. Two live traps the adapter handles — the config'sdurationfield lies (observed 4365 vs a real 3800 s, and 43970 vs a real 8825 s), so the real duration is the#EXTINFsum from the media playlist; and signed URLs are minted per request and expire (~48 h), so they are never cached.ffmpegis required here (HLS/DASH-only source); the skill fails fast with exit 7 without it. Before using this path, check for a YouTube mirror — see §1. - Fetch Skool lesson pages via a stdlib HTML scrape — public
communities work without auth; private/paid ones accept an optional
Netscape
cookies.txt. Then delegate embedded YouTube/Vimeo videos to those adapters; capture author-supplied transcript field when present. - Fall back through a configurable ladder (default for ru: manual ru -> auto ru-orig -> auto ru -> auto en).
- Clean captions to plain text — WebVTT, and for X also SRT and TTML/DFXP (
vtt/srt/ttml/bestpreference list, so a non-VTT track is no longer skipped to ASR): strip timestamps, inline timing tags, and rolling-caption overlap; decode HTML entities. TTML is parsed safely (DTD/entity declarations refused — XXE/billion-laughs guard). For X, captions-first is language-robust: if the requested--langhas no track but the post carries captions in another language, those are used (manual preferred, with a note) rather than dropping to ASR. - Preserve
>>speaker-turn markers as paragraph breaks. - Emit a JSON stat sidecar (chosen track, char count, speaker-turn count, quality flag, plus optional title/uploader/duration metadata). For X media it also records
transcript_origin(embedded-captions|macwhisper|whisper-cli|whisper-cpp|openai-api) so downstream skills know HOW the text was produced. - Optionally write
<out>.description.md(YAML frontmatter + Markdown body) when--with-descriptionis passed — gives you a ready-to-ingest description for RAG / Obsidian / human review. - Batch mode for processing multiple URLs from a text file.
- Source-agnostic architecture: each platform is one file under
scripts/sources/. The yt-dlp + ASR pipeline is shared (sources/_ytdlp_media.py+asr/), so a future TikTok/Twitch/Vimeo-ASR provider is one new file + one host entry — no pipeline changes. Zoom and podcast slots remain reserved. - Configurable + secrets-safe: a skill-local
.env(seescripts/.env.example) externalises every endpoint, model, and tool path. The cloud ASR endpoint works with any OpenAI-compatible server (Groq, self-hosted whisper). A.envholding an API key is refused unlesschmod 600(and not a symlink); the key is sent only in an HTTP header, never on argv or in logs.
3. Execution Mode
- Mode:
script-first - Why this mode: Fetching captions, parsing WebVTT, deduplicating rolling captions, and applying the fallback ladder are deterministic operations with > 5 lines of business logic. The CLI is the contract; SKILL.md is orchestration.
4. Script Contract
- Install (one-time): creates the venv + yt-dlp, then reports which optional ASR components are present:
bash skills/transcript-fetcher/scripts/install.sh # Optional ASR engines (for caption-less X media) — detect / install: ./scripts/.venv/bin/python scripts/install_components.py # status report ./scripts/.venv/bin/python scripts/install_components.py --install-whisper # pip openai-whisper into the venv ./scripts/.venv/bin/python scripts/install_components.py --system --run # brew/apt ffmpeg + whisper.cpp - Readiness check (
doctor, no network, import-free): answers "is this skill ready?" without$PATHguessing — the canonical replacement forwhich yt-dlp:
Reports the resolved interpreter, in-venv flag, yt-dlp version, ffmpeg, each ASR backend, and cloud opt-in state (key presence only, never the key)../scripts/.venv/bin/python scripts/fetch.py doctor # human-readable report ./scripts/.venv/bin/python scripts/fetch.py doctor --json # {v, interpreter, in_venv, ready, components, remediation}remediationnames a flow-blocking gap for EVERY such gap — yt-dlp missing, ffmpeg missing (it back-stops X Broadcast/Space HLS ASR and whisper/whisper.cpp at runtime even though it is itself optional), or no ASR capability at all (no local backend AND no fully-configured cloud) — but never an individual missing ALTERNATIVE local ASR engine while ASR capability already resolves elsewhere (mw present, another local engine present, or cloud configured); those show only as informational→install-hint lines in the human report. A fully-configured cloud backend (--asr-allow-cloud+ key) genuinely suppresses the no-local-ASR hint — it never appears inremediation, and the human report gets one informational note instead of a demand to install a local engine it does not need. Exit0when yt-dlp is present (regardless ofremediation),7when it is not. - Single URL:
Optional flags:cd skills/transcript-fetcher ./scripts/.venv/bin/python scripts/fetch.py <URL> --out <path/to/output.txt>--lang ru(default),--prefer manual|auto(defaultmanual),--with-description,--description-only,--cookies-file PATH,--json-errors,--debug(stage logging to stderr),--asr-allow-cloud(opt-in cloud ASR),--asr-model <id>,--asr-timeout-sec N,--max-duration-min N(X: transcribe only the first N minutes — clips the download when ffmpeg is present, the only case where the download itself is clipped; also clips the media-timeout floor below in that case),--keep-silence(X ASR: do NOT strip long silences before transcription — silence removal is ON by default to cut Whisper hallucinated filler on silent lead-in/out),--auth-map PATH(per-host cookies),--cookies-from-browser BROWSER(X: load cookies from a local browser via yt-dlp),--concurrent-fragments N(X: parallel HLS fragment downloads for the media, default 8;1= serial; values above 32 are capped at 32; CLI rejects<= 0with exit 2; envTRANSCRIPT_FETCHER_CONCURRENT_FRAGMENTSnon-positive/malformed falls back to 8 — these are three DISTINCT layers, not one shared clamp),--media-timeout-sec N(X: per-attempt budget for the media download only, separate from the probe's--timeout-sec; default duration-derivedmin(21600, max(600, duration*4))s — capped at 6h, and derived from the--max-duration-min-clipped duration when that flag is set AND ffmpeg is present (the only case where the download itself is clipped) — else1800s). - X.com / Twitter (captions-first, automatic ASR fallback; no mode switch). A Broadcast/Space usually has no captions → ASR via the first available local backend (MacWhisper, etc.):
cd skills/transcript-fetcher ./scripts/.venv/bin/python scripts/fetch.py \ "https://x.com/i/broadcasts/<id>" \ --out broadcast.txt --with-description --debug # → broadcast.txt + .stat.json (source="x", transcript_origin="macwhisper", # chosen_track_kind="asr"). A status video WITH captions skips ASR # (transcript_origin="embedded-captions"). Use --cookies-file for # protected/age-gated media, or drop a Netscape cookies.txt at # ~/.transcript-fetcher/x.com-cookies.txt for zero-flag auth # (see the X cookie contract in §5 Safety Boundaries). - Skool lesson (cookies needed ONLY for private / paid communities — public ones work without):
./scripts/.venv/bin/python scripts/fetch.py \ "https://www.skool.com/<community>/classroom/<id>?md=<lesson-id>" \ --out lesson.txt --with-description \ --cookies-file ~/.config/skool-cookies.txt - Batch:
./scripts/.venv/bin/python scripts/fetch.py --batch urls.txt --out-dir transcripts/ - Inputs: A YouTube / Vimeo / Skool-lesson URL (or a file with one URL per line). Empty lines and
#comments in the batch file are ignored. - Outputs:
<out>.txt— clean plain text, UTF-8.<out>.txt.stat.json— sidecar with the chosen track, quality flag, plus optional title / uploader / upload_date / duration_sec / embed_source / embed_url metadata.<out>.description.md— only when--with-descriptionis passed; YAML frontmatter + Markdown body.- One JSON stat record per URL on stdout.
- Failure semantics: Non-zero exit. With
--json-errors, stderr carries a single JSON line{v, error, code, type, details?}. Exit codes:2usage error (incl. malformed Skool URL, cookies file path missing on disk),3no transcript producible (no caption track in the ladder AND — for X — ASR produced nothing / every available backend failed, or the media download timed out — that case is transient/retryable and its remediation names--concurrent-fragments/--media-timeout-sec),4partial batch failure,5source-auth error (HTTP 401/403 — private Skool community needs cookies, X protected/suspended/age-gated media, or supplied cookies expired),6source rate-limit (HTTP 429),7missing dependency (yt-dlp absent, ffmpeg required-but-absent, or no ASR backend available for caption-less media —details.remediationcarries the hint),1unexpected. Whendetailscarries aremediationkey, it is ALSO printed as a second, plain stderr line (remediation: <text>) even WITHOUT--json-errors— the remedy is operator-visible either way. - Idempotency: Re-running overwrites the output file and sidecar. yt-dlp itself caches nothing the skill depends on; behaviour is reproducible given network availability.
- Dry-run support: Not currently exposed as a flag. Inspect the fallback ladder via
_build_ladderif needed.
5. Safety Boundaries
- Allowed scope: Reads from the network (yt-dlp HTTPS to YouTube/Vimeo; stdlib HTTPS to Skool). Writes ONLY to the user-specified
--out/--out-dirpath plus a.stat.jsonsidecar (and optionally a.description.mdsidecar) next to it. - Default exclusions: Never downloads the video itself (only
--skip-download+ subtitle tracks /--write-info-jsonfor metadata). Never writes to any path the user did not explicitly pass via--out/--out-dir. - Destructive actions: None. The script does not delete or modify any pre-existing files outside the chosen output paths. In batch mode, output path collisions are handled per
--on-collision={error,skip,suffix}(default:error). - Optional artifacts: The JSON stat sidecar is mandatory in single-URL mode (it is the audit trail for which track was used). The
.description.mdsidecar is written only when--with-descriptionis passed. - URL allowlist: Source dispatch is hostname-based against an explicit allowlist:
- YouTube —
youtu.be,youtube.com,m.youtube.com,music.youtube.com,youtube-nocookie.com(pluswww.variants). - Vimeo —
vimeo.com,www.vimeo.com,player.vimeo.com. - X / Twitter —
x.com,www.x.com,mobile.x.com,twitter.com,www.twitter.com,mobile.twitter.com(status…/status/<id>and…/i/broadcasts/<id>). - Skool —
skool.com,www.skool.com,app.skool.com; additionally URLs must match/<community>/classroom/<id>?md=<lesson-id>. Landing //about//calendarpages are rejected. - Yandex VH/Strm —
runtime.strm.yandex.ru,frontend.vh.yandex.ru,strm.yandex.ru(episode…/player/episode/<id>). Other*.yandex.ruhosts (e.g.music.yandex.ru) are NOT routed here. URLs that merely contain a supported host as a substring elsewhere are rejected.
- YouTube —
- Second-hop allowlist (Yandex only): the Yandex adapter resolves a stream URL out of a remote JSON document and hands it to yt-dlp, so that URL is untrusted input and is separately gated — https-only, against an exact-host set (
strm.yandex.ru,runtime.strm.yandex.ru,frontend.vh.yandex.ru,vh.yandex.ru). Deliberately not a*.yandexcloud.netsuffix rule: that is object storage where anyone can create a bucket, which would turn the allowlist into an open redirect. Every HTTP hop is fetched through the shared restricted opener, so a cross-host redirect is refused too (the pre-flight check on the original URL alone guarantees nothing about where the bytes came from), and each document read is capped at 8 MB. - Silence removal also runs on the Yandex ASR path (same default-on behaviour and
--keep-silenceopt-out documented for X below). - ASR backends (external, optional): For caption-less X media the skill shells out (argv arrays, never a shell string) to whichever local engine is present — MacWhisper
mw, Whisper CLI, or whisper.cpp. None is a pip dependency; all are probed at runtime. ffmpeg (also external) is required to turn an X Broadcast's HLS stream into a valid audio file; the skill fails fast with exit 7 (clear remediation, before any large download) when ffmpeg is absent for an HLS source.bash scripts/install.shreports which engines are available;scripts/install_components.pyguides/installs them (incl. ffmpeg). If no ASR backend is available (and cloud is not opted in), the run also fails cleanly with exit 7, never a traceback. ASR portability: the fallback chain resolves in ordermw→ Whisper CLI → whisper.cpp → (opt-in) cloud (§2 Capabilities) — a caption-less Broadcast/Space genuinely REQUIRES ffmpeg and at least one of these; a box with neither (e.g. a bare Linux/CI runner) fails hard on exit 7. Remediate withscripts/install_components.py --install-whisper(in-venv Whisper CLI;ffmpegitself needs the separate--system --run) or--asr-allow-cloud(+ an API key) to fall back to the cloud backend. Runscripts/fetch.py doctor(§4 Script Contract) before a long fetch to see which backends resolve, with zero downloads. - Cloud ASR egress (opt-in only): The OpenAI/compatible cloud backend is used only with
--asr-allow-cloud(orTRANSCRIPT_FETCHER_ASR_ALLOW_CLOUD=1) AND an API key present. When used, the audio leaves the machine to the configured endpoint — disclosed here and in the stat notes. Local backends are always tried first; cloud is the last resort. - Silence removal before ASR (X): before transcribing, the X path runs ffmpeg
silenceremoveto trim leading silence and collapse long interior/trailing gaps — this cuts Whisper-family hallucinated filler (e.g."Продолжение следует...") on silent lead-in/out. ON by default;--keep-silence(orTRANSCRIPT_FETCHER_SILENCE_REMOVAL=0) opts out;_THRESHOLD/_MIN_GAP_SEC/_KEEP_SECtune it. Only true silence is removed (music/speech survive — a music-only intro can still trigger filler, see KNOWN_ISSUES TF-X-6). Never fatal: ffmpeg absent or a filter failure transparently falls back to the original audio. The statnotesrecord what was stripped (silence-removal: stripped ~Ns ...); the original media is kept for the ffprobe duration fill. - Secrets: The API key is read from
OPENAI_API_KEY/TRANSCRIPT_FETCHER_OPENAI_API_KEYor a skill-local.env. A.envis loaded only at the CLI entry point and is refused if it is a symlink or notchmod 600(group/world-readable). The key is sent only in an HTTPAuthorizationheader — never on a command line, never logged..envis git-ignored; onlyscripts/.env.example(placeholders) is committed. - Per-host cookies (
~/.transcript-fetcher/): cookies for auth-walled media resolve (after an explicit--cookies-file) from a skill-local home folder — mirrors thehtmlskill's~/.html. Anauth-map.json(--auth-map/TRANSCRIPT_FETCHER_AUTH_MAP/~/.transcript-fetcher/auth-map.json) maps a host to its{cookies_file}, or the convention~/.transcript-fetcher/<host>-cookies.txtis used — e.g.~/.transcript-fetcher/x.com-cookies.txtforhttps://x.com/...URLs,~/.transcript-fetcher/twitter.com-cookies.txtforhttps://twitter.com/...URLs. The convention lookup tries the EXACT URL hostname first, then — ONLY for the three well-known mirror-prefix labelswww/mobile/m— the same file with that single label stripped (e.g.www.x.comandmobile.x.comboth also resolvex.com-cookies.txtwhen the exact-host file is absent); distinct domains are never aliased to each other (x.comandtwitter.comstill need separate files — no generic parent-domain walk). A custom filename (e.g.x-cookies.txt) REQUIRES anauth-map.jsonentry; the convention path only matches the literal<host>-cookies.txtname (or its single-label-stripped mirror variant). Host match is label-boundary (a keyx.commatchesx.com/*.x.com, neverevil-x.com); auth-map and convention files are hardened (symlink-reject +0600). The resolved Netscape cookies.txt feeds yt-dlp's--cookies(and Skool's opener).--cookies-from-browser BROWSERloads cookies straight from a local browser via yt-dlp (opt-in — reads the browser's cookie store). On an X auth failure (SourceAuthError, exit 5) the message names the refresh path: the resolved--cookies-filewhen one was supplied, else the convention path to create — derived from the failing URL's own host (www./mobile.labels stripped), e.g.~/.transcript-fetcher/x.com-cookies.txtfor anx.comURL or~/.transcript-fetcher/twitter.com-cookies.txtfor atwitter.comURL; the convention lookup's mirror-prefix fallback above guarantees this hinted path is actually picked up on retry, for all 6 documented X hosts. - Temp-file hygiene: For X media, all intermediates (audio, VTT, info.json,
.part,.m3u8, fragments) live under one tempdir removed in afinallyblock even on error — nothing is left behind. - Auth credentials:
--cookies-file <path>accepts a Netscapecookies.txtand is ALWAYS OPTIONAL for every source. The file is read once at startup, never copied or re-emitted. For Skool, public communities (e.g.zero-one) serve lessons without auth; private / paid communities respond with HTTP 401/403 and the user then needs to supply cookies. YouTube/Vimeo optionally forward the file to yt-dlp's--cookiesfor age-gated or unlisted videos. The skill never blocks on missing cookies up-front — it tries the fetch and surfaces aSourceAuthError(exit 5) only if the source returns 401/403.
6. Validation Evidence
- Local verification:
All offline tests must pass without network. The end-to-end network test is gated behindcd skills/transcript-fetcher ./scripts/.venv/bin/python -m unittest discover -s scripts/testsTRANSCRIPT_FETCHER_E2E=1. - Skill-validator (structural):
python3 .claude/skills/skill-creator/scripts/validate_skill.py skills/transcript-fetcher - Skill-validator (security):
python3 .claude/skills/skill-validator/scripts/validate.py skills/transcript-fetcher - Expected evidence: Both validators exit 0; unittest reports
OK.
7. Instructions
Step 1: Verify environment
If scripts/.venv/ does not exist, run bash scripts/install.sh first.
The install script is idempotent — safe to re-run.
Step 2: Choose mode
| Input | Mode | Command form |
|---|---|---|
| Single URL | single |
fetch.py <URL> --out path.txt |
| List of URLs in a file | batch |
fetch.py --batch urls.txt --out-dir dir/ |
Step 3: Pick a fallback strategy
Default is --lang ru --prefer manual. This tries:
manual:ru— user-uploaded Russian subtitles (highest quality).auto:ru-orig— YouTube auto-captions of the original Russian audio (good).auto:ru— YouTube auto-translation TO Russian (often noisy if speech was in another language).auto:en— English auto-captions as last resort (will setquality_flag = english_auto_translation).
For non-Russian content, pass --lang en (or another ISO code). For
non-Russian languages the lang-orig step is skipped — it is a YouTube
quirk that mainly matters for non-English speech.
Step 4: Run the CLI
Capture stdout (it carries the JSON stat). Read the stat to confirm which track was used:
./scripts/.venv/bin/python scripts/fetch.py \
"https://youtu.be/NSVTpCfBMK8" \
--out /tmp/talk.txt
# stdout: {"source":"youtube","url":"...","chosen_track_kind":"auto","chosen_track_lang":"ru-orig", ...}
Step 5: Inspect quality
Open the generated <out>.txt.stat.json. If quality_flag is set
(currently only "english_auto_translation"), surface a warning to
the user before passing the transcript to a downstream summarizer:
⚠️ TRANSCRIPT QUALITY: only English auto-translation was available for this URL. Idioms, proper names, and technical terms may be distorted. Consider asking the user for a manual transcription.
Step 6: Hand off
The clean .txt is now ready for summarizing-meetings or any other
downstream consumer. Pass the path; do not paste the contents inline
(transcripts are often large).
8. Workflows
- [ ] Verify scripts/.venv/ exists (run install.sh otherwise)
- [ ] Decide single vs batch mode
- [ ] Run fetch.py with the chosen language and preference
- [ ] Read the JSON stat sidecar
- [ ] Surface quality_flag warning if set
- [ ] Hand .txt path to downstream skill
9. Best Practices & Anti-Patterns
| DO THIS | DO NOT DO THIS |
|---|---|
| Always read the stat sidecar after fetching | Trust the .txt without checking which track was used |
Prefer ru-orig over ru for Russian content |
Pick the first track that returns text |
| Pass batch URLs through a file | Loop the CLI shell-side with arbitrary URLs |
Surface quality_flag in any user-visible output |
Silently downgrade to English auto-translation |
Use --json-errors in CI/automation pipelines |
Parse free-form stderr |
Rationalization Table
| Agent Excuse | Reality / Counter-Argument |
|---|---|
"yt-dlp is on $PATH so I can just call it" |
The skill invokes python -m yt_dlp from the per-skill venv. The system yt-dlp may be a different version with different output. |
"The >> markers are clutter" |
They are paragraph breaks for speaker turns. Downstream summarizers attribute statements by them. |
| "Rolling-caption dedup is overkill, just keep the longest cue" | The dedup IS keeping the longest cue. Without it, the same sentence would appear 3-4 times. |
| "I'll add a Vimeo adapter inline in fetch.py" | Add it as scripts/sources/vimeo.py. Each source is its own file. |
10. Examples
See examples/:
example_input_url.txt— batch input format.example_output_plain.txt— what a cleaned transcript looks like (excerpt).example_output_stat.json— what the stat sidecar contains.
11. Resources
scripts/fetch.py— CLI entry point.scripts/sources/youtube.py— YouTube adapter (yt-dlp orchestration + fallback ladder + description path).scripts/sources/vimeo.py— minimal Vimeo adapter (yt-dlp).scripts/sources/x.py— X.com / Twitter adapter (captions-first → ASR; theXTranscriptProvider).scripts/sources/_ytdlp_media.py— shared yt-dlp plumbing (metadata probe, caption inspection, audio-minimal download, failure classifier) reused by X and any future yt-dlp source.scripts/sources/_log.py— debug-only stage logger (stderr, gated on--debug).scripts/sources/_auth.py—~/.transcript-fetcher/per-host cookie resolution (auth-map + convention, hardened; mirrors thehtmlskill's~/.html).scripts/asr/— pluggable ASR backend package:_base.py(theASRBackendinterface),macwhisper.py,whisper_cli.py,whisper_cpp.py,openai_api.py(opt-in cloud),__init__.py(priority registry + fallback chain).scripts/_config.py— skill-local.envloader (secrets-safe) + typed config accessors (endpoints/models/tool paths).scripts/_procgroup.py— source-neutral process-group subprocess runner: kills a timed-out child AND its descendants (yt-dlp's ffmpeg / JS runtime, whisper's ffmpeg decode). Imported by bothsources/andasr/; must import from neither.scripts/_stdout.py— the machine channel's byte contract: JSON to stdout as UTF-8 regardless of the caller's locale, lone surrogates escaped, a dead reader raised to the caller instead of rewriting the exit status to 120. Stdlib-only; deliberately NOT the office skills'_errors.py(that one is proprietary, this skill is Apache-2.0 — see the module docstring).scripts/_human.py— the presentation channel's opposite contract: prose obeys the caller's codec instead of overriding it.say()(aprintdrop-in) andHumanArgumentParser(--help/usage) degrade—✓✗→⚠…to ASCII spellings per character only when the stream cannot carry them, so UTF-8 output is byte-identical; anything the table cannot spell falls back tobackslashreplacerather than raising — which is also what makes a lone surrogate from asurrogateescapefilename survivable, and it keeps a dead reader from rewriting the exit status to 120.scripts/.env.example— config/secret template (copy to.env,chmod 600).scripts/install_components.py— detect / guide / install the optional ASR components.scripts/sources/skool.py— Skool lesson adapter (cookies.txt + Next.js scrape + embed delegation).scripts/sources/_vtt_to_text.py— pure-Python WebVTT cleaner.scripts/sources/_captions.py— multi-format caption → text dispatch (SRT/TTML/DFXP build on the VTT cleaner; TTML XXE/billion-laughs guard).scripts/sources/_stat.py— sharedTranscriptStat+ sidecar writer + error classes.scripts/sources/_description.py—.description.mdwriter (YAML frontmatter + Markdown body).scripts/sources/_cookies.py— Netscape cookies.txt loader + authenticated opener.scripts/sources/_prosemirror.py— ProseMirror/TipTap v2 JSON → Markdown for Skool lesson bodies.scripts/install.sh— venv bootstrap.scripts/requirements.txt— pinned deps (single source of truth for the yt-dlp version range).scripts/tests/— offline unit tests + opt-in E2E network test.scripts/tests/_sanitize_fixture.py— utility for scrubbing PII from Skool HTML snapshots before they become fixtures.references/youtube_caption_format.md— what>>,>, rolling captions, andru-origactually mean.references/fallback_policy.md— the language ladder and why it is in this order.references/supported_sources.md— current and planned source slots.references/skool_adapter.md— Skool auth flow, schema notes, embed delegation rules.references/description_metadata.md—.description.mdformat for YouTube and Skool.docs/Manuals/transcript-fetcher_manual.md— user-facing manual with quick reference, troubleshooting, and composition recipes.
12. Composition
- Composes well with
summarizing-meetings: runtranscript-fetcherfirst to get a clean.txt, then pass that file tosummarizing-meetingsfor a structured Markdown summary. The two skills are intentionally separate — fetching is a deterministic file operation; summarizing is a prompt-first reasoning task.