Analyzing narrated recordings with talkthrough
The talkthrough MCP server turns a local video/audio file into queryable
structured data: timestamped transcript segments, scene keyframes, OCR'd
on-screen text, and wall-clock anchoring. No LLM inside — you bring the
reasoning; it brings the evidence. Everything is lazy and token-budgeted:
never ask for more than the moment you are analyzing.
Prerequisite
The talkthrough MCP server must be connected (tools like
process_media / get_transcript are visible). If not, tell the user to
install it: claude mcp add -s user talkthrough -- uvx --python ">=3.11,<3.14" "talkthrough-mcp[diarization,url]"
(see the repository README for other clients).
Core workflow
- Ingest once:
process_media(path) — idempotent by content hash;
re-calls on the same file return instantly. Given a public video/audio
URL instead of a file, call process_url(url): the source is downloaded
once (the only network step; YouTube needs the [url] extra) and kept
inside the job, then everything below is identical and local — never
download twice; a repeat call serves the stored job. Long videos take minutes and
stream progress. The summary gives you job_id, counts, wall-clock, and
a transcript preview — do NOT dump anything else eagerly. Multi-person
recording (meeting/interview)? Add diarize=true — even when the ask is
just "summarize", speaker structure is part of meeting analysis — and — whenever the
headcount is known — num_speakers=N (the main accuracy lever): segments
get S1/S2/… labels and the summary a talk-time roster. On an
already-processed job the amend re-runs ONLY diarization (no
re-transcription) — still minutes on long recordings.
- Orient:
get_transcript(job_id) (paginate via next_start_ms when
truncated) or search(job_id, "<distinctive word>") to jump straight
to the relevant moments (searches speech AND on-screen OCR text). Multi-word
search defaults to match_mode="all_words"; use "any_word" for broader
lexical recall.
- Evidence per remark:
get_moment(job_id, t0-2000, t1+2000) — one
call returns the transcript slice + up to 3 unique frames + their OCR
text + the wall-clock range. This is the workhorse; describe observed
from the returned pixels, never from imagination.
- Precision when needed:
get_frames(at_ms=...) for nearby keyframes;
extract_frame(job_id, at_ms, crop={x,y,w,h}) for an exact instant at
native resolution (keyframes capture scene changes + a 1 fps floor, so
sub-second moments can fall between them).
- Keep verified names: after proving an anonymous label's identity,
call
label_speakers(job_id, labels={"S1":"Name"}, evidence={"S1":"intro or frame proof"}). Saved names appear in later
transcript, moment, and search calls while raw S1/S2 labels remain.
If a diarization amend changes labels, those names move to
speaker_names_pending_review and stop being identities. Use the stored
old-roster context anchors to re-check them. A pending label still in the
roster can be confirmed, replaced, or removed; a stale pending label can
only be removed with labels={"Sx":null}. Never use a pending name in
minutes or search as though it were active. A full force=true rebuild of
a job with active or pending names must also use diarize=true; it rebuilds
safely and moves every old identity to pending review, while omitting
diarization is refused without changing the stored job.
- Recall across sessions:
list_jobs() — the store persists; a file
processed yesterday (even via CLI) is queryable by job_id today.
Timestamps
Every timestamped result carries t_ms (video-relative) and, when the
recording start is known, t_wall (ISO 8601 real time). Copy t_wall
VERBATIM from the payload — never compute it from t_ms yourself
(hand-derived wall-clocks drift by whole hours). Use t_wall to
correlate remarks with server/app logs (±30 s grep window). If
wall_clock is null or low-confidence, ask the user when the recording
started and re-anchor: process_media(path, recorded_at="<ISO 8601>", force=true); when the job already has speaker identities, include
diarize=true as required by the safe-rebuild contract.
Packaged workflows (server prompts)
Prefer the server prompts when the task matches — they encode the full
method: bug (one recording → evidence-backed GitHub issue draft; silent,
narration-free recordings welcome), triage-recording (screencast →
findings JSON per the contract in examples/output-contract.schema.json),
spec-from-workshop, backlog-from-demo, meeting-actions (audio-only
friendly), correlate-with-logs.
Rules of thumb
- Audio-only jobs (.m4a/.mp3/…): transcript tools work; frame tools error
by design — that error is expected, not a failure.
- Speaker labels are anonymous (
S1/S2, ordered by first voice). Mapping
them to names is YOUR job: self-introductions, vocatives, the attendees
list — and on video jobs the screen check is MANDATORY: for every label
you map, get_frames(at_ms=<that label's longest_turn_at_ms from the roster>) and read the meeting-app name plates, the recording's title
card, the active-speaker highlight BEFORE asserting the mapping. STT
homophones lie about name spellings (spoken "profit" vs on-screen
"Prophet") — trust OCR/frames over the transcript for names. State the
mapping explicitly and mark unmapped labels "unidentified".
Roster name_candidates are raw OCR hints, not identities: they may be
UI text, a job title, or somebody else's name. Inspect the cited frame and
persist only defensible mappings with label_speakers; never auto-save a
candidate.
A name_candidates_note on a pre-0.3.1 video job explains that its legacy
flat OCR may not yield hints. The job remains readable; regenerate only when
useful, with force=true, diarize=true, so old identities become pending
review instead of being lost.
diarize=true needs the [diarization] extra — its absence produces an
actionable install-hint error.
- Findings/quotes must cite the narrator's exact words +
t_ms (+ t_wall
when known) + the frame files you actually inspected.
- Low STT/vision confidence → surface a question; never silently guess.
- Any narration language works (Whisper auto-detects; the summary reports
language + language_probability). Garbled transcript or low/wrong
detection → re-call process_media(path, model="large-v3-turbo", force=true) (best multilingual quality) or pin language="…"; domain
jargon → pass vocabulary="Term1, Term2".
- Write digests/summaries for the recording author in the narrator's
language; keep quotes verbatim in the original — translate in your own
prose only, never inside a quote.
1---2name: talkthrough3description: Analyze narrated screen recordings and audio files through the talkthrough MCP server — triage feedback into findings, extract specs/backlogs/action items from recordings, and correlate spoken remarks with logs via wall-clock timestamps. Use when the user mentions a screen recording, screencast, narrated video/audio file, or asks to "watch" a recording and act on it.4license: MIT5---67# Analyzing narrated recordings with talkthrough89The talkthrough MCP server turns a local video/audio file into queryable10structured data: timestamped transcript segments, scene keyframes, OCR'd11on-screen text, and wall-clock anchoring. No LLM inside — you bring the12reasoning; it brings the evidence. Everything is lazy and token-budgeted:13never ask for more than the moment you are analyzing.1415## Prerequisite1617The `talkthrough` MCP server must be connected (tools like18`process_media` / `get_transcript` are visible). If not, tell the user to19install it: `claude mcp add -s user talkthrough -- uvx --python ">=3.11,<3.14" "talkthrough-mcp[diarization,url]"`20(see the repository README for other clients).2122## Core workflow23241. **Ingest once**: `process_media(path)` — idempotent by content hash;25 re-calls on the same file return instantly. Given a public video/audio26 URL instead of a file, call `process_url(url)`: the source is downloaded27 once (the only network step; YouTube needs the `[url]` extra) and kept28 inside the job, then everything below is identical and local — never29 download twice; a repeat call serves the stored job. Long videos take minutes and30 stream progress. The summary gives you `job_id`, counts, wall-clock, and31 a transcript preview — do NOT dump anything else eagerly. Multi-person32 recording (meeting/interview)? Add `diarize=true` — even when the ask is33 just "summarize", speaker structure is part of meeting analysis — and — whenever the34 headcount is known — `num_speakers=N` (the main accuracy lever): segments35 get `S1`/`S2`/… labels and the summary a talk-time roster. On an36 already-processed job the amend re-runs ONLY diarization (no37 re-transcription) — still minutes on long recordings.382. **Orient**: `get_transcript(job_id)` (paginate via `next_start_ms` when39 `truncated`) or `search(job_id, "<distinctive word>")` to jump straight40 to the relevant moments (searches speech AND on-screen OCR text). Multi-word41 search defaults to `match_mode="all_words"`; use `"any_word"` for broader42 lexical recall.433. **Evidence per remark**: `get_moment(job_id, t0-2000, t1+2000)` — one44 call returns the transcript slice + up to 3 unique frames + their OCR45 text + the wall-clock range. This is the workhorse; describe `observed`46 from the returned pixels, never from imagination.474. **Precision when needed**: `get_frames(at_ms=...)` for nearby keyframes;48 `extract_frame(job_id, at_ms, crop={x,y,w,h})` for an exact instant at49 native resolution (keyframes capture scene changes + a 1 fps floor, so50 sub-second moments can fall between them).515. **Keep verified names**: after proving an anonymous label's identity,52 call `label_speakers(job_id, labels={"S1":"Name"},53 evidence={"S1":"intro or frame proof"})`. Saved names appear in later54 transcript, moment, and search calls while raw `S1`/`S2` labels remain.55 If a diarization amend changes labels, those names move to56 `speaker_names_pending_review` and stop being identities. Use the stored57 old-roster context anchors to re-check them. A pending label still in the58 roster can be confirmed, replaced, or removed; a stale pending label can59 only be removed with `labels={"Sx":null}`. Never use a pending name in60 minutes or search as though it were active. A full `force=true` rebuild of61 a job with active or pending names must also use `diarize=true`; it rebuilds62 safely and moves every old identity to pending review, while omitting63 diarization is refused without changing the stored job.646. **Recall across sessions**: `list_jobs()` — the store persists; a file65 processed yesterday (even via CLI) is queryable by `job_id` today.6667## Timestamps6869Every timestamped result carries `t_ms` (video-relative) and, when the70recording start is known, `t_wall` (ISO 8601 real time). Copy `t_wall`71VERBATIM from the payload — never compute it from `t_ms` yourself72(hand-derived wall-clocks drift by whole hours). Use `t_wall` to73correlate remarks with server/app logs (±30 s grep window). If74`wall_clock` is null or low-confidence, ask the user when the recording75started and re-anchor: `process_media(path, recorded_at="<ISO 8601>",76force=true)`; when the job already has speaker identities, include77`diarize=true` as required by the safe-rebuild contract.7879## Packaged workflows (server prompts)8081Prefer the server prompts when the task matches — they encode the full82method: `bug` (one recording → evidence-backed GitHub issue draft; silent,83narration-free recordings welcome), `triage-recording` (screencast →84findings JSON per the contract in `examples/output-contract.schema.json`),85`spec-from-workshop`, `backlog-from-demo`, `meeting-actions` (audio-only86friendly), `correlate-with-logs`.8788## Rules of thumb8990- Audio-only jobs (.m4a/.mp3/…): transcript tools work; frame tools error91 by design — that error is expected, not a failure.92- Speaker labels are anonymous (`S1`/`S2`, ordered by first voice). Mapping93 them to names is YOUR job: self-introductions, vocatives, the attendees94 list — and on video jobs the screen check is MANDATORY: for every label95 you map, `get_frames(at_ms=<that label's longest_turn_at_ms from the96 roster>)` and read the meeting-app name plates, the recording's title97 card, the active-speaker highlight BEFORE asserting the mapping. STT98 homophones lie about name spellings (spoken "profit" vs on-screen99 "Prophet") — trust OCR/frames over the transcript for names. State the100 mapping explicitly and mark unmapped labels "unidentified".101 Roster `name_candidates` are raw OCR hints, not identities: they may be102 UI text, a job title, or somebody else's name. Inspect the cited frame and103 persist only defensible mappings with `label_speakers`; never auto-save a104 candidate.105 A `name_candidates_note` on a pre-0.3.1 video job explains that its legacy106 flat OCR may not yield hints. The job remains readable; regenerate only when107 useful, with `force=true, diarize=true`, so old identities become pending108 review instead of being lost.109 `diarize=true` needs the `[diarization]` extra — its absence produces an110 actionable install-hint error.111- Findings/quotes must cite the narrator's exact words + `t_ms` (+ `t_wall`112 when known) + the frame files you actually inspected.113- Low STT/vision confidence → surface a question; never silently guess.114- Any narration language works (Whisper auto-detects; the summary reports115 `language` + `language_probability`). Garbled transcript or low/wrong116 detection → re-call `process_media(path, model="large-v3-turbo",117 force=true)` (best multilingual quality) or pin `language="…"`; domain118 jargon → pass `vocabulary="Term1, Term2"`.119- Write digests/summaries for the recording author in the narrator's120 language; keep quotes verbatim in the original — translate in your own121 prose only, never inside a quote.