Media Ingest Skill
Ingest video, audio, PDF, book, screenshot, and GitHub repo content into the brain.
Filing rule: Read skills/_brain-filing-rules.md before creating any new page.
Input
| Parameter |
Required |
Description |
| source |
yes |
URL, file path, or uploaded file reference |
| title |
no |
Override title (auto-detected if omitted) |
| target_slug |
no |
Override page slug (auto-generated if omitted) |
Contract
This skill guarantees:
- Every ingested media item has a brain page with analysis (not just a transcript dump)
- Transcripts (video/audio) saved in raw and human-readable formats
- Entity extraction: every person and company mentioned gets back-linked
- Raw source files preserved via
gbrain files upload-raw
- Filing by primary subject, not by media format
Convention: See skills/conventions/quality.md for Iron Law back-linking.
Every mention of a person or company with a brain page MUST create a back-link.
Phases
Phase 1: Identify format and fetch
| Format |
Action |
| YouTube/video URL |
Fetch transcript (Whisper, transcription service, or captions) |
| Audio file |
Transcribe with available STT service |
| PDF |
Extract text (OCR if needed) |
| Book PDF |
Extract text, identify chapters/sections |
| Screenshot/image |
OCR via vision model, extract text and entities |
| GitHub repo |
Clone, read README + key files, summarize architecture |
Phase 2: Upload raw source
Save the original file for provenance: gbrain files upload-raw <file> --page <slug>
Phase 3: Create brain page
File by primary subject (not format). Use this template:
# {Title}
**Source:** {URL or file path}
**Format:** {video/audio/PDF/book/screenshot/repo}
**Created:** {date}
## Summary
{Key points, not a transcript dump}
## Key Segments / Highlights
{For video/audio: timestamped highlights. For books: chapter summaries.}
## People Mentioned
{List with links to brain pages}
## Companies Mentioned
{List with links to brain pages}
Phase 4: Entity extraction and propagation
For every person and company mentioned:
- Check brain for existing page
- Create/enrich if needed (delegate to enrich skill)
- Add back-link from entity page to this media page
- Add timeline entry on entity page
A media item is NOT fully ingested until entity propagation is complete.
Phase 5: Sync
gbrain sync to update the index.
Output Format
Brain page created with summary, highlights, and entity cross-links. Report to user:
"Ingested {title}: {N} entities detected, {N} pages updated."
Error Handling
- Transcription failure: If STT or captions are unavailable, note
[transcript unavailable] in the page and proceed with whatever metadata is available. Do NOT fabricate content.
- Duplicate detection: Before creating a page, search the brain for the source URL or file hash. If found, ask the user whether to update the existing page or skip.
- Partial OCR / audio: Mark unclear segments with
[inaudible] or [illegible]. Never guess at proper nouns.
- Large content (books > 500 pages): Summarize by chapter; do not attempt to inline the full text. Link to the raw upload.
- Retry policy: On transient API failures (network, timeout), retry once. On auth failures, abort immediately.
Known Pitfalls
- YouTube auto-captions misidentify proper nouns. Always cross-reference entity names against existing brain pages before creating new ones. A caption that garbles a name (e.g. "Alise" when the speakers are discussing alice-example) should match the existing
alice-example page, not create a new one.
- Re-running ingest on same source creates duplicates. Always check brain for existing source URL match before Phase 3.
- Book OCR quality varies wildly. Scanned PDFs often have garbled text. If OCR quality is <80% readable, flag to user rather than ingesting garbage.
- Video transcript without speaker diarization is low-value. If multiple speakers are present but no diarization is available, note this limitation prominently rather than attributing all speech to one person.
- Large audio files (>2hr) can timeout transcription services. Split into chunks before transcription if needed.
Anti-Patterns
- Dumping raw transcripts without analysis
- Skipping entity extraction ("I'll do that separately")
- Filing raw ingest by format (all videos in
media/videos/) instead of by subject. Note: format-prefixed paths under media/<format>/<slug> ARE sanctioned for synthesized one-of-one output like book-mirror's media/books/<slug>-personalized.md. The anti-pattern is for raw ingest, not for sui generis synthesis. See skills/_brain-filing-rules.md "Sanctioned exception: synthesis output is sui generis."
- Not preserving raw source files
- Creating stub pages without meaningful content
1---2name: media-ingest3description: Ingest video, audio, PDF, book, screenshot, and GitHub repo content into the brain. Multi-format handling with entity extraction and backlink propagation. Covers video-ingest, youtube-ingest, and book-ingest subtypes.4---5
6# Media Ingest Skill
7
8Ingest video, audio, PDF, book, screenshot, and GitHub repo content into the brain.
9
10> **Filing rule:** Read `skills/_brain-filing-rules.md` before creating any new page.
11
12## Input
13
14| Parameter | Required | Description |
15|-----------|----------|-------------|
16| source | yes | URL, file path, or uploaded file reference |
17| title | no | Override title (auto-detected if omitted) |
18| target_slug | no | Override page slug (auto-generated if omitted) |
19
20## Contract
21
22This skill guarantees:
23- Every ingested media item has a brain page with analysis (not just a transcript dump)
24- Transcripts (video/audio) saved in raw and human-readable formats
25- Entity extraction: every person and company mentioned gets back-linked
26- Raw source files preserved via `gbrain files upload-raw`
27- Filing by primary subject, not by media format
28
29> **Convention:** See `skills/conventions/quality.md` for Iron Law back-linking.
30
31Every mention of a person or company with a brain page MUST create a back-link.
32
33## Phases
34
35### Phase 1: Identify format and fetch
36
37| Format | Action |
38|--------|--------|
39| YouTube/video URL | Fetch transcript (Whisper, transcription service, or captions) |
40| Audio file | Transcribe with available STT service |
41| PDF | Extract text (OCR if needed) |
42| Book PDF | Extract text, identify chapters/sections |
43| Screenshot/image | OCR via vision model, extract text and entities |
44| GitHub repo | Clone, read README + key files, summarize architecture |
45
46### Phase 2: Upload raw source
47
48Save the original file for provenance: `gbrain files upload-raw <file> --page <slug>`
49
50### Phase 3: Create brain page
51
52File by primary subject (not format). Use this template:
53
54```markdown
55# {Title}
56
57**Source:** {URL or file path}
58**Format:** {video/audio/PDF/book/screenshot/repo}
59**Created:** {date}
60
61## Summary
62{Key points, not a transcript dump}
63
64## Key Segments / Highlights
65{For video/audio: timestamped highlights. For books: chapter summaries.}
66
67## People Mentioned
68{List with links to brain pages}
69
70## Companies Mentioned
71{List with links to brain pages}
72```
73
74### Phase 4: Entity extraction and propagation
75
76For every person and company mentioned:
771. Check brain for existing page
782. Create/enrich if needed (delegate to enrich skill)
793. Add back-link from entity page to this media page
804. Add timeline entry on entity page
81
82A media item is NOT fully ingested until entity propagation is complete.
83
84### Phase 5: Sync
85
86`gbrain sync` to update the index.
87
88## Output Format
89
90Brain page created with summary, highlights, and entity cross-links. Report to user:
91"Ingested {title}: {N} entities detected, {N} pages updated."
92
93## Error Handling
94
95- **Transcription failure:** If STT or captions are unavailable, note `[transcript unavailable]` in the page and proceed with whatever metadata is available. Do NOT fabricate content.
96- **Duplicate detection:** Before creating a page, search the brain for the source URL or file hash. If found, ask the user whether to update the existing page or skip.
97- **Partial OCR / audio:** Mark unclear segments with `[inaudible]` or `[illegible]`. Never guess at proper nouns.
98- **Large content (books > 500 pages):** Summarize by chapter; do not attempt to inline the full text. Link to the raw upload.
99- **Retry policy:** On transient API failures (network, timeout), retry once. On auth failures, abort immediately.
100
101## Known Pitfalls
102
1031. **YouTube auto-captions misidentify proper nouns.** Always cross-reference entity names against existing brain pages before creating new ones. A caption that garbles a name (e.g. "Alise" when the speakers are discussing alice-example) should match the existing `alice-example` page, not create a new one.
1042. **Re-running ingest on same source creates duplicates.** Always check brain for existing source URL match before Phase 3.
1053. **Book OCR quality varies wildly.** Scanned PDFs often have garbled text. If OCR quality is <80% readable, flag to user rather than ingesting garbage.
1064. **Video transcript without speaker diarization is low-value.** If multiple speakers are present but no diarization is available, note this limitation prominently rather than attributing all speech to one person.
1075. **Large audio files (>2hr) can timeout transcription services.** Split into chunks before transcription if needed.
108
109## Anti-Patterns
110
111- Dumping raw transcripts without analysis
112- Skipping entity extraction ("I'll do that separately")
113- Filing **raw ingest** by format (all videos in `media/videos/`) instead of by subject. Note: format-prefixed paths under `media/<format>/<slug>` ARE sanctioned for **synthesized one-of-one output** like book-mirror's `media/books/<slug>-personalized.md`. The anti-pattern is for raw ingest, not for sui generis synthesis. See `skills/_brain-filing-rules.md` "Sanctioned exception: synthesis output is sui generis."
114- Not preserving raw source files
115- Creating stub pages without meaningful content