Book Pipeline — Full End-to-End Orchestrator
Purpose
Run a single book all the way from a PDF or EPUB in _inbox/ to live across pgvector, OpenAI Vector Stores, Gemini File Search, and Anthropic Files. Enforce quality gates so a broken phase blocks advancement. Skip work that's already done (idempotent on re-run).
Pipeline graph
/book-convert (Phase 01)
│
▼
/book-validate (Phase 01b — 5-layer harness)
│
▼
/book-audit (BookQuality v1.0)
│
┌────────┴────────┐
GREEN YELLOW or RED
│ │
│ ▼
│ /book-fix (writes fix-<slug>.md;
│ optionally executes)
│ │
│ ▼
│ re-run /book-audit
│ │
└────────┬───────────┘
│
▼
/book-frontmatter (Phase 02 — Zod validation)
│
▼
/book-ingest (Phase 03 — Supabase upsert)
│
▼
/book-chunk (Phase 04a — chunks.jsonl)
│
▼
/book-rag-push (Phase 04b — fan out + smoke test)
│
▼
✓
Invocation
$ARGUMENTS:
<input-file-or-book-slug>— Either a path to a PDF/EPUB in_inbox/(for a brand-new book) OR a<book-slug>whose corpus already exists (for re-processing). Required.--from-phase <phase>— Resume from a phase. Default: auto-detect (skip phases whose outputs already exist and are unchanged).- Values:
convert | validate | audit | fix | frontmatter | ingest | chunk | rag-push
- Values:
--stop-after <phase>— Stop after this phase completes. Default:rag-push.--auto-fix— On YELLOW/RED audit, automatically run/book-fix --mode=executeand re-audit. Default:false(ask the user).--targets <list>— Pass-through to/book-rag-push. Default:pgvector,openai,gemini,claude.--dry-run— Print the planned phase sequence without running.
Process
Phase routing
For each phase, decide: skip (output up to date), run, or fail. Use these signals:
| Phase | Skip if | Run if |
|---|---|---|
| Convert | corpus/alan_hirsch/<slug>/*.md exists AND _inbox/<slug>.<ext> mtime older than chapter mtime |
otherwise |
| Validate | <book-dir>/.ingest/validation-*.json exists, verdict PASS, all chapter SHAs match |
otherwise |
| Audit | docs/build/audits/<slug>-*.md exists from current week |
otherwise |
| Fix | only if audit verdict ≠ GREEN | n/a |
| Frontmatter | <book-dir>/.ingest/frontmatter-*.json PASS, all SHAs match |
otherwise |
| Ingest | DB row's manifest:version matches book.json:version AND no chapter SHA changed |
otherwise |
| Chunk | <book-dir>/.ingest/chunks.jsonl exists, all chunk SHAs match |
chapter SHA changed |
| RAG-push | <book-dir>/.ingest/rag-state.json records each target's last upload at current chunk SHA set |
chunks changed |
Step 1 — Detect input + slug
If $ARGUMENTS is a file path, infer slug from filename. If it's a slug, look up _inbox/<slug>.pdf or _inbox/<slug>.epub.
Step 2 — Run phases (sequential, with gates)
For each phase:
- Print
▶ Phase 01 — Convert(or whatever). - Invoke the corresponding skill via
Skilltool with the right arguments. - Read the skill's structured summary output.
- Gate: if the phase failed, stop and report. Do not advance.
- Skip: if the skill reports "no work to do, output current", note it and advance.
- Quality gate after audit: branch on verdict.
- GREEN → advance to frontmatter.
- YELLOW / RED → if
--auto-fix, run/book-fix --mode=executeand re-audit; otherwise stop and ask.
Step 3 — Final report
Pipeline complete: <book-slug>
✓ Phase 01 Convert (skipped — up to date)
✓ Phase 01b Validate (skipped — up to date)
✓ BookQuality Audit GREEN (4 minor low-severity findings)
✓ Phase 02 Frontmatter 18/18 chapters validated
✓ Phase 03 Ingest 3 chapters changed; 47 sections regenerated
✓ Phase 04a Chunk +12 added, ~7 modified, 4218 unchanged
✓ Phase 04b RAG push pgvector +12 · openai +12 · gemini +12 · claude +0
smoke test: top-3 Jaccard 0.62 ✓
Verdict: ✓ Live on all four stores
Total wall time: 8m 42s
Total API cost: $1.34
Next: monitor `book_chunks` for query coverage; watch
`corpus/alan_hirsch/_archive/<slug>-*` for any pre-fix snapshots
that can be deleted after a successful 30-day soak.
Operating Principles
- Each phase is its own skill. This skill is a thin orchestrator — never duplicate phase logic. If a phase needs a behavior change, change the phase skill, not this orchestrator.
- Quality gates are non-negotiable. A book that fails validation does not get a frontmatter pass; a RED audit does not get ingested. The pipeline halts and surfaces the problem.
- Idempotent re-runs. Running the pipeline twice on an unchanged book should produce the same artifacts and skip every phase. Running it on a book where Ch.4 was hand-edited should re-validate, re-audit, re-chunk, and re-push only the affected chunks.
- Pre-archive on every destructive step.
/book-convertand/book-fixarchive the current state before overwriting; this skill verifies those archives exist before allowing those phases to run. - Cost transparency. Every phase reports its API spend; this skill aggregates and prints a total.
Anti-patterns
- ❌ Running
/book-rag-pushwithout/book-auditfirst — that's how fabricated content lands in production retrieval. - ❌ Skipping
/book-validatebecause Phase 01 "looked fine" — the harness exists because eyeballing doesn't catch alternate-line drops. - ❌ Auto-fixing a RED audit without human review — content rewrites need eyes; the prompt-based approach in
/book-fix --mode=generateis safer. - ❌ Marking a book "live" without the smoke test in
/book-rag-pushconfirming all four stores agree on the top results.
References
- docs/html/books-pipeline.html#runbook — End-to-end runbook brief
- All sibling skill SKILL.md files
- docs/build/prompts/ — Remediation prompt exemplars
- docs/build/audits/ — Past audit reports (created on demand)