Corpus Ingestion

Convert heterogeneous source files (PDFs, EPUBs, DOCX, PPTX, HTML, audio, images, ZIPs, web pages) into normalized markdown using MarkItDown, then prepare them for Movemental-style corpus pipelines. Use this skill whenever ingesting an author's books, sermons, transcripts, slide decks, or archive materials for downstream theme/voice/concept analysis, RAG pipeline preparation, course content generation, or vector store upload. Also trigger on: "ingest this book", "convert and add to corpus", "prepare source material", "add to the knowledge base", "batch convert documents for analysis", or any time raw heterogeneous files need to become clean, curated markdown assets.

JoshuaShepherd Updated 1 repo stars

File contents

JoshuaShepherd/my-skills/tree/main/claude/content/corpus-ingestion commit 339508e498

Frequently asked questions

npx skillmds@latest add joshuashepherd/corpus-ingestion