Video RAG

Video ingestion for multimodal RAG. Covers keyframe extraction (ffmpeg, PySceneDetect), audio-track transcription via Whisper, visual embeddings per keyframe (CLIP, SigLIP, VoyageAI multimodal), multi-modal retrieval combining visual + transcript hits, timestamp-anchored citations, and YouTube transcript ingestion. USE WHEN: user mentions "video RAG", "video transcription", "keyframe extraction", "PySceneDetect", "ffmpeg scenes", "CLIP embeddings", "SigLIP", "VoyageAI multimodal", "YouTube transcript", "multimodal retrieval", "video chunking" DO NOT USE FOR: audio-only content (podcasts, meetings) - use `audio-transcription`; screen recordings that are effectively slide decks - use `office-docs` on the source PPTX; static images - use `ocr` or pure vision embeddings

claude-dev-suite Updated 28 repo stars

File contents

claude-dev-suite/claude-dev-suite/tree/main/skills/document-processing/video-rag commit e83bdb9b06

Frequently asked questions

npx skillmds@latest add claude-dev-suite/video-rag