Streaming RAG

Streaming LLM responses with inline citations. Token-level source attribution, SSE vs WebSocket, TTFT optimization, progressive disclosure (retrieval status then tokens), Python async generators, Vercel AI SDK streaming with sources, LangChain streaming callbacks, client-side citation rendering. USE WHEN: user mentions "streaming RAG", "streaming citations", "SSE RAG", "TTFT", "progressive disclosure", "AI SDK streaming", "token streaming", "inline citations" DO NOT USE FOR: generic RAG pipelines - use `rag-architecture`; citation correctness evaluation - use `rag-evaluation`; conversational memory - use `conversational-rag`

claude-dev-suite Updated 28 repo stars

File contents

claude-dev-suite/claude-dev-suite/tree/main/skills/rag/streaming-rag commit 925bc8fdae

Frequently asked questions

npx skillmds@latest add claude-dev-suite/streaming-rag