Inference Serving Topology

LLM/model inference serving architecture: the engine → serving → orchestration layering (vLLM/SGLang/TensorRT-LLM, Triton, KServe/Ray Serve), KV-cache & continuous batching, prefill-decode disaggregation, and scaling. Architect-level topology, not model training. USE WHEN: designing model/LLM serving infra, "vLLM", "SGLang", "TensorRT-LLM", "Triton", "KServe", "Ray Serve", "continuous batching", "KV cache", "prefill decode", "TTFT", multi-GPU/multi-model serving, inference autoscaling. DO NOT USE FOR: on-device (use `edge-inference`); provider routing (use `model-gateway-routing`); RAG app logic (use rag/rag-frameworks skills).

claude-dev-suite Updated 28 repo stars

File contents

claude-dev-suite/claude-dev-suite/tree/main/skills/ai-systems/inference-serving-topology commit 896d8a3878

Frequently asked questions

npx skillmds@latest add claude-dev-suite/inference-serving-topology