Inference Performance

Use when serving or optimizing LLM inference in production — diagnosing or improving TTFT/TPOT/throughput, choosing batching strategy, sizing GPUs, picking vLLM/TensorRT-LLM, or debugging low GPU utilization, TTFT spikes, and OOM. Covers prefill vs decode, the roofline, continuous batching, PagedAttention, chunked prefill, disaggregation, and FlashAttention.

jpoindexter Updated

File contents

jpoindexter/design-and-ai-skills/tree/main/ai-engineering-skills/inference-performance commit 0da4d2f0b7

Frequently asked questions

npx skillmds@latest add jpoindexter/inference-performance