Latency Budgeting

Budget and engineer latency for an LLM agent — TTFT, tokens-per-second, tool round-trips, parallelism, streaming. Use when the user is building a user-facing or real-time agent and mentions latency, p50, p95, p99, TTFT, streaming, throughput, time-to-first-token, slow agent, or asks "why is my agent slow?" / "how do I hit a 2-second latency target?".

cobusgreyling Updated

File contents

cobusgreyling/agent-skills/tree/main/skills/latency-budgeting commit 0258450aa0

Frequently asked questions

npx skillmds@latest add cobusgreyling/latency-budgeting