LLM Cost Latency Budget

Model the cost and latency of an LLM feature before it ships and surprises the bill. Use when asked to estimate LLM API costs, set a latency/token budget, decide which model tier to use, or bring down the cost of an AI feature. Produces a cost & latency budget — token math per request, monthly cost projection, model tiering, caching/streaming levers, p95 latency targets, and a guardrail/alert plan.

Mohit Aggarwal 05980d0 3.8 KB Updated

File contents

mohitagw15856/pm-claude-skills/tree/main/skills/llm-cost-latency-budget commit 05980d0967

Frequently asked questions

npx skillmds add mohitagw15856/llm-cost-latency-budget