LLM Inference Optimization

Diagnose and optimize self-hosted LLM runtime latency, goodput, cost, and capacity through prefill/decode profiling, KV cache, batching, admission, quantization, and inference parallelism. Use for execution bottlenecks and runtime scaling signals; replica provisioning and fleet placement belong to ai-platform-llmops.

BejeweledMe Updated

File contents

BejeweledMe/codex-pro-agent-skills/tree/main/skills/llm-inference-optimization commit f26da69232

Frequently asked questions

npx skillmds@latest add bejeweledme/llm-inference-optimization