LLM Caching

Optimize LLM costs and latency through KV caching and prompt caching. Use when (1) structuring prompts for cache hits, (2) configuring API cache_control for Anthropic/Cohere/OpenAI/Gemini, (3) setting up self-hosted inference with vLLM/SGLang/Ollama, (4) building agentic workflows with prefix reuse, (5) designing batch processing pipelines, or (6) understanding cache pricing and tradeoffs.

diegosouzapw Updated 54 repo stars

File contents

diegosouzapw/awesome-omni-skill/tree/main/skills/data-ai/llm-caching commit d944f86a05

Frequently asked questions

npx skillmds@latest add diegosouzapw/llm-caching