LLM Gateway

LLM gateways in front of RAG stacks. Covers Portkey (caching, fallbacks, retries, observability), OpenRouter (300+ model routing), LiteLLM Proxy, Kong AI Gateway, semantic caching at gateway layer, cost-based routing (cheap model for easy queries), rate-limit handling, and unified API across providers. Config examples. USE WHEN: user mentions "LLM gateway", "Portkey", "OpenRouter", "LiteLLM", "Kong AI Gateway", "AI gateway", "semantic cache gateway", "provider fallback", "unified LLM API" DO NOT USE FOR: in-app caching - use `rag-caching`; multi-region routing - use `multi-region`; cost tracking/dashboarding - use `cost-allocation`

claude-dev-suite Updated 28 repo stars

File contents

claude-dev-suite/claude-dev-suite/tree/main/skills/rag-ops/llm-gateway commit e99340af4e

Frequently asked questions

npx skillmds@latest add claude-dev-suite/llm-gateway