Portkey — AI Gateway for Production LLM Apps
You are an expert in Portkey, the AI gateway that sits between your app and LLM providers. You help developers add caching, fallbacks, load balancing, request retries, guardrails, semantic caching, budget limits, and observability to LLM calls — using a single unified API that works with 200+ models from OpenAI, Anthropic, Google, and open-source providers.
Core Capabilities
import Portkey from "portkey-ai";
const portkey = new Portkey({
apiKey: process.env.PORTKEY_API_KEY,
config: {
strategy: { mode: "fallback" }, // Auto-fallback on errors
targets: [
{
provider: "openai", api_key: process.env.OPENAI_KEY,
override_params: { model: "gpt-4o" },
weight: 0.7,
},
{
provider: "anthropic", api_key: process.env.ANTHROPIC_KEY,
override_params: { model: "claude-sonnet-4-20250514" },
weight: 0.3,
},
],
cache: { mode: "semantic", max_age: 3600 }, // Semantic caching
retry: { attempts: 3, on_status_codes: [429, 500, 503] },
},
});
// Use like OpenAI SDK — Portkey handles routing, caching, fallbacks
const response = await portkey.chat.completions.create({
messages: [{ role: "user", content: "Explain microservices" }],
max_tokens: 1024,
});
// Guardrails
const guarded = new Portkey({
apiKey: process.env.PORTKEY_API_KEY,
config: {
before_request_hooks: [{ type: "guardrail", id: "no-pii" }],
after_request_hooks: [{ type: "guardrail", id: "no-hallucination" }],
},
});
// Budget limits
// Set in Portkey dashboard: max $100/day per API key
Installation
npm install portkey-ai
# or
pip install portkey-ai
Best Practices
- OpenAI SDK compatible — Drop-in replacement; change import and add config; existing code works
- Fallbacks — Route to backup provider when primary fails; 99.99% effective uptime
- Semantic caching — Cache similar (not just identical) queries; 40-60% cache hit rate typical
- Load balancing — Split traffic across providers by weight; optimize cost vs quality
- Retry with backoff — Auto-retry on 429/500/503; configurable attempts and status codes
- Guardrails — PII detection, content moderation, hallucination checks; pre and post request
- Budget limits — Set per-key spending caps; prevent runaway costs from bugs or abuse
- Observability — Dashboard shows latency, cost, tokens, errors per provider; no additional SDK
1---2name: portkey3description: You are an expert in Portkey, the AI gateway that sits between your app and LLM providers. You help developers add caching, fallbacks, load balancing, request retries, guardrails, semantic caching, budget limits, and observability to LLM calls — using a single unified API that works with 200+ models from OpenAI, Anthropic, Google, and open-source providers.4license: Apache-2.05---67# Portkey — AI Gateway for Production LLM Apps89You are an expert in Portkey, the AI gateway that sits between your app and LLM providers. You help developers add caching, fallbacks, load balancing, request retries, guardrails, semantic caching, budget limits, and observability to LLM calls — using a single unified API that works with 200+ models from OpenAI, Anthropic, Google, and open-source providers.1011## Core Capabilities1213```typescript14import Portkey from "portkey-ai";1516const portkey = new Portkey({17 apiKey: process.env.PORTKEY_API_KEY,18 config: {19 strategy: { mode: "fallback" }, // Auto-fallback on errors20 targets: [21 {22 provider: "openai", api_key: process.env.OPENAI_KEY,23 override_params: { model: "gpt-4o" },24 weight: 0.7,25 },26 {27 provider: "anthropic", api_key: process.env.ANTHROPIC_KEY,28 override_params: { model: "claude-sonnet-4-20250514" },29 weight: 0.3,30 },31 ],32 cache: { mode: "semantic", max_age: 3600 }, // Semantic caching33 retry: { attempts: 3, on_status_codes: [429, 500, 503] },34 },35});3637// Use like OpenAI SDK — Portkey handles routing, caching, fallbacks38const response = await portkey.chat.completions.create({39 messages: [{ role: "user", content: "Explain microservices" }],40 max_tokens: 1024,41});4243// Guardrails44const guarded = new Portkey({45 apiKey: process.env.PORTKEY_API_KEY,46 config: {47 before_request_hooks: [{ type: "guardrail", id: "no-pii" }],48 after_request_hooks: [{ type: "guardrail", id: "no-hallucination" }],49 },50});5152// Budget limits53// Set in Portkey dashboard: max $100/day per API key54```5556## Installation5758```bash59npm install portkey-ai60# or61pip install portkey-ai62```6364## Best Practices65661. **OpenAI SDK compatible** — Drop-in replacement; change import and add config; existing code works672. **Fallbacks** — Route to backup provider when primary fails; 99.99% effective uptime683. **Semantic caching** — Cache similar (not just identical) queries; 40-60% cache hit rate typical694. **Load balancing** — Split traffic across providers by weight; optimize cost vs quality705. **Retry with backoff** — Auto-retry on 429/500/503; configurable attempts and status codes716. **Guardrails** — PII detection, content moderation, hallucination checks; pre and post request727. **Budget limits** — Set per-key spending caps; prevent runaway costs from bugs or abuse738. **Observability** — Dashboard shows latency, cost, tokens, errors per provider; no additional SDK