1---2name: deploy-this-version-353description: import Image from '@theme/IdealImage'; import Tabs from '@theme/Tabs'; import TabItem from '@theme/TabItem';4---56import Image from '@theme/IdealImage';7import Tabs from '@theme/Tabs';8import TabItem from '@theme/TabItem';910## Deploy this version1112<Tabs>13<TabItem value="docker" label="Docker">1415``` showLineNumbers title="docker run litellm"16docker run \17-e STORE_MODEL_IN_DB=True \18-p 4000:4000 \19docker.litellm.ai/berriai/litellm:v1.78.5-stable20```2122</TabItem>2324<TabItem value="pip" label="Pip">2526``` showLineNumbers title="pip install litellm"27pip install litellm==1.78.528```2930</TabItem>31</Tabs>3233---3435## Key Highlights3637- **Native OCR Endpoints** - Native `/v1/ocr` endpoint support with cost tracking for Mistral OCR and Azure AI OCR38- **Global Vendor Discounts** - Specify global vendor discount percentages for accurate cost tracking and reporting39- **Team Spending Reports** - Team admins can now export detailed spending reports for their teams40- **Claude Haiku 4.5** - Day 0 support for Claude Haiku 4.5 across Bedrock, Vertex AI, and OpenRouter with 200K context window41- **GPT-5-Codex** - Support for GPT-5-Codex via Responses API on OpenAI and Azure42- **Performance Improvements** - Major router optimizations: O(1) model lookups, 10-100x faster shallow copy, 30-40% faster timing calls, and O(n) to O(1) hash generation4344---4546## New Models / Updated Models4748#### New Model Support4950| Provider | Model | Context Window | Input ($/1M tokens) | Output ($/1M tokens) | Features |51| -------- | ----- | -------------- | ------------------- | -------------------- | -------- |52| Anthropic | `claude-haiku-4-5` | 200K | $1.00 | $5.00 | Chat, reasoning, vision, function calling, prompt caching, computer use |53| Anthropic | `claude-haiku-4-5-20251001` | 200K | $1.00 | $5.00 | Chat, reasoning, vision, function calling, prompt caching, computer use |54| Bedrock | `anthropic.claude-haiku-4-5-20251001-v1:0` | 200K | $1.00 | $5.00 | Chat, reasoning, vision, function calling, prompt caching |55| Bedrock | `global.anthropic.claude-haiku-4-5-20251001-v1:0` | 200K | $1.00 | $5.00 | Chat, reasoning, vision, function calling, prompt caching |56| Bedrock | `jp.anthropic.claude-haiku-4-5-20251001-v1:0` | 200K | $1.10 | $5.50 | Chat, reasoning, vision, function calling, prompt caching (JP Cross-Region) |57| Bedrock | `us.anthropic.claude-haiku-4-5-20251001-v1:0` | 200K | $1.10 | $5.50 | Chat, reasoning, vision, function calling, prompt caching (US region) |58| Bedrock | `eu.anthropic.claude-haiku-4-5-20251001-v1:0` | 200K | $1.10 | $5.50 | Chat, reasoning, vision, function calling, prompt caching (EU region) |59| Bedrock | `apac.anthropic.claude-haiku-4-5-20251001-v1:0` | 200K | $1.10 | $5.50 | Chat, reasoning, vision, function calling, prompt caching (APAC region) |60| Bedrock | `au.anthropic.claude-haiku-4-5-20251001-v1:0` | 200K | $1.10 | $5.50 | Chat, reasoning, vision, function calling, prompt caching (AU region) |61| Vertex AI | `vertex_ai/claude-haiku-4-5@20251001` | 200K | $1.00 | $5.00 | Chat, reasoning, vision, function calling, prompt caching |62| OpenAI | `gpt-5` | 272K | $1.25 | $10.00 | Chat, responses API, reasoning, vision, function calling, prompt caching |63| OpenAI | `gpt-5-codex` | 272K | $1.25 | $10.00 | Responses API mode |64| Azure | `azure/gpt-5-codex` | 272K | $1.25 | $10.00 | Responses API mode |65| Gemini | `gemini-2.5-flash-image` | 32K | $0.30 | $2.50 | Image generation (GA - Nano Banana) - $0.039/image |66| ZhipuAI | `glm-4.6` | - | - | - | Chat completions |6768#### Features6970- **[OpenAI](../../docs/providers/openai)**71 - GPT-5 return reasoning content via /chat/completions + GPT-5-Codex working on Claude Code - [PR #15441](https://github.com/BerriAI/litellm/pull/15441)7273- **[Anthropic](../../docs/providers/anthropic)**74 - Reduce claude-4-sonnet max_output_tokens to 64k - [PR #15409](https://github.com/BerriAI/litellm/pull/15409)75 - Added claude-haiku-4.5 - [PR #15579](https://github.com/BerriAI/litellm/pull/15579)76 - Add support for thinking blocks and redacted thinking blocks in Anthropic v1/messages API - [PR #15501](https://github.com/BerriAI/litellm/pull/15501)7778- **[Bedrock](../../docs/providers/bedrock)**79 - Add anthropic.claude-haiku-4-5-20251001-v1:0 on Bedrock, VertexAI - [PR #15581](https://github.com/BerriAI/litellm/pull/15581)80 - Add Claude Haiku 4.5 support for Bedrock global and US regions - [PR #15650](https://github.com/BerriAI/litellm/pull/15650)81 - Add Claude Haiku 4.5 support for Bedrock Other regions - [PR #15653](https://github.com/BerriAI/litellm/pull/15653)82 - Add JP Cross-Region Inference jp.anthropic.claude-haiku-4-5-20251001 - [PR #15598](https://github.com/BerriAI/litellm/pull/15598)83 - Fix: bedrock-pricing-geo-inregion-cross-region / add Global Cross-Region Inference - [PR #15685](https://github.com/BerriAI/litellm/pull/15685)84 - Fix: Support us-gov prefix for AWS GovCloud Bedrock models - [PR #15626](https://github.com/BerriAI/litellm/pull/15626)85 - Fix GPT-OSS in Bedrock now supports streaming. Revert fake streaming - [PR #15668](https://github.com/BerriAI/litellm/pull/15668)8687- **[Gemini](../../docs/providers/gemini)**88 - Feat(pricing): Add Gemini 2.5 Flash Image (Nano Banana) in GA - [PR #15557](https://github.com/BerriAI/litellm/pull/15557)89 - Fix: Gemini 2.5 Flash Image should not have supports_web_search=true - [PR #15642](https://github.com/BerriAI/litellm/pull/15642)90 - Remove penalty params as supported params for gemini preview model - [PR #15503](https://github.com/BerriAI/litellm/pull/15503)9192- **[Ollama](../../docs/providers/ollama)**93 - Fix(ollama/chat): correctly map reasoning_effort to think in requests - [PR #15465](https://github.com/BerriAI/litellm/pull/15465)9495- **[OpenRouter](../../docs/providers/openrouter)**96 - Add anthropic/claude-sonnet-4.5 to OpenRouter cost map - [PR #15472](https://github.com/BerriAI/litellm/pull/15472)97 - Prompt caching for anthropic models with OpenRouter - [PR #15535](https://github.com/BerriAI/litellm/pull/15535)98 - Get completion cost directly from OpenRouter - [PR #15448](https://github.com/BerriAI/litellm/pull/15448)99 - Fix OpenRouter Claude Opus 4 model naming - [PR #15495](https://github.com/BerriAI/litellm/pull/15495)100101- **[CometAPI](../../docs/providers/comet)**102 - Fix(cometapi): improve CometAPI provider support (embeddings, image generation, docs) - [PR #15591](https://github.com/BerriAI/litellm/pull/15591)103104- **[Lemonade](../../docs/providers/lemonade)**105 - Adding new models to the lemonade provider - [PR #15554](https://github.com/BerriAI/litellm/pull/15554)106107- **[Watson X](../../docs/providers/watsonx)**108 - Fix (pricing): Fix pricing for watsonx model family for various models - [PR #15670](https://github.com/BerriAI/litellm/pull/15670)109110- **[Vercel AI Gateway](../../docs/providers/vercel_ai_gateway)**111 - Add glm-4.6 model to pricing configuration - [PR #15679](https://github.com/BerriAI/litellm/pull/15679)112113- **[Vertex AI](../../docs/providers/vertex)**114 - Add Vertex AI Discovery Engine Rerank Support - [PR #15532](https://github.com/BerriAI/litellm/pull/15532)115116### Bug Fixes117118- **[Anthropic](../../docs/providers/anthropic)**119 - Fix: Pricing for Claude Sonnet 4.5 in US regions is 10x too high - [PR #15374](https://github.com/BerriAI/litellm/pull/15374)120121- **[OpenRouter](../../docs/providers/openrouter)**122 - Change gpt-5-codex support in model_price json - [PR #15540](https://github.com/BerriAI/litellm/pull/15540)123124- **[Bedrock](../../docs/providers/bedrock)**125 - Fix filtering headers for signature calcs - [PR #15590](https://github.com/BerriAI/litellm/pull/15590)126127- **General**128 - Add native reasoning and streaming support flag for gpt-5-codex - [PR #15569](https://github.com/BerriAI/litellm/pull/15569)129130---131132## LLM API Endpoints133134#### Features135136- **[Responses API](../../docs/response_api)**137 - Responses API - enable calling anthropic/gemini models in Responses API streaming in openai ruby sdk + DB - sanity check pending migrations before startup - [PR #15432](https://github.com/BerriAI/litellm/pull/15432)138 - Add support for responses mode in health check - [PR #15658](https://github.com/BerriAI/litellm/pull/15658)139140- **[OCR API](../../docs/ocr)**141 - Feat: Add native litellm.ocr() functions - [PR #15567](https://github.com/BerriAI/litellm/pull/15567)142 - Feat: Add /ocr route on LiteLLM AI Gateway - Adds support for native Mistral OCR calling - [PR #15571](https://github.com/BerriAI/litellm/pull/15571)143 - Feat: Add Azure AI Mistral OCR Integration - [PR #15572](https://github.com/BerriAI/litellm/pull/15572)144 - Feat: Native /ocr endpoint support - [PR #15573](https://github.com/BerriAI/litellm/pull/15573)145 - Feat: Add Cost Tracking for /ocr endpoints - [PR #15678](https://github.com/BerriAI/litellm/pull/15678)146147- **[/generateContent](../../docs/providers/gemini)**148 - Fix: GEMINI - CLI - add google_routes to llm_api_routes - [PR #15500](https://github.com/BerriAI/litellm/pull/15500)149 - Fix Pydantic validation error for citationMetadata.citationSources in Google GenAI responses - [PR #15592](https://github.com/BerriAI/litellm/pull/15592)150151- **[Images API](../../docs/image_generation)**152 - Fix: Dall-e-2 for Image Edits API - [PR #15604](https://github.com/BerriAI/litellm/pull/15604)153154- **[Bedrock Passthrough](../../docs/pass_through/bedrock)**155 - Feat: Allow calling /invoke, /converse routes through AI Gateway + models on config.yaml - [PR #15618](https://github.com/BerriAI/litellm/pull/15618)156157#### Bugs158159- **General**160 - Fix: Convert object to a correct type - [PR #15634](https://github.com/BerriAI/litellm/pull/15634)161 - Bug Fix: Tags as metadata dicts were raising exceptions - [PR #15625](https://github.com/BerriAI/litellm/pull/15625)162 - Add type hint to function_to_dict and fix typo - [PR #15580](https://github.com/BerriAI/litellm/pull/15580)163164---165166## Management Endpoints / UI167168#### Features169170- **Virtual Keys**171 - Docs: Key Rotations - [PR #15455](https://github.com/BerriAI/litellm/pull/15455)172 - Fix: UI - Key Max Budget Removal Error Fix - [PR #15672](https://github.com/BerriAI/litellm/pull/15672)173 - litellm_Key Settings Max Budget Removal Error Fix - [PR #15669](https://github.com/BerriAI/litellm/pull/15669)174175- **Teams**176 - Feat: Allow Team Admins to export a report of the team spending - [PR #15542](https://github.com/BerriAI/litellm/pull/15542)177178- **Passthrough**179 - Feat: Passthrough - allow admin to give access to specific passthrough endpoints - [PR #15401](https://github.com/BerriAI/litellm/pull/15401)180181- **SCIM v2**182 - Feat(scim_v2.py): if group.id doesn't exist, use external id + Passthrough - ensure updates and deletions persist across instances - [PR #15276](https://github.com/BerriAI/litellm/pull/15276)183184- **SSO**185 - Feat: UI SSO - Add PKCE for OKTA SSO - [PR #15608](https://github.com/BerriAI/litellm/pull/15608)186 - Fix: Separate OAuth M2M authentication from UI SSO + Handle Introspection endpoint for Oauth2 - [PR #15667](https://github.com/BerriAI/litellm/pull/15667)187 - Fix/entraid app roles jwt claim clean - [PR #15583](https://github.com/BerriAI/litellm/pull/15583)188189---190191## Logging / Guardrail / Prompt Management Integrations192193#### Guardrails194195- **General**196 - Fix apply_guardrail endpoint returning raw string instead of ApplyGuardrailResponse - [PR #15436](https://github.com/BerriAI/litellm/pull/15436)197 - Fix: Ensure guardrail memory sync after database updates - [PR #15633](https://github.com/BerriAI/litellm/pull/15633)198 - Feat: add guardrail for image generation - [PR #15619](https://github.com/BerriAI/litellm/pull/15619)199 - Feat: Add Guardrails for /v1/messages and /v1/responses API - [PR #15686](https://github.com/BerriAI/litellm/pull/15686)200201- **[Pillar Security](../../docs/proxy/guardrails)**202 - Feature: update pillar security integration to support no persistence mode in litellm proxy - [PR #15599](https://github.com/BerriAI/litellm/pull/15599)203204#### Prompt Management205206- **General**207 - Small fix code snippet custom_prompt_management.md - [PR #15544](https://github.com/BerriAI/litellm/pull/15544)208209---210211## Spend Tracking, Budgets and Rate Limiting212213- **Cost Tracking**214 - Feat: Cost Tracking - specify a global vendor discount for costs - [PR #15546](https://github.com/BerriAI/litellm/pull/15546)215 - Feat: UI - Allow setting Provider Discounts on UI - [PR #15550](https://github.com/BerriAI/litellm/pull/15550)216217- **Budgets**218 - Fix: improve budget clarity - [PR #15682](https://github.com/BerriAI/litellm/pull/15682)219220---221222## Performance / Loadbalancing / Reliability improvements223224- **Router Optimizations**225 - Perf(router): use shallow copy instead of deepcopy for model aliases - 10-100x faster than deepcopy on nested dict structures - [PR #15576](https://github.com/BerriAI/litellm/pull/15576)226 - Perf(router): optimize string concatenation in hash generation - Improves time complexity from O(n²) to O(n) - [PR #15575](https://github.com/BerriAI/litellm/pull/15575)227 - Perf(router): optimize model lookups with O(1) data structures - Replace O(n) scans with index map lookups - [PR #15578](https://github.com/BerriAI/litellm/pull/15578)228 - Perf(router): optimize model lookups with O(1) index maps - Use model_id_to_deployment_index_map and model_name_to_deployment_indices for instant lookups - [PR #15574](https://github.com/BerriAI/litellm/pull/15574)229 - Perf(router): optimize timing functions in completion hot path - Use time.perf_counter() for duration measurements and time.monotonic() for timeout calculations, providing 30-40% faster timing calls - [PR #15617](https://github.com/BerriAI/litellm/pull/15617)230231- **SSL/TLS Performance**232 - Feat(ssl): add configurable ECDH curve for TLS performance - Configure via ssl_ecdh_curve setting to disable PQC on OpenSSL 3.x for better performance - [PR #15617](https://github.com/BerriAI/litellm/pull/15617)233234- **Token Counter**235 - Fix(token-counter): extract model_info from deployment for custom_tokenizer - [PR #15680](https://github.com/BerriAI/litellm/pull/15680)236237- **Performance Metrics**238 - Add: perf summary - [PR #15458](https://github.com/BerriAI/litellm/pull/15458)239240- **CI/CD**241 - Fix: CI/CD - Missing env key & Linter type error - [PR #15606](https://github.com/BerriAI/litellm/pull/15606)242243---244245## Documentation Updates246247- **Provider Documentation**248 - Litellm docs 10 11 2025 - [PR #15457](https://github.com/BerriAI/litellm/pull/15457)249 - Docs: add ecs deployment guide - [PR #15468](https://github.com/BerriAI/litellm/pull/15468)250 - Docs: Update benchmark results - [PR #15461](https://github.com/BerriAI/litellm/pull/15461)251 - Fix: add missing context to benchmark docs - [PR #15688](https://github.com/BerriAI/litellm/pull/15688)252253- **General**254 - Fixed a few typos - [PR #15267](https://github.com/BerriAI/litellm/pull/15267)255256---257258## New Contributors259260* @jlan-nl made their first contribution in [PR #15374](https://github.com/BerriAI/litellm/pull/15374)261* @ImadSaddik made their first contribution in [PR #15267](https://github.com/BerriAI/litellm/pull/15267)262* @huangyafei made their first contribution in [PR #15472](https://github.com/BerriAI/litellm/pull/15472)263* @mubashir1osmani made their first contribution in [PR #15468](https://github.com/BerriAI/litellm/pull/15468)264* @kowyo made their first contribution in [PR #15465](https://github.com/BerriAI/litellm/pull/15465)265* @dhruvyad made their first contribution in [PR #15448](https://github.com/BerriAI/litellm/pull/15448)266* @davizucon made their first contribution in [PR #15544](https://github.com/BerriAI/litellm/pull/15544)267* @FelipeRodriguesGare made their first contribution in [PR #15540](https://github.com/BerriAI/litellm/pull/15540)268* @ndrsfel made their first contribution in [PR #15557](https://github.com/BerriAI/litellm/pull/15557)269* @shinharaguchi made their first contribution in [PR #15598](https://github.com/BerriAI/litellm/pull/15598)270* @TensorNull made their first contribution in [PR #15591](https://github.com/BerriAI/litellm/pull/15591)271* @TeddyAmkie made their first contribution in [PR #15583](https://github.com/BerriAI/litellm/pull/15583)272* @aniketmaurya made their first contribution in [PR #15580](https://github.com/BerriAI/litellm/pull/15580)273* @eddierichter-amd made their first contribution in [PR #15554](https://github.com/BerriAI/litellm/pull/15554)274* @konekohana made their first contribution in [PR #15535](https://github.com/BerriAI/litellm/pull/15535)275* @Classic298 made their first contribution in [PR #15495](https://github.com/BerriAI/litellm/pull/15495)276* @afogel made their first contribution in [PR #15599](https://github.com/BerriAI/litellm/pull/15599)277* @orolega made their first contribution in [PR #15633](https://github.com/BerriAI/litellm/pull/15633)278* @LucasSugi made their first contribution in [PR #15634](https://github.com/BerriAI/litellm/pull/15634)279* @uc4w6c made their first contribution in [PR #15619](https://github.com/BerriAI/litellm/pull/15619)280* @Sameerlite made their first contribution in [PR #15658](https://github.com/BerriAI/litellm/pull/15658)281* @yuneng-jiang made their first contribution in [PR #15672](https://github.com/BerriAI/litellm/pull/15672)282* @Nikro made their first contribution in [PR #15680](https://github.com/BerriAI/litellm/pull/15680)283284---285286## Full Changelog287288**[View complete changelog on GitHub](https://github.com/BerriAI/litellm/compare/v1.78.0-stable...v1.78.4-stable)**289