import Tabs from '@theme/Tabs';
import TabItem from '@theme/TabItem';
import Image from '@theme/IdealImage';
📈 Prometheus metrics
LiteLLM Exposes a /metrics endpoint for Prometheus to Poll
Quick Start
If you're using the LiteLLM CLI with litellm --config proxy_config.yaml then you need to pip install prometheus_client==0.20.0. This is already pre-installed on the litellm Docker image
Add this to your proxy config.yaml
model_list:
- model_name: gpt-4o
litellm_params:
model: gpt-4o
litellm_settings:
callbacks:
- prometheus
Start the proxy
litellm --config config.yaml --debug
Test Request
curl --location 'http://0.0.0.0:4000/chat/completions' \
--header 'Content-Type: application/json' \
--data '{
"model": "gpt-4o",
"messages": [
{
"role": "user",
"content": "what llm are you"
}
]
}'
View Metrics on /metrics, Visit http://localhost:4000/metrics
http://localhost:4000/metrics
# <proxy_base_url>/metrics
Multiple Workers
When using LiteLLM with multiple workers, you need to set the PROMETHEUS_MULTIPROC_DIR environment variable to enable aggregated metric collection across worker processes.
export PROMETHEUS_MULTIPROC_DIR="/prometheus_multiproc"
This directory is used by the Prometheus client library to store metric files that can be shared across multiple worker processes. Make sure the directory exists and is writable by your LiteLLM process.
Virtual Keys, Teams, Internal Users
Use this for for tracking per user, key, team, etc.
| Metric Name |
Description |
litellm_spend_metric |
Total Spend, per "end_user", "hashed_api_key", "api_key_alias", "model", "team", "team_alias", "user" |
litellm_total_tokens_metric |
input + output tokens per "end_user", "hashed_api_key", "api_key_alias", "requested_model", "team", "team_alias", "user", "model" |
litellm_input_tokens_metric |
input tokens per "end_user", "hashed_api_key", "api_key_alias", "requested_model", "team", "team_alias", "user", "model" |
litellm_output_tokens_metric |
output tokens per "end_user", "hashed_api_key", "api_key_alias", "requested_model", "team", "team_alias", "user", "model" |
Team - Budget
| Metric Name |
Description |
litellm_team_max_budget_metric |
Max Budget for Team Labels: "team", "team_alias" |
litellm_remaining_team_budget_metric |
Remaining Budget for Team (A team created on LiteLLM) Labels: "team", "team_alias" |
litellm_team_budget_remaining_hours_metric |
Hours before the team budget is reset Labels: "team", "team_alias" |
Virtual Key - Budget
| Metric Name |
Description |
litellm_api_key_max_budget_metric |
Max Budget for API Key Labels: "hashed_api_key", "api_key_alias" |
litellm_remaining_api_key_budget_metric |
Remaining Budget for API Key (A key Created on LiteLLM) Labels: "hashed_api_key", "api_key_alias" |
litellm_api_key_budget_remaining_hours_metric |
Hours before the API Key budget is reset Labels: "hashed_api_key", "api_key_alias" |
Virtual Key - Rate Limit
| Metric Name |
Description |
litellm_remaining_api_key_requests_for_model |
Remaining Requests for a LiteLLM virtual API key, only if a model-specific rate limit (rpm) has been set for that virtual key. Labels: "hashed_api_key", "api_key_alias", "model" |
litellm_remaining_api_key_tokens_for_model |
Remaining Tokens for a LiteLLM virtual API key, only if a model-specific token limit (tpm) has been set for that virtual key. Labels: "hashed_api_key", "api_key_alias", "model" |
Initialize Budget Metrics on Startup
If you want litellm to emit the budget metrics for all keys, teams irrespective of whether they are getting requests or not, set prometheus_initialize_budget_metrics to true in the config.yaml
How this works:
- If the
prometheus_initialize_budget_metrics is set to true
- Every 5 minutes litellm runs a cron job to read all keys, teams from the database
- It then emits the budget metrics for each key, team
- This is used to populate the budget metrics on the
/metrics endpoint
litellm_settings:
callbacks: ["prometheus"]
prometheus_initialize_budget_metrics: true
Proxy Level Tracking Metrics
Use this to track overall LiteLLM Proxy usage.
- Track Actual traffic rate to proxy
- Number of client side requests and failures for requests made to proxy
| Metric Name |
Description |
litellm_proxy_failed_requests_metric |
Total number of failed responses from proxy - the client did not get a success response from litellm proxy. Labels: "end_user", "hashed_api_key", "api_key_alias", "requested_model", "team", "team_alias", "user", "user_email", "exception_status", "exception_class", "route", "model_id" |
litellm_proxy_total_requests_metric |
Total number of requests made to the proxy server - track number of client side requests. Labels: "end_user", "hashed_api_key", "api_key_alias", "requested_model", "team", "team_alias", "user", "status_code", "user_email", "route", "model_id" |
Callback Logging Metrics
Monitor failures while shipping logs to downstream callbacks like s3_v3 cold storage
| Metric Name |
Description |
litellm_callback_logging_failures_metric |
Total number of failed attempts to emit logs to a configured callback. Labels: "callback_name". Use this to alert on callback delivery issues such as repeated failures when writing to s3_v3, langfuse, or langfuse_otel and other otel providers |
Supported Callbacks:
S3Logger - S3 v2 cold storage failures
langfuse - Langfuse logging failures
otel - OpenTelemetry logging failures
LLM Provider Metrics
Use this for LLM API Error monitoring and tracking remaining rate limits and token limits
Labels Tracked
| Label |
Description |
| litellm_model_name |
The name of the LLM model used by LiteLLM |
| requested_model |
The model sent in the request |
| model_id |
The model_id of the deployment. Autogenerated by LiteLLM, each deployment has a unique model_id |
| api_base |
The API Base of the deployment |
| api_provider |
The LLM API provider, used for the provider. Example (azure, openai, vertex_ai) |
| hashed_api_key |
The hashed api key of the request |
| api_key_alias |
The alias of the api key used |
| team |
The team of the request |
| team_alias |
The alias of the team used |
| exception_status |
The status of the exception, if any |
| exception_class |
The class of the exception, if any |
Success and Failure
| Metric Name |
Description |
litellm_deployment_success_responses |
Total number of successful LLM API calls for deployment. Labels: "requested_model", "litellm_model_name", "model_id", "api_base", "api_provider", "hashed_api_key", "api_key_alias", "team", "team_alias" |
litellm_deployment_failure_responses |
Total number of failed LLM API calls for a specific LLM deployment. Labels: "requested_model", "litellm_model_name", "model_id", "api_base", "api_provider", "hashed_api_key", "api_key_alias", "team", "team_alias", "exception_status", "exception_class" |
litellm_deployment_total_requests |
Total number of LLM API calls for deployment - success + failure. Labels: "requested_model", "litellm_model_name", "model_id", "api_base", "api_provider", "hashed_api_key", "api_key_alias", "team", "team_alias" |
Remaining Requests and Tokens
| Metric Name |
Description |
litellm_remaining_requests_metric |
Track x-ratelimit-remaining-requests returned from LLM API Deployment. Labels: "model_group", "api_provider", "api_base", "litellm_model_name", "hashed_api_key", "api_key_alias" |
litellm_remaining_tokens_metric |
Track x-ratelimit-remaining-tokens return from LLM API Deployment. Labels: "model_group", "api_provider", "api_base", "litellm_model_name", "hashed_api_key", "api_key_alias" |
Deployment State
| Metric Name |
Description |
litellm_deployment_state |
The state of the deployment: 0 = healthy, 1 = partial outage, 2 = complete outage. Labels: "litellm_model_name", "model_id", "api_base", "api_provider" |
litellm_deployment_latency_per_output_token |
Latency per output token for deployment. Labels: "litellm_model_name", "model_id", "api_base", "api_provider", "hashed_api_key", "api_key_alias", "team", "team_alias" |
Fallback (Failover) Metrics
| Metric Name |
Description |
litellm_deployment_cooled_down |
Number of times a deployment has been cooled down by LiteLLM load balancing logic. Labels: "litellm_model_name", "model_id", "api_base", "api_provider" |
litellm_deployment_successful_fallbacks |
Number of successful fallback requests from primary model -> fallback model. Labels: "requested_model", "fallback_model", "hashed_api_key", "api_key_alias", "team", "team_alias", "exception_status", "exception_class" |
litellm_deployment_failed_fallbacks |
Number of failed fallback requests from primary model -> fallback model. Labels: "requested_model", "fallback_model", "hashed_api_key", "api_key_alias", "team", "team_alias", "exception_status", "exception_class" |
Request Counting Metrics
| Metric Name |
Description |
litellm_requests_metric |
Total number of requests tracked per endpoint. Labels: "end_user", "hashed_api_key", "api_key_alias", "model", "team", "team_alias", "user", "user_email" |
Request Latency Metrics
| Metric Name |
Description |
litellm_request_total_latency_metric |
Total latency (seconds) for a request to LiteLLM Proxy Server - tracked for labels "end_user", "hashed_api_key", "api_key_alias", "requested_model", "team", "team_alias", "user", "model", "model_id" |
litellm_overhead_latency_metric |
Latency overhead (seconds) added by LiteLLM processing - tracked for labels "model_group", "api_provider", "api_base", "litellm_model_name", "hashed_api_key", "api_key_alias" |
litellm_llm_api_latency_metric |
Latency (seconds) for just the LLM API call - tracked for labels "model", "hashed_api_key", "api_key_alias", "team", "team_alias", "requested_model", "end_user", "user" |
litellm_llm_api_time_to_first_token_metric |
Time to first token for LLM API call - tracked for labels model, hashed_api_key, api_key_alias, team, team_alias, requested_model, end_user, user, model_id [Note: only emitted for streaming requests] |
Tracking end_user on Prometheus
By default LiteLLM does not track end_user on Prometheus. This is done to reduce the cardinality of the metrics from LiteLLM Proxy.
If you want to track end_user on Prometheus, you can do the following:
litellm_settings:
callbacks: ["prometheus"]
enable_end_user_cost_tracking_prometheus_only: true
[BETA] Custom Metrics
Track custom metrics on prometheus on all events mentioned above.
Custom Metadata Labels
- Define the custom metadata labels in the
config.yaml
model_list:
- model_name: openai/gpt-4o
litellm_params:
model: openai/gpt-4o
api_key: os.environ/OPENAI_API_KEY
litellm_settings:
callbacks: ["prometheus"]
custom_prometheus_metadata_labels: ["metadata.foo", "metadata.bar"]
- Make a request with the custom metadata labels
curl -L -X POST 'http://0.0.0.0:4000/key/generate' \
-H 'Authorization: Bearer sk-1234' \
-H 'Content-Type: application/json' \
-d '{
"metadata": {
"foo": "hello world"
}
}'
curl -L -X POST 'http://0.0.0.0:4000/team/new' \
-H 'Authorization: Bearer sk-1234' \
-H 'Content-Type: application/json' \
-d '{
"metadata": {
"foo": "hello world"
}
}'
- Check your
/metrics endpoint for the custom metrics
... "metadata_foo": "hello world" ...
Custom Tags
Track specific tags as prometheus labels for better filtering and monitoring.
- Define the custom tags in the
config.yaml
model_list:
- model_name: openai/gpt-4o
litellm_params:
model: openai/gpt-4o
api_key: os.environ/OPENAI_API_KEY
litellm_settings:
callbacks: ["prometheus"]
custom_prometheus_metadata_labels: ["metadata.foo", "metadata.bar"]
custom_prometheus_tags:
- "prod"
- "staging"
- "batch-job"
- "User-Agent: RooCode/*"
- "User-Agent: claude-cli/*"
- Make a request with tags
curl -L -X POST 'http://0.0.0.0:4000/v1/chat/completions' \
-H 'Content-Type: application/json' \
-H 'Authorization: Bearer <LITELLM_API_KEY>' \
-d '{
"model": "openai/gpt-4o",
"messages": [
{
"role": "user",
"content": [
{
"type": "text",
"text": "What's in this image?"
}
]
}
],
"max_tokens": 300,
"metadata": {
"tags": ["prod", "user-facing"]
}
}'
- Check your
/metrics endpoint for the custom tag metrics
... "tag_prod": "true", "tag_staging": "false", "tag_batch_job": "false" ...
How Custom Tags Work:
- Each configured tag becomes a boolean label in prometheus metrics
- If a tag matches (exact or wildcard), the label value is
"true", otherwise "false"
- Tag names are sanitized for prometheus compatibility (e.g.,
"batch-job" becomes "tag_batch_job")
- Wildcard patterns supported using
* (e.g., "User-Agent: RooCode/*" matches "User-Agent: RooCode/1.0.0")
Example with wildcards:
litellm_settings:
callbacks: ["prometheus"]
custom_prometheus_tags:
- "User-Agent: RooCode/*"
- "User-Agent: claude-cli/*"
Use Cases:
- Environment tracking (
prod, staging, dev)
- Request type classification (
batch-job, user-facing, background)
- Feature flags (
new-feature, beta-users)
- Team or service identification (
team-a, service-xyz)
- User-Agent Tracking - use this to track how much Roo Code, Claude Code, Gemini CLI are used (
User-Agent: RooCode/*, User-Agent: claude-cli/*, User-Agent: gemini-cli/*)
Configuring Metrics and Labels
You can selectively enable specific metrics and control which labels are included to optimize performance and reduce cardinality.
Enable Specific Metrics and Labels
Configure which metrics to emit by specifying them in prometheus_metrics_config. Each configuration group needs a group name (for organization) and a list of metrics to enable. You can optionally include a list of include_labels to filter the labels for the metrics.
model_list:
- model_name: gpt-4o
litellm_params:
model: gpt-4o
litellm_settings:
callbacks: ["prometheus"]
prometheus_metrics_config:
# High-cardinality metrics with minimal labels
- group: "proxy_metrics"
metrics:
- "litellm_proxy_total_requests_metric"
- "litellm_proxy_failed_requests_metric"
include_labels:
- "hashed_api_key"
- "requested_model"
- "model_group"
On starting up LiteLLM if your metrics were correctly configured, you should see the following on your container logs
<Image
img={require('../../img/prom_config.png')}
style={{width: '100%', display: 'block', margin: '2rem auto'}}
/>
Filter Labels Per Metric
Control which labels are included for each metric to reduce cardinality:
litellm_settings:
callbacks: ["prometheus"]
prometheus_metrics_config:
- group: "token_consumption"
metrics:
- "litellm_input_tokens_metric"
- "litellm_output_tokens_metric"
- "litellm_total_tokens_metric"
include_labels:
- "model"
- "team"
- "hashed_api_key"
- group: "request_tracking"
metrics:
- "litellm_proxy_total_requests_metric"
include_labels:
- "status_code"
- "requested_model"
Advanced Configuration
You can create multiple configuration groups with different label sets:
litellm_settings:
callbacks: ["prometheus"]
prometheus_metrics_config:
# High-cardinality metrics with minimal labels
- group: "deployment_health"
metrics:
- "litellm_deployment_success_responses"
- "litellm_deployment_failure_responses"
include_labels:
- "api_provider"
- "requested_model"
# Budget metrics with full label set
- group: "budget_tracking"
metrics:
- "litellm_remaining_team_budget_metric"
include_labels:
- "team"
- "team_alias"
- "hashed_api_key"
- "api_key_alias"
- "model"
- "end_user"
# Latency metrics with performance-focused labels
- group: "performance"
metrics:
- "litellm_request_total_latency_metric"
- "litellm_llm_api_latency_metric"
include_labels:
- "model"
- "api_provider"
- "requested_model"
Configuration Structure:
group: A descriptive name for organizing related metrics
metrics: List of metric names to include in this group
include_labels: (Optional) List of labels to include for these metrics
Default Behavior: If no prometheus_metrics_config is specified, all metrics are enabled with their default labels (backward compatible).
Monitor System Health
To monitor the health of litellm adjacent services (redis / postgres), do:
model_list:
- model_name: gpt-4o
litellm_params:
model: gpt-4o
litellm_settings:
service_callback: ["prometheus_system"]
| Metric Name |
Description |
litellm_redis_latency |
histogram latency for redis calls |
litellm_redis_fails |
Number of failed redis calls |
litellm_self_latency |
Histogram latency for successful litellm api call |
DB Transaction Queue Health Metrics
Use these metrics to monitor the health of the DB Transaction Queue. Eg. Monitoring the size of the in-memory and redis buffers.
| Metric Name |
Description |
Storage Type |
litellm_pod_lock_manager_size |
Indicates which pod has the lock to write updates to the database. |
Redis |
litellm_in_memory_daily_spend_update_queue_size |
Number of items in the in-memory daily spend update queue. These are the aggregate spend logs for each user. |
In-Memory |
litellm_redis_daily_spend_update_queue_size |
Number of items in the Redis daily spend update queue. These are the aggregate spend logs for each user. |
Redis |
litellm_in_memory_spend_update_queue_size |
In-memory aggregate spend values for keys, users, teams, team members, etc. |
In-Memory |
litellm_redis_spend_update_queue_size |
Redis aggregate spend values for keys, users, teams, etc. |
Redis |
🔥 LiteLLM Maintained Grafana Dashboards
Link to Grafana Dashboards maintained by LiteLLM
https://github.com/BerriAI/litellm/tree/main/cookbook/litellm_proxy_server/grafana_dashboard
Here is a screenshot of the metrics you can monitor with the LiteLLM Grafana Dashboard
<Image img={require('../../img/grafana_1.png')} />
<Image img={require('../../img/grafana_2.png')} />
<Image img={require('../../img/grafana_3.png')} />
Deprecated Metrics
| Metric Name |
Description |
litellm_llm_api_failed_requests_metric |
deprecated use litellm_proxy_failed_requests_metric |
Add authentication on /metrics endpoint
By default /metrics endpoint is unauthenticated.
You can opt into running litellm authentication on the /metrics endpoint by setting the following on the config
litellm_settings:
require_auth_for_metrics_endpoint: true
FAQ
What are _created vs. _total metrics?
_created metrics are metrics that are created when the proxy starts
_total metrics are metrics that are incremented for each request
You should consume the _total metrics for your counting purposes
1---2name: prometheus-metrics-33description: import Tabs from '@theme/Tabs'; import TabItem from '@theme/TabItem'; import Image from '@theme/IdealImage';4---5import Tabs from '@theme/Tabs';6import TabItem from '@theme/TabItem';7import Image from '@theme/IdealImage';89# 📈 Prometheus metrics101112LiteLLM Exposes a `/metrics` endpoint for Prometheus to Poll1314## Quick Start1516If you're using the LiteLLM CLI with `litellm --config proxy_config.yaml` then you need to `pip install prometheus_client==0.20.0`. **This is already pre-installed on the litellm Docker image**1718Add this to your proxy config.yaml 19```yaml20model_list:21 - model_name: gpt-4o22 litellm_params:23 model: gpt-4o24litellm_settings:25 callbacks:26 - prometheus27```2829Start the proxy30```shell31litellm --config config.yaml --debug32```3334Test Request35```shell36curl --location 'http://0.0.0.0:4000/chat/completions' \37 --header 'Content-Type: application/json' \38 --data '{39 "model": "gpt-4o",40 "messages": [41 {42 "role": "user",43 "content": "what llm are you"44 }45 ]46}'47```4849View Metrics on `/metrics`, Visit `http://localhost:4000/metrics` 50```shell51http://localhost:4000/metrics5253# <proxy_base_url>/metrics54```5556### Multiple Workers5758When using LiteLLM with multiple workers, you need to set the `PROMETHEUS_MULTIPROC_DIR` environment variable to enable aggregated metric collection across worker processes.5960```shell61export PROMETHEUS_MULTIPROC_DIR="/prometheus_multiproc"62```6364This directory is used by the Prometheus client library to store metric files that can be shared across multiple worker processes. Make sure the directory exists and is writable by your LiteLLM process.6566## Virtual Keys, Teams, Internal Users6768Use this for for tracking per [user, key, team, etc.](virtual_keys)6970| Metric Name | Description |71|----------------------|--------------------------------------|72| `litellm_spend_metric` | Total Spend, per `"end_user", "hashed_api_key", "api_key_alias", "model", "team", "team_alias", "user"` |73| `litellm_total_tokens_metric` | input + output tokens per `"end_user", "hashed_api_key", "api_key_alias", "requested_model", "team", "team_alias", "user", "model"` |74| `litellm_input_tokens_metric` | input tokens per `"end_user", "hashed_api_key", "api_key_alias", "requested_model", "team", "team_alias", "user", "model"` |75| `litellm_output_tokens_metric` | output tokens per `"end_user", "hashed_api_key", "api_key_alias", "requested_model", "team", "team_alias", "user", "model"` |7677### Team - Budget787980| Metric Name | Description |81|----------------------|--------------------------------------|82| `litellm_team_max_budget_metric` | Max Budget for Team Labels: `"team", "team_alias"`|83| `litellm_remaining_team_budget_metric` | Remaining Budget for Team (A team created on LiteLLM) Labels: `"team", "team_alias"`|84| `litellm_team_budget_remaining_hours_metric` | Hours before the team budget is reset Labels: `"team", "team_alias"`|8586### Virtual Key - Budget8788| Metric Name | Description |89|----------------------|--------------------------------------|90| `litellm_api_key_max_budget_metric` | Max Budget for API Key Labels: `"hashed_api_key", "api_key_alias"`|91| `litellm_remaining_api_key_budget_metric` | Remaining Budget for API Key (A key Created on LiteLLM) Labels: `"hashed_api_key", "api_key_alias"`|92| `litellm_api_key_budget_remaining_hours_metric` | Hours before the API Key budget is reset Labels: `"hashed_api_key", "api_key_alias"`|9394### Virtual Key - Rate Limit9596| Metric Name | Description |97|----------------------|--------------------------------------|98| `litellm_remaining_api_key_requests_for_model` | Remaining Requests for a LiteLLM virtual API key, only if a model-specific rate limit (rpm) has been set for that virtual key. Labels: `"hashed_api_key", "api_key_alias", "model"`|99| `litellm_remaining_api_key_tokens_for_model` | Remaining Tokens for a LiteLLM virtual API key, only if a model-specific token limit (tpm) has been set for that virtual key. Labels: `"hashed_api_key", "api_key_alias", "model"`|100101102### Initialize Budget Metrics on Startup103104If you want litellm to emit the budget metrics for all keys, teams irrespective of whether they are getting requests or not, set `prometheus_initialize_budget_metrics` to `true` in the `config.yaml`105106**How this works:**107108- If the `prometheus_initialize_budget_metrics` is set to `true`109 - Every 5 minutes litellm runs a cron job to read all keys, teams from the database110 - It then emits the budget metrics for each key, team111 - This is used to populate the budget metrics on the `/metrics` endpoint112113```yaml114litellm_settings:115 callbacks: ["prometheus"]116 prometheus_initialize_budget_metrics: true117```118119120## Proxy Level Tracking Metrics121122Use this to track overall LiteLLM Proxy usage.123- Track Actual traffic rate to proxy 124- Number of **client side** requests and failures for requests made to proxy 125126| Metric Name | Description |127|----------------------|--------------------------------------|128| `litellm_proxy_failed_requests_metric` | Total number of failed responses from proxy - the client did not get a success response from litellm proxy. Labels: `"end_user", "hashed_api_key", "api_key_alias", "requested_model", "team", "team_alias", "user", "user_email", "exception_status", "exception_class", "route", "model_id"` |129| `litellm_proxy_total_requests_metric` | Total number of requests made to the proxy server - track number of client side requests. Labels: `"end_user", "hashed_api_key", "api_key_alias", "requested_model", "team", "team_alias", "user", "status_code", "user_email", "route", "model_id"` |130131### Callback Logging Metrics132133Monitor failures while shipping logs to downstream callbacks like `s3_v3` cold storage134135| Metric Name | Description |136|----------------------|--------------------------------------|137| `litellm_callback_logging_failures_metric` | Total number of failed attempts to emit logs to a configured callback. Labels: `"callback_name"`. Use this to alert on callback delivery issues such as repeated failures when writing to `s3_v3`, `langfuse`, or `langfuse_otel` and other otel providers |138139**Supported Callbacks:**140- `S3Logger` - S3 v2 cold storage failures141- `langfuse` - Langfuse logging failures142- `otel` - OpenTelemetry logging failures143144## LLM Provider Metrics145146Use this for LLM API Error monitoring and tracking remaining rate limits and token limits147148### Labels Tracked149150| Label | Description |151|-------|-------------|152| litellm_model_name | The name of the LLM model used by LiteLLM |153| requested_model | The model sent in the request |154| model_id | The model_id of the deployment. Autogenerated by LiteLLM, each deployment has a unique model_id |155| api_base | The API Base of the deployment |156| api_provider | The LLM API provider, used for the provider. Example (azure, openai, vertex_ai) |157| hashed_api_key | The hashed api key of the request |158| api_key_alias | The alias of the api key used |159| team | The team of the request |160| team_alias | The alias of the team used |161| exception_status | The status of the exception, if any |162| exception_class | The class of the exception, if any |163164### Success and Failure165166| Metric Name | Description |167|----------------------|--------------------------------------|168 `litellm_deployment_success_responses` | Total number of successful LLM API calls for deployment. Labels: `"requested_model", "litellm_model_name", "model_id", "api_base", "api_provider", "hashed_api_key", "api_key_alias", "team", "team_alias"` |169| `litellm_deployment_failure_responses` | Total number of failed LLM API calls for a specific LLM deployment. Labels: `"requested_model", "litellm_model_name", "model_id", "api_base", "api_provider", "hashed_api_key", "api_key_alias", "team", "team_alias", "exception_status", "exception_class"` |170| `litellm_deployment_total_requests` | Total number of LLM API calls for deployment - success + failure. Labels: `"requested_model", "litellm_model_name", "model_id", "api_base", "api_provider", "hashed_api_key", "api_key_alias", "team", "team_alias"` |171172### Remaining Requests and Tokens173174| Metric Name | Description |175|----------------------|--------------------------------------|176| `litellm_remaining_requests_metric` | Track `x-ratelimit-remaining-requests` returned from LLM API Deployment. Labels: `"model_group", "api_provider", "api_base", "litellm_model_name", "hashed_api_key", "api_key_alias"` |177| `litellm_remaining_tokens_metric` | Track `x-ratelimit-remaining-tokens` return from LLM API Deployment. Labels: `"model_group", "api_provider", "api_base", "litellm_model_name", "hashed_api_key", "api_key_alias"` |178179### Deployment State 180| Metric Name | Description |181|----------------------|--------------------------------------|182| `litellm_deployment_state` | The state of the deployment: 0 = healthy, 1 = partial outage, 2 = complete outage. Labels: `"litellm_model_name", "model_id", "api_base", "api_provider"` |183| `litellm_deployment_latency_per_output_token` | Latency per output token for deployment. Labels: `"litellm_model_name", "model_id", "api_base", "api_provider", "hashed_api_key", "api_key_alias", "team", "team_alias"` |184185#### Fallback (Failover) Metrics186187| Metric Name | Description |188|----------------------|--------------------------------------|189| `litellm_deployment_cooled_down` | Number of times a deployment has been cooled down by LiteLLM load balancing logic. Labels: `"litellm_model_name", "model_id", "api_base", "api_provider"` |190| `litellm_deployment_successful_fallbacks` | Number of successful fallback requests from primary model -> fallback model. Labels: `"requested_model", "fallback_model", "hashed_api_key", "api_key_alias", "team", "team_alias", "exception_status", "exception_class"` |191| `litellm_deployment_failed_fallbacks` | Number of failed fallback requests from primary model -> fallback model. Labels: `"requested_model", "fallback_model", "hashed_api_key", "api_key_alias", "team", "team_alias", "exception_status", "exception_class"` |192193## Request Counting Metrics194195| Metric Name | Description |196|----------------------|--------------------------------------|197| `litellm_requests_metric` | Total number of requests tracked per endpoint. Labels: `"end_user", "hashed_api_key", "api_key_alias", "model", "team", "team_alias", "user", "user_email"` |198199## Request Latency Metrics 200201| Metric Name | Description |202|----------------------|--------------------------------------|203| `litellm_request_total_latency_metric` | Total latency (seconds) for a request to LiteLLM Proxy Server - tracked for labels "end_user", "hashed_api_key", "api_key_alias", "requested_model", "team", "team_alias", "user", "model", "model_id" |204| `litellm_overhead_latency_metric` | Latency overhead (seconds) added by LiteLLM processing - tracked for labels "model_group", "api_provider", "api_base", "litellm_model_name", "hashed_api_key", "api_key_alias" |205| `litellm_llm_api_latency_metric` | Latency (seconds) for just the LLM API call - tracked for labels "model", "hashed_api_key", "api_key_alias", "team", "team_alias", "requested_model", "end_user", "user" |206| `litellm_llm_api_time_to_first_token_metric` | Time to first token for LLM API call - tracked for labels `model`, `hashed_api_key`, `api_key_alias`, `team`, `team_alias`, `requested_model`, `end_user`, `user`, `model_id` [Note: only emitted for streaming requests] |207208## Tracking `end_user` on Prometheus209210By default LiteLLM does not track `end_user` on Prometheus. This is done to reduce the cardinality of the metrics from LiteLLM Proxy.211212If you want to track `end_user` on Prometheus, you can do the following:213214```yaml showLineNumbers title="config.yaml"215litellm_settings:216 callbacks: ["prometheus"]217 enable_end_user_cost_tracking_prometheus_only: true218```219220221## [BETA] Custom Metrics222223Track custom metrics on prometheus on all events mentioned above. 224225### Custom Metadata Labels2262271. Define the custom metadata labels in the `config.yaml`228229```yaml230model_list:231 - model_name: openai/gpt-4o232 litellm_params:233 model: openai/gpt-4o234 api_key: os.environ/OPENAI_API_KEY235236litellm_settings:237 callbacks: ["prometheus"]238 custom_prometheus_metadata_labels: ["metadata.foo", "metadata.bar"]239```2402412. Make a request with the custom metadata labels242243<Tabs>244<TabItem value="Curl" label="Curl Request">245```bash246curl -L -X POST 'http://0.0.0.0:4000/v1/chat/completions' \247-H 'Content-Type: application/json' \248-H 'Authorization: Bearer <LITELLM_API_KEY>' \249-d '{250 "model": "openai/gpt-4o",251 "messages": [252 {253 "role": "user",254 "content": [255 {256 "type": "text",257 "text": "What's in this image?"258 }259 ]260 }261 ],262 "max_tokens": 300,263 "metadata": {264 "foo": "hello world"265 }266}'267```268</TabItem>269<TabItem value="key" label="on Key">270271```bash272curl -L -X POST 'http://0.0.0.0:4000/key/generate' \273-H 'Authorization: Bearer sk-1234' \274-H 'Content-Type: application/json' \275-d '{276 "metadata": {277 "foo": "hello world"278 }279}'280```281</TabItem>282<TabItem value="team" label="on Team">283284```bash285curl -L -X POST 'http://0.0.0.0:4000/team/new' \286-H 'Authorization: Bearer sk-1234' \287-H 'Content-Type: application/json' \288-d '{289 "metadata": {290 "foo": "hello world"291 }292}'293```294</TabItem>295</Tabs>2962973. Check your `/metrics` endpoint for the custom metrics 298299```300... "metadata_foo": "hello world" ...301```302303### Custom Tags304305Track specific tags as prometheus labels for better filtering and monitoring.3063071. Define the custom tags in the `config.yaml`308309```yaml310model_list:311 - model_name: openai/gpt-4o312 litellm_params:313 model: openai/gpt-4o314 api_key: os.environ/OPENAI_API_KEY315316litellm_settings:317 callbacks: ["prometheus"]318 custom_prometheus_metadata_labels: ["metadata.foo", "metadata.bar"]319 custom_prometheus_tags: 320 - "prod"321 - "staging"322 - "batch-job"323 - "User-Agent: RooCode/*"324 - "User-Agent: claude-cli/*"325```3263272. Make a request with tags328329```bash330curl -L -X POST 'http://0.0.0.0:4000/v1/chat/completions' \331-H 'Content-Type: application/json' \332-H 'Authorization: Bearer <LITELLM_API_KEY>' \333-d '{334 "model": "openai/gpt-4o",335 "messages": [336 {337 "role": "user",338 "content": [339 {340 "type": "text",341 "text": "What's in this image?"342 }343 ]344 }345 ],346 "max_tokens": 300,347 "metadata": {348 "tags": ["prod", "user-facing"]349 }350}'351```3523533. Check your `/metrics` endpoint for the custom tag metrics354355```356... "tag_prod": "true", "tag_staging": "false", "tag_batch_job": "false" ...357```358359**How Custom Tags Work:**360- Each configured tag becomes a boolean label in prometheus metrics 361- If a tag matches (exact or wildcard), the label value is `"true"`, otherwise `"false"`362- Tag names are sanitized for prometheus compatibility (e.g., `"batch-job"` becomes `"tag_batch_job"`)363- **Wildcard patterns** supported using `*` (e.g., `"User-Agent: RooCode/*"` matches `"User-Agent: RooCode/1.0.0"`)364365**Example with wildcards:**366```yaml367litellm_settings:368 callbacks: ["prometheus"]369 custom_prometheus_tags:370 - "User-Agent: RooCode/*"371 - "User-Agent: claude-cli/*"372``` 373374**Use Cases:**375- Environment tracking (`prod`, `staging`, `dev`)376- Request type classification (`batch-job`, `user-facing`, `background`)377- Feature flags (`new-feature`, `beta-users`)378- Team or service identification (`team-a`, `service-xyz`)379- User-Agent Tracking - use this to track how much Roo Code, Claude Code, Gemini CLI are used (`User-Agent: RooCode/*`, `User-Agent: claude-cli/*`, `User-Agent: gemini-cli/*`)380381382## Configuring Metrics and Labels383384You can selectively enable specific metrics and control which labels are included to optimize performance and reduce cardinality.385386### Enable Specific Metrics and Labels387388Configure which metrics to emit by specifying them in `prometheus_metrics_config`. Each configuration group needs a `group` name (for organization) and a list of `metrics` to enable. You can optionally include a list of `include_labels` to filter the labels for the metrics.389390```yaml391model_list:392 - model_name: gpt-4o393 litellm_params:394 model: gpt-4o395396litellm_settings:397 callbacks: ["prometheus"]398 prometheus_metrics_config:399 # High-cardinality metrics with minimal labels400 - group: "proxy_metrics"401 metrics:402 - "litellm_proxy_total_requests_metric"403 - "litellm_proxy_failed_requests_metric"404 include_labels:405 - "hashed_api_key"406 - "requested_model"407 - "model_group"408```409410On starting up LiteLLM if your metrics were correctly configured, you should see the following on your container logs411412<Image 413 img={require('../../img/prom_config.png')}414 style={{width: '100%', display: 'block', margin: '2rem auto'}}415/>416417418### Filter Labels Per Metric419420Control which labels are included for each metric to reduce cardinality:421422```yaml423litellm_settings:424 callbacks: ["prometheus"]425 prometheus_metrics_config:426 - group: "token_consumption"427 metrics:428 - "litellm_input_tokens_metric"429 - "litellm_output_tokens_metric"430 - "litellm_total_tokens_metric"431 include_labels:432 - "model"433 - "team"434 - "hashed_api_key"435 - group: "request_tracking"436 metrics:437 - "litellm_proxy_total_requests_metric"438 include_labels:439 - "status_code"440 - "requested_model"441```442443### Advanced Configuration444445You can create multiple configuration groups with different label sets:446447```yaml448litellm_settings:449 callbacks: ["prometheus"]450 prometheus_metrics_config:451 # High-cardinality metrics with minimal labels452 - group: "deployment_health"453 metrics:454 - "litellm_deployment_success_responses"455 - "litellm_deployment_failure_responses"456 include_labels:457 - "api_provider"458 - "requested_model"459 460 # Budget metrics with full label set461 - group: "budget_tracking"462 metrics:463 - "litellm_remaining_team_budget_metric"464 include_labels:465 - "team"466 - "team_alias"467 - "hashed_api_key"468 - "api_key_alias"469 - "model"470 - "end_user"471 472 # Latency metrics with performance-focused labels473 - group: "performance"474 metrics:475 - "litellm_request_total_latency_metric"476 - "litellm_llm_api_latency_metric"477 include_labels:478 - "model"479 - "api_provider"480 - "requested_model"481```482483**Configuration Structure:**484- `group`: A descriptive name for organizing related metrics485- `metrics`: List of metric names to include in this group 486- `include_labels`: (Optional) List of labels to include for these metrics487488**Default Behavior**: If no `prometheus_metrics_config` is specified, all metrics are enabled with their default labels (backward compatible).489490## Monitor System Health491492To monitor the health of litellm adjacent services (redis / postgres), do:493494```yaml495model_list:496 - model_name: gpt-4o497 litellm_params:498 model: gpt-4o499litellm_settings:500 service_callback: ["prometheus_system"]501```502503| Metric Name | Description |504|----------------------|--------------------------------------|505| `litellm_redis_latency` | histogram latency for redis calls |506| `litellm_redis_fails` | Number of failed redis calls |507| `litellm_self_latency` | Histogram latency for successful litellm api call |508509#### DB Transaction Queue Health Metrics510511Use these metrics to monitor the health of the DB Transaction Queue. Eg. Monitoring the size of the in-memory and redis buffers. 512513| Metric Name | Description | Storage Type |514|-----------------------------------------------------|-----------------------------------------------------------------------------|--------------|515| `litellm_pod_lock_manager_size` | Indicates which pod has the lock to write updates to the database. | Redis |516| `litellm_in_memory_daily_spend_update_queue_size` | Number of items in the in-memory daily spend update queue. These are the aggregate spend logs for each user. | In-Memory |517| `litellm_redis_daily_spend_update_queue_size` | Number of items in the Redis daily spend update queue. These are the aggregate spend logs for each user. | Redis |518| `litellm_in_memory_spend_update_queue_size` | In-memory aggregate spend values for keys, users, teams, team members, etc.| In-Memory |519| `litellm_redis_spend_update_queue_size` | Redis aggregate spend values for keys, users, teams, etc. | Redis |520521522523## 🔥 LiteLLM Maintained Grafana Dashboards 524525Link to Grafana Dashboards maintained by LiteLLM526527https://github.com/BerriAI/litellm/tree/main/cookbook/litellm_proxy_server/grafana_dashboard528529Here is a screenshot of the metrics you can monitor with the LiteLLM Grafana Dashboard530531532<Image img={require('../../img/grafana_1.png')} />533534<Image img={require('../../img/grafana_2.png')} />535536<Image img={require('../../img/grafana_3.png')} />537538539## Deprecated Metrics 540541| Metric Name | Description |542|----------------------|--------------------------------------|543| `litellm_llm_api_failed_requests_metric` | **deprecated** use `litellm_proxy_failed_requests_metric` |544545546547## Add authentication on /metrics endpoint548549**By default /metrics endpoint is unauthenticated.** 550551You can opt into running litellm authentication on the /metrics endpoint by setting the following on the config 552553```yaml554litellm_settings:555 require_auth_for_metrics_endpoint: true556```557558## FAQ 559560### What are `_created` vs. `_total` metrics?561562- `_created` metrics are metrics that are created when the proxy starts563- `_total` metrics are metrics that are incremented for each request564565You should consume the `_total` metrics for your counting purposes