# Scaling ContextForge

> ContextForge is designed to scale from single-container development environments to distributed multi-node production deployments.

- Skill: `tools-only/scaling-contextforge` (Agent Skill, multi-file: 3 files)
- Install (CLI): `npx skillmds@latest add tools-only/scaling-contextforge`
- Raw SKILL.md: https://api.skillmd.com/api/skills/tools-only/scaling-contextforge/raw
- Safety review: pending (external: skill-scanner WARNING, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: tools-only (https://skillmd.com/u/tools-only)
- Updated: 2026-09-29
- Page: https://skillmd.com/skills/tools-only/scaling-contextforge

---

# Scaling ContextForge

> Comprehensive guide to scaling ContextForge from development to production, covering vertical scaling, horizontal scaling, connection pooling, performance tuning, and Kubernetes deployment strategies.

## Overview

ContextForge is designed to scale from single-container development environments to distributed multi-node production deployments. For a visual overview of the high-performance architecture including Rust-powered components, see the [Performance Architecture Diagram](../architecture/performance-architecture.md).

This guide covers:

- **Vertical Scaling**: Optimizing single-instance performance with Gunicorn workers
- **Horizontal Scaling**: Multi-container deployments with shared state
- **Database Optimization**: PostgreSQL connection pooling, PgBouncer, and indexing
- **Cache Architecture**: Redis with hiredis, multi-level application caching
- **Performance Tuning**: orjson, compression, and configuration
- **Kubernetes Deployment**: HPA, resource limits, and best practices

---

## Table of Contents

1. [Understanding the GIL and Worker Architecture](#1-understanding-the-gil-and-worker-architecture)
2. [Vertical Scaling with Gunicorn](#2-vertical-scaling-with-gunicorn)
3. [Python 3.14 Free-Threading and PostgreSQL 18](#3-python-314-free-threading-and-postgresql-18)
4. [Horizontal Scaling with Kubernetes](#4-horizontal-scaling-with-kubernetes)
5. [Database Connection Pooling](#5-database-connection-pooling)
6. [Redis for Distributed Caching](#6-redis-for-distributed-caching)
7. [Performance Tuning](#7-performance-tuning)
8. [Benchmarking and Load Testing](#8-benchmarking-and-load-testing)
9. [Health Checks and Readiness](#9-health-checks-and-readiness)
10. [Stateless Architecture and Long-Running Connections](#10-stateless-architecture-and-long-running-connections)
11. [Kubernetes Production Deployment](#11-kubernetes-production-deployment)
12. [Monitoring and Observability](#12-monitoring-and-observability)

---

## 1. Understanding the GIL and Worker Architecture

### The Python Global Interpreter Lock (GIL)

Python's Global Interpreter Lock (GIL) prevents multiple native threads from executing Python bytecode simultaneously. This means:

- **Single worker** = Single CPU core usage (even on multi-core systems)
- **I/O-bound workloads** (API calls, database queries) benefit from async/await
- **CPU-bound workloads** (JSON parsing, encryption) require multiple processes

### Pydantic v2: Rust-Powered Performance

ContextForge leverages **Pydantic v2.11+** for all request/response validation and schema definitions. Unlike pure Python libraries, Pydantic v2 includes a **Rust-based core** (`pydantic-core`) that significantly improves performance:

**Performance benefits:**

- **5-50x faster validation** compared to Pydantic v1
- **JSON parsing** in Rust (bypasses GIL for serialization/deserialization)
- **Schema validation** runs in compiled Rust code
- **Reduced CPU overhead** for request processing

**Impact on scaling:**

- 5,463 lines of Pydantic schemas (`mcpgateway/schemas.py`)
- Every API request validated through Rust-optimized code
- Lower CPU usage per request = higher throughput per worker
- Rust components release the GIL during execution

This means that even within a single worker process, Pydantic's Rust core can run concurrently with Python code for validation-heavy workloads.

### ContextForge's Solution: Gunicorn with Multiple Workers

ContextForge uses **Gunicorn with UvicornWorker** to spawn multiple worker processes:

```python
# gunicorn.config.py
workers = 8                    # Multiple processes bypass the GIL
worker_class = "uvicorn.workers.UvicornWorker"  # Async support
timeout = 600                  # 10-minute timeout for long-running operations
preload_app = True            # Load app once, then fork (memory efficient)
```

**Key benefits:**

- Each worker is a separate process with its own GIL
- 8 workers = ability to use 8 CPU cores
- UvicornWorker enables async I/O within each worker
- Preloading reduces memory footprint (shared code segments)

The trade-off is that you are running multiple Python interpreter instances, and each consumes additional memory.

This also requires having shared state (e.g. Redis or a Database).
---

## 2. Vertical Scaling with Gunicorn

### Worker Count Calculation

**Formula**: `workers = (2 × CPU_cores) + 1`

**Examples:**

| CPU Cores | Recommended Workers | Use Case |
|-----------|---------------------|----------|
| 1 | 2-3 | Development/testing |
| 2 | 4-5 | Small production |
| 4 | 8-9 | Medium production |
| 8 | 16-17 | Large production |

### Configuration Methods

#### Environment Variables

```bash
# Automatic detection based on CPU cores
export GUNICORN_WORKERS=auto

# Manual override
export GUNICORN_WORKERS=16
export GUNICORN_TIMEOUT=600
export GUNICORN_MAX_REQUESTS=100000
export GUNICORN_MAX_REQUESTS_JITTER=100
export GUNICORN_PRELOAD_APP=true
```

#### Kubernetes ConfigMap

```yaml
# charts/mcp-stack/values.yaml
mcpContextForge:
  config:
    GUNICORN_WORKERS: "16"               # Number of worker processes
    GUNICORN_TIMEOUT: "600"              # Worker timeout (seconds)
    GUNICORN_MAX_REQUESTS: "100000"      # Requests before worker restart
    GUNICORN_MAX_REQUESTS_JITTER: "100"  # Prevents thundering herd
    GUNICORN_PRELOAD_APP: "true"         # Memory optimization
```

### Resource Allocation

**CPU**: Allocate 1 CPU core per 2 workers (allows for I/O wait)

**Memory**:

- Base: 256MB
- Per worker: 128-256MB (depending on workload)
- Formula: `memory = 256 + (workers × 200)` MB

**Example for 16 workers:**

- CPU: `8-10 cores` (allows headroom)
- Memory: `3.5-4 GB` (256 + 16×200 = 3.5GB)

```yaml
# Kubernetes resource limits
resources:
  limits:
    cpu: 10000m        # 10 cores
    memory: 4Gi
  requests:
    cpu: 8000m         # 8 cores
    memory: 3584Mi     # 3.5GB
```

---

## 3. Python 3.14 Free-Threading and PostgreSQL 18

### PostgreSQL 18 (Current)

**Status**: Production-ready - ContextForge's default Docker Compose configuration uses PostgreSQL 18.

PostgreSQL 18 provides significant performance improvements:

- **Improved async I/O**: Better non-blocking query performance
- **Reduced latency**: Optimized connection handling
- **Enhanced parallelism**: Better parallel query execution
- **Connection multiplexing**: More efficient connection reuse

**Docker Compose Configuration** (default):

```yaml
postgres:
  image: postgres:18
  command:
    - "postgres"
    - "-c"
    - "max_connections=500"       # With PgBouncer (4000 without)
    - "-c"
    - "shared_buffers=512MB"      # 25% of available RAM
    - "-c"
    - "work_mem=16MB"             # Per-operation memory
    - "-c"
    - "effective_cache_size=1536MB"  # 75% of RAM
    - "-c"
    - "maintenance_work_mem=128MB"
    - "-c"
    - "checkpoint_completion_target=0.9"
    - "-c"
    - "wal_buffers=16MB"
    - "-c"
    - "random_page_cost=1.1"      # SSD optimization
    - "-c"
    - "effective_io_concurrency=200"  # SSD parallel I/O
    - "-c"
    - "max_worker_processes=4"
    - "-c"
    - "max_parallel_workers_per_gather=2"
    - "-c"
    - "max_parallel_workers=4"
```

**Connection URL** (psycopg3 required):

```bash
# Via PgBouncer (recommended for high concurrency)
DATABASE_URL=postgresql+psycopg://postgres:password@pgbouncer:6432/mcp

# Direct connection (for development or low concurrency)
DATABASE_URL=postgresql+psycopg://postgres:password@postgres:5432/mcp
```

### Python 3.14 (Free-Threaded Mode)

**Status**: Beta (as of July 2025) - [PEP 703](https://peps.python.org/pep-0703/)

Python 3.14 introduces **optional free-threading** (GIL removal), a groundbreaking change that enables true parallel multi-threading:

```bash
# Enable free-threading mode
python3.14 -X gil=0 -m gunicorn ...

# Or use PYTHON_GIL environment variable
PYTHON_GIL=0 python3.14 -m gunicorn ...
```

**Performance characteristics:**

| Workload Type | Expected Impact |
|---------------|----------------|
| Single-threaded | **3-15% slower** (overhead from thread-safety mechanisms) |
| Multi-threaded (I/O-bound) | **Minimal impact** (already benefits from async/await) |
| Multi-threaded (CPU-bound) | **Near-linear scaling** with CPU cores |
| Multi-process (current) | **No change** (already bypasses GIL) |

**Benefits when available:**

- **True parallel threads**: Multiple threads execute Python code simultaneously
- **Lower memory overhead**: Threads share memory (vs. separate processes)
- **Faster inter-thread communication**: Shared memory, no IPC overhead
- **Better resource efficiency**: One interpreter instance instead of multiple processes

**Trade-offs:**

- **Single-threaded penalty**: 3-15% slower due to fine-grained locking
- **Library compatibility**: Some C extensions need updates (most popular libraries already compatible)
- **Different scaling model**: Move from `workers=16` to `workers=2 --threads=32`

**Migration strategy:**

1. **Now (Python 3.11-3.13)**: Continue using multi-process Gunicorn
   ```python
   workers = 16                    # Multiple processes
   worker_class = "uvicorn.workers.UvicornWorker"
   ```

2. **Python 3.14 beta**: Test in staging environment
   ```bash
   # Build free-threaded Python
   ./configure --enable-experimental-jit --with-pydebug
   make

   # Test with free-threading
   PYTHON_GIL=0 python3.14 -m pytest tests/
   ```

3. **Python 3.14 stable**: Evaluate hybrid approach
   ```python
   workers = 4                     # Fewer processes
   threads = 8                     # More threads per process
   worker_class = "uvicorn.workers.UvicornWorker"
   ```

4. **Post-migration**: Thread-based scaling
   ```python
   workers = 2                     # Minimal processes
   threads = 32                    # Scale with threads
   preload_app = True              # Single app load
   ```

**Current recommendation**:

- **Production**: Use Python 3.11-3.13 with multi-process Gunicorn (proven, stable)
- **Testing**: Experiment with Python 3.14 beta in non-production environments
- **Monitoring**: Watch for library compatibility announcements

**Why ContextForge is well-positioned for free-threading:**

ContextForge's architecture already benefits from components that will perform even better with Python 3.14:

1. **Pydantic v2 Rust core**: Already bypasses GIL for validation - will work seamlessly with free-threading
2. **FastAPI/Uvicorn**: Built for async I/O - natural fit for thread-based concurrency
3. **SQLAlchemy async**: Database operations already non-blocking
4. **Stateless design**: No shared mutable state between requests

**Resources:**

- [Python 3.14 Free-Threading Guide](https://www.pythoncheatsheet.org/blog/python-3-14-breaking-free-from-gil)
- [PEP 703: Making the GIL Optional](https://peps.python.org/pep-0703/)
- [Python 3.14 Release Schedule](https://peps.python.org/pep-0745/)
- [Pydantic v2 Performance](https://docs.pydantic.dev/latest/blog/pydantic-v2/)

---

## 4. Horizontal Scaling with Kubernetes

### Architecture Overview

```
+------------------------------------------------------------------------------+
|                              Load Balancer                                    |
|                        (Kubernetes Ingress / Service)                         |
+----------------------------------+-------------------------------------------+
                                   |
+----------------------------------v-------------------------------------------+
|                           Nginx Caching Layer                                 |
|   +---------------------+                  +---------------------+            |
|   |   Nginx Cache 1     |                  |   Nginx Cache 2     |            |
|   | - Brotli/Gzip/Zstd  |                  | - Brotli/Gzip/Zstd  |            |
|   | - Static caching    |                  | - Static caching    |            |
|   | - Rate limiting     |                  | - Rate limiting     |            |
|   +----------+----------+                  +----------+----------+            |
+--------------|-----------------------------------------|---------------------+
               |                                         |
+--------------v-----------------------------------------v---------------------+
|                         Gateway Application Layer                             |
|  +------------------+  +------------------+  +------------------+             |
|  |  Gateway Pod 1   |  |  Gateway Pod 2   |  |  Gateway Pod N   |             |
|  |  (16 workers)    |  |  (16 workers)    |  |  (16 workers)    |             |
|  | Gunicorn/Granian |  | Gunicorn/Granian |  | Gunicorn/Granian |             |
|  | orjson, psycopg3 |  | orjson, psycopg3 |  | orjson, psycopg3 |             |
|  +--------+---------+  +--------+---------+  +--------+---------+             |
+-----------|-----------------------|----------------------|-------------------+
            |                       |                      |
            +-----------+-----------+-----------+----------+
                        |                       |
+-------------------------------------------------------------------+
|                           Data Layer                               |
|                                                                    |
|  +---------------------------+    +---------------------------+   |
|  |        PgBouncer          |    |          Redis            |   |
|  | - Connection multiplexing |    | - Distributed cache       |   |
|  | - Transaction pooling     |    | - Session storage         |   |
|  | - 3000 client connections |    | - hiredis parser (83x)    |   |
|  | - 200 server connections  |    | - Leader election         |   |
|  +-------------+-------------+    +---------------------------+   |
|                |                                                   |
|  +-------------v-------------+                                    |
|  |      PostgreSQL 18        |                                    |
|  | - Async I/O               |                                    |
|  | - Auto-prepared stmts     |                                    |
|  | - 500 max_connections     |                                    |
|  | - Parallel query exec     |                                    |
|  +---------------------------+                                    |
+-------------------------------------------------------------------+
```

**Layer Summary:**

| Layer | Component | Purpose | Key Performance Features |
|-------|-----------|---------|--------------------------|
| Edge | Load Balancer | Traffic distribution | SSL termination, health checks |
| Proxy | Nginx | Caching, compression | Brotli/Gzip/Zstd (30-70% bandwidth reduction) |
| App | Gateway Pods | Request processing | Gunicorn/Granian, orjson, multi-level caching |
| Pool | PgBouncer | Connection multiplexing | 3000 client → 200 server connections |
| Cache | Redis | Distributed state | hiredis parser (up to 83x faster) |
| DB | PostgreSQL 18 | Persistent storage | psycopg3 COPY/pipeline, async I/O |

### Shared State Requirements

For multi-pod deployments:

1. **Shared PostgreSQL 18**: All persistent data (servers, tools, users, teams)
2. **PgBouncer**: Connection pooling and multiplexing between gateway pods and PostgreSQL
3. **Shared Redis**: Distributed caching, session storage, and leader election
4. **Stateless pods**: No local state, can be killed/restarted anytime

### Kubernetes Deployment

#### Helm Chart Configuration

```yaml
# charts/mcp-stack/values.yaml
mcpContextForge:
  replicaCount: 3                   # Start with 3 pods

  # Horizontal Pod Autoscaler
  hpa:
    enabled: true
    minReplicas: 3                  # Never scale below 3
    maxReplicas: 20                 # Scale up to 20 pods
    targetCPUUtilizationPercentage: 70    # Scale at 70% CPU
    targetMemoryUtilizationPercentage: 80 # Scale at 80% memory

  # Pod resources
  resources:
    limits:
      cpu: 2000m                    # 2 cores per pod
      memory: 4Gi
    requests:
      cpu: 1000m                    # 1 core per pod
      memory: 2Gi

  # Environment configuration
  config:
    GUNICORN_WORKERS: "8"           # 8 workers per pod
    CACHE_TYPE: redis               # Shared cache
    DB_POOL_SIZE: "50"              # Per-pod pool size

# Shared PostgreSQL
postgres:
  enabled: true
  resources:
    limits:
      cpu: 4000m                    # 4 cores
      memory: 8Gi
    requests:
      cpu: 2000m
      memory: 4Gi

  # Important: Set max_connections
  # Formula: (num_pods × DB_POOL_SIZE × 1.2) + 20
  # Example: (20 pods × 50 pool × 1.2) + 20 = 1220
  config:
    max_connections: 1500           # Adjust based on scale

# Shared Redis
redis:
  enabled: true
  resources:
    limits:
      cpu: 2000m
      memory: 4Gi
    requests:
      cpu: 1000m
      memory: 2Gi
```

#### Deploy with Helm

```bash
# Install/upgrade with custom values
helm upgrade --install mcp-stack ./charts/mcp-stack \
  --namespace mcp-gateway \
  --create-namespace \
  --values production-values.yaml

# Verify HPA
kubectl get hpa -n mcp-gateway
```

### Horizontal Scaling Calculation

**Total capacity** = `pods × workers × requests_per_second`

**Example:**

- 10 pods × 8 workers × 100 RPS = **8,000 RPS**

**Database connections needed:**

- 10 pods × 50 pool size = **500 connections**
- Add 20% overhead = **600 connections**
- Set `max_connections=1000` (buffer for maintenance)

### Docker Compose Reference Architecture

The `docker-compose.yml` provides a production-ready reference architecture with all performance optimizations pre-configured:

```yaml
services:
  # Nginx caching proxy (port 8080)
  nginx:
    image: mcpgateway/nginx-cache:latest
    ports: ["8080:80"]
    volumes:
      - nginx_cache:/var/cache/nginx
    # Brotli/Gzip compression, static caching, rate limiting

  # Gateway application (replicas: 2)
  gateway:
    image: mcpgateway/mcpgateway:latest
    environment:
      # HTTP Server: gunicorn (stable) or granian (faster)
      - HTTP_SERVER=gunicorn
      - GUNICORN_WORKERS=16

      # Database: via PgBouncer
      - DATABASE_URL=postgresql+psycopg://postgres:password@pgbouncer:6432/mcp
      - DB_POOL_SIZE=15              # Smaller with PgBouncer
      - DB_MAX_OVERFLOW=30

      # Redis with hiredis
      - CACHE_TYPE=redis
      - REDIS_URL=redis://redis:6379/0
      - REDIS_PARSER=hiredis
      - REDIS_MAX_CONNECTIONS=150

      # Multi-level caching
      - AUTH_CACHE_ENABLED=true
      - REGISTRY_CACHE_ENABLED=true
      - ADMIN_STATS_CACHE_ENABLED=true

      # Performance: disable overhead
      - LOG_LEVEL=ERROR
      - DISABLE_ACCESS_LOG=true
      - AUDIT_TRAIL_ENABLED=false
      - COMPRESSION_ENABLED=false    # Nginx handles this
    deploy:
      replicas: 2
      resources:
        limits: { cpus: '8', memory: 8G }

  # PgBouncer connection pooler
  pgbouncer:
    image: edoburu/pgbouncer:latest
    environment:
      - DATABASE_URL=postgres://postgres:password@postgres:5432/mcp
      - POOL_MODE=transaction
      - MAX_CLIENT_CONN=3000
      - DEFAULT_POOL_SIZE=120
      - MAX_DB_CONNECTIONS=200

  # PostgreSQL 18
  postgres:
    image: postgres:18
    command:
      - "postgres"
      - "-c" - "max_connections=500"
      - "-c" - "shared_buffers=512MB"
      - "-c" - "effective_cache_size=1536MB"
      - "-c" - "random_page_cost=1.1"
      - "-c" - "effective_io_concurrency=200"

  # Redis with performance tuning
  redis:
    image: redis:latest
    command:
      - "redis-server"
      - "--maxmemory" - "1gb"
      - "--maxmemory-policy" - "allkeys-lru"
      - "--tcp-backlog" - "2048"
      - "--maxclients" - "10000"
```

**Access Points:**

| Port | Service | Use |
|------|---------|-----|
| 8080 | Nginx | Production access (caching, compression) |
| 4444 | Gateway | Direct access (debugging, internal) |
| 6432 | PgBouncer | Database connection pooling |
| 5433 | PostgreSQL | Direct DB access (admin) |
| 6379 | Redis | Cache access |

**Quick Start:**

```bash
# Start full stack
docker-compose up -d

# Access via caching proxy
curl http://localhost:8080/health

# View logs
docker-compose logs -f gateway
```

---

## 5. Database Connection Pooling

### Connection Pool Architecture

ContextForge uses a two-tier connection pooling architecture:

**Without PgBouncer** (direct connection):
```
+-------------------+     +-------------------+     +-------------------+
| Pod 1 (16 workers)|     | Pod 2 (16 workers)|     | Pod N (16 workers)|
| 16 SQLAlchemy     |     | 16 SQLAlchemy     |     | 16 SQLAlchemy     |
| pools × 50 conns  |     | pools × 50 conns  |     | pools × 50 conns  |
+--------+----------+     +--------+----------+     +--------+----------+
         |                         |                         |
         +------------+------------+------------+------------+
                      |
         +------------v------------+
         |      PostgreSQL 18      |
         | max_connections = 4000  |
         +-------------------------+
```

**With PgBouncer** (recommended for high concurrency):
```
+-------------------+     +-------------------+     +-------------------+
| Pod 1 (16 workers)|     | Pod 2 (16 workers)|     | Pod N (16 workers)|
| 16 SQLAlchemy     |     | 16 SQLAlchemy     |     | 16 SQLAlchemy     |
| pools × 15 conns  |     | pools × 15 conns  |     | pools × 15 conns  |
+--------+----------+     +--------+----------+     +--------+----------+
         |                         |                         |
         +------------+------------+------------+------------+
                      |
         +------------v------------+
         |       PgBouncer         |
         | MAX_CLIENT_CONN = 3000  |  <-- Application connections
         | DEFAULT_POOL_SIZE = 120 |
         | MAX_DB_CONNECTIONS = 200|  <-- PostgreSQL connections
         +------------+------------+
                      |
         +------------v------------+
         |      PostgreSQL 18      |
         | max_connections = 500   |  <-- Much lower requirement
         +-------------------------+
```

**Connection Multiplexing Benefits**:

| Metric | Without PgBouncer | With PgBouncer | Improvement |
|--------|-------------------|----------------|-------------|
| App connections | 2 pods × 16 × 50 = 1,600 | 2 pods × 16 × 15 = 480 | App-level reduction |
| PostgreSQL connections | 1,600+ | 200 | **8x reduction** |
| `max_connections` needed | 4,000 | 500 | **8x lower** |
| Memory per connection | ~10MB | ~10MB | Same |
| PostgreSQL memory | ~40GB | ~5GB | **8x reduction** |
| Connection setup time | Per request | Reused | Near-zero |

### Pool Configuration

#### Environment Variables

```bash
# Connection pool settings
DB_POOL_SIZE=50              # Persistent connections per worker
DB_MAX_OVERFLOW=10           # Additional connections allowed
DB_POOL_TIMEOUT=60           # Wait time before timeout (seconds)
DB_POOL_RECYCLE=3600         # Recycle connections after 1 hour
DB_MAX_RETRIES=30            # Retry attempts on failure (exponential backoff)
DB_RETRY_INTERVAL_MS=2000    # Base retry interval (doubles each attempt, max 30s)

# psycopg3-specific optimizations
DB_PREPARE_THRESHOLD=5       # Auto-prepare queries after N executions (0=disable)
```

#### psycopg3 Performance Features

ContextForge uses **psycopg3** (`psycopg[c,binary]`) instead of psycopg2, providing significant performance improvements and modern features.

**Why psycopg3:**

| Feature | psycopg2 | psycopg3 |
|---------|----------|----------|
| Parameter binding | Client-side | Server-side (more secure) |
| Prepared statements | Manual | Automatic after N executions |
| Binary protocol | Optional | Native support |
| Async support | Wrapper | First-class built-in |
| Maintenance | Maintenance mode | Active development |

**Performance Benchmarks:**

| Operation | psycopg2 | psycopg3 | Improvement |
|-----------|----------|----------|-------------|
| Bulk INSERT (1000+ rows) | Standard | COPY protocol | 5-10x |
| Repeated queries | Parsed each time | Auto-prepared | 2-3x |
| Batch queries | Sequential | Pipelined | 2-5x |

**Automatic Prepared Statements:**

```bash
# Number of executions before auto-preparing a query server-side
# Default: 5 (balance between memory and performance)
# Set to 0 to disable, 1 to prepare immediately
DB_PREPARE_THRESHOLD=5
```

After N executions of the same query pattern, psycopg3 creates a server-side prepared statement, reducing:

- Query parsing overhead (5-15% per query)
- Network round-trips for query plans
- PostgreSQL CPU usage for repeated queries

**COPY Protocol for Bulk Inserts:**

The utility module `mcpgateway/utils/psycopg3_optimizations.py` provides COPY protocol support:

```python
from mcpgateway.utils.psycopg3_optimizations import bulk_insert_with_copy

# 5-10x faster for 1000+ rows
bulk_insert_with_copy(db, "tool_metrics", columns, rows)
```

Note: COPY is only faster for large batches (1000+ rows); for small batches (<100 rows), SQLAlchemy's `bulk_insert_mappings()` is actually faster due to lower protocol overhead.

**Pipeline Mode for Batch Queries:**

Execute multiple queries with reduced round-trips:

```python
from mcpgateway.utils.psycopg3_optimizations import execute_pipelined

# Send multiple queries without waiting for individual responses
results = execute_pipelined(db, [
    ("SELECT * FROM tools WHERE id = %s", {"id": tool_id}),
    ("SELECT * FROM gateways WHERE id = %s", {"id": gateway_id}),
])
```

**Connection URL Format:**

```bash
# IMPORTANT: Required format for psycopg3 (NOT postgresql://)
DATABASE_URL=postgresql+psycopg://user:pass@host:5432/db
```

See [ADR-027: Migrate to psycopg3](../architecture/adr/027-migrate-psycopg3.md) for details.

#### Configuration in Code

```python
# mcpgateway/config.py
@property
def database_settings(self) -> dict:
    return {
        "pool_size": self.db_pool_size,          # 50
        "max_overflow": self.db_max_overflow,    # 10
        "pool_timeout": self.db_pool_timeout,    # 60s
        "pool_recycle": self.db_pool_recycle,    # 3600s
    }
```

### PostgreSQL Configuration

#### Calculate max_connections

```bash
# Formula
max_connections = (num_pods × num_workers × pool_size × 1.2) + buffer

# Example: 10 pods, 8 workers, 50 pool size
max_connections = (10 × 8 × 50 × 1.2) + 200 = 5000 connections
```

#### PostgreSQL Configuration File

```ini
# postgresql.conf
max_connections = 5000
shared_buffers = 16GB              # 25% of RAM
effective_cache_size = 48GB        # 75% of RAM
work_mem = 16MB                    # Per operation
maintenance_work_mem = 2GB
```

#### Managed Services

**IBM Cloud Databases for PostgreSQL:**
```bash
# Increase max_connections via CLI
ibmcloud cdb deployment-configuration postgres \
  --configuration max_connections=5000
```

**AWS RDS:**
```bash
# Via parameter group
max_connections = {DBInstanceClassMemory/9531392}
```

**Google Cloud SQL:**
```bash
# Auto-scales based on instance size
# 4 vCPU = 400 connections
# 8 vCPU = 800 connections
```

### PgBouncer Connection Pooling

For high-concurrency deployments, PgBouncer provides connection multiplexing between the gateway and PostgreSQL, dramatically reducing connection overhead.

**Architecture:**

```
Without PgBouncer:
  Gateway (2 replicas × 16 workers) → PostgreSQL (max_connections=4000)
  Each worker maintains its own pool → High connection churn

With PgBouncer:
  Gateway (2 replicas × 16 workers) → PgBouncer → PostgreSQL (max_connections=500)
  Connections multiplexed → Efficient reuse, lower PostgreSQL overhead
```

**Benefits:**

- **Connection multiplexing**: Many app connections share fewer database connections
- **Reduced PostgreSQL overhead**: Lower `max_connections` reduces memory per connection
- **Connection reuse**: PgBouncer maintains persistent connections to PostgreSQL
- **Graceful handling of connection storms**: Queues requests instead of rejecting

**Docker Compose Setup:**

```yaml
pgbouncer:
  image: edoburu/pgbouncer:latest
  restart: unless-stopped
  ports:
    - "6432:6432"
  environment:
    - DATABASE_URL=postgres://postgres:password@postgres:5432/mcp
    - POOL_MODE=transaction
    - MAX_CLIENT_CONN=2000
    - DEFAULT_POOL_SIZE=100
    - MIN_POOL_SIZE=10
    - RESERVE_POOL_SIZE=25
    - MAX_DB_CONNECTIONS=200
    - SERVER_LIFETIME=3600
    - SERVER_IDLE_TIMEOUT=600
    - AUTH_TYPE=scram-sha-256
  depends_on:
    postgres:
      condition: service_healthy
  healthcheck:
    test: ["CMD", "pg_isready", "-h", "localhost", "-p", "6432"]
    interval: 10s
    timeout: 5s
    retries: 3
```

**Gateway Configuration with PgBouncer:**

```bash
# Connect via PgBouncer instead of direct PostgreSQL
DATABASE_URL=postgresql+psycopg://postgres:password@pgbouncer:6432/mcp

# Smaller pool since PgBouncer handles pooling
DB_POOL_SIZE=10
DB_MAX_OVERFLOW=20
```

**Pool Modes:**

| Mode | Description | Use Case |
|------|-------------|----------|
| `transaction` | Connection returned after transaction commit | **Recommended** for most workloads |
| `session` | Connection held for entire session | Legacy apps requiring session state |
| `statement` | Connection returned after each statement | Limited use cases |

**Key Tuning Parameters:**

| Parameter | Description | Suggested Value |
|-----------|-------------|-----------------|
| `MAX_CLIENT_CONN` | Max app connections | 2000-5000 |
| `DEFAULT_POOL_SIZE` | Connections per user/db pair | 100 |
| `MAX_DB_CONNECTIONS` | Max connections to PostgreSQL | 200-500 |
| `POOL_MODE` | When to return connections | `transaction` |

**Monitoring PgBouncer:**

```bash
# Connect to PgBouncer admin console
psql -h localhost -p 6432 -U pgbouncer pgbouncer

# Show connection pool statistics
SHOW STATS;
SHOW POOLS;
SHOW CLIENTS;
SHOW SERVERS;
```

### Connection Pool Monitoring

```python
# Health endpoint checks pool status
@app.get("/health")
async def healthcheck(db: Session = Depends(get_db)):
    try:
        db.execute(text("SELECT 1"))
        return {"status": "healthy"}
    except Exception as e:
        return {"status": "unhealthy", "error": str(e)}
```

```bash
# Check PostgreSQL connections
kubectl exec -it postgres-pod -- psql -U admin -d postgresdb \
  -c "SELECT count(*) FROM pg_stat_activity;"
```

### Database Session Management

To prevent connection pool exhaustion under high load, ContextForge releases database sessions before making upstream HTTP calls:

**Problem:** Database sessions held during slow upstream calls (100ms - 4+ minutes) exhaust the connection pool even when the database is lightly loaded.

**Solution:** "Fetch-Then-Release" pattern:

1. **Eager load** required data with `joinedload()` in single query
2. **Extract data** to local variables before network I/O
3. **Release session** with `db.expunge()` + `db.close()` before HTTP calls
4. **Fresh session** for metrics recording after the call

This reduces connection hold time from minutes to <50ms and enables 10x+ higher concurrency.

---

## 6. Redis for Distributed Caching

### Architecture

Redis provides shared state across all Gateway pods:

- **Session storage**: User sessions (TTL: 3600s)
- **Message cache**: Ephemeral data (TTL: 600s)
- **Federation cache**: Gateway peer discovery

### Configuration

#### Enable Redis Caching

```bash
# .env or Kubernetes ConfigMap
CACHE_TYPE=redis
REDIS_URL=redis://redis-service:6379/0
CACHE_PREFIX=mcpgw:
SESSION_TTL=3600
MESSAGE_TTL=600
REDIS_MAX_RETRIES=30             # Retry attempts on failure (exponential backoff)
REDIS_RETRY_INTERVAL_MS=2000     # Base retry interval (doubles each attempt, max 30s)

# Connection pool (standard)
REDIS_MAX_CONNECTIONS=50
REDIS_SOCKET_TIMEOUT=2.0
REDIS_SOCKET_CONNECT_TIMEOUT=2.0
REDIS_RETRY_ON_TIMEOUT=true
REDIS_HEALTH_CHECK_INTERVAL=30

# Leader election (multi-node)
REDIS_LEADER_TTL=15
REDIS_LEADER_HEARTBEAT_INTERVAL=5
```

#### High-Concurrency Redis Tuning

For 1000+ concurrent users:

```bash
# Connection pool - Formula: (concurrent_requests / workers) * 1.5
# Example: 32 workers × 150 = 4800 < Redis maxclients (10000)
REDIS_MAX_CONNECTIONS=150

# Timeouts - keep low for fast failure detection
REDIS_SOCKET_TIMEOUT=5.0
REDIS_SOCKET_CONNECT_TIMEOUT=5.0
REDIS_HEALTH_CHECK_INTERVAL=30
```

**Redis Server Tuning** (docker-compose.yml):

```yaml
redis:
  command:
    - "redis-server"
    - "--maxmemory"
    - "1gb"
    - "--maxmemory-policy"
    - "allkeys-lru"
    - "--tcp-backlog"
    - "2048"                    # Higher for pending connections
    - "--maxclients"
    - "10000"                   # Max client connections
```

#### Kubernetes Deployment

```yaml
# charts/mcp-stack/values.yaml
redis:
  enabled: true

  resources:
    limits:
      cpu: 2000m
      memory: 4Gi
    requests:
      cpu: 1000m
      memory: 2Gi

  # Enable persistence
  persistence:
    enabled: true
    size: 10Gi
```

### Redis Sizing

**Memory calculation:**

- Sessions: `concurrent_users × 50KB`
- Messages: `messages_per_minute × 100KB × (TTL/60)`

**Example:**

- 10,000 users × 50KB = 500MB
- 1,000 msg/min × 100KB × 10min = 1GB
- **Total: 1.5GB + 50% overhead = 2.5GB**

**Connection pool sizing:**

- Formula: `REDIS_MAX_CONNECTIONS = (concurrent_requests / workers) × 1.5`
- Default 50 handles ~500 concurrent requests with 10 workers
- High-concurrency: increase to 100 and lower timeouts

```bash
# High-concurrency production overrides
REDIS_MAX_CONNECTIONS=100
REDIS_SOCKET_TIMEOUT=1.0
REDIS_SOCKET_CONNECT_TIMEOUT=1.0
REDIS_HEALTH_CHECK_INTERVAL=15
REDIS_LEADER_TTL=10
REDIS_LEADER_HEARTBEAT_INTERVAL=3
```

### Hiredis High-Performance Parser

ContextForge uses `redis[hiredis]` for significantly faster Redis protocol parsing, especially beneficial for large responses.

**Performance Impact:**

| Operation | Pure Python | With Hiredis | Improvement |
|-----------|-------------|--------------|-------------|
| Simple SET/GET | Baseline | +10% | 1.1x |
| LRANGE (10 items) | Baseline | +170% | 2.7x |
| LRANGE (100 items) | Baseline | +1000% | ~10x |
| LRANGE (999 items) | Baseline | +8220% | **83x** |

The larger the response, the greater the improvement. This is critical for:

- Tool registry queries returning many tools
- Bulk operations and federation
- Cached response retrieval
- Metrics aggregation

**Configuration:**

```bash
# Redis parser selection (hiredis is default for performance)
# Options: auto (default - uses hiredis if available), hiredis, python
# Use "python" to force pure-Python parser (useful for debugging)
REDIS_PARSER=auto
```

**When to use each parser:**

| Parser | Use Case |
|--------|----------|
| `hiredis` (default) | Production, high throughput |
| `python` | Debugging Redis protocol issues, restricted environments |

### High Availability

**Redis Sentinel** (3+ nodes):
```yaml
redis:
  sentinel:
    enabled: true
    quorum: 2

  replicas: 3  # 1 primary + 2 replicas
```

**Redis Cluster** (6+ nodes):
```bash
REDIS_URL=redis://redis-cluster:6379/0?cluster=true
```

---

## 7. Performance Tuning

### Application Architecture Performance

ContextForge's technology stack is optimized for high performance:

**Rust-Powered Components:**

- **Pydantic v2** (5-50x faster validation via Rust core)
- **Uvicorn with [standard] extras** (ASGI server with high-performance components)

**Async-First Design:**

- **FastAPI** (async request handling)
- **SQLAlchemy 2.0** (async database operations)
- **asyncio** event loop per worker

**Performance characteristics:**

- Request validation: **< 1ms** (Pydantic v2 Rust core)
- JSON serialization: **3-5x faster** than pure Python
- Database queries: Non-blocking async I/O
- Concurrent requests per worker: **1000+** (async event loop)

### System-Level Optimization

#### Kernel Parameters

```bash
# /etc/sysctl.conf
net.core.somaxconn=4096
net.ipv4.tcp_max_syn_backlog=4096
net.ipv4.ip_local_port_range=1024 65535
net.ipv4.tcp_tw_reuse=1
fs.file-max=2097152

# Apply changes
sysctl -p
```

#### File Descriptors

```bash
# /etc/security/limits.conf
* soft nofile 1048576
* hard nofile 1048576

# Verify
ulimit -n
```

### HTTP Server Selection

ContextForge supports two production HTTP servers:

| Server | Description | Best For |
|--------|-------------|----------|
| **Gunicorn** (default) | Python-based with Uvicorn workers | Stable, well-tested |
| **Granian** | Rust-based HTTP server | Maximum performance (+20-50%) |

```bash
# Select HTTP server (in containers)
HTTP_SERVER=gunicorn    # Default, stable
HTTP_SERVER=granian     # Alternative, Rust-based
```

### Gunicorn Configuration

```bash
# Number of worker processes
# Options: "auto" (default, 2*CPU+1 capped at 16), or positive integer
GUNICORN_WORKERS=auto

# Worker timeout in seconds (increase for long-running LLM requests)
GUNICORN_TIMEOUT=600

# Maximum requests per worker before automatic restart (prevents memory leaks)
GUNICORN_MAX_REQUESTS=100000

# Random jitter added to max requests (prevents thundering herd)
GUNICORN_MAX_REQUESTS_JITTER=100

# Preload application before forking workers
# true: Saves memory (shared code), runs migrations once before forking
# false: Each worker loads app independently (more memory, better isolation)
GUNICORN_PRELOAD_APP=true

# Development mode with hot reload (NOT for production)
GUNICORN_DEV_MODE=false
```

**Worker Class**: UvicornWorker (default) for async support.

### Granian Configuration (Alternative)

Granian is a Rust-based HTTP server with native backpressure for overload protection:

```bash
# HTTP server selection
HTTP_SERVER=granian

# Worker count (auto = CPU count, max 16)
GRANIAN_WORKERS=16

# TCP backlog for pending connections
GRANIAN_BACKLOG=4096

# Backpressure: max concurrent requests per worker before 503 rejection
# Total capacity = WORKERS × BACKPRESSURE = 16 × 64 = 1024 concurrent requests
GRANIAN_BACKPRESSURE=64

# HTTP/1.1 buffer size (bytes)
GRANIAN_HTTP1_BUFFER_SIZE=524288

# Auto-restart failed workers
GRANIAN_RESPAWN_FAILED=true
```

**Backpressure behavior:**

- Requests within capacity (≤1024): Processed normally
- Requests over capacity (>1024): Immediate 503 Service Unavailable
- No queuing, no memory growth, no cascading timeouts

**When to use Granian:**

- Load spike protection (backpressure rejects excess gracefully)
- Bursty or unpredictable traffic patterns
- High-concurrency deployments (1000+ concurrent users)

**When to use Gunicorn:**

- Memory-constrained environments (32% less RAM)
- Maximum stability and compatibility
- Standard deployments with predictable traffic

### Application Tuning

```bash
# Resource limits
TOOL_TIMEOUT=60
TOOL_CONCURRENT_LIMIT=10
RESOURCE_CACHE_SIZE=1000
RESOURCE_CACHE_TTL=3600

# Retry configuration
RETRY_MAX_ATTEMPTS=3
RETRY_BASE_DELAY=1.0
RETRY_MAX_DELAY=60

# Health check intervals
HEALTH_CHECK_INTERVAL=60
HEALTH_CHECK_TIMEOUT=5
UNHEALTHY_THRESHOLD=3

# Gateway health check timeout (seconds)
GATEWAY_HEALTH_CHECK_TIMEOUT=5.0

# Auto-refresh tools during health checks
# When enabled, tools/resources/prompts are fetched and synced during health checks
AUTO_REFRESH_SERVERS=false

```

### Logging for Performance

Logging can significantly impact performance under high load:

```bash
# Log level - ERROR recommended for production
# DEBUG/INFO create massive I/O overhead
LOG_LEVEL=ERROR

# Disable access logging (massive I/O overhead under high concurrency)
DISABLE_ACCESS_LOG=true

# Disable database logging for performance
STRUCTURED_LOGGING_DATABASE_ENABLED=false
```

**Impact of logging settings:**

| Setting | I/O Overhead | Use Case |
|---------|--------------|----------|
| `LOG_LEVEL=DEBUG` | Very High | Development only |
| `LOG_LEVEL=INFO` | High | Light load, debugging |
| `LOG_LEVEL=ERROR` | Low | Production (recommended) |
| `DISABLE_ACCESS_LOG=true` | None | Production (recommended) |
| `STRUCTURED_LOGGING_DATABASE_ENABLED=true` | Very High | Compliance (use external aggregator) |

### Metrics Buffer Configuration

Bat

…(truncated)
