Scaling ContextForge
Comprehensive guide to scaling ContextForge from development to production, covering vertical scaling, horizontal scaling, connection pooling, performance tuning, and Kubernetes deployment strategies.
Overview
ContextForge is designed to scale from single-container development environments to distributed multi-node production deployments. For a visual overview of the high-performance architecture including Rust-powered components, see the Performance Architecture Diagram.
This guide covers:
- Vertical Scaling: Optimizing single-instance performance with Gunicorn workers
- Horizontal Scaling: Multi-container deployments with shared state
- Database Optimization: PostgreSQL connection pooling, PgBouncer, and indexing
- Cache Architecture: Redis with hiredis, multi-level application caching
- Performance Tuning: orjson, compression, and configuration
- Kubernetes Deployment: HPA, resource limits, and best practices
Table of Contents
- Understanding the GIL and Worker Architecture
- Vertical Scaling with Gunicorn
- Python 3.14 Free-Threading and PostgreSQL 18
- Horizontal Scaling with Kubernetes
- Database Connection Pooling
- Redis for Distributed Caching
- Performance Tuning
- Benchmarking and Load Testing
- Health Checks and Readiness
- Stateless Architecture and Long-Running Connections
- Kubernetes Production Deployment
- Monitoring and Observability
1. Understanding the GIL and Worker Architecture
The Python Global Interpreter Lock (GIL)
Python's Global Interpreter Lock (GIL) prevents multiple native threads from executing Python bytecode simultaneously. This means:
- Single worker = Single CPU core usage (even on multi-core systems)
- I/O-bound workloads (API calls, database queries) benefit from async/await
- CPU-bound workloads (JSON parsing, encryption) require multiple processes
Pydantic v2: Rust-Powered Performance
ContextForge leverages Pydantic v2.11+ for all request/response validation and schema definitions. Unlike pure Python libraries, Pydantic v2 includes a Rust-based core (pydantic-core) that significantly improves performance:
Performance benefits:
- 5-50x faster validation compared to Pydantic v1
- JSON parsing in Rust (bypasses GIL for serialization/deserialization)
- Schema validation runs in compiled Rust code
- Reduced CPU overhead for request processing
Impact on scaling:
- 5,463 lines of Pydantic schemas (
mcpgateway/schemas.py) - Every API request validated through Rust-optimized code
- Lower CPU usage per request = higher throughput per worker
- Rust components release the GIL during execution
This means that even within a single worker process, Pydantic's Rust core can run concurrently with Python code for validation-heavy workloads.
ContextForge's Solution: Gunicorn with Multiple Workers
ContextForge uses Gunicorn with UvicornWorker to spawn multiple worker processes:
# gunicorn.config.py
workers = 8 # Multiple processes bypass the GIL
worker_class = "uvicorn.workers.UvicornWorker" # Async support
timeout = 600 # 10-minute timeout for long-running operations
preload_app = True # Load app once, then fork (memory efficient)
Key benefits:
- Each worker is a separate process with its own GIL
- 8 workers = ability to use 8 CPU cores
- UvicornWorker enables async I/O within each worker
- Preloading reduces memory footprint (shared code segments)
The trade-off is that you are running multiple Python interpreter instances, and each consumes additional memory.
This also requires having shared state (e.g. Redis or a Database).
2. Vertical Scaling with Gunicorn
Worker Count Calculation
Formula: workers = (2 × CPU_cores) + 1
Examples:
| CPU Cores | Recommended Workers | Use Case |
|---|---|---|
| 1 | 2-3 | Development/testing |
| 2 | 4-5 | Small production |
| 4 | 8-9 | Medium production |
| 8 | 16-17 | Large production |
Configuration Methods
Environment Variables
# Automatic detection based on CPU cores
export GUNICORN_WORKERS=auto
# Manual override
export GUNICORN_WORKERS=16
export GUNICORN_TIMEOUT=600
export GUNICORN_MAX_REQUESTS=100000
export GUNICORN_MAX_REQUESTS_JITTER=100
export GUNICORN_PRELOAD_APP=true
Kubernetes ConfigMap
# charts/mcp-stack/values.yaml
mcpContextForge:
config:
GUNICORN_WORKERS: "16" # Number of worker processes
GUNICORN_TIMEOUT: "600" # Worker timeout (seconds)
GUNICORN_MAX_REQUESTS: "100000" # Requests before worker restart
GUNICORN_MAX_REQUESTS_JITTER: "100" # Prevents thundering herd
GUNICORN_PRELOAD_APP: "true" # Memory optimization
Resource Allocation
CPU: Allocate 1 CPU core per 2 workers (allows for I/O wait)
Memory:
- Base: 256MB
- Per worker: 128-256MB (depending on workload)
- Formula:
memory = 256 + (workers × 200)MB
Example for 16 workers:
- CPU:
8-10 cores(allows headroom) - Memory:
3.5-4 GB(256 + 16×200 = 3.5GB)
# Kubernetes resource limits
resources:
limits:
cpu: 10000m # 10 cores
memory: 4Gi
requests:
cpu: 8000m # 8 cores
memory: 3584Mi # 3.5GB
3. Python 3.14 Free-Threading and PostgreSQL 18
PostgreSQL 18 (Current)
Status: Production-ready - ContextForge's default Docker Compose configuration uses PostgreSQL 18.
PostgreSQL 18 provides significant performance improvements:
- Improved async I/O: Better non-blocking query performance
- Reduced latency: Optimized connection handling
- Enhanced parallelism: Better parallel query execution
- Connection multiplexing: More efficient connection reuse
Docker Compose Configuration (default):
postgres:
image: postgres:18
command:
- "postgres"
- "-c"
- "max_connections=500" # With PgBouncer (4000 without)
- "-c"
- "shared_buffers=512MB" # 25% of available RAM
- "-c"
- "work_mem=16MB" # Per-operation memory
- "-c"
- "effective_cache_size=1536MB" # 75% of RAM
- "-c"
- "maintenance_work_mem=128MB"
- "-c"
- "checkpoint_completion_target=0.9"
- "-c"
- "wal_buffers=16MB"
- "-c"
- "random_page_cost=1.1" # SSD optimization
- "-c"
- "effective_io_concurrency=200" # SSD parallel I/O
- "-c"
- "max_worker_processes=4"
- "-c"
- "max_parallel_workers_per_gather=2"
- "-c"
- "max_parallel_workers=4"
Connection URL (psycopg3 required):
# Via PgBouncer (recommended for high concurrency)
DATABASE_URL=postgresql+psycopg://postgres:password@pgbouncer:6432/mcp
# Direct connection (for development or low concurrency)
DATABASE_URL=postgresql+psycopg://postgres:password@postgres:5432/mcp
Python 3.14 (Free-Threaded Mode)
Status: Beta (as of July 2025) - PEP 703
Python 3.14 introduces optional free-threading (GIL removal), a groundbreaking change that enables true parallel multi-threading:
# Enable free-threading mode
python3.14 -X gil=0 -m gunicorn ...
# Or use PYTHON_GIL environment variable
PYTHON_GIL=0 python3.14 -m gunicorn ...
Performance characteristics:
| Workload Type | Expected Impact |
|---|---|
| Single-threaded | 3-15% slower (overhead from thread-safety mechanisms) |
| Multi-threaded (I/O-bound) | Minimal impact (already benefits from async/await) |
| Multi-threaded (CPU-bound) | Near-linear scaling with CPU cores |
| Multi-process (current) | No change (already bypasses GIL) |
Benefits when available:
- True parallel threads: Multiple threads execute Python code simultaneously
- Lower memory overhead: Threads share memory (vs. separate processes)
- Faster inter-thread communication: Shared memory, no IPC overhead
- Better resource efficiency: One interpreter instance instead of multiple processes
Trade-offs:
- Single-threaded penalty: 3-15% slower due to fine-grained locking
- Library compatibility: Some C extensions need updates (most popular libraries already compatible)
- Different scaling model: Move from
workers=16toworkers=2 --threads=32
Migration strategy:
Now (Python 3.11-3.13): Continue using multi-process Gunicorn
workers = 16 # Multiple processes worker_class = "uvicorn.workers.UvicornWorker"Python 3.14 beta: Test in staging environment
# Build free-threaded Python ./configure --enable-experimental-jit --with-pydebug make # Test with free-threading PYTHON_GIL=0 python3.14 -m pytest tests/Python 3.14 stable: Evaluate hybrid approach
workers = 4 # Fewer processes threads = 8 # More threads per process worker_class = "uvicorn.workers.UvicornWorker"Post-migration: Thread-based scaling
workers = 2 # Minimal processes threads = 32 # Scale with threads preload_app = True # Single app load
Current recommendation:
- Production: Use Python 3.11-3.13 with multi-process Gunicorn (proven, stable)
- Testing: Experiment with Python 3.14 beta in non-production environments
- Monitoring: Watch for library compatibility announcements
Why ContextForge is well-positioned for free-threading:
ContextForge's architecture already benefits from components that will perform even better with Python 3.14:
- Pydantic v2 Rust core: Already bypasses GIL for validation - will work seamlessly with free-threading
- FastAPI/Uvicorn: Built for async I/O - natural fit for thread-based concurrency
- SQLAlchemy async: Database operations already non-blocking
- Stateless design: No shared mutable state between requests
Resources:
- Python 3.14 Free-Threading Guide
- PEP 703: Making the GIL Optional
- Python 3.14 Release Schedule
- Pydantic v2 Performance
4. Horizontal Scaling with Kubernetes
Architecture Overview
+------------------------------------------------------------------------------+
| Load Balancer |
| (Kubernetes Ingress / Service) |
+----------------------------------+-------------------------------------------+
|
+----------------------------------v-------------------------------------------+
| Nginx Caching Layer |
| +---------------------+ +---------------------+ |
| | Nginx Cache 1 | | Nginx Cache 2 | |
| | - Brotli/Gzip/Zstd | | - Brotli/Gzip/Zstd | |
| | - Static caching | | - Static caching | |
| | - Rate limiting | | - Rate limiting | |
| +----------+----------+ +----------+----------+ |
+--------------|-----------------------------------------|---------------------+
| |
+--------------v-----------------------------------------v---------------------+
| Gateway Application Layer |
| +------------------+ +------------------+ +------------------+ |
| | Gateway Pod 1 | | Gateway Pod 2 | | Gateway Pod N | |
| | (16 workers) | | (16 workers) | | (16 workers) | |
| | Gunicorn/Granian | | Gunicorn/Granian | | Gunicorn/Granian | |
| | orjson, psycopg3 | | orjson, psycopg3 | | orjson, psycopg3 | |
| +--------+---------+ +--------+---------+ +--------+---------+ |
+-----------|-----------------------|----------------------|-------------------+
| | |
+-----------+-----------+-----------+----------+
| |
+-------------------------------------------------------------------+
| Data Layer |
| |
| +---------------------------+ +---------------------------+ |
| | PgBouncer | | Redis | |
| | - Connection multiplexing | | - Distributed cache | |
| | - Transaction pooling | | - Session storage | |
| | - 3000 client connections | | - hiredis parser (83x) | |
| | - 200 server connections | | - Leader election | |
| +-------------+-------------+ +---------------------------+ |
| | |
| +-------------v-------------+ |
| | PostgreSQL 18 | |
| | - Async I/O | |
| | - Auto-prepared stmts | |
| | - 500 max_connections | |
| | - Parallel query exec | |
| +---------------------------+ |
+-------------------------------------------------------------------+
Layer Summary:
| Layer | Component | Purpose | Key Performance Features |
|---|---|---|---|
| Edge | Load Balancer | Traffic distribution | SSL termination, health checks |
| Proxy | Nginx | Caching, compression | Brotli/Gzip/Zstd (30-70% bandwidth reduction) |
| App | Gateway Pods | Request processing | Gunicorn/Granian, orjson, multi-level caching |
| Pool | PgBouncer | Connection multiplexing | 3000 client → 200 server connections |
| Cache | Redis | Distributed state | hiredis parser (up to 83x faster) |
| DB | PostgreSQL 18 | Persistent storage | psycopg3 COPY/pipeline, async I/O |
Shared State Requirements
For multi-pod deployments:
- Shared PostgreSQL 18: All persistent data (servers, tools, users, teams)
- PgBouncer: Connection pooling and multiplexing between gateway pods and PostgreSQL
- Shared Redis: Distributed caching, session storage, and leader election
- Stateless pods: No local state, can be killed/restarted anytime
Kubernetes Deployment
Helm Chart Configuration
# charts/mcp-stack/values.yaml
mcpContextForge:
replicaCount: 3 # Start with 3 pods
# Horizontal Pod Autoscaler
hpa:
enabled: true
minReplicas: 3 # Never scale below 3
maxReplicas: 20 # Scale up to 20 pods
targetCPUUtilizationPercentage: 70 # Scale at 70% CPU
targetMemoryUtilizationPercentage: 80 # Scale at 80% memory
# Pod resources
resources:
limits:
cpu: 2000m # 2 cores per pod
memory: 4Gi
requests:
cpu: 1000m # 1 core per pod
memory: 2Gi
# Environment configuration
config:
GUNICORN_WORKERS: "8" # 8 workers per pod
CACHE_TYPE: redis # Shared cache
DB_POOL_SIZE: "50" # Per-pod pool size
# Shared PostgreSQL
postgres:
enabled: true
resources:
limits:
cpu: 4000m # 4 cores
memory: 8Gi
requests:
cpu: 2000m
memory: 4Gi
# Important: Set max_connections
# Formula: (num_pods × DB_POOL_SIZE × 1.2) + 20
# Example: (20 pods × 50 pool × 1.2) + 20 = 1220
config:
max_connections: 1500 # Adjust based on scale
# Shared Redis
redis:
enabled: true
resources:
limits:
cpu: 2000m
memory: 4Gi
requests:
cpu: 1000m
memory: 2Gi
Deploy with Helm
# Install/upgrade with custom values
helm upgrade --install mcp-stack ./charts/mcp-stack \
--namespace mcp-gateway \
--create-namespace \
--values production-values.yaml
# Verify HPA
kubectl get hpa -n mcp-gateway
Horizontal Scaling Calculation
Total capacity = pods × workers × requests_per_second
Example:
- 10 pods × 8 workers × 100 RPS = 8,000 RPS
Database connections needed:
- 10 pods × 50 pool size = 500 connections
- Add 20% overhead = 600 connections
- Set
max_connections=1000(buffer for maintenance)
Docker Compose Reference Architecture
The docker-compose.yml provides a production-ready reference architecture with all performance optimizations pre-configured:
services:
# Nginx caching proxy (port 8080)
nginx:
image: mcpgateway/nginx-cache:latest
ports: ["8080:80"]
volumes:
- nginx_cache:/var/cache/nginx
# Brotli/Gzip compression, static caching, rate limiting
# Gateway application (replicas: 2)
gateway:
image: mcpgateway/mcpgateway:latest
environment:
# HTTP Server: gunicorn (stable) or granian (faster)
- HTTP_SERVER=gunicorn
- GUNICORN_WORKERS=16
# Database: via PgBouncer
- DATABASE_URL=postgresql+psycopg://postgres:password@pgbouncer:6432/mcp
- DB_POOL_SIZE=15 # Smaller with PgBouncer
- DB_MAX_OVERFLOW=30
# Redis with hiredis
- CACHE_TYPE=redis
- REDIS_URL=redis://redis:6379/0
- REDIS_PARSER=hiredis
- REDIS_MAX_CONNECTIONS=150
# Multi-level caching
- AUTH_CACHE_ENABLED=true
- REGISTRY_CACHE_ENABLED=true
- ADMIN_STATS_CACHE_ENABLED=true
# Performance: disable overhead
- LOG_LEVEL=ERROR
- DISABLE_ACCESS_LOG=true
- AUDIT_TRAIL_ENABLED=false
- COMPRESSION_ENABLED=false # Nginx handles this
deploy:
replicas: 2
resources:
limits: { cpus: '8', memory: 8G }
# PgBouncer connection pooler
pgbouncer:
image: edoburu/pgbouncer:latest
environment:
- DATABASE_URL=postgres://postgres:password@postgres:5432/mcp
- POOL_MODE=transaction
- MAX_CLIENT_CONN=3000
- DEFAULT_POOL_SIZE=120
- MAX_DB_CONNECTIONS=200
# PostgreSQL 18
postgres:
image: postgres:18
command:
- "postgres"
- "-c" - "max_connections=500"
- "-c" - "shared_buffers=512MB"
- "-c" - "effective_cache_size=1536MB"
- "-c" - "random_page_cost=1.1"
- "-c" - "effective_io_concurrency=200"
# Redis with performance tuning
redis:
image: redis:latest
command:
- "redis-server"
- "--maxmemory" - "1gb"
- "--maxmemory-policy" - "allkeys-lru"
- "--tcp-backlog" - "2048"
- "--maxclients" - "10000"
Access Points:
| Port | Service | Use |
|---|---|---|
| 8080 | Nginx | Production access (caching, compression) |
| 4444 | Gateway | Direct access (debugging, internal) |
| 6432 | PgBouncer | Database connection pooling |
| 5433 | PostgreSQL | Direct DB access (admin) |
| 6379 | Redis | Cache access |
Quick Start:
# Start full stack
docker-compose up -d
# Access via caching proxy
curl http://localhost:8080/health
# View logs
docker-compose logs -f gateway
5. Database Connection Pooling
Connection Pool Architecture
ContextForge uses a two-tier connection pooling architecture:
Without PgBouncer (direct connection):
+-------------------+ +-------------------+ +-------------------+
| Pod 1 (16 workers)| | Pod 2 (16 workers)| | Pod N (16 workers)|
| 16 SQLAlchemy | | 16 SQLAlchemy | | 16 SQLAlchemy |
| pools × 50 conns | | pools × 50 conns | | pools × 50 conns |
+--------+----------+ +--------+----------+ +--------+----------+
| | |
+------------+------------+------------+------------+
|
+------------v------------+
| PostgreSQL 18 |
| max_connections = 4000 |
+-------------------------+
With PgBouncer (recommended for high concurrency):
+-------------------+ +-------------------+ +-------------------+
| Pod 1 (16 workers)| | Pod 2 (16 workers)| | Pod N (16 workers)|
| 16 SQLAlchemy | | 16 SQLAlchemy | | 16 SQLAlchemy |
| pools × 15 conns | | pools × 15 conns | | pools × 15 conns |
+--------+----------+ +--------+----------+ +--------+----------+
| | |
+------------+------------+------------+------------+
|
+------------v------------+
| PgBouncer |
| MAX_CLIENT_CONN = 3000 | <-- Application connections
| DEFAULT_POOL_SIZE = 120 |
| MAX_DB_CONNECTIONS = 200| <-- PostgreSQL connections
+------------+------------+
|
+------------v------------+
| PostgreSQL 18 |
| max_connections = 500 | <-- Much lower requirement
+-------------------------+
Connection Multiplexing Benefits:
| Metric | Without PgBouncer | With PgBouncer | Improvement |
|---|---|---|---|
| App connections | 2 pods × 16 × 50 = 1,600 | 2 pods × 16 × 15 = 480 | App-level reduction |
| PostgreSQL connections | 1,600+ | 200 | 8x reduction |
max_connections needed |
4,000 | 500 | 8x lower |
| Memory per connection | ~10MB | ~10MB | Same |
| PostgreSQL memory | ~40GB | ~5GB | 8x reduction |
| Connection setup time | Per request | Reused | Near-zero |
Pool Configuration
Environment Variables
# Connection pool settings
DB_POOL_SIZE=50 # Persistent connections per worker
DB_MAX_OVERFLOW=10 # Additional connections allowed
DB_POOL_TIMEOUT=60 # Wait time before timeout (seconds)
DB_POOL_RECYCLE=3600 # Recycle connections after 1 hour
DB_MAX_RETRIES=30 # Retry attempts on failure (exponential backoff)
DB_RETRY_INTERVAL_MS=2000 # Base retry interval (doubles each attempt, max 30s)
# psycopg3-specific optimizations
DB_PREPARE_THRESHOLD=5 # Auto-prepare queries after N executions (0=disable)
psycopg3 Performance Features
ContextForge uses psycopg3 (psycopg[c,binary]) instead of psycopg2, providing significant performance improvements and modern features.
Why psycopg3:
| Feature | psycopg2 | psycopg3 |
|---|---|---|
| Parameter binding | Client-side | Server-side (more secure) |
| Prepared statements | Manual | Automatic after N executions |
| Binary protocol | Optional | Native support |
| Async support | Wrapper | First-class built-in |
| Maintenance | Maintenance mode | Active development |
Performance Benchmarks:
| Operation | psycopg2 | psycopg3 | Improvement |
|---|---|---|---|
| Bulk INSERT (1000+ rows) | Standard | COPY protocol | 5-10x |
| Repeated queries | Parsed each time | Auto-prepared | 2-3x |
| Batch queries | Sequential | Pipelined | 2-5x |
Automatic Prepared Statements:
# Number of executions before auto-preparing a query server-side
# Default: 5 (balance between memory and performance)
# Set to 0 to disable, 1 to prepare immediately
DB_PREPARE_THRESHOLD=5
After N executions of the same query pattern, psycopg3 creates a server-side prepared statement, reducing:
- Query parsing overhead (5-15% per query)
- Network round-trips for query plans
- PostgreSQL CPU usage for repeated queries
COPY Protocol for Bulk Inserts:
The utility module mcpgateway/utils/psycopg3_optimizations.py provides COPY protocol support:
from mcpgateway.utils.psycopg3_optimizations import bulk_insert_with_copy
# 5-10x faster for 1000+ rows
bulk_insert_with_copy(db, "tool_metrics", columns, rows)
Note: COPY is only faster for large batches (1000+ rows); for small batches (<100 rows), SQLAlchemy's bulk_insert_mappings() is actually faster due to lower protocol overhead.
Pipeline Mode for Batch Queries:
Execute multiple queries with reduced round-trips:
from mcpgateway.utils.psycopg3_optimizations import execute_pipelined
# Send multiple queries without waiting for individual responses
results = execute_pipelined(db, [
("SELECT * FROM tools WHERE id = %s", {"id": tool_id}),
("SELECT * FROM gateways WHERE id = %s", {"id": gateway_id}),
])
Connection URL Format:
# IMPORTANT: Required format for psycopg3 (NOT postgresql://)
DATABASE_URL=postgresql+psycopg://user:pass@host:5432/db
See ADR-027: Migrate to psycopg3 for details.
Configuration in Code
# mcpgateway/config.py
@property
def database_settings(self) -> dict:
return {
"pool_size": self.db_pool_size, # 50
"max_overflow": self.db_max_overflow, # 10
"pool_timeout": self.db_pool_timeout, # 60s
"pool_recycle": self.db_pool_recycle, # 3600s
}
PostgreSQL Configuration
Calculate max_connections
# Formula
max_connections = (num_pods × num_workers × pool_size × 1.2) + buffer
# Example: 10 pods, 8 workers, 50 pool size
max_connections = (10 × 8 × 50 × 1.2) + 200 = 5000 connections
PostgreSQL Configuration File
# postgresql.conf
max_connections = 5000
shared_buffers = 16GB # 25% of RAM
effective_cache_size = 48GB # 75% of RAM
work_mem = 16MB # Per operation
maintenance_work_mem = 2GB
Managed Services
IBM Cloud Databases for PostgreSQL:
# Increase max_connections via CLI
ibmcloud cdb deployment-configuration postgres \
--configuration max_connections=5000
AWS RDS:
# Via parameter group
max_connections = {DBInstanceClassMemory/9531392}
Google Cloud SQL:
# Auto-scales based on instance size
# 4 vCPU = 400 connections
# 8 vCPU = 800 connections
PgBouncer Connection Pooling
For high-concurrency deployments, PgBouncer provides connection multiplexing between the gateway and PostgreSQL, dramatically reducing connection overhead.
Architecture:
Without PgBouncer:
Gateway (2 replicas × 16 workers) → PostgreSQL (max_connections=4000)
Each worker maintains its own pool → High connection churn
With PgBouncer:
Gateway (2 replicas × 16 workers) → PgBouncer → PostgreSQL (max_connections=500)
Connections multiplexed → Efficient reuse, lower PostgreSQL overhead
Benefits:
- Connection multiplexing: Many app connections share fewer database connections
- Reduced PostgreSQL overhead: Lower
max_connectionsreduces memory per connection - Connection reuse: PgBouncer maintains persistent connections to PostgreSQL
- Graceful handling of connection storms: Queues requests instead of rejecting
Docker Compose Setup:
pgbouncer:
image: edoburu/pgbouncer:latest
restart: unless-stopped
ports:
- "6432:6432"
environment:
- DATABASE_URL=postgres://postgres:password@postgres:5432/mcp
- POOL_MODE=transaction
- MAX_CLIENT_CONN=2000
- DEFAULT_POOL_SIZE=100
- MIN_POOL_SIZE=10
- RESERVE_POOL_SIZE=25
- MAX_DB_CONNECTIONS=200
- SERVER_LIFETIME=3600
- SERVER_IDLE_TIMEOUT=600
- AUTH_TYPE=scram-sha-256
depends_on:
postgres:
condition: service_healthy
healthcheck:
test: ["CMD", "pg_isready", "-h", "localhost", "-p", "6432"]
interval: 10s
timeout: 5s
retries: 3
Gateway Configuration with PgBouncer:
# Connect via PgBouncer instead of direct PostgreSQL
DATABASE_URL=postgresql+psycopg://postgres:password@pgbouncer:6432/mcp
# Smaller pool since PgBouncer handles pooling
DB_POOL_SIZE=10
DB_MAX_OVERFLOW=20
Pool Modes:
| Mode | Description | Use Case |
|---|---|---|
transaction |
Connection returned after transaction commit | Recommended for most workloads |
session |
Connection held for entire session | Legacy apps requiring session state |
statement |
Connection returned after each statement | Limited use cases |
Key Tuning Parameters:
| Parameter | Description | Suggested Value |
|---|---|---|
MAX_CLIENT_CONN |
Max app connections | 2000-5000 |
DEFAULT_POOL_SIZE |
Connections per user/db pair | 100 |
MAX_DB_CONNECTIONS |
Max connections to PostgreSQL | 200-500 |
POOL_MODE |
When to return connections | transaction |
Monitoring PgBouncer:
# Connect to PgBouncer admin console
psql -h localhost -p 6432 -U pgbouncer pgbouncer
# Show connection pool statistics
SHOW STATS;
SHOW POOLS;
SHOW CLIENTS;
SHOW SERVERS;
Connection Pool Monitoring
# Health endpoint checks pool status
@app.get("/health")
async def healthcheck(db: Session = Depends(get_db)):
try:
db.execute(text("SELECT 1"))
return {"status": "healthy"}
except Exception as e:
return {"status": "unhealthy", "error": str(e)}
# Check PostgreSQL connections
kubectl exec -it postgres-pod -- psql -U admin -d postgresdb \
-c "SELECT count(*) FROM pg_stat_activity;"
Database Session Management
To prevent connection pool exhaustion under high load, ContextForge releases database sessions before making upstream HTTP calls:
Problem: Database sessions held during slow upstream calls (100ms - 4+ minutes) exhaust the connection pool even when the database is lightly loaded.
Solution: "Fetch-Then-Release" pattern:
- Eager load required data with
joinedload()in single query - Extract data to local variables before network I/O
- Release session with
db.expunge()+db.close()before HTTP calls - Fresh session for metrics recording after the call
This reduces connection hold time from minutes to <50ms and enables 10x+ higher concurrency.
6. Redis for Distributed Caching
Architecture
Redis provides shared state across all Gateway pods:
- Session storage: User sessions (TTL: 3600s)
- Message cache: Ephemeral data (TTL: 600s)
- Federation cache: Gateway peer discovery
Configuration
Enable Redis Caching
# .env or Kubernetes ConfigMap
CACHE_TYPE=redis
REDIS_URL=redis://redis-service:6379/0
CACHE_PREFIX=mcpgw:
SESSION_TTL=3600
MESSAGE_TTL=600
REDIS_MAX_RETRIES=30 # Retry attempts on failure (exponential backoff)
REDIS_RETRY_INTERVAL_MS=2000 # Base retry interval (doubles each attempt, max 30s)
# Connection pool (standard)
REDIS_MAX_CONNECTIONS=50
REDIS_SOCKET_TIMEOUT=2.0
REDIS_SOCKET_CONNECT_TIMEOUT=2.0
REDIS_RETRY_ON_TIMEOUT=true
REDIS_HEALTH_CHECK_INTERVAL=30
# Leader election (multi-node)
REDIS_LEADER_TTL=15
REDIS_LEADER_HEARTBEAT_INTERVAL=5
High-Concurrency Redis Tuning
For 1000+ concurrent users:
# Connection pool - Formula: (concurrent_requests / workers) * 1.5
# Example: 32 workers × 150 = 4800 < Redis maxclients (10000)
REDIS_MAX_CONNECTIONS=150
# Timeouts - keep low for fast failure detection
REDIS_SOCKET_TIMEOUT=5.0
REDIS_SOCKET_CONNECT_TIMEOUT=5.0
REDIS_HEALTH_CHECK_INTERVAL=30
Redis Server Tuning (docker-compose.yml):
redis:
command:
- "redis-server"
- "--maxmemory"
- "1gb"
- "--maxmemory-policy"
- "allkeys-lru"
- "--tcp-backlog"
- "2048" # Higher for pending connections
- "--maxclients"
- "10000" # Max client connections
Kubernetes Deployment
# charts/mcp-stack/values.yaml
redis:
enabled: true
resources:
limits:
cpu: 2000m
memory: 4Gi
requests:
cpu: 1000m
memory: 2Gi
# Enable persistence
persistence:
enabled: true
size: 10Gi
Redis Sizing
Memory calculation:
- Sessions:
concurrent_users × 50KB - Messages:
messages_per_minute × 100KB × (TTL/60)
Example:
- 10,000 users × 50KB = 500MB
- 1,000 msg/min × 100KB × 10min = 1GB
- Total: 1.5GB + 50% overhead = 2.5GB
Connection pool sizing:
- Formula:
REDIS_MAX_CONNECTIONS = (concurrent_requests / workers) × 1.5 - Default 50 handles ~500 concurrent requests with 10 workers
- High-concurrency: increase to 100 and lower timeouts
# High-concurrency production overrides
REDIS_MAX_CONNECTIONS=100
REDIS_SOCKET_TIMEOUT=1.0
REDIS_SOCKET_CONNECT_TIMEOUT=1.0
REDIS_HEALTH_CHECK_INTERVAL=15
REDIS_LEADER_TTL=10
REDIS_LEADER_HEARTBEAT_INTERVAL=3
Hiredis High-Performance Parser
ContextForge uses redis[hiredis] for significantly faster Redis protocol parsing, especially beneficial for large responses.
Performance Impact:
| Operation | Pure Python | With Hiredis | Improvement |
|---|---|---|---|
| Simple SET/GET | Baseline | +10% | 1.1x |
| LRANGE (10 items) | Baseline | +170% | 2.7x |
| LRANGE (100 items) | Baseline | +1000% | ~10x |
| LRANGE (999 items) | Baseline | +8220% | 83x |
The larger the response, the greater the improvement. This is critical for:
- Tool registry queries returning many tools
- Bulk operations and federation
- Cached response retrieval
- Metrics aggregation
Configuration:
# Redis parser selection (hiredis is default for performance)
# Options: auto (default - uses hiredis if available), hiredis, python
# Use "python" to force pure-Python parser (useful for debugging)
REDIS_PARSER=auto
When to use each parser:
| Parser | Use Case |
|---|---|
hiredis (default) |
Production, high throughput |
python |
Debugging Redis protocol issues, restricted environments |
High Availability
Redis Sentinel (3+ nodes):
redis:
sentinel:
enabled: true
quorum: 2
replicas: 3 # 1 primary + 2 replicas
Redis Cluster (6+ nodes):
REDIS_URL=redis://redis-cluster:6379/0?cluster=true
7. Performance Tuning
Application Architecture Performance
ContextForge's technology stack is optimized for high performance:
Rust-Powered Components:
- Pydantic v2 (5-50x faster validation via Rust core)
- Uvicorn with [standard] extras (ASGI server with high-performance components)
Async-First Design:
- FastAPI (async request handling)
- SQLAlchemy 2.0 (async database operations)
- asyncio event loop per worker
Performance characteristics:
- Request validation: < 1ms (Pydantic v2 Rust core)
- JSON serialization: 3-5x faster than pure Python
- Database queries: Non-blocking async I/O
- Concurrent requests per worker: 1000+ (async event loop)
System-Level Optimization
Kernel Parameters
# /etc/sysctl.conf
net.core.somaxconn=4096
net.ipv4.tcp_max_syn_backlog=4096
net.ipv4.ip_local_port_range=1024 65535
net.ipv4.tcp_tw_reuse=1
fs.file-max=2097152
# Apply changes
sysctl -p
File Descriptors
# /etc/security/limits.conf
* soft nofile 1048576
* hard nofile 1048576
# Verify
ulimit -n
HTTP Server Selection
ContextForge supports two production HTTP servers:
| Server | Description | Best For |
|---|---|---|
| Gunicorn (default) | Python-based with Uvicorn workers | Stable, well-tested |
| Granian | Rust-based HTTP server | Maximum performance (+20-50%) |
# Select HTTP server (in containers)
HTTP_SERVER=gunicorn # Default, stable
HTTP_SERVER=granian # Alternative, Rust-based
Gunicorn Configuration
# Number of worker processes
# Options: "auto" (default, 2*CPU+1 capped at 16), or positive integer
GUNICORN_WORKERS=auto
# Worker timeout in seconds (increase for long-running LLM requests)
GUNICORN_TIMEOUT=600
# Maximum requests per worker before automatic restart (prevents memory leaks)
GUNICORN_MAX_REQUESTS=100000
# Random jitter added to max requests (prevents thundering herd)
GUNICORN_MAX_REQUESTS_JITTER=100
# Preload application before forking workers
# true: Saves memory (shared code), runs migrations once before forking
# false: Each worker loads app independently (more memory, better isolation)
GUNICORN_PRELOAD_APP=true
# Development mode with hot reload (NOT for production)
GUNICORN_DEV_MODE=false
Worker Class: UvicornWorker (default) for async support.
Granian Configuration (Alternative)
Granian is a Rust-based HTTP server with native backpressure for overload protection:
# HTTP server selection
HTTP_SERVER=granian
# Worker count (auto = CPU count, max 16)
GRANIAN_WORKERS=16
# TCP backlog for pending connections
GRANIAN_BACKLOG=4096
# Backpressure: max concurrent requests per worker before 503 rejection
# Total capacity = WORKERS × BACKPRESSURE = 16 × 64 = 1024 concurrent requests
GRANIAN_BACKPRESSURE=64
# HTTP/1.1 buffer size (bytes)
GRANIAN_HTTP1_BUFFER_SIZE=524288
# Auto-restart failed workers
GRANIAN_RESPAWN_FAILED=true
Backpressure behavior:
- Requests within capacity (≤1024): Processed normally
- Requests over capacity (>1024): Immediate 503 Service Unavailable
- No queuing, no memory growth, no cascading timeouts
When to use Granian:
- Load spike protection (backpressure rejects excess gracefully)
- Bursty or unpredictable traffic patterns
- High-concurrency deployments (1000+ concurrent users)
When to use Gunicorn:
- Memory-constrained environments (32% less RAM)
- Maximum stability and compatibility
- Standard deployments with predictable traffic
Application Tuning
# Resource limits
TOOL_TIMEOUT=60
TOOL_CONCURRENT_LIMIT=10
RESOURCE_CACHE_SIZE=1000
RESOURCE_CACHE_TTL=3600
# Retry configuration
RETRY_MAX_ATTEMPTS=3
RETRY_BASE_DELAY=1.0
RETRY_MAX_DELAY=60
# Health check intervals
HEALTH_CHECK_INTERVAL=60
HEALTH_CHECK_TIMEOUT=5
UNHEALTHY_THRESHOLD=3
# Gateway health check timeout (seconds)
GATEWAY_HEALTH_CHECK_TIMEOUT=5.0
# Auto-refresh tools during health checks
# When enabled, tools/resources/prompts are fetched and synced during health checks
AUTO_REFRESH_SERVERS=false
Logging for Performance
Logging can significantly impact performance under high load:
# Log level - ERROR recommended for production
# DEBUG/INFO create massive I/O overhead
LOG_LEVEL=ERROR
# Disable access logging (massive I/O overhead under high concurrency)
DISABLE_ACCESS_LOG=true
# Disable database logging for performance
STRUCTURED_LOGGING_DATABASE_ENABLED=false
Impact of logging settings:
| Setting | I/O Overhead | Use Case |
|---|---|---|
LOG_LEVEL=DEBUG |
Very High | Development only |
LOG_LEVEL=INFO |
High | Light load, debugging |
LOG_LEVEL=ERROR |
Low | Production (recommended) |
DISABLE_ACCESS_LOG=true |
None | Production (recommended) |
STRUCTURED_LOGGING_DATABASE_ENABLED=true |
Very High | Compliance (use external aggregator) |
Metrics Buffer Configuration
Bat
…(truncated)