MLOps Production Patterns
Production ML infrastructure patterns for model deployment, monitoring, and lifecycle management.
Table of Contents
Model Deployment Pipeline
Deployment Workflow
- Export trained model to standardized format (ONNX, TorchScript, SavedModel)
- Package model with dependencies in Docker container
- Deploy to staging environment
- Run integration tests against staging
- Deploy canary (5% traffic) to production
- Monitor latency and error rates for 1 hour
- Promote to full production if metrics pass
- Validation: p95 latency < 100ms, error rate < 0.1%
Container Structure
FROM python:3.11-slim
# Install dependencies
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
# Copy model artifacts
COPY model/ /app/model/
COPY src/ /app/src/
# Health check endpoint
HEALTHCHECK CMD curl -f http://localhost:8080/health || exit 1
EXPOSE 8080
CMD ["uvicorn", "src.server:app", "--host", "0.0.0.0", "--port", "8080"]
Model Serving Options
| Option |
Latency |
Throughput |
Use Case |
| FastAPI + Uvicorn |
Low |
Medium |
REST APIs, small models |
| Triton Inference Server |
Very Low |
Very High |
GPU inference, batching |
| TensorFlow Serving |
Low |
High |
TensorFlow models |
| TorchServe |
Low |
High |
PyTorch models |
| Ray Serve |
Medium |
High |
Complex pipelines, multi-model |
Kubernetes Deployment
apiVersion: apps/v1
kind: Deployment
metadata:
name: model-serving
spec:
replicas: 3
selector:
matchLabels:
app: model-serving
template:
spec:
containers:
- name: model
image: model:v1.0.0
resources:
requests:
memory: "2Gi"
cpu: "1"
limits:
memory: "4Gi"
cpu: "2"
readinessProbe:
httpGet:
path: /health
port: 8080
initialDelaySeconds: 10
periodSeconds: 5
Feature Store Architecture
Feature Store Components
| Component |
Purpose |
Tools |
| Offline Store |
Training data, batch features |
BigQuery, Snowflake, S3 |
| Online Store |
Low-latency serving |
Redis, DynamoDB, Feast |
| Feature Registry |
Metadata, lineage |
Feast, Tecton, Hopsworks |
| Transformation |
Feature engineering |
Spark, Flink, dbt |
Feature Pipeline Workflow
- Define feature schema in registry
- Implement transformation logic (SQL or Python)
- Backfill historical features to offline store
- Schedule incremental updates
- Materialize to online store for serving
- Monitor feature freshness and quality
- Validation: Feature values within expected ranges, no nulls in required fields
Feature Definition Example
from feast import Entity, Feature, FeatureView, FileSource
user = Entity(name="user_id", value_type=ValueType.INT64)
user_features = FeatureView(
name="user_features",
entities=["user_id"],
ttl=timedelta(days=1),
features=[
Feature(name="purchase_count_30d", dtype=ValueType.INT64),
Feature(name="avg_order_value", dtype=ValueType.FLOAT),
Feature(name="days_since_last_purchase", dtype=ValueType.INT64),
],
source=FileSource(path="data/user_features.parquet"),
)
Model Monitoring
Monitoring Dimensions
| Dimension |
Metrics |
Alert Threshold |
| Latency |
p50, p95, p99 |
p95 > 100ms |
| Throughput |
requests/sec |
< 80% baseline |
| Errors |
error rate, 5xx count |
> 0.1% |
| Data Drift |
PSI, KS statistic |
PSI > 0.2 |
| Model Drift |
accuracy, AUC decay |
> 5% drop |
Data Drift Detection
from scipy.stats import ks_2samp
import numpy as np
def detect_drift(reference: np.array, current: np.array, threshold: float = 0.05):
"""Detect distribution drift using Kolmogorov-Smirnov test."""
statistic, p_value = ks_2samp(reference, current)
drift_detected = p_value < threshold
return {
"drift_detected": drift_detected,
"ks_statistic": statistic,
"p_value": p_value,
"threshold": threshold
}
Monitoring Dashboard Metrics
Infrastructure:
- Request latency (p50, p95, p99)
- Requests per second
- Error rate by type
- CPU/memory utilization
- GPU utilization (if applicable)
Model Performance:
- Prediction distribution
- Feature value distributions
- Model output confidence
- Ground truth vs predictions (when available)
A/B Testing Infrastructure
Experiment Workflow
- Define experiment hypothesis and success metrics
- Calculate required sample size for statistical power
- Configure traffic split (control vs treatment)
- Deploy treatment model alongside control
- Route traffic based on user/session hash
- Collect metrics for both variants
- Run statistical significance test
- Validation: p-value < 0.05, minimum sample size reached
Traffic Splitting
import hashlib
def get_variant(user_id: str, experiment: str, control_pct: float = 0.5) -> str:
"""Deterministic traffic splitting based on user ID."""
hash_input = f"{user_id}:{experiment}"
hash_value = int(hashlib.md5(hash_input.encode()).hexdigest(), 16)
bucket = (hash_value % 100) / 100.0
return "control" if bucket < control_pct else "treatment"
Metrics Collection
| Metric Type |
Examples |
Collection Method |
| Primary |
Conversion rate, revenue |
Event logging |
| Secondary |
Latency, engagement |
Request logs |
| Guardrail |
Error rate, crashes |
Monitoring system |
Automated Retraining
Retraining Triggers
| Trigger |
Detection Method |
Action |
| Scheduled |
Cron (weekly/monthly) |
Full retrain |
| Performance drop |
Accuracy < threshold |
Immediate retrain |
| Data drift |
PSI > 0.2 |
Evaluate, then retrain |
| New data volume |
X new samples |
Incremental update |
Retraining Pipeline
- Trigger detection (schedule, drift, performance)
- Fetch latest training data from feature store
- Run training job with hyperparameter config
- Evaluate model on holdout set
- Compare against production model
- If improved: register new model version
- Deploy to staging for validation
- Promote to production via canary
- Validation: New model outperforms baseline on key metrics
MLflow Model Registry Integration
import mlflow
def register_model(model, metrics: dict, model_name: str):
"""Register trained model with MLflow."""
with mlflow.start_run():
# Log metrics
for name, value in metrics.items():
mlflow.log_metric(name, value)
# Log model
mlflow.sklearn.log_model(model, "model")
# Register in model registry
model_uri = f"runs:/{mlflow.active_run().info.run_id}/model"
mlflow.register_model(model_uri, model_name)
1---2name: mlops-production-patterns3description: Production ML infrastructure patterns for model deployment, monitoring, and lifecycle management.4---5# MLOps Production Patterns67Production ML infrastructure patterns for model deployment, monitoring, and lifecycle management.89---1011## Table of Contents1213- [Model Deployment Pipeline](#model-deployment-pipeline)14- [Feature Store Architecture](#feature-store-architecture)15- [Model Monitoring](#model-monitoring)16- [A/B Testing Infrastructure](#ab-testing-infrastructure)17- [Automated Retraining](#automated-retraining)1819---2021## Model Deployment Pipeline2223### Deployment Workflow24251. Export trained model to standardized format (ONNX, TorchScript, SavedModel)262. Package model with dependencies in Docker container273. Deploy to staging environment284. Run integration tests against staging295. Deploy canary (5% traffic) to production306. Monitor latency and error rates for 1 hour317. Promote to full production if metrics pass328. **Validation:** p95 latency < 100ms, error rate < 0.1%3334### Container Structure3536```dockerfile37FROM python:3.11-slim3839# Install dependencies40COPY requirements.txt .41RUN pip install --no-cache-dir -r requirements.txt4243# Copy model artifacts44COPY model/ /app/model/45COPY src/ /app/src/4647# Health check endpoint48HEALTHCHECK CMD curl -f http://localhost:8080/health || exit 14950EXPOSE 808051CMD ["uvicorn", "src.server:app", "--host", "0.0.0.0", "--port", "8080"]52```5354### Model Serving Options5556| Option | Latency | Throughput | Use Case |57|--------|---------|------------|----------|58| FastAPI + Uvicorn | Low | Medium | REST APIs, small models |59| Triton Inference Server | Very Low | Very High | GPU inference, batching |60| TensorFlow Serving | Low | High | TensorFlow models |61| TorchServe | Low | High | PyTorch models |62| Ray Serve | Medium | High | Complex pipelines, multi-model |6364### Kubernetes Deployment6566```yaml67apiVersion: apps/v168kind: Deployment69metadata:70 name: model-serving71spec:72 replicas: 373 selector:74 matchLabels:75 app: model-serving76 template:77 spec:78 containers:79 - name: model80 image: model:v1.0.081 resources:82 requests:83 memory: "2Gi"84 cpu: "1"85 limits:86 memory: "4Gi"87 cpu: "2"88 readinessProbe:89 httpGet:90 path: /health91 port: 808092 initialDelaySeconds: 1093 periodSeconds: 594```9596---9798## Feature Store Architecture99100### Feature Store Components101102| Component | Purpose | Tools |103|-----------|---------|-------|104| Offline Store | Training data, batch features | BigQuery, Snowflake, S3 |105| Online Store | Low-latency serving | Redis, DynamoDB, Feast |106| Feature Registry | Metadata, lineage | Feast, Tecton, Hopsworks |107| Transformation | Feature engineering | Spark, Flink, dbt |108109### Feature Pipeline Workflow1101111. Define feature schema in registry1122. Implement transformation logic (SQL or Python)1133. Backfill historical features to offline store1144. Schedule incremental updates1155. Materialize to online store for serving1166. Monitor feature freshness and quality1177. **Validation:** Feature values within expected ranges, no nulls in required fields118119### Feature Definition Example120121```python122from feast import Entity, Feature, FeatureView, FileSource123124user = Entity(name="user_id", value_type=ValueType.INT64)125126user_features = FeatureView(127 name="user_features",128 entities=["user_id"],129 ttl=timedelta(days=1),130 features=[131 Feature(name="purchase_count_30d", dtype=ValueType.INT64),132 Feature(name="avg_order_value", dtype=ValueType.FLOAT),133 Feature(name="days_since_last_purchase", dtype=ValueType.INT64),134 ],135 online=True,136 source=FileSource(path="data/user_features.parquet"),137)138```139140---141142## Model Monitoring143144### Monitoring Dimensions145146| Dimension | Metrics | Alert Threshold |147|-----------|---------|-----------------|148| Latency | p50, p95, p99 | p95 > 100ms |149| Throughput | requests/sec | < 80% baseline |150| Errors | error rate, 5xx count | > 0.1% |151| Data Drift | PSI, KS statistic | PSI > 0.2 |152| Model Drift | accuracy, AUC decay | > 5% drop |153154### Data Drift Detection155156```python157from scipy.stats import ks_2samp158import numpy as np159160def detect_drift(reference: np.array, current: np.array, threshold: float = 0.05):161 """Detect distribution drift using Kolmogorov-Smirnov test."""162 statistic, p_value = ks_2samp(reference, current)163164 drift_detected = p_value < threshold165166 return {167 "drift_detected": drift_detected,168 "ks_statistic": statistic,169 "p_value": p_value,170 "threshold": threshold171 }172```173174### Monitoring Dashboard Metrics175176**Infrastructure:**177- Request latency (p50, p95, p99)178- Requests per second179- Error rate by type180- CPU/memory utilization181- GPU utilization (if applicable)182183**Model Performance:**184- Prediction distribution185- Feature value distributions186- Model output confidence187- Ground truth vs predictions (when available)188189---190191## A/B Testing Infrastructure192193### Experiment Workflow1941951. Define experiment hypothesis and success metrics1962. Calculate required sample size for statistical power1973. Configure traffic split (control vs treatment)1984. Deploy treatment model alongside control1995. Route traffic based on user/session hash2006. Collect metrics for both variants2017. Run statistical significance test2028. **Validation:** p-value < 0.05, minimum sample size reached203204### Traffic Splitting205206```python207import hashlib208209def get_variant(user_id: str, experiment: str, control_pct: float = 0.5) -> str:210 """Deterministic traffic splitting based on user ID."""211 hash_input = f"{user_id}:{experiment}"212 hash_value = int(hashlib.md5(hash_input.encode()).hexdigest(), 16)213 bucket = (hash_value % 100) / 100.0214215 return "control" if bucket < control_pct else "treatment"216```217218### Metrics Collection219220| Metric Type | Examples | Collection Method |221|-------------|----------|-------------------|222| Primary | Conversion rate, revenue | Event logging |223| Secondary | Latency, engagement | Request logs |224| Guardrail | Error rate, crashes | Monitoring system |225226---227228## Automated Retraining229230### Retraining Triggers231232| Trigger | Detection Method | Action |233|---------|------------------|--------|234| Scheduled | Cron (weekly/monthly) | Full retrain |235| Performance drop | Accuracy < threshold | Immediate retrain |236| Data drift | PSI > 0.2 | Evaluate, then retrain |237| New data volume | X new samples | Incremental update |238239### Retraining Pipeline2402411. Trigger detection (schedule, drift, performance)2422. Fetch latest training data from feature store2433. Run training job with hyperparameter config2444. Evaluate model on holdout set2455. Compare against production model2466. If improved: register new model version2477. Deploy to staging for validation2488. Promote to production via canary2499. **Validation:** New model outperforms baseline on key metrics250251### MLflow Model Registry Integration252253```python254import mlflow255256def register_model(model, metrics: dict, model_name: str):257 """Register trained model with MLflow."""258 with mlflow.start_run():259 # Log metrics260 for name, value in metrics.items():261 mlflow.log_metric(name, value)262263 # Log model264 mlflow.sklearn.log_model(model, "model")265266 # Register in model registry267 model_uri = f"runs:/{mlflow.active_run().info.run_id}/model"268 mlflow.register_model(model_uri, model_name)269```