LLMOps Platform Engineering
Design and operate an internal LLM platform that supports rapid experimentation without compromising reliability, cost, or compliance.
When to Use This Skill
- Building an internal platform for teams to deploy and manage LLM-powered features
- Designing CI/CD pipelines that include model evaluation gates
- Setting up A/B testing infrastructure for model versions
- Creating Kubernetes-based model serving infrastructure
- Establishing governance workflows for model promotion
Prerequisites
- Kubernetes cluster with GPU node pools (or cloud inference API access)
- Container registry (Harbor, ECR, GCR, or ACR)
- CI/CD system (GitHub Actions, GitLab CI, or Argo Workflows)
- Observability stack (Prometheus + Grafana + OpenTelemetry)
- Model registry (MLflow or custom metadata store)
Outcomes
- Standardized path from experiment to production
- Safe model rollout with quality and safety gates
- Repeatable infra modules for inference, vector DB, and observability
- Clear ownership model across platform, app, and security teams
Reference Architecture
- Control Plane: model registry, prompt/version catalog, policy checks, eval pipeline.
- Data Plane: inference gateway, vector database, cache, feature store.
- Ops Plane: telemetry, alerting, SLO dashboards, cost analytics.
- Security Plane: IAM boundaries, secret rotation, content filters, audit logs.
Model Promotion Pipeline
# .github/workflows/model-promotion.yaml
name: Model Promotion Pipeline
on:
workflow_dispatch:
inputs:
model_name:
description: "Model identifier"
required: true
model_version:
description: "Model version to promote"
required: true
target_env:
description: "Target environment"
required: true
type: choice
options: [staging, production]
jobs:
evaluate:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Run quality evaluation suite
run: |
python -m evals.run \
--model "${{ inputs.model_name }}:${{ inputs.model_version }}" \
--suite quality \
--output results/quality.json
- name: Run safety evaluation suite
run: |
python -m evals.run \
--model "${{ inputs.model_name }}:${{ inputs.model_version }}" \
--suite safety \
--output results/safety.json
- name: Run latency benchmark
run: |
python -m evals.benchmark \
--model "${{ inputs.model_name }}:${{ inputs.model_version }}" \
--concurrent-users 50 \
--duration 300 \
--output results/latency.json
- name: Gate check - quality
run: |
python -m evals.gate_check \
--results results/quality.json \
--threshold-file thresholds/quality.yaml
- name: Gate check - safety
run: |
python -m evals.gate_check \
--results results/safety.json \
--threshold-file thresholds/safety.yaml
- name: Gate check - latency
run: |
python -m evals.gate_check \
--results results/latency.json \
--threshold-file thresholds/latency.yaml
- name: Upload eval evidence
uses: actions/upload-artifact@v4
with:
name: eval-results-${{ inputs.model_version }}
path: results/
approve:
needs: evaluate
runs-on: ubuntu-latest
environment: ${{ inputs.target_env }}
steps:
- name: Record approval
run: |
echo "Approved by: ${{ github.actor }}"
echo "Model: ${{ inputs.model_name }}:${{ inputs.model_version }}"
echo "Target: ${{ inputs.target_env }}"
echo "Time: $(date -u +%Y-%m-%dT%H:%M:%SZ)"
deploy:
needs: approve
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Deploy canary
run: |
kubectl set image deployment/${{ inputs.model_name }}-canary \
model=${{ inputs.model_name }}:${{ inputs.model_version }} \
-n ai-${{ inputs.target_env }}
- name: Wait for canary validation (15 min)
run: |
python -m canary.validate \
--deployment ${{ inputs.model_name }}-canary \
--namespace ai-${{ inputs.target_env }} \
--duration 900 \
--quality-threshold 0.85 \
--error-rate-threshold 0.02
- name: Promote to full rollout
run: |
kubectl set image deployment/${{ inputs.model_name }} \
model=${{ inputs.model_name }}:${{ inputs.model_version }} \
-n ai-${{ inputs.target_env }}
kubectl rollout status deployment/${{ inputs.model_name }} \
-n ai-${{ inputs.target_env }} --timeout=300s
Evaluation Gate Thresholds
# thresholds/quality.yaml
gates:
groundedness:
metric: groundedness_score
min: 0.85
comparison: gte
task_success:
metric: task_success_rate
min: 0.90
comparison: gte
hallucination:
metric: hallucination_rate
max: 0.08
comparison: lte
regression:
metric: quality_delta_vs_baseline
min: -0.02
comparison: gte
description: "Must not regress more than 2% vs current production"
# thresholds/latency.yaml
gates:
p50_latency:
metric: latency_p50_ms
max: 800
comparison: lte
p95_latency:
metric: latency_p95_ms
max: 2000
comparison: lte
p99_latency:
metric: latency_p99_ms
max: 5000
comparison: lte
throughput:
metric: requests_per_second
min: 50
comparison: gte
A/B Testing Configuration
# ab-test-config.yaml
apiVersion: gateway.ai/v1
kind: ABTest
metadata:
name: model-comparison-q1
namespace: ai-production
spec:
duration: 7d
traffic_split:
control:
model: gpt-4o-2024-08-06
weight: 70
treatment:
model: gpt-4o-2025-01-15
weight: 30
metrics:
primary:
- task_success_rate
- user_satisfaction_score
secondary:
- latency_p95
- cost_per_request
- hallucination_rate
guardrails:
auto_rollback_if:
- metric: task_success_rate
threshold: 0.80
window: 1h
- metric: hallucination_rate
threshold: 0.15
window: 30m
assignment:
strategy: sticky_user
hash_key: user_id
Kubernetes Model Serving Deployment
# model-serving-deployment.yaml
apiVersion: apps/v1
kind: Deployment
metadata:
name: llm-inference
namespace: ai-production
labels:
app: llm-inference
model: gpt-4o
version: "2025-01"
spec:
replicas: 3
strategy:
type: RollingUpdate
rollingUpdate:
maxSurge: 1
maxUnavailable: 0
selector:
matchLabels:
app: llm-inference
template:
metadata:
labels:
app: llm-inference
model: gpt-4o
annotations:
prometheus.io/scrape: "true"
prometheus.io/port: "8080"
prometheus.io/path: "/metrics"
spec:
topologySpreadConstraints:
- maxSkew: 1
topologyKey: topology.kubernetes.io/zone
whenUnsatisfiable: DoNotSchedule
labelSelector:
matchLabels:
app: llm-inference
containers:
- name: model
image: registry.internal/vllm-server:0.4.1
args:
- "--model=/models/current"
- "--tensor-parallel-size=1"
- "--max-model-len=8192"
- "--gpu-memory-utilization=0.90"
ports:
- containerPort: 8000
name: inference
- containerPort: 8080
name: metrics
resources:
requests:
cpu: "4"
memory: "16Gi"
nvidia.com/gpu: "1"
limits:
cpu: "8"
memory: "32Gi"
nvidia.com/gpu: "1"
readinessProbe:
httpGet:
path: /health
port: 8000
initialDelaySeconds: 60
periodSeconds: 10
livenessProbe:
httpGet:
path: /health
port: 8000
initialDelaySeconds: 120
periodSeconds: 30
volumeMounts:
- name: model-weights
mountPath: /models
readOnly: true
- name: config
mountPath: /etc/vllm
volumes:
- name: model-weights
persistentVolumeClaim:
claimName: model-weights-pvc
- name: config
configMap:
name: vllm-config
tolerations:
- key: nvidia.com/gpu
operator: Exists
effect: NoSchedule
nodeSelector:
gpu-type: a100
---
apiVersion: v1
kind: Service
metadata:
name: llm-inference
namespace: ai-production
spec:
selector:
app: llm-inference
ports:
- name: inference
port: 8000
targetPort: 8000
- name: metrics
port: 8080
targetPort: 8080
---
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: llm-inference-hpa
namespace: ai-production
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: llm-inference
minReplicas: 2
maxReplicas: 10
metrics:
- type: Pods
pods:
metric:
name: llm_queue_depth
target:
type: AverageValue
averageValue: "5"
- type: Pods
pods:
metric:
name: gpu_utilization_percent
target:
type: AverageValue
averageValue: "75"
behavior:
scaleUp:
stabilizationWindowSeconds: 60
policies:
- type: Pods
value: 2
periodSeconds: 120
scaleDown:
stabilizationWindowSeconds: 300
policies:
- type: Pods
value: 1
periodSeconds: 300
CI/CD Design for AI Services
- Build immutable containers with pinned dependencies and model hashes.
- Use environment promotion:
dev -> stage -> prod.
- Fail deployment if:
- regression evals drop below baseline,
- safety tests exceed risk threshold,
- p95 latency exceeds SLO budget.
- Store deployment evidence for audits (commit SHA, eval report, approver).
Operational SLOs
| Signal |
Target |
Measurement Window |
| Availability |
99.9% |
30-day rolling |
| p95 Latency |
< 1200ms |
5-min buckets |
| Cost per request |
< $0.05 |
1-hour average |
| Task success rate |
> 90% |
24-hour rolling |
| Groundedness |
> 85% |
24-hour rolling |
Platform Guardrails
- Enforce tenant quotas and model allow-lists.
- Require structured output contracts for automation paths.
- Default to low-risk model settings for critical workflows.
- Disable unconstrained tool execution in production.
Tooling Stack (Example)
| Layer |
Tools |
| Orchestration |
Argo Workflows, GitHub Actions, Airflow |
| Model Registry |
MLflow, custom metadata DB |
| Gateway |
LiteLLM, Envoy-based API gateway |
| Observability |
OpenTelemetry + Prometheus + Grafana + Langfuse |
| Policy |
OPA/Rego for deployment and runtime checks |
| Evaluation |
RAGAS, custom eval harness, Promptfoo |
| Serving |
vLLM, TGI, Triton Inference Server |
Troubleshooting
| Issue |
Diagnosis |
Resolution |
| Canary fails quality gate |
Compare eval results with baseline |
Adjust model config or revert version |
| Deployment stuck in rollout |
Check pod events and resource quotas |
Fix resource limits or node availability |
| A/B test shows no significant difference |
Verify traffic split and sample size |
Extend test duration or increase treatment weight |
| Model cold start too slow |
Large model weight download |
Use pre-cached PVCs or init containers |
| Eval pipeline flaky |
Non-deterministic model outputs |
Set temperature=0 for evals, increase sample size |
Related Skills
1---2name: llmops-platform-engineering3description: Build production LLMOps platforms with CI/CD, model promotion workflows, evaluation gates, rollback, and governance across cloud and self-hosted inference.4license: MIT5---6
7# LLMOps Platform Engineering
8
9Design and operate an internal LLM platform that supports rapid experimentation without compromising reliability, cost, or compliance.
10
11## When to Use This Skill
12
13- Building an internal platform for teams to deploy and manage LLM-powered features
14- Designing CI/CD pipelines that include model evaluation gates
15- Setting up A/B testing infrastructure for model versions
16- Creating Kubernetes-based model serving infrastructure
17- Establishing governance workflows for model promotion
18
19## Prerequisites
20
21- Kubernetes cluster with GPU node pools (or cloud inference API access)
22- Container registry (Harbor, ECR, GCR, or ACR)
23- CI/CD system (GitHub Actions, GitLab CI, or Argo Workflows)
24- Observability stack (Prometheus + Grafana + OpenTelemetry)
25- Model registry (MLflow or custom metadata store)
26
27## Outcomes
28
29- Standardized path from experiment to production
30- Safe model rollout with quality and safety gates
31- Repeatable infra modules for inference, vector DB, and observability
32- Clear ownership model across platform, app, and security teams
33
34## Reference Architecture
35
361. **Control Plane**: model registry, prompt/version catalog, policy checks, eval pipeline.
372. **Data Plane**: inference gateway, vector database, cache, feature store.
383. **Ops Plane**: telemetry, alerting, SLO dashboards, cost analytics.
394. **Security Plane**: IAM boundaries, secret rotation, content filters, audit logs.
40
41## Model Promotion Pipeline
42
43```yaml
44# .github/workflows/model-promotion.yaml
45name: Model Promotion Pipeline
46on:
47 workflow_dispatch:
48 inputs:
49 model_name:
50 description: "Model identifier"
51 required: true
52 model_version:
53 description: "Model version to promote"
54 required: true
55 target_env:
56 description: "Target environment"
57 required: true
58 type: choice
59 options: [staging, production]
60
61jobs:
62 evaluate:
63 runs-on: ubuntu-latest
64 steps:
65 - uses: actions/checkout@v4
66
67 - name: Run quality evaluation suite
68 run: |
69 python -m evals.run \
70 --model "${{ inputs.model_name }}:${{ inputs.model_version }}" \
71 --suite quality \
72 --output results/quality.json
73
74 - name: Run safety evaluation suite
75 run: |
76 python -m evals.run \
77 --model "${{ inputs.model_name }}:${{ inputs.model_version }}" \
78 --suite safety \
79 --output results/safety.json
80
81 - name: Run latency benchmark
82 run: |
83 python -m evals.benchmark \
84 --model "${{ inputs.model_name }}:${{ inputs.model_version }}" \
85 --concurrent-users 50 \
86 --duration 300 \
87 --output results/latency.json
88
89 - name: Gate check - quality
90 run: |
91 python -m evals.gate_check \
92 --results results/quality.json \
93 --threshold-file thresholds/quality.yaml
94
95 - name: Gate check - safety
96 run: |
97 python -m evals.gate_check \
98 --results results/safety.json \
99 --threshold-file thresholds/safety.yaml
100
101 - name: Gate check - latency
102 run: |
103 python -m evals.gate_check \
104 --results results/latency.json \
105 --threshold-file thresholds/latency.yaml
106
107 - name: Upload eval evidence
108 uses: actions/upload-artifact@v4
109 with:
110 name: eval-results-${{ inputs.model_version }}
111 path: results/
112
113 approve:
114 needs: evaluate
115 runs-on: ubuntu-latest
116 environment: ${{ inputs.target_env }}
117 steps:
118 - name: Record approval
119 run: |
120 echo "Approved by: ${{ github.actor }}"
121 echo "Model: ${{ inputs.model_name }}:${{ inputs.model_version }}"
122 echo "Target: ${{ inputs.target_env }}"
123 echo "Time: $(date -u +%Y-%m-%dT%H:%M:%SZ)"
124
125 deploy:
126 needs: approve
127 runs-on: ubuntu-latest
128 steps:
129 - uses: actions/checkout@v4
130
131 - name: Deploy canary
132 run: |
133 kubectl set image deployment/${{ inputs.model_name }}-canary \
134 model=${{ inputs.model_name }}:${{ inputs.model_version }} \
135 -n ai-${{ inputs.target_env }}
136
137 - name: Wait for canary validation (15 min)
138 run: |
139 python -m canary.validate \
140 --deployment ${{ inputs.model_name }}-canary \
141 --namespace ai-${{ inputs.target_env }} \
142 --duration 900 \
143 --quality-threshold 0.85 \
144 --error-rate-threshold 0.02
145
146 - name: Promote to full rollout
147 run: |
148 kubectl set image deployment/${{ inputs.model_name }} \
149 model=${{ inputs.model_name }}:${{ inputs.model_version }} \
150 -n ai-${{ inputs.target_env }}
151 kubectl rollout status deployment/${{ inputs.model_name }} \
152 -n ai-${{ inputs.target_env }} --timeout=300s
153```
154
155## Evaluation Gate Thresholds
156
157```yaml
158# thresholds/quality.yaml
159gates:
160 groundedness:
161 metric: groundedness_score
162 min: 0.85
163 comparison: gte
164 task_success:
165 metric: task_success_rate
166 min: 0.90
167 comparison: gte
168 hallucination:
169 metric: hallucination_rate
170 max: 0.08
171 comparison: lte
172 regression:
173 metric: quality_delta_vs_baseline
174 min: -0.02
175 comparison: gte
176 description: "Must not regress more than 2% vs current production"
177
178# thresholds/latency.yaml
179gates:
180 p50_latency:
181 metric: latency_p50_ms
182 max: 800
183 comparison: lte
184 p95_latency:
185 metric: latency_p95_ms
186 max: 2000
187 comparison: lte
188 p99_latency:
189 metric: latency_p99_ms
190 max: 5000
191 comparison: lte
192 throughput:
193 metric: requests_per_second
194 min: 50
195 comparison: gte
196```
197
198## A/B Testing Configuration
199
200```yaml
201# ab-test-config.yaml
202apiVersion: gateway.ai/v1
203kind: ABTest
204metadata:
205 name: model-comparison-q1
206 namespace: ai-production
207spec:
208 duration: 7d
209 traffic_split:
210 control:
211 model: gpt-4o-2024-08-06
212 weight: 70
213 treatment:
214 model: gpt-4o-2025-01-15
215 weight: 30
216 metrics:
217 primary:
218 - task_success_rate
219 - user_satisfaction_score
220 secondary:
221 - latency_p95
222 - cost_per_request
223 - hallucination_rate
224 guardrails:
225 auto_rollback_if:
226 - metric: task_success_rate
227 threshold: 0.80
228 window: 1h
229 - metric: hallucination_rate
230 threshold: 0.15
231 window: 30m
232 assignment:
233 strategy: sticky_user
234 hash_key: user_id
235```
236
237## Kubernetes Model Serving Deployment
238
239```yaml
240# model-serving-deployment.yaml
241apiVersion: apps/v1
242kind: Deployment
243metadata:
244 name: llm-inference
245 namespace: ai-production
246 labels:
247 app: llm-inference
248 model: gpt-4o
249 version: "2025-01"
250spec:
251 replicas: 3
252 strategy:
253 type: RollingUpdate
254 rollingUpdate:
255 maxSurge: 1
256 maxUnavailable: 0
257 selector:
258 matchLabels:
259 app: llm-inference
260 template:
261 metadata:
262 labels:
263 app: llm-inference
264 model: gpt-4o
265 annotations:
266 prometheus.io/scrape: "true"
267 prometheus.io/port: "8080"
268 prometheus.io/path: "/metrics"
269 spec:
270 topologySpreadConstraints:
271 - maxSkew: 1
272 topologyKey: topology.kubernetes.io/zone
273 whenUnsatisfiable: DoNotSchedule
274 labelSelector:
275 matchLabels:
276 app: llm-inference
277 containers:
278 - name: model
279 image: registry.internal/vllm-server:0.4.1
280 args:
281 - "--model=/models/current"
282 - "--tensor-parallel-size=1"
283 - "--max-model-len=8192"
284 - "--gpu-memory-utilization=0.90"
285 ports:
286 - containerPort: 8000
287 name: inference
288 - containerPort: 8080
289 name: metrics
290 resources:
291 requests:
292 cpu: "4"
293 memory: "16Gi"
294 nvidia.com/gpu: "1"
295 limits:
296 cpu: "8"
297 memory: "32Gi"
298 nvidia.com/gpu: "1"
299 readinessProbe:
300 httpGet:
301 path: /health
302 port: 8000
303 initialDelaySeconds: 60
304 periodSeconds: 10
305 livenessProbe:
306 httpGet:
307 path: /health
308 port: 8000
309 initialDelaySeconds: 120
310 periodSeconds: 30
311 volumeMounts:
312 - name: model-weights
313 mountPath: /models
314 readOnly: true
315 - name: config
316 mountPath: /etc/vllm
317 volumes:
318 - name: model-weights
319 persistentVolumeClaim:
320 claimName: model-weights-pvc
321 - name: config
322 configMap:
323 name: vllm-config
324 tolerations:
325 - key: nvidia.com/gpu
326 operator: Exists
327 effect: NoSchedule
328 nodeSelector:
329 gpu-type: a100
330---
331apiVersion: v1
332kind: Service
333metadata:
334 name: llm-inference
335 namespace: ai-production
336spec:
337 selector:
338 app: llm-inference
339 ports:
340 - name: inference
341 port: 8000
342 targetPort: 8000
343 - name: metrics
344 port: 8080
345 targetPort: 8080
346---
347apiVersion: autoscaling/v2
348kind: HorizontalPodAutoscaler
349metadata:
350 name: llm-inference-hpa
351 namespace: ai-production
352spec:
353 scaleTargetRef:
354 apiVersion: apps/v1
355 kind: Deployment
356 name: llm-inference
357 minReplicas: 2
358 maxReplicas: 10
359 metrics:
360 - type: Pods
361 pods:
362 metric:
363 name: llm_queue_depth
364 target:
365 type: AverageValue
366 averageValue: "5"
367 - type: Pods
368 pods:
369 metric:
370 name: gpu_utilization_percent
371 target:
372 type: AverageValue
373 averageValue: "75"
374 behavior:
375 scaleUp:
376 stabilizationWindowSeconds: 60
377 policies:
378 - type: Pods
379 value: 2
380 periodSeconds: 120
381 scaleDown:
382 stabilizationWindowSeconds: 300
383 policies:
384 - type: Pods
385 value: 1
386 periodSeconds: 300
387```
388
389## CI/CD Design for AI Services
390
391- Build immutable containers with pinned dependencies and model hashes.
392- Use environment promotion: `dev -> stage -> prod`.
393- Fail deployment if:
394 - regression evals drop below baseline,
395 - safety tests exceed risk threshold,
396 - p95 latency exceeds SLO budget.
397- Store deployment evidence for audits (commit SHA, eval report, approver).
398
399## Operational SLOs
400
401| Signal | Target | Measurement Window |
402|--------|--------|--------------------|
403| Availability | 99.9% | 30-day rolling |
404| p95 Latency | < 1200ms | 5-min buckets |
405| Cost per request | < $0.05 | 1-hour average |
406| Task success rate | > 90% | 24-hour rolling |
407| Groundedness | > 85% | 24-hour rolling |
408
409## Platform Guardrails
410
411- Enforce tenant quotas and model allow-lists.
412- Require structured output contracts for automation paths.
413- Default to low-risk model settings for critical workflows.
414- Disable unconstrained tool execution in production.
415
416## Tooling Stack (Example)
417
418| Layer | Tools |
419|-------|-------|
420| Orchestration | Argo Workflows, GitHub Actions, Airflow |
421| Model Registry | MLflow, custom metadata DB |
422| Gateway | LiteLLM, Envoy-based API gateway |
423| Observability | OpenTelemetry + Prometheus + Grafana + Langfuse |
424| Policy | OPA/Rego for deployment and runtime checks |
425| Evaluation | RAGAS, custom eval harness, Promptfoo |
426| Serving | vLLM, TGI, Triton Inference Server |
427
428## Troubleshooting
429
430| Issue | Diagnosis | Resolution |
431|-------|-----------|------------|
432| Canary fails quality gate | Compare eval results with baseline | Adjust model config or revert version |
433| Deployment stuck in rollout | Check pod events and resource quotas | Fix resource limits or node availability |
434| A/B test shows no significant difference | Verify traffic split and sample size | Extend test duration or increase treatment weight |
435| Model cold start too slow | Large model weight download | Use pre-cached PVCs or init containers |
436| Eval pipeline flaky | Non-deterministic model outputs | Set temperature=0 for evals, increase sample size |
437
438## Related Skills
439
440- [ai-pipeline-orchestration](../ai-pipeline-orchestration/) - Orchestrate ingestion and inference workflows
441- [agent-evals](../agent-evals/) - Build evaluation gates for releases
442- [llm-gateway](../../../infrastructure/networking/llm-gateway/) - Route and control LLM traffic
443- [model-registry-governance](../model-registry-governance/) - Model lifecycle and approval workflows
444- [ai-sre-incident-response](../ai-sre-incident-response/) - AI-specific incident response