Kubernetes Platform Engineering
Overview
Platform engineering on Kubernetes is about making the “golden path” easy: secure-by-default workloads, consistent delivery via GitOps, and predictable operations (capacity, upgrades, incident response).
Why This Matters
- Reliability: self-healing + controlled rollouts reduce incidents
- Security: least privilege and isolation by default
- Developer velocity: standard templates + paved roads
- Cost control: right-sizing and autoscaling without surprises
Core Concepts
1. Cluster Architecture
- Separate node pools by workload class (stateless, stateful, GPU, batch, system).
- Define upgrade strategy (surge capacity, maintenance windows, version skew policy).
- Decide multi-cluster approach (per env/region/tenant) and shared services (ingress, monitoring, secrets).
2. Workload Patterns
- Deployment for stateless services; use
PodDisruptionBudget, readinessProbe, and topologySpreadConstraints.
- StatefulSet for stable identity/storage (databases, queues) with explicit backup/restore.
- Job/CronJob for batch; enforce concurrency policies and deadlines.
- DaemonSet for node-level agents (logging, CNI, security).
3. Networking
- Use
Ingress/Gateway with TLS termination and standardized auth/rate limiting at the edge.
- Apply NetworkPolicies (default-deny + explicit allow) for zero-trust inside the cluster.
- Consider service mesh only when you need it (mTLS, traffic shifting, retries, L7 telemetry); keep complexity intentional.
4. Storage
- Standardize
StorageClass per tier (fast SSD, standard, archival) and backup policies.
- For stateful apps: set anti-affinity/zone spread; verify recovery time objectives (RTO/RPO).
- Avoid “pet” PVs by having a documented restore path and regular restore drills.
5. Security Baseline
- Enforce Pod Security (restricted baseline): non-root, read-only root filesystem, drop capabilities.
- RBAC least privilege: namespace-scoped roles, separate human vs automation identities.
- Secrets: prefer external secret managers (e.g., External Secrets + AWS/GCP/Azure secret store); avoid long-lived static secrets in Git.
- Admission policies: Kyverno / Gatekeeper for guardrails (image registries, resource limits, required labels).
6. Observability & Ops
- Metrics: Prometheus + dashboards per service (latency, errors, saturation, restarts).
- Logs: structured JSON with trace IDs; centralized retention policies.
- Traces: OpenTelemetry for request graphs across services.
- SLOs: define error budgets and link alerts to user impact.
7. GitOps & Delivery
- Git is source of truth: Argo CD / Flux reconciles desired state.
- Use Helm/Kustomize overlays per environment; keep secrets out-of-band.
- Progressive delivery: canary/blue-green via Argo Rollouts / Flagger (or mesh-based traffic shifting).
8. Cost & Capacity Management
- Require resource requests/limits; use VPA recommendations carefully (watch for restarts).
- Use HPA/KEDA for demand-based scaling; use cluster autoscaler/Karpenter for nodes.
- Separate spot/preemptible pools for tolerant workloads; enforce pod placement policies.
Quick Start (Minimal “Production-ish” Deployment)
apiVersion: apps/v1
kind: Deployment
metadata:
name: app
spec:
replicas: 2
strategy:
type: RollingUpdate
rollingUpdate: { maxUnavailable: 0, maxSurge: 1 }
selector: { matchLabels: { app: app } }
template:
metadata: { labels: { app: app } }
spec:
securityContext: { runAsNonRoot: true }
containers:
- name: app
image: your-registry/app:1.0.0
ports: [{ containerPort: 3000 }]
resources:
requests: { cpu: "100m", memory: "256Mi" }
limits: { cpu: "500m", memory: "512Mi" }
readinessProbe:
httpGet: { path: /readyz, port: 3000 }
initialDelaySeconds: 5
livenessProbe:
httpGet: { path: /healthz, port: 3000 }
initialDelaySeconds: 10
---
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
name: app
spec:
minAvailable: 1
selector: { matchLabels: { app: app } }
Production Checklist
Tools & Libraries
| Tool |
Purpose |
| Argo CD / Flux |
GitOps continuous delivery |
| Helm / Kustomize |
Packaging and environment overlays |
| Prometheus + Grafana |
Metrics + dashboards |
| Loki/ELK |
Central logging |
| OpenTelemetry |
Distributed tracing |
| External Secrets |
Secret manager integration |
| Kyverno / Gatekeeper |
Policy enforcement |
Anti-patterns
- No requests/limits: unpredictable scheduling and noisy-neighbor incidents
- Hand-applied changes:
kubectl apply drift instead of GitOps reconciliation
- Flat network: no NetworkPolicies; lateral movement is trivial
- Single replica: planned/unplanned disruption becomes downtime
Real-World Examples
Example: Progressive Delivery (Canary)
- Deploy new version at 5% traffic; watch SLOs; ramp to 25/50/100%; auto-rollback on regressions.
Example: Autoscaling
- HPA on CPU/requests-per-second + cluster autoscaler for nodes; cap max replicas to protect dependencies.
Example: Zero-Trust Networking
- Default-deny in each namespace; only allow traffic from ingress controller to app and app to explicit dependencies.
Common Mistakes
- Missing readiness probes (traffic hits pods before they’re ready)
- Running containers as root (wider blast radius on compromise)
- Unbounded egress (data exfiltration paths and surprise costs)
- Treating the cluster as the product (no SLOs, no runbooks, no ownership)
Integration Points
- CI (image build, SBOM, scanning, signing)
- CD (GitOps reconciler + progressive delivery)
- Cloud provider (EKS/GKE/AKS, IAM, load balancers, DNS)
- Observability + incident management (paging, runbooks, postmortems)
Further Reading
1---2name: kubernetes-platform-engineering3description: Production Kubernetes platform patterns covering cluster architecture, security, GitOps, observability, autoscaling, and operational guardrails4---5
6# Kubernetes Platform Engineering
7
8## Overview
9
10Platform engineering on Kubernetes is about making the “golden path” easy: secure-by-default workloads, consistent delivery via GitOps, and predictable operations (capacity, upgrades, incident response).
11
12## Why This Matters
13
14- **Reliability**: self-healing + controlled rollouts reduce incidents
15- **Security**: least privilege and isolation by default
16- **Developer velocity**: standard templates + paved roads
17- **Cost control**: right-sizing and autoscaling without surprises
18
19---
20
21## Core Concepts
22
23### 1. Cluster Architecture
24
25- Separate node pools by workload class (stateless, stateful, GPU, batch, system).
26- Define upgrade strategy (surge capacity, maintenance windows, version skew policy).
27- Decide multi-cluster approach (per env/region/tenant) and shared services (ingress, monitoring, secrets).
28
29### 2. Workload Patterns
30
31- **Deployment** for stateless services; use `PodDisruptionBudget`, `readinessProbe`, and `topologySpreadConstraints`.
32- **StatefulSet** for stable identity/storage (databases, queues) with explicit backup/restore.
33- **Job/CronJob** for batch; enforce concurrency policies and deadlines.
34- **DaemonSet** for node-level agents (logging, CNI, security).
35
36### 3. Networking
37
38- Use `Ingress`/Gateway with TLS termination and standardized auth/rate limiting at the edge.
39- Apply **NetworkPolicies** (default-deny + explicit allow) for zero-trust inside the cluster.
40- Consider service mesh only when you need it (mTLS, traffic shifting, retries, L7 telemetry); keep complexity intentional.
41
42### 4. Storage
43
44- Standardize `StorageClass` per tier (fast SSD, standard, archival) and backup policies.
45- For stateful apps: set anti-affinity/zone spread; verify recovery time objectives (RTO/RPO).
46- Avoid “pet” PVs by having a documented restore path and regular restore drills.
47
48### 5. Security Baseline
49
50- Enforce Pod Security (restricted baseline): non-root, read-only root filesystem, drop capabilities.
51- RBAC least privilege: namespace-scoped roles, separate human vs automation identities.
52- Secrets: prefer external secret managers (e.g., External Secrets + AWS/GCP/Azure secret store); avoid long-lived static secrets in Git.
53- Admission policies: Kyverno / Gatekeeper for guardrails (image registries, resource limits, required labels).
54
55### 6. Observability & Ops
56
57- Metrics: Prometheus + dashboards per service (latency, errors, saturation, restarts).
58- Logs: structured JSON with trace IDs; centralized retention policies.
59- Traces: OpenTelemetry for request graphs across services.
60- SLOs: define error budgets and link alerts to user impact.
61
62### 7. GitOps & Delivery
63
64- Git is source of truth: Argo CD / Flux reconciles desired state.
65- Use Helm/Kustomize overlays per environment; keep secrets out-of-band.
66- Progressive delivery: canary/blue-green via Argo Rollouts / Flagger (or mesh-based traffic shifting).
67
68### 8. Cost & Capacity Management
69
70- Require resource requests/limits; use VPA recommendations carefully (watch for restarts).
71- Use HPA/KEDA for demand-based scaling; use cluster autoscaler/Karpenter for nodes.
72- Separate spot/preemptible pools for tolerant workloads; enforce pod placement policies.
73
74## Quick Start (Minimal “Production-ish” Deployment)
75
76```yaml
77apiVersion: apps/v1
78kind: Deployment
79metadata:
80 name: app
81spec:
82 replicas: 2
83 strategy:
84 type: RollingUpdate
85 rollingUpdate: { maxUnavailable: 0, maxSurge: 1 }
86 selector: { matchLabels: { app: app } }
87 template:
88 metadata: { labels: { app: app } }
89 spec:
90 securityContext: { runAsNonRoot: true }
91 containers:
92 - name: app
93 image: your-registry/app:1.0.0
94 ports: [{ containerPort: 3000 }]
95 resources:
96 requests: { cpu: "100m", memory: "256Mi" }
97 limits: { cpu: "500m", memory: "512Mi" }
98 readinessProbe:
99 httpGet: { path: /readyz, port: 3000 }
100 initialDelaySeconds: 5
101 livenessProbe:
102 httpGet: { path: /healthz, port: 3000 }
103 initialDelaySeconds: 10
104---
105apiVersion: policy/v1
106kind: PodDisruptionBudget
107metadata:
108 name: app
109spec:
110 minAvailable: 1
111 selector: { matchLabels: { app: app } }
112```
113
114## Production Checklist
115
116- [ ] Requests/limits set; namespaces have `LimitRange`/`ResourceQuota`
117- [ ] `readinessProbe` + `livenessProbe` + graceful shutdown configured
118- [ ] PDB + topology spread/anti-affinity to survive node/zone events
119- [ ] NetworkPolicies enforce default-deny and least access
120- [ ] RBAC least privilege; service accounts scoped per app
121- [ ] Metrics/logs/traces exported; alerts tied to SLOs
122- [ ] Backup/restore drills for stateful components
123- [ ] Upgrade playbook exists (cluster + addons + workload compatibility)
124
125## Tools & Libraries
126
127| Tool | Purpose |
128|------|---------|
129| Argo CD / Flux | GitOps continuous delivery |
130| Helm / Kustomize | Packaging and environment overlays |
131| Prometheus + Grafana | Metrics + dashboards |
132| Loki/ELK | Central logging |
133| OpenTelemetry | Distributed tracing |
134| External Secrets | Secret manager integration |
135| Kyverno / Gatekeeper | Policy enforcement |
136
137## Anti-patterns
138
1391. **No requests/limits**: unpredictable scheduling and noisy-neighbor incidents
1402. **Hand-applied changes**: `kubectl apply` drift instead of GitOps reconciliation
1413. **Flat network**: no NetworkPolicies; lateral movement is trivial
1424. **Single replica**: planned/unplanned disruption becomes downtime
143
144## Real-World Examples
145
146### Example: Progressive Delivery (Canary)
147
148- Deploy new version at 5% traffic; watch SLOs; ramp to 25/50/100%; auto-rollback on regressions.
149
150### Example: Autoscaling
151
152- HPA on CPU/requests-per-second + cluster autoscaler for nodes; cap max replicas to protect dependencies.
153
154### Example: Zero-Trust Networking
155
156- Default-deny in each namespace; only allow traffic from ingress controller to app and app to explicit dependencies.
157
158## Common Mistakes
159
1601. Missing readiness probes (traffic hits pods before they’re ready)
1612. Running containers as root (wider blast radius on compromise)
1623. Unbounded egress (data exfiltration paths and surprise costs)
1634. Treating the cluster as the product (no SLOs, no runbooks, no ownership)
164
165## Integration Points
166
167- CI (image build, SBOM, scanning, signing)
168- CD (GitOps reconciler + progressive delivery)
169- Cloud provider (EKS/GKE/AKS, IAM, load balancers, DNS)
170- Observability + incident management (paging, runbooks, postmortems)
171
172## Further Reading
173
174- [Kubernetes Documentation](https://kubernetes.io/docs/)
175- [Kubernetes Production Best Practices](https://learnk8s.io/production-best-practices)
176- [Kubernetes Patterns](https://k8spatterns.io/)