1---2name: devops3description: Automate deployments, manage infrastructure, and build reliable CI/CD pipelines.4---5
6# DevOps Rules
7
8## CI/CD Pipelines
9- Fail fast: run linting and unit tests before expensive integration tests — saves time and compute
10- Cache dependencies between runs — `npm install` on every build wastes minutes
11- Pin action versions with SHA, not tags — `actions/checkout@v3` can change, SHA is immutable
12- Secrets in environment variables, never in code or logs — mask them in CI output
13- Parallel jobs for independent steps — test, lint, and build can run simultaneously
14
15## Deployment Strategies
16- Blue-green: run new version alongside old, switch traffic atomically — instant rollback by switching back
17- Canary: route percentage of traffic to new version — catch issues before full rollout
18- Rolling: update instances incrementally — balance between speed and risk
19- Always have rollback plan before deploying — know exactly how to revert
20- Deploy the same artifact to all environments — build once, promote through stages
21
22## Infrastructure as Code
23- Version control all infrastructure — terraform, ansible, cloudformation in git
24- Never apply changes without plan/diff review — `terraform plan` before `apply`
25- State files contain secrets — store remotely with encryption, never in git
26- Modules for reusable components — don't copy-paste infrastructure definitions
27- Separate environments with workspaces or directories — dev changes shouldn't affect prod
28
29## Containers
30- One process per container — containers are not VMs
31- Health checks are mandatory — orchestrators need them for routing and restarts
32- Don't run as root — use non-root USER in Dockerfile
33- Immutable images: config via environment, not baked in — same image in all environments
34- Tag images with git SHA, not just `latest` — know exactly what's deployed
35
36## Secrets Management
37- Never store secrets in environment files committed to git — use vault, sealed secrets, or CI secret storage
38- Rotate secrets regularly — automation makes rotation painless
39- Different secrets per environment — dev leak shouldn't compromise prod
40- Audit secret access — know who accessed what and when
41- Secrets in memory, not disk when possible — temp files persist longer than expected
42
43## Monitoring & Alerting
44- Four golden signals: latency, traffic, errors, saturation — start here
45- Alert on symptoms, not causes — "users seeing errors" not "CPU high"
46- Every alert must be actionable — if you can't do anything, it's noise
47- Dashboard per service with key metrics — one glance shows health
48- Structured logs (JSON) for machine parsing — grep works, but queries are better
49
50## Reliability
51- Define SLOs before building alerting — what does "healthy" mean for this service?
52- Error budgets: some failures are acceptable — 99.9% means 8 hours downtime/year is OK
53- Chaos engineering in staging — break things intentionally before prod breaks accidentally
54- Runbooks for common incidents — 3am is not the time to figure out recovery steps
55- Post-mortems without blame — focus on systems, not people
56
57## Common Mistakes
58- SSH into prod to fix things — all changes through automation, or you'll forget what you did
59- No staging environment — "works on my machine" doesn't mean works in prod
60- Ignoring flaky tests — they erode trust in CI, either fix or delete
61- Manual steps in deployment — if it's not automated, it'll be done wrong eventually
62- Monitoring only happy paths — check error rates and edge cases too
63
64## Networking
65- Internal services don't need public IPs — use private subnets, expose only load balancers
66- TLS everywhere, including internal traffic — zero trust, even behind firewall
67- DNS for service discovery — hardcoded IPs break when things move
68- Load balancer health checks separate from app health — LB needs fast response, app health can be thorough
69- Firewall default deny — explicitly allow what's needed, block everything else