1---2name: devops3description: Automate deployments, manage infrastructure, and build reliable CI/CD pipelines.4---56# DevOps Rules78## CI/CD Pipelines9- Fail fast: run linting and unit tests before expensive integration tests — saves time and compute10- Cache dependencies between runs — `npm install` on every build wastes minutes11- Pin action versions with SHA, not tags — `actions/checkout@v3` can change, SHA is immutable12- Secrets in environment variables, never in code or logs — mask them in CI output13- Parallel jobs for independent steps — test, lint, and build can run simultaneously1415## Deployment Strategies16- Blue-green: run new version alongside old, switch traffic atomically — instant rollback by switching back17- Canary: route percentage of traffic to new version — catch issues before full rollout18- Rolling: update instances incrementally — balance between speed and risk19- Always have rollback plan before deploying — know exactly how to revert20- Deploy the same artifact to all environments — build once, promote through stages2122## Infrastructure as Code23- Version control all infrastructure — terraform, ansible, cloudformation in git24- Never apply changes without plan/diff review — `terraform plan` before `apply`25- State files contain secrets — store remotely with encryption, never in git26- Modules for reusable components — don't copy-paste infrastructure definitions27- Separate environments with workspaces or directories — dev changes shouldn't affect prod2829## Containers30- One process per container — containers are not VMs31- Health checks are mandatory — orchestrators need them for routing and restarts32- Don't run as root — use non-root USER in Dockerfile33- Immutable images: config via environment, not baked in — same image in all environments34- Tag images with git SHA, not just `latest` — know exactly what's deployed3536## Secrets Management37- Never store secrets in environment files committed to git — use vault, sealed secrets, or CI secret storage38- Rotate secrets regularly — automation makes rotation painless39- Different secrets per environment — dev leak shouldn't compromise prod40- Audit secret access — know who accessed what and when41- Secrets in memory, not disk when possible — temp files persist longer than expected4243## Monitoring & Alerting44- Four golden signals: latency, traffic, errors, saturation — start here45- Alert on symptoms, not causes — "users seeing errors" not "CPU high"46- Every alert must be actionable — if you can't do anything, it's noise47- Dashboard per service with key metrics — one glance shows health48- Structured logs (JSON) for machine parsing — grep works, but queries are better4950## Reliability51- Define SLOs before building alerting — what does "healthy" mean for this service?52- Error budgets: some failures are acceptable — 99.9% means 8 hours downtime/year is OK53- Chaos engineering in staging — break things intentionally before prod breaks accidentally54- Runbooks for common incidents — 3am is not the time to figure out recovery steps55- Post-mortems without blame — focus on systems, not people5657## Common Mistakes58- SSH into prod to fix things — all changes through automation, or you'll forget what you did59- No staging environment — "works on my machine" doesn't mean works in prod60- Ignoring flaky tests — they erode trust in CI, either fix or delete61- Manual steps in deployment — if it's not automated, it'll be done wrong eventually62- Monitoring only happy paths — check error rates and edge cases too6364## Networking65- Internal services don't need public IPs — use private subnets, expose only load balancers66- TLS everywhere, including internal traffic — zero trust, even behind firewall67- DNS for service discovery — hardcoded IPs break when things move68- Load balancer health checks separate from app health — LB needs fast response, app health can be thorough69- Firewall default deny — explicitly allow what's needed, block everything else