DevOps Guidelines
Apply only the sections relevant to the workload, provider, environment, and acceptance criteria. Skip Docker, Kubernetes, mobile, production, rollback, health, feature-flag, and security checks when they do not apply.
Deployment strategy
- Rolling (default): gradual, zero-downtime replacement.
- Blue-green: duplicate environments, atomic cutover, instant rollback, 2× infrastructure.
- Canary: route a small percentage first; requires traffic splitting.
Docker
- Pin specific base-image tags (for example
node:22-alpine); NEVER use :latest.
- Use multi-stage builds and a non-root user. Copy dependencies first for caching.
.dockerignore: node_modules, .git, tests. Define HEALTHCHECK and resource limits.
Kubernetes
Configure startup, readiness, and liveness probes with workload-appropriate initial delays and thresholds.
CI/CD
- PR: lint -> typecheck -> unit -> integration -> preview.
- Main: build -> staging -> smoke -> production.
Health and shutdown
- Simple:
GET /health -> { "status": "ok" }.
- Detailed: dependencies, uptime, version.
- Services MUST expose meaningful health and gracefully handle
SIGTERM when the workload requires it.
Configuration
Use environment variables (Twelve-Factor), separated by environment. Validate at startup and fail fast. NEVER commit secrets or hard-code NODE_ENV=production.
Rollback
Kubernetes: kubectl rollout undo. Vercel: vercel rollback. Docker: redeploy the previous pinned image.
Feature Flags
- Lifecycle: create -> enable -> 5% -> 25% -> 50% -> 100% -> remove flag and dead code.
- Every flag MUST have an owner, expiration, and rollback trigger. Remove within two weeks.
Checklists
- Pre-deploy, when applicable: passing tests, code review, environment variables, migrations, rollback plan.
- Post-deploy services: healthy, monitored, old pods terminated, outcome documented.
- Production services: passing tests; no hardcoded secrets; JSON logs; meaningful health; pinned versions; validated environment variables; resource limits; TLS; CVE scan; CORS; rate limiting; CSP/HSTS/X-Frame-Options; tested rollback; runbook; on-call.
- Apply security/CVE checks to executable or security-sensitive workloads.
Mobile Deployment
- EAS:
eas build:configure; eas build -p ios|android --profile preview; eas update --branch production; --auto-submit.
- Fastlane: iOS
match/cert/sigh/pilot; Android Gradle/supply.
- Keep credentials in environment/secret storage, never Git. Automate iOS development/distribution signing with
fastlane match; use keytool and Google Play App Signing for Android.
- TestFlight: internal instant; external 90 days/100 testers. Google Play: internal/beta/production. Expect 1–7 days for review.
- Rollback: EAS
eas update:rollback; native release -> revert build; store release -> reduce phased rollout.
1---2name: gem-devops-guidelines3description: Design or review infrastructure, deployment, CI/CD, Docker, Kubernetes, health checks, rollback, feature flags, production readiness, and mobile release workflows. Use for DevOps, platform, container, pipeline, or release tasks.4---5
6# DevOps Guidelines
7
8Apply only the sections relevant to the workload, provider, environment, and acceptance criteria. Skip Docker, Kubernetes, mobile, production, rollback, health, feature-flag, and security checks when they do not apply.
9
10## Deployment strategy
11
12- Rolling (default): gradual, zero-downtime replacement.
13- Blue-green: duplicate environments, atomic cutover, instant rollback, 2× infrastructure.
14- Canary: route a small percentage first; requires traffic splitting.
15
16## Docker
17
18- Pin specific base-image tags (for example `node:22-alpine`); NEVER use `:latest`.
19- Use multi-stage builds and a non-root user. Copy dependencies first for caching.
20- `.dockerignore`: `node_modules`, `.git`, tests. Define `HEALTHCHECK` and resource limits.
21
22## Kubernetes
23
24Configure startup, readiness, and liveness probes with workload-appropriate initial delays and thresholds.
25
26## CI/CD
27
28- PR: lint -> typecheck -> unit -> integration -> preview.
29- Main: build -> staging -> smoke -> production.
30
31## Health and shutdown
32
33- Simple: `GET /health` -> `{ "status": "ok" }`.
34- Detailed: dependencies, uptime, version.
35- Services MUST expose meaningful health and gracefully handle `SIGTERM` when the workload requires it.
36
37## Configuration
38
39Use environment variables (Twelve-Factor), separated by environment. Validate at startup and fail fast. NEVER commit secrets or hard-code `NODE_ENV=production`.
40
41## Rollback
42
43Kubernetes: `kubectl rollout undo`. Vercel: `vercel rollback`. Docker: redeploy the previous pinned image.
44
45## Feature Flags
46
47- Lifecycle: create -> enable -> 5% -> 25% -> 50% -> 100% -> remove flag and dead code.
48- Every flag MUST have an owner, expiration, and rollback trigger. Remove within two weeks.
49
50## Checklists
51
52- Pre-deploy, when applicable: passing tests, code review, environment variables, migrations, rollback plan.
53- Post-deploy services: healthy, monitored, old pods terminated, outcome documented.
54- Production services: passing tests; no hardcoded secrets; JSON logs; meaningful health; pinned versions; validated environment variables; resource limits; TLS; CVE scan; CORS; rate limiting; CSP/HSTS/X-Frame-Options; tested rollback; runbook; on-call.
55- Apply security/CVE checks to executable or security-sensitive workloads.
56
57## Mobile Deployment
58
59- EAS: `eas build:configure`; `eas build -p ios|android --profile preview`; `eas update --branch production`; `--auto-submit`.
60- Fastlane: iOS `match`/`cert`/`sigh`/`pilot`; Android Gradle/`supply`.
61- Keep credentials in environment/secret storage, never Git. Automate iOS development/distribution signing with `fastlane match`; use `keytool` and Google Play App Signing for Android.
62- TestFlight: internal instant; external 90 days/100 testers. Google Play: internal/beta/production. Expect 1–7 days for review.
63- Rollback: EAS `eas update:rollback`; native release -> revert build; store release -> reduce phased rollout.