1---2name: system-design-resilience-ops3description: Make a design survivable and operable: eliminate single points of failure, pick failover topology, set RPO/RTO, define readiness signals and observability, choose a deployment and rollback strategy. Use when reviewing availability, planning DR, or deciding rollout mechanics.4---5
6# Resilience and Operations
7
8## **Priority: P1 (HIGH)**
9
10A design is not done until its failure and its rollout are designed.
11
12## SPOF Elimination
13
14- Walk every component and ask what happens when exactly one instance dies, then when the whole zone dies.
15- Any component with one instance, one writer, or one shared config plane is a single point of failure. Name it or remove it.
16- Redundancy only helps when failure modes are independent: shared credentials, shared config, and a shared control plane cancel the benefit.
17- Blast radius: state which users or flows are affected per component failure, and cap it with cells, bulkheads, or per-tenant quotas.
18
19## Failover and Recovery
20
21| Topology | Recovery time | Cost | Fits |
22| --- | --- | --- | --- |
23| Single region, multi-AZ | Minutes, automatic | Low | Most products |
24| Active-passive across regions | Minutes to hours, drill-dependent | Medium | Regulated or high-value flows |
25| Active-active across regions | Seconds | High | Global low-latency, conflict-tolerant data |
26
27- Set **RPO** (tolerable data loss) and **RTO** (tolerable downtime) as numbers before choosing a topology; the numbers pick the topology, not the reverse.
28- Untested failover is a hypothesis. Schedule a drill and record the measured RTO against the target.
29- Backups need a restore test. A backup that has never been restored is not a backup.
30
31## Observability
32
33- Instrument the four signals per service: traffic, error rate, latency percentiles, saturation.
34- Alert on user-visible symptoms and on error-budget burn rate, not on raw CPU.
35- Propagate a trace and correlation id across every hop, including queue messages.
36- Every alert needs an owner, a runbook link, and a defined next action; an alert nobody acts on is noise.
37
38## Rollout
39
40| Strategy | Blast radius | Rollback | Cost |
41| --- | --- | --- | --- |
42| Rolling | Grows during the roll | Roll forward or back, slow | Low |
43| Blue-green | Full switch at cutover | Instant switch back | Double capacity |
44| Canary | Small cohort first | Stop and drain the cohort | Needs routing plus metrics |
45| Feature flag | Per user or tenant | Instant, no redeploy | Flag lifecycle debt |
46
47- Schema and code deploy separately: expand, migrate, contract. Never ship a migration that only the new code can read.
48- Define the rollback trigger as a metric threshold and a time box before the deploy starts.
49
50## Anti-Patterns
51
52- **No untested failover**: no DR claim without a drill date and a measured RTO.
53- **No unbounded retry**: retries need budget, backoff with jitter, and a stop condition, or they amplify an outage.
54- **No liveness probe on dependencies**: a downstream outage must not restart the fleet.
55- **No deploy without rollback**: irreversible releases are outages waiting for a bad build.
56- **No autoscaling without a floor and ceiling**: unbounded scaling turns a bug into a bill.
57
58## References
59
60- [Reliability Operations](references/reliability-operations.md) - failure drills, health check design, DR runbook shape, scaling policy notes