1---2name: systems-architect3description: Design infrastructure, networks, and cloud systems with integration, reliability, and security patterns.4---56# Systems Architecture Rules78## Infrastructure Design9- Design for failure at every layer — hardware fails, networks partition, regions go down10- Redundancy costs money, downtime costs more — calculate acceptable risk11- Prefer managed services for undifferentiated work — run less, build more12- Infrastructure as code from day one — manual changes drift and break13- Immutable infrastructure beats patching — replace, don't repair1415## Cloud Architecture16- Multi-AZ minimum, multi-region for critical systems — availability zones fail together sometimes17- Right-size first, auto-scale second — baseline must be correct18- Reserved capacity for steady load, spot/preemptible for bursts — cost optimization requires planning19- Egress costs add up — keep traffic within regions when possible20- Cloud vendor lock-in is real — abstract where escape matters, accept where it doesn't2122## Networking23- Private subnets for workloads, public only for load balancers — minimize attack surface24- VPC peering and transit gateways for multi-account — plan topology before scaling25- DNS for service discovery — hardcoded IPs break migrations26- Zero trust: authenticate and encrypt internal traffic — perimeter security isn't enough27- Network segmentation limits blast radius — flat networks let attackers roam2829## Integration Patterns30- APIs for synchronous, queues for asynchronous — match pattern to requirements31- Event-driven for loose coupling — producers don't know consumers32- Service mesh for complex microservices — observability and security at network layer33- Rate limiting and backpressure protect systems — don't let slow consumers crash fast producers34- Dead letter queues for failed messages — don't lose data, process later3536## Reliability37- Define SLOs before building — what does "up" mean for this system?38- Error budgets allow controlled risk — 99.9% means 8 hours downtime per year is acceptable39- Blast radius reduction: cell-based architecture — limit how many users one failure affects40- Chaos engineering in staging first — break things intentionally before production breaks accidentally41- Runbooks for every alert — 3 AM isn't debugging time4243## Disaster Recovery44- RTO (recovery time) and RPO (data loss) are business decisions — architect for the requirement45- Backups aren't recovery until tested — restore regularly46- Hot/warm/cold standby each have trade-offs — cost vs speed of recovery47- Cross-region replication for critical data — single region is single point of failure48- DR drills reveal real problems — plan meets reality4950## Security51- Defense in depth: multiple barriers — one layer will fail52- Least privilege for services too — not just users53- Secrets management centralized — no secrets in code, config files, or environment variables in images54- Audit logging for compliance and forensics — you'll need it after a breach55- Patch aggressively — known vulnerabilities are actively exploited5657## Monitoring and Observability58- Metrics, logs, and traces together — each tells part of the story59- Alerting on symptoms, not causes — users down matters, CPU high might not60- Dashboards for each service with golden signals — latency, traffic, errors, saturation61- Distributed tracing across services — follow requests end to end62- Log aggregation with retention policy — balance cost and forensic needs6364## Capacity Planning65- Measure current baseline before projecting — can't scale what you don't measure66- Load test to find breaking points — theory differs from reality67- Capacity leads demand — scaling takes time, be ahead68- Cost modeling for growth scenarios — 10x users is rarely 10x cost69- Review quarterly at minimum — patterns change7071## Migration and Evolution72- Strangler fig pattern for legacy replacement — route traffic gradually73- Blue-green or canary for infrastructure changes — test in production safely74- Database migrations are hardest — plan data migration separately75- Rollback plans before rollout — assume failure, prepare for it76- Communicate maintenance windows — surprises damage trust