AWS Architecture Reference
Comprehensive guide for AWS services, patterns, and Well-Architected Framework implementation.
Well-Architected Framework
Six Pillars
Operational Excellence
- Infrastructure as Code (CloudFormation, CDK, Terraform)
- Continuous integration/deployment
- Observability (CloudWatch, X-Ray)
- Runbooks and playbooks
- Game days and failure injection
Security
- Identity and Access Management (IAM)
- Detective controls (GuardDuty, Security Hub)
- Infrastructure protection (VPC, security groups, NACLs)
- Data protection (KMS, encryption at rest/transit)
- Incident response automation
Reliability
- Multi-AZ deployments
- Auto Scaling groups
- Route 53 health checks and failover
- Backup and restore (AWS Backup)
- Chaos engineering (AWS FIS)
Performance Efficiency
- Right-sizing with Compute Optimizer
- Caching strategies (CloudFront, ElastiCache)
- Database optimization (RDS Performance Insights)
- Serverless architectures
- Global content delivery
Cost Optimization
- Reserved Instances and Savings Plans
- Spot Instances for fault-tolerant workloads
- S3 Intelligent-Tiering and lifecycle policies
- Right-sizing recommendations
- Cost allocation tags and budgets
Sustainability
- Region selection for renewable energy
- Serverless to minimize idle resources
- Efficient data storage patterns
- Resource utilization optimization
Core Services Architecture
Compute
EC2 (Elastic Compute Cloud)
- Instance families: General (t3, m5), Compute (c5), Memory (r5), GPU (p3, g4)
- Auto Scaling: Target tracking, step scaling, scheduled scaling
- Placement groups: Cluster, partition, spread
- Best practices: Use latest generation, right-size, enable detailed monitoring
Lambda
- Invocation models: Synchronous, asynchronous, event source mapping
- Concurrency: Reserved, provisioned, burst limits
- Layers for shared dependencies
- Best practices: Keep functions small, use environment variables, set timeouts
ECS/EKS (Container Services)
- ECS: Fargate for serverless, EC2 for control
- EKS: Managed Kubernetes with AWS integration
- Service mesh: App Mesh for observability
- Best practices: Use Fargate for simplicity, EKS for portability
Elastic Beanstalk
- Managed platform for web apps
- Auto-scaling and load balancing included
- Support for multiple languages and Docker
Storage
S3 (Simple Storage Service)
- Storage classes: Standard, IA, One Zone-IA, Glacier, Deep Archive
- Lifecycle policies for automatic tiering
- Versioning and MFA delete for protection
- Cross-region replication for DR
- Best practices: Enable versioning, use lifecycle policies, block public access
EBS (Elastic Block Store)
- Volume types: gp3 (general), io2 (IOPS), st1 (throughput), sc1 (cold)
- Snapshots to S3 for backup
- Encryption by default
- Best practices: Use gp3 for most workloads, enable encryption
EFS (Elastic File System)
- NFSv4 file system for shared access
- Performance modes: General purpose, Max I/O
- Throughput modes: Bursting, provisioned
- Best practices: Use lifecycle management, enable encryption
FSx
- FSx for Windows File Server (SMB)
- FSx for Lustre (HPC workloads)
- FSx for NetApp ONTAP
- FSx for OpenZFS
Database
RDS (Relational Database Service)
- Engines: MySQL, PostgreSQL, MariaDB, Oracle, SQL Server, Aurora
- Multi-AZ for high availability
- Read replicas for scalability
- Automated backups and point-in-time recovery
- Best practices: Use Aurora for performance, enable Multi-AZ, use read replicas
Aurora
- MySQL and PostgreSQL compatible
- 5x MySQL, 3x PostgreSQL performance
- Global databases for cross-region DR
- Serverless v2 for variable workloads
- Best practices: Use Aurora Serverless for unpredictable workloads
DynamoDB
- NoSQL key-value and document database
- On-demand or provisioned capacity
- Global tables for multi-region replication
- DynamoDB Streams for change data capture
- Best practices: Use on-demand for unpredictable traffic, implement GSI carefully
ElastiCache
- Redis or Memcached in-memory caching
- Cluster mode for Redis scalability
- Best practices: Use for session storage, API caching
Networking
VPC (Virtual Private Cloud)
- CIDR planning: Avoid overlaps, plan for growth
- Subnets: Public (IGW), private (NAT), isolated (no internet)
- Route tables and routing decisions
- Security groups (stateful) and NACLs (stateless)
- Best practices: Use /16 for VPC, /24 for subnets, plan IP space
Route 53
- DNS service with health checks
- Routing policies: Simple, weighted, latency, failover, geolocation
- Best practices: Use alias records, enable DNSSEC
CloudFront
- Global CDN with edge locations
- Origin types: S3, ALB, custom origins
- Lambda@Edge for request/response manipulation
- Best practices: Enable compression, use field-level encryption
VPN and Direct Connect
- Site-to-Site VPN for encrypted tunnels
- Direct Connect for dedicated bandwidth
- Transit Gateway for hub-and-spoke topology
- Best practices: Use Direct Connect for high bandwidth, Transit Gateway for complex routing
API Gateway
- REST APIs, HTTP APIs, WebSocket APIs
- Throttling and quotas
- Integration with Lambda, HTTP endpoints, AWS services
- Best practices: Use HTTP APIs for lower cost, implement caching
Security
IAM (Identity and Access Management)
- Principle of least privilege
- Roles for applications, not access keys
- MFA for privileged users
- Service Control Policies (SCPs) for organization-wide controls
- Best practices: Use roles, enable MFA, rotate credentials
KMS (Key Management Service)
- Customer managed keys (CMKs)
- Automatic key rotation
- Envelope encryption pattern
- Best practices: Enable automatic rotation, use grants for temporary access
Secrets Manager
- Automatic rotation for RDS credentials
- Versioning and rollback
- Best practices: Rotate secrets regularly, use VPC endpoints
Security Hub
- Centralized security findings
- CIS AWS Foundations Benchmark
- Integration with GuardDuty, Inspector, Macie
GuardDuty
- Threat detection using ML
- Monitors CloudTrail, VPC Flow Logs, DNS logs
Architecture Patterns
High Availability
Multi-AZ Pattern
- Application Load Balancer across 3 AZs
- Auto Scaling group with instances in each AZ
- RDS Multi-AZ for database
- S3 for static assets (11 9's durability)
Multi-Region Pattern
- Route 53 with health checks and failover
- CloudFront for global distribution
- Aurora Global Database for <1s RPO
- S3 Cross-Region Replication
Serverless Architecture
API-Driven Pattern
API Gateway -> Lambda -> DynamoDB
|
v
EventBridge -> Lambda (async processing)
Event-Driven Pattern
S3 Event -> Lambda -> Process -> SNS
|
v
Multiple subscribers
Microservices on AWS
Container-Based
ALB -> ECS Fargate (multiple services)
|
v
Service Discovery (Cloud Map)
|
v
RDS/DynamoDB per service
Service Mesh
App Mesh for traffic management
X-Ray for distributed tracing
CloudWatch Container Insights
Data Lake Architecture
Data Sources -> Kinesis Data Streams
|
v
Kinesis Firehose
|
v
S3 (raw bucket)
|
v
Glue ETL or Lambda processing
|
v
S3 (processed bucket)
|
v
Athena/Redshift Spectrum
|
v
QuickSight dashboards
Migration Strategies (6Rs)
Rehost (Lift-and-Shift)
- AWS Application Migration Service (MGN)
- Minimal changes, quick migration
- Use for legacy apps with compliance constraints
Replatform (Lift-Tinker-and-Shift)
- Migrate to RDS instead of self-managed databases
- Use Elastic Beanstalk instead of custom app servers
- Small optimizations during migration
Repurchase (Drop-and-Shop)
- Move to SaaS (e.g., Salesforce, Workday)
- Reduce maintenance burden
Refactor/Re-architect
- Modernize to serverless or containers
- Highest effort, highest benefit
- Use for competitive advantage applications
Retire
- Decommission unused applications
- Reduce attack surface and costs
Retain
- Keep on-premises temporarily
- Migrate later or keep for regulatory reasons
Landing Zone Design
AWS Control Tower
- Multi-account strategy (AWS Organizations)
- Account factory for standardization
- Guardrails for governance (SCPs)
- Centralized logging (CloudTrail, Config)
Account Structure
Root
├── Security OU
│ ├── Log Archive Account
│ └── Security Tooling Account
├── Infrastructure OU
│ ├── Network Account (Transit Gateway, VPN)
│ └── Shared Services Account
└── Workloads OU
├── Production Account
├── Staging Account
└── Development Account
Network Design
Transit Gateway (hub)
|
├── Production VPC
├── Staging VPC
├── Development VPC
└── On-premises (Direct Connect/VPN)
Cost Optimization Strategies
Compute Savings
- Compute Savings Plans (up to 66% savings)
- EC2 Reserved Instances (1-year or 3-year)
- Spot Instances for batch/fault-tolerant workloads
- Lambda: Reduce memory if possible, use reserved concurrency
Storage Savings
- S3 Intelligent-Tiering for unpredictable access
- Lifecycle policies to Glacier/Deep Archive
- EBS gp3 instead of gp2 (20% cheaper, better performance)
- Delete unused snapshots and volumes
Database Savings
- Aurora Serverless v2 for variable workloads
- RDS Reserved Instances
- DynamoDB on-demand for unpredictable workloads
- Read replicas in same region to reduce cross-AZ data transfer
Monitoring and Alerting
- AWS Cost Explorer for analysis
- AWS Budgets for alerts
- Cost Anomaly Detection
- Trusted Advisor for recommendations
Disaster Recovery
RPO and RTO Targets
- Backup and Restore: Hours RPO/RTO (lowest cost)
- Pilot Light: Minutes RPO, hours RTO
- Warm Standby: Seconds RPO, minutes RTO
- Multi-Site Active/Active: Near-zero RPO/RTO (highest cost)
Implementation
- AWS Backup for centralized backup management
- Aurora Global Database for cross-region replication
- S3 Cross-Region Replication
- Route 53 health checks and failover routing
- Regular DR testing with CloudFormation/Terraform
Monitoring and Observability
CloudWatch
- Metrics: Standard (5 min) and detailed (1 min)
- Alarms with SNS notifications
- Logs Insights for log analysis
- Dashboards for visualization
X-Ray
- Distributed tracing for microservices
- Service map visualization
- Trace annotations and metadata
AWS Config
- Resource inventory and change tracking
- Compliance rules evaluation
- Relationship tracking between resources
1---2name: aws-architecture-reference3description: Comprehensive guide for AWS services, patterns, and Well-Architected Framework implementation.4---5# AWS Architecture Reference67Comprehensive guide for AWS services, patterns, and Well-Architected Framework implementation.89## Well-Architected Framework1011### Six Pillars12131. **Operational Excellence**14 - Infrastructure as Code (CloudFormation, CDK, Terraform)15 - Continuous integration/deployment16 - Observability (CloudWatch, X-Ray)17 - Runbooks and playbooks18 - Game days and failure injection19202. **Security**21 - Identity and Access Management (IAM)22 - Detective controls (GuardDuty, Security Hub)23 - Infrastructure protection (VPC, security groups, NACLs)24 - Data protection (KMS, encryption at rest/transit)25 - Incident response automation26273. **Reliability**28 - Multi-AZ deployments29 - Auto Scaling groups30 - Route 53 health checks and failover31 - Backup and restore (AWS Backup)32 - Chaos engineering (AWS FIS)33344. **Performance Efficiency**35 - Right-sizing with Compute Optimizer36 - Caching strategies (CloudFront, ElastiCache)37 - Database optimization (RDS Performance Insights)38 - Serverless architectures39 - Global content delivery40415. **Cost Optimization**42 - Reserved Instances and Savings Plans43 - Spot Instances for fault-tolerant workloads44 - S3 Intelligent-Tiering and lifecycle policies45 - Right-sizing recommendations46 - Cost allocation tags and budgets47486. **Sustainability**49 - Region selection for renewable energy50 - Serverless to minimize idle resources51 - Efficient data storage patterns52 - Resource utilization optimization5354## Core Services Architecture5556### Compute5758**EC2 (Elastic Compute Cloud)**59- Instance families: General (t3, m5), Compute (c5), Memory (r5), GPU (p3, g4)60- Auto Scaling: Target tracking, step scaling, scheduled scaling61- Placement groups: Cluster, partition, spread62- Best practices: Use latest generation, right-size, enable detailed monitoring6364**Lambda**65- Invocation models: Synchronous, asynchronous, event source mapping66- Concurrency: Reserved, provisioned, burst limits67- Layers for shared dependencies68- Best practices: Keep functions small, use environment variables, set timeouts6970**ECS/EKS (Container Services)**71- ECS: Fargate for serverless, EC2 for control72- EKS: Managed Kubernetes with AWS integration73- Service mesh: App Mesh for observability74- Best practices: Use Fargate for simplicity, EKS for portability7576**Elastic Beanstalk**77- Managed platform for web apps78- Auto-scaling and load balancing included79- Support for multiple languages and Docker8081### Storage8283**S3 (Simple Storage Service)**84- Storage classes: Standard, IA, One Zone-IA, Glacier, Deep Archive85- Lifecycle policies for automatic tiering86- Versioning and MFA delete for protection87- Cross-region replication for DR88- Best practices: Enable versioning, use lifecycle policies, block public access8990**EBS (Elastic Block Store)**91- Volume types: gp3 (general), io2 (IOPS), st1 (throughput), sc1 (cold)92- Snapshots to S3 for backup93- Encryption by default94- Best practices: Use gp3 for most workloads, enable encryption9596**EFS (Elastic File System)**97- NFSv4 file system for shared access98- Performance modes: General purpose, Max I/O99- Throughput modes: Bursting, provisioned100- Best practices: Use lifecycle management, enable encryption101102**FSx**103- FSx for Windows File Server (SMB)104- FSx for Lustre (HPC workloads)105- FSx for NetApp ONTAP106- FSx for OpenZFS107108### Database109110**RDS (Relational Database Service)**111- Engines: MySQL, PostgreSQL, MariaDB, Oracle, SQL Server, Aurora112- Multi-AZ for high availability113- Read replicas for scalability114- Automated backups and point-in-time recovery115- Best practices: Use Aurora for performance, enable Multi-AZ, use read replicas116117**Aurora**118- MySQL and PostgreSQL compatible119- 5x MySQL, 3x PostgreSQL performance120- Global databases for cross-region DR121- Serverless v2 for variable workloads122- Best practices: Use Aurora Serverless for unpredictable workloads123124**DynamoDB**125- NoSQL key-value and document database126- On-demand or provisioned capacity127- Global tables for multi-region replication128- DynamoDB Streams for change data capture129- Best practices: Use on-demand for unpredictable traffic, implement GSI carefully130131**ElastiCache**132- Redis or Memcached in-memory caching133- Cluster mode for Redis scalability134- Best practices: Use for session storage, API caching135136### Networking137138**VPC (Virtual Private Cloud)**139- CIDR planning: Avoid overlaps, plan for growth140- Subnets: Public (IGW), private (NAT), isolated (no internet)141- Route tables and routing decisions142- Security groups (stateful) and NACLs (stateless)143- Best practices: Use /16 for VPC, /24 for subnets, plan IP space144145**Route 53**146- DNS service with health checks147- Routing policies: Simple, weighted, latency, failover, geolocation148- Best practices: Use alias records, enable DNSSEC149150**CloudFront**151- Global CDN with edge locations152- Origin types: S3, ALB, custom origins153- Lambda@Edge for request/response manipulation154- Best practices: Enable compression, use field-level encryption155156**VPN and Direct Connect**157- Site-to-Site VPN for encrypted tunnels158- Direct Connect for dedicated bandwidth159- Transit Gateway for hub-and-spoke topology160- Best practices: Use Direct Connect for high bandwidth, Transit Gateway for complex routing161162**API Gateway**163- REST APIs, HTTP APIs, WebSocket APIs164- Throttling and quotas165- Integration with Lambda, HTTP endpoints, AWS services166- Best practices: Use HTTP APIs for lower cost, implement caching167168### Security169170**IAM (Identity and Access Management)**171- Principle of least privilege172- Roles for applications, not access keys173- MFA for privileged users174- Service Control Policies (SCPs) for organization-wide controls175- Best practices: Use roles, enable MFA, rotate credentials176177**KMS (Key Management Service)**178- Customer managed keys (CMKs)179- Automatic key rotation180- Envelope encryption pattern181- Best practices: Enable automatic rotation, use grants for temporary access182183**Secrets Manager**184- Automatic rotation for RDS credentials185- Versioning and rollback186- Best practices: Rotate secrets regularly, use VPC endpoints187188**Security Hub**189- Centralized security findings190- CIS AWS Foundations Benchmark191- Integration with GuardDuty, Inspector, Macie192193**GuardDuty**194- Threat detection using ML195- Monitors CloudTrail, VPC Flow Logs, DNS logs196197## Architecture Patterns198199### High Availability200201**Multi-AZ Pattern**202```203- Application Load Balancer across 3 AZs204- Auto Scaling group with instances in each AZ205- RDS Multi-AZ for database206- S3 for static assets (11 9's durability)207```208209**Multi-Region Pattern**210```211- Route 53 with health checks and failover212- CloudFront for global distribution213- Aurora Global Database for <1s RPO214- S3 Cross-Region Replication215```216217### Serverless Architecture218219**API-Driven Pattern**220```221API Gateway -> Lambda -> DynamoDB222 |223 v224 EventBridge -> Lambda (async processing)225```226227**Event-Driven Pattern**228```229S3 Event -> Lambda -> Process -> SNS230 |231 v232 Multiple subscribers233```234235### Microservices on AWS236237**Container-Based**238```239ALB -> ECS Fargate (multiple services)240 |241 v242Service Discovery (Cloud Map)243 |244 v245RDS/DynamoDB per service246```247248**Service Mesh**249```250App Mesh for traffic management251X-Ray for distributed tracing252CloudWatch Container Insights253```254255### Data Lake Architecture256257```258Data Sources -> Kinesis Data Streams259 |260 v261 Kinesis Firehose262 |263 v264 S3 (raw bucket)265 |266 v267 Glue ETL or Lambda processing268 |269 v270 S3 (processed bucket)271 |272 v273 Athena/Redshift Spectrum274 |275 v276 QuickSight dashboards277```278279## Migration Strategies (6Rs)2802811. **Rehost (Lift-and-Shift)**282 - AWS Application Migration Service (MGN)283 - Minimal changes, quick migration284 - Use for legacy apps with compliance constraints2852862. **Replatform (Lift-Tinker-and-Shift)**287 - Migrate to RDS instead of self-managed databases288 - Use Elastic Beanstalk instead of custom app servers289 - Small optimizations during migration2902913. **Repurchase (Drop-and-Shop)**292 - Move to SaaS (e.g., Salesforce, Workday)293 - Reduce maintenance burden2942954. **Refactor/Re-architect**296 - Modernize to serverless or containers297 - Highest effort, highest benefit298 - Use for competitive advantage applications2993005. **Retire**301 - Decommission unused applications302 - Reduce attack surface and costs3033046. **Retain**305 - Keep on-premises temporarily306 - Migrate later or keep for regulatory reasons307308## Landing Zone Design309310**AWS Control Tower**311- Multi-account strategy (AWS Organizations)312- Account factory for standardization313- Guardrails for governance (SCPs)314- Centralized logging (CloudTrail, Config)315316**Account Structure**317```318Root319├── Security OU320│ ├── Log Archive Account321│ └── Security Tooling Account322├── Infrastructure OU323│ ├── Network Account (Transit Gateway, VPN)324│ └── Shared Services Account325└── Workloads OU326 ├── Production Account327 ├── Staging Account328 └── Development Account329```330331**Network Design**332```333Transit Gateway (hub)334 |335 ├── Production VPC336 ├── Staging VPC337 ├── Development VPC338 └── On-premises (Direct Connect/VPN)339```340341## Cost Optimization Strategies342343**Compute Savings**344- Compute Savings Plans (up to 66% savings)345- EC2 Reserved Instances (1-year or 3-year)346- Spot Instances for batch/fault-tolerant workloads347- Lambda: Reduce memory if possible, use reserved concurrency348349**Storage Savings**350- S3 Intelligent-Tiering for unpredictable access351- Lifecycle policies to Glacier/Deep Archive352- EBS gp3 instead of gp2 (20% cheaper, better performance)353- Delete unused snapshots and volumes354355**Database Savings**356- Aurora Serverless v2 for variable workloads357- RDS Reserved Instances358- DynamoDB on-demand for unpredictable workloads359- Read replicas in same region to reduce cross-AZ data transfer360361**Monitoring and Alerting**362- AWS Cost Explorer for analysis363- AWS Budgets for alerts364- Cost Anomaly Detection365- Trusted Advisor for recommendations366367## Disaster Recovery368369**RPO and RTO Targets**370- Backup and Restore: Hours RPO/RTO (lowest cost)371- Pilot Light: Minutes RPO, hours RTO372- Warm Standby: Seconds RPO, minutes RTO373- Multi-Site Active/Active: Near-zero RPO/RTO (highest cost)374375**Implementation**376- AWS Backup for centralized backup management377- Aurora Global Database for cross-region replication378- S3 Cross-Region Replication379- Route 53 health checks and failover routing380- Regular DR testing with CloudFormation/Terraform381382## Monitoring and Observability383384**CloudWatch**385- Metrics: Standard (5 min) and detailed (1 min)386- Alarms with SNS notifications387- Logs Insights for log analysis388- Dashboards for visualization389390**X-Ray**391- Distributed tracing for microservices392- Service map visualization393- Trace annotations and metadata394395**AWS Config**396- Resource inventory and change tracking397- Compliance rules evaluation398- Relationship tracking between resources