1---2name: cloud-solution-architect3description: Design well-architected Azure cloud systems — 10 design principles, 6 architecture styles, 44 design patterns, technology choices, WAF pillars4---5
6# Cloud Solution Architect
7
8
9> Design well-architected, production-grade cloud systems following Azure Architecture Center best practices
10
11---
12
13## Ten Design Principles for Azure Applications
14
15| # | Principle | Key Tactics |
16|---|-----------|-------------|
17| 1 | **Design for self-healing** | Retry with backoff, circuit breaker, bulkhead isolation, health endpoint monitoring, graceful degradation |
18| 2 | **Make all things redundant** | Eliminate single points of failure, use availability zones, deploy multi-region, replicate data |
19| 3 | **Minimize coordination** | Decouple services, use async messaging, embrace eventual consistency, use domain events |
20| 4 | **Design to scale out** | Horizontal scaling, autoscaling rules, stateless services, avoid session stickiness, partition workloads |
21| 5 | **Partition around limits** | Data partitioning (shard/hash/range), respect compute & network limits, use CDNs for static content |
22| 6 | **Design for operations** | Structured logging, distributed tracing, metrics & dashboards, runbook automation, infrastructure as code |
23| 7 | **Use managed services** | Prefer PaaS over IaaS, reduce operational burden, leverage built-in HA/DR/scaling |
24| 8 | **Use an identity service** | Microsoft Entra ID, managed identity, RBAC, avoid storing credentials, zero-trust principles |
25| 9 | **Design for evolution** | Loose coupling, versioned APIs, backward compatibility, async messaging for integration, feature flags |
26| 10 | **Build for business needs** | Define SLAs/SLOs, establish RTO/RPO targets, domain-driven design, cost modeling, composite SLAs |
27
28---
29
30## Architecture Styles
31
32| Style | Description | When to Use | Key Services |
33|-------|-------------|-------------|--------------|
34| **N-tier** | Horizontal layers (presentation, business, data) | Traditional enterprise apps, lift-and-shift | App Service, SQL Database, VNets |
35| **Web-Queue-Worker** | Web frontend → message queue → backend worker | Moderate-complexity apps with long-running tasks | App Service, Service Bus, Functions |
36| **Microservices** | Small autonomous services, bounded contexts, independent deploy | Complex domains, independent team scaling | AKS, Container Apps, API Management |
37| **Event-driven** | Pub/sub model, event producers/consumers | Real-time processing, IoT, reactive systems | Event Hubs, Event Grid, Functions |
38| **Big data** | Batch + stream processing pipeline | Analytics, ML pipelines, large-scale data | Synapse, Data Factory, Databricks |
39| **Big compute** | HPC, parallel processing | Simulations, modeling, rendering, genomics | Batch, CycleCloud, HPC VMs |
40
41### Selection Criteria
42
43- **Domain complexity** → Microservices (high), N-tier (low-medium)
44- **Team autonomy** → Microservices (independent teams), N-tier (single team)
45- **Data volume** → Big data (TB+), others (GB)
46- **Latency requirements** → Event-driven (real-time), Web-Queue-Worker (tolerant)
47
48---
49
50## Cloud Design Patterns (44 Patterns)
51
52WAF pillar mapping: **R**=Reliability, **S**=Security, **CO**=Cost Optimization, **OE**=Operational Excellence, **PE**=Performance Efficiency.
53
54### Messaging & Communication
55
56| Pattern | Summary | Pillars |
57|---------|---------|---------|
58| **Asynchronous Request-Reply** | Decouple request/response with polling or callbacks | R, PE |
59| **Claim Check** | Split large messages; store payload separately, pass reference | R, PE |
60| **Choreography** | Services coordinate via events without central orchestrator | R, OE |
61| **Competing Consumers** | Multiple consumers process messages from shared queue concurrently | R, PE |
62| **Pipes and Filters** | Decompose complex processing into reusable filter stages | R, OE |
63| **Priority Queue** | Prioritize requests so higher-priority work is processed first | R, PE |
64| **Publisher/Subscriber** | Decouple senders from receivers via topics/subscriptions | R, PE |
65| **Queue-Based Load Leveling** | Buffer requests with a queue to smooth intermittent loads | R, PE |
66| **Sequential Convoy** | Process related messages in order while allowing parallel groups | R, PE |
67
68### Reliability & Resilience
69
70| Pattern | Summary | Pillars |
71|---------|---------|---------|
72| **Bulkhead** | Isolate resources per workload to prevent cascading failure | R |
73| **Circuit Breaker** | Stop calling a failing service; fail fast to protect resources | R |
74| **Compensating Transaction** | Undo previously committed steps when a later step fails | R |
75| **Health Endpoint Monitoring** | Expose health checks for load balancers and orchestrators | R, OE |
76| **Leader Election** | Coordinate distributed instances by electing a leader | R |
77| **Retry** | Handle transient faults by retrying with exponential backoff | R |
78| **Saga** | Manage data consistency across microservices with compensating transactions | R |
79| **Scheduler Agent Supervisor** | Coordinate distributed actions with retry and failure handling | R |
80
81### Data Management
82
83| Pattern | Summary | Pillars |
84|---------|---------|---------|
85| **Cache-Aside** | Load data on demand into cache from data store | PE |
86| **CQRS** | Separate read and write models for independent scaling | PE, R |
87| **Event Sourcing** | Store state as append-only sequence of domain events | R, OE |
88| **Index Table** | Create indexes over frequently queried fields in data stores | PE |
89| **Materialized View** | Pre-compute views over data for efficient queries | PE |
90| **Sharding** | Distribute data across partitions for scale and performance | PE, R |
91| **Static Content Hosting** | Serve static content from cloud storage/CDN directly | PE, CO |
92| **Valet Key** | Grant clients limited direct access to storage resources | S, PE |
93
94### Design & Structure
95
96| Pattern | Summary | Pillars |
97|---------|---------|---------|
98| **Ambassador** | Offload cross-cutting concerns to a helper sidecar proxy | OE |
99| **Anti-Corruption Layer** | Translate between new and legacy system models | OE, R |
100| **Backends for Frontends** | Create separate backends per frontend type (mobile, web, etc.) | OE, PE |
101| **Compute Resource Consolidation** | Combine multiple workloads into fewer compute instances | CO |
102| **External Configuration Store** | Externalize configuration from deployment packages | OE |
103| **Sidecar** | Deploy helper components alongside the main service | OE |
104| **Strangler Fig** | Incrementally migrate legacy systems by replacing pieces | OE, R |
105
106### Security & Access
107
108| Pattern | Summary | Pillars |
109|---------|---------|---------|
110| **Federated Identity** | Delegate authentication to an external identity provider | S |
111| **Gatekeeper** | Protect services using a dedicated broker that validates requests | S |
112| **Quarantine** | Isolate and validate external assets before allowing use | S |
113| **Rate Limiting** | Control consumption rate of resources by consumers | R, S |
114| **Throttling** | Control resource consumption to sustain SLAs under load | R, PE |
115
116### Deployment & Scaling
117
118| Pattern | Summary | Pillars |
119|---------|---------|---------|
120| **Deployment Stamps** | Deploy multiple independent copies of application components | R, PE |
121| **Gateway Aggregation** | Aggregate multiple backend calls into a single client request | PE |
122| **Gateway Offloading** | Offload shared functionality (SSL, auth) to a gateway | OE, S |
123| **Gateway Routing** | Route requests to multiple backends using a single endpoint | OE |
124| **Geode** | Deploy backends to multiple regions for active-active serving | R, PE |
125
126---
127
128## Technology Choices
129
130### Decision Framework
131
132For each technology area, evaluate: **requirements → constraints → tradeoffs → select**.
133
134| Area | Key Options | Selection Criteria |
135|------|-------------|-------------------|
136| **Compute** | App Service, Functions, Container Apps, AKS, VMs, Batch | Hosting model, scaling, cost, team skills |
137| **Storage** | Blob Storage, Data Lake, Files, Disks, Managed Lustre | Access patterns, throughput, cost tier |
138| **Data stores** | SQL Database, Cosmos DB, PostgreSQL, Redis, Table Storage | Consistency model, query patterns, scale |
139| **Messaging** | Service Bus, Event Hubs, Event Grid, Queue Storage | Ordering, throughput, pub/sub vs queue |
140| **Networking** | Front Door, Application Gateway, Load Balancer, Traffic Manager | Global vs regional, L4 vs L7, WAF |
141| **AI services** | Azure OpenAI, AI Search, AI Foundry, Document Intelligence | Model needs, data grounding, orchestration |
142| **Containers** | Container Apps, AKS, Container Instances | Operational control vs simplicity |
143
144---
145
146## Well-Architected Framework (WAF) Pillars
147
148| Pillar | Focus | Key Questions |
149|--------|-------|---------------|
150| **Reliability** | Resiliency, availability, disaster recovery | What is the RTO/RPO? How does it handle failures? Is there redundancy? |
151| **Security** | Threat protection, identity, data protection | Is identity managed? Is data encrypted? Are there network controls? |
152| **Cost Optimization** | Cost management, efficiency, right-sizing | Is compute right-sized? Are there reserved instances? Is there waste? |
153| **Operational Excellence** | Monitoring, deployment, automation | Is deployment automated? Is there observability? Are there runbooks? |
154| **Performance Efficiency** | Scaling, load testing, performance targets | Can it scale horizontally? Are there performance baselines? Is caching used? |
155
156### WAF Tradeoff Matrix
157
158| Optimizing for... | May impact... |
159|-------------------|---------------|
160| Reliability (redundancy) | Cost (more resources) |
161| Security (isolation) | Performance (added latency) |
162| Cost (consolidation) | Reliability (shared failure domains) |
163| Performance (caching) | Cost (cache infrastructure), Reliability (stale data) |
164
165---
166
167## Performance Antipatterns
168
169| Antipattern | Problem | Fix |
170|-------------|---------|-----|
171| **Busy Database** | Offloading too much processing to the database | Move logic to application tier, use caching |
172| **Busy Front End** | Resource-intensive work on frontend request threads | Offload to background workers/queues |
173| **Chatty I/O** | Many small I/O requests instead of fewer large ones | Batch requests, use bulk APIs, buffer writes |
174| **Extraneous Fetching** | Retrieving more data than needed | Project only required fields, paginate, filter server-side |
175| **Improper Instantiation** | Recreating expensive objects per request | Use singletons, connection pooling, HttpClientFactory |
176| **Monolithic Persistence** | Single data store for all data types | Polyglot persistence — right store for each workload |
177| **No Caching** | Repeatedly fetching unchanged data | Cache-aside pattern, CDN, output caching, Redis |
178| **Noisy Neighbor** | One tenant consuming all shared resources | Bulkhead isolation, per-tenant quotas, throttling |
179| **Retry Storm** | Aggressive retries overwhelming a recovering service | Exponential backoff + jitter, circuit breaker, retry budgets |
180| **Synchronous I/O** | Blocking threads on I/O operations | Async/await, non-blocking I/O, reactive streams |
181
182---
183
184## Mission-Critical Design (99.99%+ SLO)
185
186| Design Area | Key Considerations |
187|-------------|-------------------|
188| **Application platform** | Multi-region active-active, availability zones, Container Apps or AKS with zone redundancy |
189| **Application design** | Stateless services, idempotent operations, graceful degradation, bulkhead isolation |
190| **Networking** | Azure Front Door (global LB), DDoS Protection, private endpoints, redundant connectivity |
191| **Data platform** | Multi-region Cosmos DB, zone-redundant SQL, async replication, conflict resolution |
192| **Deployment & testing** | Blue-green deployments, canary releases, chaos engineering, automated rollback |
193| **Health modeling** | Composite health scores, dependency health tracking, automated remediation, SLI dashboards |
194| **Security** | Zero-trust, managed identity everywhere, key rotation, WAF policies, threat modeling |
195| **Operational procedures** | Automated runbooks, incident response playbooks, game days, postmortems |
196
197---
198
199## Architecture Review Workflow
200
201### Step 1: Identify Requirements
202
203```text
204Functional: What must the system do?
205Non-functional:
206 - Availability target (e.g., 99.9%, 99.99%)
207 - Latency requirements (p50, p95, p99)
208 - Throughput (requests/sec, messages/sec)
209 - Data residency and compliance
210 - Recovery targets (RTO, RPO)
211 - Cost constraints
212```
213
214### Step 2: Select Architecture Style
215
216Match requirements to architecture style using the selection criteria table above.
217
218### Step 3: Choose Technology Stack
219
220Use the technology choices decision framework. Prefer managed services (PaaS) over IaaS.
221
222### Step 4: Apply Design Patterns
223
224Select relevant patterns from the 44 cloud design patterns based on identified concerns.
225
226### Step 5: Address Cross-Cutting Concerns
227
228- **Identity & access** — Microsoft Entra ID, managed identity, RBAC
229- **Monitoring** — Application Insights, Azure Monitor, Log Analytics
230- **Security** — Network segmentation, encryption at rest/in transit, Key Vault
231- **CI/CD** — GitHub Actions, Azure DevOps Pipelines, infrastructure as code
232
233### Step 6: Validate Against WAF Pillars
234
235Review each pillar systematically. Document tradeoffs explicitly.
236
237### Step 7: Document Decisions
238
239Use Architecture Decision Records (ADRs):
240
241```markdown
242# ADR-NNN: [Decision Title]
243
244## Status: [Proposed | Accepted | Deprecated]
245
246## Context
247[What is the issue we're addressing?]
248
249## Decision
250[What did we decide and why?]
251
252## Consequences
253[What are the positive and negative impacts?]
254```
255
256---
257
258## Related Skills
259
260This skill complements:
261
262- **azure-architecture-patterns** — Detailed Azure-specific patterns
263- **bicep-avm-mastery** — Infrastructure as code for Azure
264- **security-review** — STRIDE, threat modeling
265- **observability-monitoring** — Production visibility
266
267---
268
269## References
270
271- [Azure Architecture Center](https://learn.microsoft.com/en-us/azure/architecture/)
272- [Well-Architected Framework](https://learn.microsoft.com/en-us/azure/well-architected/)
273- [Cloud Design Patterns](https://learn.microsoft.com/en-us/azure/architecture/patterns/)
274- [Mission-Critical Workloads](https://learn.microsoft.com/en-us/azure/architecture/framework/mission-critical/)