Architecture Discipline
When to Use
Use when decisions affect:
- Data models or schema changes
- Service boundaries or new services
- Deployment topology or infrastructure
- Scale characteristics (10x growth implications)
- Technology stack choices
- API contracts or external integrations
Skip for:
- Parameter tweaks or configuration changes
- UI/styling changes
- Localization or copy changes
- Bug fixes within existing architecture
- Adding fields to existing models (unless schema migration)
Threshold: If the decision could cause a 3+ month re-architecture project if wrong, use this skill.
CRITICAL: This Is Reasoning Discipline, Not a Checklist
The 7 sections are a reasoning sequence, not boxes to check:
- Alternatives BEFORE choosing
- Scale requirements BEFORE design
- Failure modes built in, not bolted on
🚨 If you wrote architecture without starting with all 7 sections: DELETE and restart. Retrofitting analysis is rationalization, not evaluation.
MANDATORY FIRST STEP
CREATE TodoWrite with these 7 sections (22+ items total):
| Section |
Minimum Items |
| Scale Analysis |
4+ |
| Architectural Options |
3+ |
| Ripple Effect Analysis |
5+ |
| Failure Modes |
3+ |
| Observability |
3+ |
| Documentation |
2+ |
| Migration/Compatibility |
2+ |
Do not design, propose solutions, or implement until TodoWrite is verified.
Verification Checkpoint
After creating TodoWrite, verify 3 random items pass this test:
Each item must have ALL THREE:
- ✓ Concrete numbers/thresholds ("100K users", "$500/mo", "P95 < 500ms")
- ✓ Specific tools/technologies ("PostgreSQL", "Redis", "CloudWatch")
- ✓ Measurable outcome ("handles 1M req/sec", "costs $X at 10x")
| ❌ FAILS |
✅ PASSES |
| "Add monitoring" |
"CloudWatch: websocket.connections.active, alert if >5% error rate via PagerDuty" |
| "Evaluate caching" |
"Compare: Redis (1ms, $300/mo) vs In-memory LRU (0.1ms, $0) vs No cache (100ms)" |
| "Analyze scale" |
"Current: 100K DAU, 50 req/sec. 10x: 1M users, 500 req/sec. Bottleneck: PostgreSQL connection pool" |
DO NOT PROCEED until 22+ items AND quality check passes.
Section Requirements
1. Scale Analysis (4+ items)
NEVER design for current scale only. Before proposing any solution:
- Current scale: Users (DAU/MAU), requests/sec, data volume, read/write ratio
- 10x scale: What numbers at 10x? When expected?
- Bottlenecks: What breaks at 10x? (DB connections, API limits, memory)
- Mitigation: Specific solution for each bottleneck
2. Architectural Options (3+ items)
NEVER present single solution. Minimum 3 distinct options, each with:
- Performance: Latency (P50/P95/P99), throughput, scale limit
- Complexity: LOC estimate, services involved, operational burden
- Cost: Infrastructure ($X/mo current, $Y/mo at 10x), development (engineer-weeks)
- Trade-offs: Specific advantages (✅) and disadvantages (❌)
If stakeholder suggests solution: Add as Option A, evaluate with SAME rigor as alternatives.
3. Ripple Effect Analysis (5+ items)
Changes propagate across layers. Analyze ALL:
- Data layer: Schema changes, migrations, indexes, query performance
- Services: Which need updates? API contracts changed?
- API: Breaking changes? Version bump? Backward compatibility?
- Clients: Mobile updates? Web UI changes?
- Operations: Deployment changes? New monitoring? Cost changes?
4. Failure Modes (3+ items)
For each mode:
- Scenario: [Component] fails because [reason]
- Detection: How we know (metrics drop, error rate spike)
- Impact: What breaks (user features, data integrity)
- Mitigation: Circuit breaker, fallback, redundancy
5. Observability (3+ items)
- Metrics: Specific (latency P95, error rate %, throughput)
- Alerts: Conditions (error rate > 5%, latency P95 > 500ms)
- Dashboards: Key visualizations
6. Documentation (2+ items)
- ADR: Chosen option, rejected alternatives, trade-offs, constraints
- Diagram: New components, data flows, failure paths
7. Migration/Compatibility (2+ items)
- Backward compatibility: Old clients work? API versioning?
- Migration path: Phased rollout, feature flags, rollback procedure
Red Flags - STOP When You Think:
| Thought |
Reality |
| "Analysis paralysis" |
This IS the analysis that prevents expensive mistakes |
| "We'll add scale/alternatives/failure modes later" |
Retrofitting costs 5-10x more |
| "CTO already decided" |
Still needs independent evaluation |
| "Being pragmatic not dogmatic" |
These requirements ARE pragmatic |
| "Just a simple feature" |
Simple becomes complex at scale |
| "We already know the solution" |
Compare 3 alternatives first |
| "Keep it simple" |
Simple for current scale = complex re-architecture at 10x |
| "I can add missing sections to existing work" |
DELETE and restart |
Override Requirements
To skip ANY requirement, you MUST provide ALL 4:
- Specific retrofit date (not "later")
- Budget allocated (engineer-weeks)
- Risk acceptance signed by decision maker
- Interim mitigation plan
| Skipped |
Risk |
Cost |
| Scale Analysis |
Re-architecture in 6-12 months |
3-6 month project, 5-10x cost |
| Alternatives |
Optimize wrong dimension |
2-4 month migration |
| Failure Modes |
Production incidents |
$5-50K per incident |
| Ripple Effects |
Broken clients, data issues |
Deployment failures |
Verification Before Complete
| Category |
Requirements |
| Scale |
✓ Current + 10x projected ✓ Bottlenecks ✓ Mitigations |
| Trade-offs |
✓ 3+ options ✓ Performance/complexity/cost ✓ Rationale |
| Impact |
✓ All layers analyzed ✓ Breaking changes identified |
| Failure |
✓ Specific modes ✓ Detection ✓ Mitigation ✓ Rollback |
| Documentation |
✓ ADR ✓ Diagram updated |
If any item missing, do not proceed to implementation.
1---2name: architecture-discipline3description: Architecture Discipline4---56# Architecture Discipline78## When to Use910**Use when decisions affect:**11- Data models or schema changes12- Service boundaries or new services13- Deployment topology or infrastructure14- Scale characteristics (10x growth implications)15- Technology stack choices16- API contracts or external integrations1718**Skip for:**19- Parameter tweaks or configuration changes20- UI/styling changes21- Localization or copy changes22- Bug fixes within existing architecture23- Adding fields to existing models (unless schema migration)2425**Threshold:** If the decision could cause a 3+ month re-architecture project if wrong, use this skill.2627## CRITICAL: This Is Reasoning Discipline, Not a Checklist2829The 7 sections are a **reasoning sequence**, not boxes to check:30- Alternatives BEFORE choosing31- Scale requirements BEFORE design32- Failure modes built in, not bolted on3334🚨 **If you wrote architecture without starting with all 7 sections:** DELETE and restart. Retrofitting analysis is rationalization, not evaluation.3536---3738## MANDATORY FIRST STEP3940**CREATE TodoWrite** with these 7 sections (22+ items total):4142| Section | Minimum Items |43|---------|---------------|44| Scale Analysis | 4+ |45| Architectural Options | 3+ |46| Ripple Effect Analysis | 5+ |47| Failure Modes | 3+ |48| Observability | 3+ |49| Documentation | 2+ |50| Migration/Compatibility | 2+ |5152**Do not design, propose solutions, or implement until TodoWrite is verified.**5354---5556## Verification Checkpoint5758After creating TodoWrite, verify 3 random items pass this test:5960**Each item must have ALL THREE:**61- ✓ Concrete numbers/thresholds ("100K users", "$500/mo", "P95 < 500ms")62- ✓ Specific tools/technologies ("PostgreSQL", "Redis", "CloudWatch")63- ✓ Measurable outcome ("handles 1M req/sec", "costs $X at 10x")6465| ❌ FAILS | ✅ PASSES |66|----------|-----------|67| "Add monitoring" | "CloudWatch: `websocket.connections.active`, alert if >5% error rate via PagerDuty" |68| "Evaluate caching" | "Compare: Redis (1ms, $300/mo) vs In-memory LRU (0.1ms, $0) vs No cache (100ms)" |69| "Analyze scale" | "Current: 100K DAU, 50 req/sec. 10x: 1M users, 500 req/sec. Bottleneck: PostgreSQL connection pool" |7071**DO NOT PROCEED until 22+ items AND quality check passes.**7273---7475## Section Requirements7677### 1. Scale Analysis (4+ items)7879**NEVER design for current scale only.** Before proposing any solution:80- Current scale: Users (DAU/MAU), requests/sec, data volume, read/write ratio81- 10x scale: What numbers at 10x? When expected?82- Bottlenecks: What breaks at 10x? (DB connections, API limits, memory)83- Mitigation: Specific solution for each bottleneck8485### 2. Architectural Options (3+ items)8687**NEVER present single solution.** Minimum 3 distinct options, each with:88- Performance: Latency (P50/P95/P99), throughput, scale limit89- Complexity: LOC estimate, services involved, operational burden90- Cost: Infrastructure ($X/mo current, $Y/mo at 10x), development (engineer-weeks)91- Trade-offs: Specific advantages (✅) and disadvantages (❌)9293**If stakeholder suggests solution:** Add as Option A, evaluate with SAME rigor as alternatives.9495### 3. Ripple Effect Analysis (5+ items)9697Changes propagate across layers. Analyze ALL:98- Data layer: Schema changes, migrations, indexes, query performance99- Services: Which need updates? API contracts changed?100- API: Breaking changes? Version bump? Backward compatibility?101- Clients: Mobile updates? Web UI changes?102- Operations: Deployment changes? New monitoring? Cost changes?103104### 4. Failure Modes (3+ items)105106For each mode:107- Scenario: [Component] fails because [reason]108- Detection: How we know (metrics drop, error rate spike)109- Impact: What breaks (user features, data integrity)110- Mitigation: Circuit breaker, fallback, redundancy111112### 5. Observability (3+ items)113114- Metrics: Specific (latency P95, error rate %, throughput)115- Alerts: Conditions (error rate > 5%, latency P95 > 500ms)116- Dashboards: Key visualizations117118### 6. Documentation (2+ items)119120- ADR: Chosen option, rejected alternatives, trade-offs, constraints121- Diagram: New components, data flows, failure paths122123### 7. Migration/Compatibility (2+ items)124125- Backward compatibility: Old clients work? API versioning?126- Migration path: Phased rollout, feature flags, rollback procedure127128---129130## Red Flags - STOP When You Think:131132| Thought | Reality |133|---------|---------|134| "Analysis paralysis" | This IS the analysis that prevents expensive mistakes |135| "We'll add scale/alternatives/failure modes later" | Retrofitting costs 5-10x more |136| "CTO already decided" | Still needs independent evaluation |137| "Being pragmatic not dogmatic" | These requirements ARE pragmatic |138| "Just a simple feature" | Simple becomes complex at scale |139| "We already know the solution" | Compare 3 alternatives first |140| "Keep it simple" | Simple for current scale = complex re-architecture at 10x |141| "I can add missing sections to existing work" | DELETE and restart |142143---144145## Override Requirements146147To skip ANY requirement, you MUST provide ALL 4:1481. Specific retrofit date (not "later")1492. Budget allocated (engineer-weeks)1503. Risk acceptance signed by decision maker1514. Interim mitigation plan152153| Skipped | Risk | Cost |154|---------|------|------|155| Scale Analysis | Re-architecture in 6-12 months | 3-6 month project, 5-10x cost |156| Alternatives | Optimize wrong dimension | 2-4 month migration |157| Failure Modes | Production incidents | $5-50K per incident |158| Ripple Effects | Broken clients, data issues | Deployment failures |159160---161162## Verification Before Complete163164| Category | Requirements |165|----------|-------------|166| Scale | ✓ Current + 10x projected ✓ Bottlenecks ✓ Mitigations |167| Trade-offs | ✓ 3+ options ✓ Performance/complexity/cost ✓ Rationale |168| Impact | ✓ All layers analyzed ✓ Breaking changes identified |169| Failure | ✓ Specific modes ✓ Detection ✓ Mitigation ✓ Rollback |170| Documentation | ✓ ADR ✓ Diagram updated |171172**If any item missing, do not proceed to implementation.**