Legacy System Modernization Engine
Complete methodology for assessing, planning, and executing legacy system modernization — from monolith decomposition to cloud migration. Works for any tech stack, any scale.
Phase 1: System Assessment
Modernization Brief
system_name: "[Name]"
age_years: 0
primary_language: ""
framework: ""
database: ""
deployment: "on-prem | VM | container | serverless"
lines_of_code: 0
team_size: 0
monthly_users: 0
annual_revenue_supported: "$0"
compliance_requirements: []
known_pain_points: []
business_driver: "cost | speed | talent | risk | compliance | scale"
timeline_pressure: "low | medium | high | critical"
budget_range: "$0-$0"
sponsor: ""
Technical Debt Inventory
Score each dimension 1-5 (1=critical, 5=healthy):
| Dimension |
Score |
Evidence |
| Code quality — test coverage, complexity, duplication |
|
|
| Architecture — coupling, modularity, clear boundaries |
|
|
| Infrastructure — deployment automation, monitoring, scaling |
|
|
| Dependencies — outdated libraries, EOL frameworks, security vulns |
|
|
| Data — schema quality, migration history, backup/recovery |
|
|
| Documentation — accuracy, coverage, onboarding effectiveness |
|
|
| Operations — deployment frequency, MTTR, incident rate |
|
|
| Security — auth patterns, encryption, audit trail, compliance gaps |
|
|
| Developer experience — build time, local setup, debugging tools |
|
|
| Business logic clarity — documented rules, test coverage of logic |
|
|
Total: /50
- 40-50: Healthy — incremental improvement
- 30-39: Aging — targeted modernization
- 20-29: Legacy — systematic modernization needed
- 10-19: Critical — modernize or replace
Dependency Risk Matrix
For each major dependency:
dependency: ""
current_version: ""
latest_version: ""
eol_date: "" # End of life
security_vulns: 0 # Known CVEs
upgrade_difficulty: "trivial | moderate | hard | rewrite"
business_risk: "low | medium | high | critical"
alternatives: []
Priority rules:
- EOL within 12 months → P0
- Known unpatched CVEs → P0
- 3+ major versions behind → P1
- No active maintainer → P1
- Everything else → P2
Phase 2: Strategy Selection
Modernization Strategy Decision Matrix
| Strategy |
When to Use |
Risk |
Cost |
Speed |
Disruption |
| Rehost (lift & shift) |
Datacenter exit, minimal change |
Low |
Low |
Fast |
Low |
| Replatform (lift & optimize) |
Cloud benefits without rewrite |
Low-Med |
Medium |
Medium |
Low-Med |
| Refactor (restructure) |
Good code, bad architecture |
Medium |
Medium |
Medium |
Medium |
| Re-architect (rebuild patterns) |
Monolith→services, new patterns |
High |
High |
Slow |
High |
| Rebuild (rewrite) |
Small system, clear requirements |
Very High |
Very High |
Very Slow |
Very High |
| Replace (buy/SaaS) |
Commodity functionality |
Medium |
Variable |
Fast |
High |
| Retire |
No longer needed |
None |
Negative |
Instant |
Low |
| Retain (do nothing) |
Working fine, other priorities |
None |
Ongoing |
N/A |
None |
Strategy Selection Decision Tree
Is the system still needed?
├─ No → RETIRE
├─ Yes → Is it a commodity (CRM, email, etc.)?
│ ├─ Yes → REPLACE (buy SaaS)
│ └─ No → Is the code maintainable?
│ ├─ Yes → Is the architecture the problem?
│ │ ├─ Yes → RE-ARCHITECT (strangler fig)
│ │ └─ No → Is the infrastructure the problem?
│ │ ├─ Yes → REPLATFORM
│ │ └─ No → REFACTOR incrementally
│ └─ No → Is the system small (<50K LOC)?
│ ├─ Yes → Can requirements be clearly defined?
│ │ ├─ Yes → REBUILD
│ │ └─ No → REFACTOR + RE-ARCHITECT
│ └─ No → STRANGLER FIG (never big-bang rewrite)
The Big Rewrite Anti-Pattern
NEVER do a full rewrite of a large system. It fails 70%+ of the time because:
- The old system keeps getting features — moving target
- Hidden business rules only exist in code — they get lost
- Timeline always 2-3x longer than estimated
- Two systems to maintain during transition
- Team burns out before completion
Always use Strangler Fig instead. Replace piece by piece.
Phase 3: Strangler Fig Pattern
How It Works
- Identify a boundary — a feature, page, or API endpoint
- Build the replacement — new stack, new patterns
- Route traffic — proxy/facade sends requests to new or old
- Verify parity — same behavior, same data
- Cut over — remove the proxy, retire the old code
- Repeat — next boundary
Strangler Facade YAML
facade_name: "[API Gateway / Reverse Proxy / BFF]"
routing_rules:
- path: "/api/users/*"
target: "new-service"
status: "migrated"
migrated_date: "2025-01-15"
- path: "/api/orders/*"
target: "legacy"
status: "planned"
target_date: "2025-Q2"
- path: "/api/reports/*"
target: "legacy"
status: "not-planned"
notes: "Low priority, rarely used"
Migration Sequence Rules
- Start with the easiest, most isolated module — build confidence
- Then the highest-value business capability — prove ROI early
- Leave the hardest, most coupled parts for last — team learns patterns first
- Never migrate auth/identity early — it touches everything
- Migrate data access layer before business logic — clean data = clean migration
- Always keep the old system as fallback until new is proven
Dual-Write / Data Sync Patterns
| Pattern |
When |
Complexity |
Risk |
| Dual write |
Both systems write simultaneously |
High |
Data inconsistency |
| CDC (Change Data Capture) |
Stream changes from old→new DB |
Medium |
Lag, ordering |
| ETL batch sync |
Periodic bulk sync |
Low |
Stale data |
| Event sourcing bridge |
Events from old, replay in new |
High |
Schema mapping |
| Read from new, write to old |
Transition period |
Medium |
Routing complexity |
Golden rule: Pick ONE source of truth. Never let both systems own the same data simultaneously.
Phase 4: Monolith Decomposition
Domain Discovery
Before splitting a monolith, identify bounded contexts:
- Event Storming (preferred) — sticky notes for domain events, commands, aggregates
- Code analysis — find clusters of related classes/tables
- Team analysis — which teams own which features?
- Data coupling analysis — which tables are joined together?
Bounded Context YAML
context_name: ""
description: ""
team: ""
entities: []
commands: []
events_published: []
events_consumed: []
database_tables: []
external_integrations: []
coupling_score: 0 # 0=independent, 10=deeply coupled
extraction_difficulty: "easy | moderate | hard | very-hard"
business_value: "low | medium | high | critical"
Extraction Priority Matrix
Plot contexts on: Business Value (Y) × Extraction Difficulty (X)
|
Easy |
Moderate |
Hard |
| High value |
🟢 Do first |
🟡 Do second |
🟠 Plan carefully |
| Medium value |
🟢 Quick win |
🟡 Evaluate ROI |
🔴 Probably not worth it |
| Low value |
🟡 If easy, why not |
🔴 Skip |
🔴 Definitely skip |
Service Extraction Checklist
For each service being extracted:
Phase 5: Database Modernization
Database Migration Strategies
| Strategy |
Description |
Downtime |
Risk |
| Parallel run |
New DB alongside old, sync both |
Zero |
High complexity |
| Blue-green |
Full copy, switch DNS |
Minutes |
Medium |
| Rolling |
Migrate table by table |
Zero per table |
Medium |
| Big bang |
Stop, migrate, start |
Hours |
High |
Schema Evolution Rules
- Always additive — add columns/tables, never remove in the same release
- Two-phase removal — Release 1: stop writing. Release 2: drop column (after backfill verified)
- Default values always — every new column gets a default
- Backward compatible — old code must work with new schema during rollout
- Index concurrently — never lock production tables
- Test with production-scale data — 100 rows ≠ 100M rows
Data Quality Gates
Before migrating data:
table: ""
row_count_source: 0
row_count_target: 0
count_match: false
checksum_match: false
null_analysis: "pass | fail"
referential_integrity: "pass | fail"
business_rule_validation: "pass | fail"
sample_manual_review: "pass | fail"
performance_benchmark: "pass | fail"
rollback_tested: false
Rule: All gates must pass before cutover. No exceptions.
Phase 6: Cloud Migration
Cloud Readiness Assessment
Score each workload:
| Factor |
Score (1-5) |
Notes |
| Stateless design |
|
|
| Configuration externalized |
|
|
| Logging to stdout |
|
|
| Health check endpoint |
|
|
| Graceful shutdown |
|
|
| Horizontal scalability |
|
|
| Secret management |
|
|
| 12-factor compliance |
|
|
35-40: Cloud-native ready
25-34: Minor modifications needed
15-24: Significant refactoring
8-14: Major redesign required
Cloud Migration Checklist
Cost Optimization from Day 1
- Right-size instances — start small, scale up with data
- Reserved/committed use — only after 3 months of usage data
- Spot/preemptible — for batch jobs, CI/CD, dev/test
- Auto-scaling — scale down at night, weekends
- Storage tiers — hot/warm/cold/archive based on access patterns
- Tag everything — cost allocation by team, service, environment
- Monthly review — unused resources, oversized instances
Phase 7: API Modernization
API Wrapping Pattern
For legacy systems without APIs:
- Screen scraping adapter — parse HTML/mainframe screens
- Database tap — read directly from legacy DB (read-only!)
- File-based integration — watch folders, parse files
- Message queue bridge — legacy writes to queue, new reads
- RPC wrapper — expose existing functions via REST/gRPC
API Contract-First Migration
endpoint: "/api/v2/orders"
legacy_source: "stored_procedure: sp_GetOrders"
new_implementation: "orders-service"
migration_status: "legacy | dual-run | new-only"
contract_changes:
- field: "order_date"
old_format: "MM/DD/YYYY string"
new_format: "ISO 8601"
adapter: "date_format_adapter"
- field: "status"
old_values: ["A", "C", "P"]
new_values: ["active", "completed", "pending"]
adapter: "status_code_mapper"
parity_tests: 47
parity_passing: 47
Phase 8: Testing Strategy
Migration Testing Pyramid
/ Smoke Tests \ ← Whole system alive?
/ Parity Tests \ ← Same behavior old vs new?
/ Integration Tests \ ← Services work together?
/ Contract Tests \ ← API contracts honored?
/ Performance Tests \ ← Not slower than before?
/ Data Validation Tests \ ← Data migrated correctly?
/ Unit Tests \ ← New code works?
Parity Testing Framework
For EVERY migrated feature:
feature: ""
test_type: "api_parity | ui_parity | data_parity"
method: "shadow traffic | replay | parallel run"
sample_size: 0
match_rate: "0%" # Target: 99.9%+
mismatches_investigated: 0
mismatches_accepted: 0 # Known intentional differences
mismatches_bugs: 0
sign_off: false
Shadow traffic — copy production requests to new system, compare responses (don't serve new responses to users yet).
Performance Regression Rules
- P95 latency must not increase >10% vs legacy
- Throughput must meet or exceed legacy under same load
- Database query count must not increase per request
- Memory usage must not increase >20%
- If ANY metric regresses → investigate before proceeding
Phase 9: Team & Process
Modernization Team Structure
| Role |
Responsibility |
When Needed |
| Modernization Lead |
Strategy, sequencing, blockers |
Full-time |
| Legacy Expert |
Knows where the bodies are buried |
Part-time, on-call |
| New Platform Engineer |
Builds target architecture |
Full-time |
| Data Engineer |
Migration, sync, validation |
Phase-dependent |
| QA/Test Engineer |
Parity testing, automation |
Full-time |
| DevOps/Platform |
CI/CD, infrastructure |
Part-time |
| Product Owner |
Business priority, acceptance |
Part-time |
Knowledge Mining from Legacy
The most dangerous part of modernization is losing undocumented business rules.
- Code archaeology — git blame, find oldest unchanged code, understand why
- Interview stakeholders — "What would break if we changed X?"
- Production log analysis — what edge cases actually occur?
- Error handling review — each catch block is a documented business rule
- Test suite review — tests describe expected behavior
- Configuration review — magic numbers, feature flags, overrides
Communication Plan
| Audience |
Frequency |
Content |
| Executive sponsor |
Bi-weekly |
Progress, risks, budget, timeline |
| Engineering team |
Weekly |
Sprint goals, technical decisions, blockers |
| Dependent teams |
Monthly |
Upcoming changes, migration dates, API changes |
| End users |
Per migration |
What's changing, when, how it affects them |
Phase 10: Risk Management
Top 10 Modernization Risks (Pre-Built)
| # |
Risk |
Likelihood |
Impact |
Mitigation |
| 1 |
Undocumented business rules lost |
High |
Critical |
Code archaeology + stakeholder interviews + parity tests |
| 2 |
Timeline underestimation |
Very High |
High |
2x initial estimate, phase-gated checkpoints |
| 3 |
Data migration corruption |
Medium |
Critical |
Checksums, parallel runs, rollback plans |
| 4 |
Feature parity gaps |
High |
High |
Shadow traffic testing, user acceptance testing |
| 5 |
Team knowledge loss (people leave) |
Medium |
High |
Document everything, pair programming, knowledge sharing |
| 6 |
Legacy system changes during migration |
High |
Medium |
Feature freeze or dual-write contract |
| 7 |
Performance regression |
Medium |
High |
Load testing at every phase, performance budgets |
| 8 |
Scope creep (improve while migrating) |
Very High |
Medium |
Strict "migrate, don't improve" rule for Phase 1 |
| 9 |
Integration failures |
Medium |
High |
Contract testing, circuit breakers, fallback routing |
| 10 |
Stakeholder fatigue |
High |
Medium |
Quick wins early, visible progress dashboard |
Kill Criteria
Stop the modernization if:
- Budget exceeds 2x initial estimate with <50% complete
- Key business rules can't be verified after migration
- Team attrition >30% during project
- Legacy system stability degrades due to migration work
- Business context changes (M&A, pivot, sunset)
If kill criteria triggered: Stabilize what's done, document learnings, reassess in 6 months.
Phase 11: Patterns & Playbooks
Language/Framework Migration Patterns
Java → Modern Java (8→17+)
- Records, sealed classes, pattern matching
- Virtual threads (Project Loom) for thread-per-request
- Migrate build: Maven→Gradle or update Maven plugins
- Spring Boot 2→3: javax→jakarta namespace
Python 2→3
- Use
2to3 tool for automated conversion
- Fix: print(), division, unicode, dict methods
- Upgrade dependencies (check py3 compat)
jQuery→React/Vue
- Extract components from page sections
- State management replaces DOM manipulation
- Event handlers become component methods
- Ajax calls become API service layer
Monolith→Microservices
- Strangler fig (see Phase 3)
- Start with read models (reporting, search)
- Extract stateless services first
- Shared database → database-per-service last
On-Prem→Cloud
- Rehost first (lift & shift)
- Then replatform (managed services)
- Then re-architect (cloud-native patterns)
- Never skip steps — each proves value
COBOL/Mainframe Modernization
- API wrapping — expose CICS/IMS transactions as REST APIs
- Screen scraping — automate 3270 terminal interactions
- Gradual extraction — one transaction at a time
- Data replication — DB2/VSAM → PostgreSQL/cloud DB
- Rule extraction — COBOL paragraphs → business rule engine
- Never rewrite all at once — decades of business logic = decades of edge cases
Microservices Anti-Patterns to Avoid
| Anti-Pattern |
Symptom |
Fix |
| Distributed monolith |
Services must deploy together |
Identify and break coupling |
| Shared database |
Multiple services write same tables |
Database-per-service |
| Synchronous chains |
A calls B calls C calls D |
Async events, choreography |
| Nano-services |
Hundreds of tiny services |
Merge related services |
| Shared libraries for business logic |
Library update breaks consumers |
Duplicate code > shared coupling |
| No API versioning |
Breaking changes cascade |
Semantic versioning, deprecation policy |
Phase 12: Metrics & Reporting
Modernization Health Dashboard
project: ""
assessment_date: ""
overall_health: "green | yellow | red"
progress:
modules_total: 0
modules_migrated: 0
modules_in_progress: 0
percent_complete: "0%"
velocity:
modules_per_sprint: 0
estimated_completion: ""
on_track: true
quality:
parity_test_pass_rate: "0%"
production_incidents_from_migration: 0
rollbacks: 0
risk:
open_risks: 0
p0_risks: 0
blocked_items: 0
cost:
budget_total: "$0"
budget_spent: "$0"
budget_remaining: "$0"
burn_rate_monthly: "$0"
100-Point Modernization Quality Rubric
| Dimension |
Weight |
Score (0-10) |
Weighted |
| Strategy clarity |
15% |
|
|
| Risk management |
15% |
|
|
| Testing rigor |
15% |
|
|
| Data integrity |
15% |
|
|
| Architecture quality |
10% |
|
|
| Team capability |
10% |
|
|
| Stakeholder alignment |
10% |
|
|
| Documentation |
10% |
|
|
| Total |
100% |
|
/100 |
90-100: Exemplary — reference project
70-89: Strong — minor improvements
50-69: Adequate — address gaps
Below 50: At risk — pause and reassess
Weekly Status Template
## Modernization Status — Week of [DATE]
### Progress
- Modules migrated this week: [N]
- Total migrated: [N]/[TOTAL] ([X]%)
- On track for [TARGET DATE]: [Yes/No]
### Completed
- [What shipped this week]
### In Progress
- [What's being worked on]
### Blockers
- [What's stuck and what's needed]
### Risks
- [New or changed risks]
### Next Week
- [Plan for next sprint]
Edge Cases
"We need to modernize but can't stop adding features"
- Strangler fig — modernize around new features
- Feature freeze on legacy module ONLY when that module is being migrated
- New features build in new stack from day 1
"We don't know what the system does"
- Start with observability: instrument logging, tracing, metrics
- Run for 2-4 weeks to understand actual usage patterns
- Code coverage analysis shows what code is actually executed
- Interview longest-tenured team members
"Multiple systems need modernizing simultaneously"
- Sequence by dependency order — downstream first
- Shared services (auth, data) get modernized once, reused
- Never parallelize more than 2 modernization streams
"The original developers are gone"
- Treat code as the documentation
- Invest 2-4 weeks in code archaeology before any migration work
- Pair new developers with business stakeholders
- Write tests for existing behavior before changing anything
"We're being acquired / merging systems"
- Map overlapping functionality first
- Pick "winner" system per domain — don't merge codebases
- API integration layer between systems
- 18-month realistic timeline for full consolidation
"Compliance requires the old system"
- Maintain compliance evidence chain during migration
- Dual-audit period with both systems
- Get compliance team involved in migration planning from Day 1
- Document control mapping: old control → new implementation
Natural Language Commands
| Command |
Action |
| "Assess this system for modernization" |
Run full Technical Debt Inventory |
| "Which modernization strategy should we use?" |
Walk through Strategy Decision Tree |
| "Plan a strangler fig migration" |
Generate Strangler Facade YAML + sequence |
| "Decompose this monolith" |
Domain discovery + Bounded Context mapping |
| "Migrate this database" |
Data Quality Gates + migration strategy |
| "Check cloud readiness" |
Run Cloud Readiness Assessment |
| "Create a migration testing plan" |
Build Testing Pyramid with parity tests |
| "What are the risks?" |
Generate Top 10 risk register |
| "How do we migrate from [X] to [Y]?" |
Pattern-specific playbook |
| "Status update for modernization" |
Generate Weekly Status Template |
| "Score this modernization project" |
Run 100-Point Quality Rubric |
| "Should we kill this modernization?" |
Evaluate Kill Criteria |
1---2name: legacy-system-modernization-engine3description: Complete methodology for assessing, planning, and executing legacy system modernization — from monolith decomposition to cloud migration. Works for any tech stack, any scale.4---5
6# Legacy System Modernization Engine
7
8Complete methodology for assessing, planning, and executing legacy system modernization — from monolith decomposition to cloud migration. Works for any tech stack, any scale.
9
10---
11
12## Phase 1: System Assessment
13
14### Modernization Brief
15
16```yaml
17system_name: "[Name]"
18age_years: 0
19primary_language: ""
20framework: ""
21database: ""
22deployment: "on-prem | VM | container | serverless"
23lines_of_code: 0
24team_size: 0
25monthly_users: 0
26annual_revenue_supported: "$0"
27compliance_requirements: []
28known_pain_points: []
29business_driver: "cost | speed | talent | risk | compliance | scale"
30timeline_pressure: "low | medium | high | critical"
31budget_range: "$0-$0"
32sponsor: ""
33```
34
35### Technical Debt Inventory
36
37Score each dimension 1-5 (1=critical, 5=healthy):
38
39| Dimension | Score | Evidence |
40|-----------|-------|----------|
41| **Code quality** — test coverage, complexity, duplication | | |
42| **Architecture** — coupling, modularity, clear boundaries | | |
43| **Infrastructure** — deployment automation, monitoring, scaling | | |
44| **Dependencies** — outdated libraries, EOL frameworks, security vulns | | |
45| **Data** — schema quality, migration history, backup/recovery | | |
46| **Documentation** — accuracy, coverage, onboarding effectiveness | | |
47| **Operations** — deployment frequency, MTTR, incident rate | | |
48| **Security** — auth patterns, encryption, audit trail, compliance gaps | | |
49| **Developer experience** — build time, local setup, debugging tools | | |
50| **Business logic clarity** — documented rules, test coverage of logic | | |
51
52**Total: /50**
53
54- **40-50**: Healthy — incremental improvement
55- **30-39**: Aging — targeted modernization
56- **20-29**: Legacy — systematic modernization needed
57- **10-19**: Critical — modernize or replace
58
59### Dependency Risk Matrix
60
61For each major dependency:
62
63```yaml
64dependency: ""
65current_version: ""
66latest_version: ""
67eol_date: "" # End of life
68security_vulns: 0 # Known CVEs
69upgrade_difficulty: "trivial | moderate | hard | rewrite"
70business_risk: "low | medium | high | critical"
71alternatives: []
72```
73
74**Priority rules:**
75- EOL within 12 months → P0
76- Known unpatched CVEs → P0
77- 3+ major versions behind → P1
78- No active maintainer → P1
79- Everything else → P2
80
81---
82
83## Phase 2: Strategy Selection
84
85### Modernization Strategy Decision Matrix
86
87| Strategy | When to Use | Risk | Cost | Speed | Disruption |
88|----------|-------------|------|------|-------|------------|
89| **Rehost** (lift & shift) | Datacenter exit, minimal change | Low | Low | Fast | Low |
90| **Replatform** (lift & optimize) | Cloud benefits without rewrite | Low-Med | Medium | Medium | Low-Med |
91| **Refactor** (restructure) | Good code, bad architecture | Medium | Medium | Medium | Medium |
92| **Re-architect** (rebuild patterns) | Monolith→services, new patterns | High | High | Slow | High |
93| **Rebuild** (rewrite) | Small system, clear requirements | Very High | Very High | Very Slow | Very High |
94| **Replace** (buy/SaaS) | Commodity functionality | Medium | Variable | Fast | High |
95| **Retire** | No longer needed | None | Negative | Instant | Low |
96| **Retain** (do nothing) | Working fine, other priorities | None | Ongoing | N/A | None |
97
98### Strategy Selection Decision Tree
99
100```
101Is the system still needed?
102├─ No → RETIRE
103├─ Yes → Is it a commodity (CRM, email, etc.)?
104│ ├─ Yes → REPLACE (buy SaaS)
105│ └─ No → Is the code maintainable?
106│ ├─ Yes → Is the architecture the problem?
107│ │ ├─ Yes → RE-ARCHITECT (strangler fig)
108│ │ └─ No → Is the infrastructure the problem?
109│ │ ├─ Yes → REPLATFORM
110│ │ └─ No → REFACTOR incrementally
111│ └─ No → Is the system small (<50K LOC)?
112│ ├─ Yes → Can requirements be clearly defined?
113│ │ ├─ Yes → REBUILD
114│ │ └─ No → REFACTOR + RE-ARCHITECT
115│ └─ No → STRANGLER FIG (never big-bang rewrite)
116```
117
118### The Big Rewrite Anti-Pattern
119
120**NEVER do a full rewrite of a large system.** It fails 70%+ of the time because:
1211. The old system keeps getting features — moving target
1222. Hidden business rules only exist in code — they get lost
1233. Timeline always 2-3x longer than estimated
1244. Two systems to maintain during transition
1255. Team burns out before completion
126
127**Always use Strangler Fig instead.** Replace piece by piece.
128
129---
130
131## Phase 3: Strangler Fig Pattern
132
133### How It Works
134
1351. **Identify a boundary** — a feature, page, or API endpoint
1362. **Build the replacement** — new stack, new patterns
1373. **Route traffic** — proxy/facade sends requests to new or old
1384. **Verify parity** — same behavior, same data
1395. **Cut over** — remove the proxy, retire the old code
1406. **Repeat** — next boundary
141
142### Strangler Facade YAML
143
144```yaml
145facade_name: "[API Gateway / Reverse Proxy / BFF]"
146routing_rules:
147 - path: "/api/users/*"
148 target: "new-service"
149 status: "migrated"
150 migrated_date: "2025-01-15"
151 - path: "/api/orders/*"
152 target: "legacy"
153 status: "planned"
154 target_date: "2025-Q2"
155 - path: "/api/reports/*"
156 target: "legacy"
157 status: "not-planned"
158 notes: "Low priority, rarely used"
159```
160
161### Migration Sequence Rules
162
1631. **Start with the easiest, most isolated module** — build confidence
1642. **Then the highest-value business capability** — prove ROI early
1653. **Leave the hardest, most coupled parts for last** — team learns patterns first
1664. **Never migrate auth/identity early** — it touches everything
1675. **Migrate data access layer before business logic** — clean data = clean migration
1686. **Always keep the old system as fallback** until new is proven
169
170### Dual-Write / Data Sync Patterns
171
172| Pattern | When | Complexity | Risk |
173|---------|------|-----------|------|
174| **Dual write** | Both systems write simultaneously | High | Data inconsistency |
175| **CDC (Change Data Capture)** | Stream changes from old→new DB | Medium | Lag, ordering |
176| **ETL batch sync** | Periodic bulk sync | Low | Stale data |
177| **Event sourcing bridge** | Events from old, replay in new | High | Schema mapping |
178| **Read from new, write to old** | Transition period | Medium | Routing complexity |
179
180**Golden rule:** Pick ONE source of truth. Never let both systems own the same data simultaneously.
181
182---
183
184## Phase 4: Monolith Decomposition
185
186### Domain Discovery
187
188Before splitting a monolith, identify bounded contexts:
189
1901. **Event Storming** (preferred) — sticky notes for domain events, commands, aggregates
1912. **Code analysis** — find clusters of related classes/tables
1923. **Team analysis** — which teams own which features?
1934. **Data coupling analysis** — which tables are joined together?
194
195### Bounded Context YAML
196
197```yaml
198context_name: ""
199description: ""
200team: ""
201entities: []
202commands: []
203events_published: []
204events_consumed: []
205database_tables: []
206external_integrations: []
207coupling_score: 0 # 0=independent, 10=deeply coupled
208extraction_difficulty: "easy | moderate | hard | very-hard"
209business_value: "low | medium | high | critical"
210```
211
212### Extraction Priority Matrix
213
214Plot contexts on: **Business Value** (Y) × **Extraction Difficulty** (X)
215
216| | Easy | Moderate | Hard |
217|---|---|---|---|
218| **High value** | 🟢 Do first | 🟡 Do second | 🟠 Plan carefully |
219| **Medium value** | 🟢 Quick win | 🟡 Evaluate ROI | 🔴 Probably not worth it |
220| **Low value** | 🟡 If easy, why not | 🔴 Skip | 🔴 Definitely skip |
221
222### Service Extraction Checklist
223
224For each service being extracted:
225
226- [ ] Bounded context clearly defined
227- [ ] API contract designed (OpenAPI spec)
228- [ ] Database separated (no shared tables)
229- [ ] Authentication/authorization integrated
230- [ ] Event publishing for cross-service communication
231- [ ] Circuit breaker for calls back to monolith
232- [ ] Monitoring and alerting configured
233- [ ] Deployment pipeline independent
234- [ ] Feature flag for traffic routing
235- [ ] Rollback plan documented
236- [ ] Performance baseline captured (before/after)
237- [ ] Data migration script tested
238- [ ] Integration tests with monolith passing
239- [ ] Runbook for on-call written
240
241---
242
243## Phase 5: Database Modernization
244
245### Database Migration Strategies
246
247| Strategy | Description | Downtime | Risk |
248|----------|-------------|----------|------|
249| **Parallel run** | New DB alongside old, sync both | Zero | High complexity |
250| **Blue-green** | Full copy, switch DNS | Minutes | Medium |
251| **Rolling** | Migrate table by table | Zero per table | Medium |
252| **Big bang** | Stop, migrate, start | Hours | High |
253
254### Schema Evolution Rules
255
2561. **Always additive** — add columns/tables, never remove in the same release
2572. **Two-phase removal** — Release 1: stop writing. Release 2: drop column (after backfill verified)
2583. **Default values always** — every new column gets a default
2594. **Backward compatible** — old code must work with new schema during rollout
2605. **Index concurrently** — never lock production tables
2616. **Test with production-scale data** — 100 rows ≠ 100M rows
262
263### Data Quality Gates
264
265Before migrating data:
266
267```yaml
268table: ""
269row_count_source: 0
270row_count_target: 0
271count_match: false
272checksum_match: false
273null_analysis: "pass | fail"
274referential_integrity: "pass | fail"
275business_rule_validation: "pass | fail"
276sample_manual_review: "pass | fail"
277performance_benchmark: "pass | fail"
278rollback_tested: false
279```
280
281**Rule: All gates must pass before cutover.** No exceptions.
282
283---
284
285## Phase 6: Cloud Migration
286
287### Cloud Readiness Assessment
288
289Score each workload:
290
291| Factor | Score (1-5) | Notes |
292|--------|-------------|-------|
293| Stateless design | | |
294| Configuration externalized | | |
295| Logging to stdout | | |
296| Health check endpoint | | |
297| Graceful shutdown | | |
298| Horizontal scalability | | |
299| Secret management | | |
300| 12-factor compliance | | |
301
302**35-40**: Cloud-native ready
303**25-34**: Minor modifications needed
304**15-24**: Significant refactoring
305**8-14**: Major redesign required
306
307### Cloud Migration Checklist
308
309- [ ] Network architecture designed (VPC, subnets, security groups)
310- [ ] Identity and access management configured
311- [ ] Data residency requirements verified
312- [ ] Compliance mapping (cloud controls ↔ requirements)
313- [ ] Cost estimation completed (TCO comparison)
314- [ ] Disaster recovery plan updated
315- [ ] Monitoring and alerting migrated
316- [ ] DNS and certificate management planned
317- [ ] CDN configuration
318- [ ] Load testing in cloud environment
319- [ ] Security scanning pipeline
320- [ ] Backup and restore verified
321- [ ] Runbooks updated for cloud operations
322- [ ] Team trained on cloud platform
323- [ ] Vendor lock-in assessment
324
325### Cost Optimization from Day 1
326
327- **Right-size instances** — start small, scale up with data
328- **Reserved/committed use** — only after 3 months of usage data
329- **Spot/preemptible** — for batch jobs, CI/CD, dev/test
330- **Auto-scaling** — scale down at night, weekends
331- **Storage tiers** — hot/warm/cold/archive based on access patterns
332- **Tag everything** — cost allocation by team, service, environment
333- **Monthly review** — unused resources, oversized instances
334
335---
336
337## Phase 7: API Modernization
338
339### API Wrapping Pattern
340
341For legacy systems without APIs:
342
3431. **Screen scraping adapter** — parse HTML/mainframe screens
3442. **Database tap** — read directly from legacy DB (read-only!)
3453. **File-based integration** — watch folders, parse files
3464. **Message queue bridge** — legacy writes to queue, new reads
3475. **RPC wrapper** — expose existing functions via REST/gRPC
348
349### API Contract-First Migration
350
351```yaml
352endpoint: "/api/v2/orders"
353legacy_source: "stored_procedure: sp_GetOrders"
354new_implementation: "orders-service"
355migration_status: "legacy | dual-run | new-only"
356contract_changes:
357 - field: "order_date"
358 old_format: "MM/DD/YYYY string"
359 new_format: "ISO 8601"
360 adapter: "date_format_adapter"
361 - field: "status"
362 old_values: ["A", "C", "P"]
363 new_values: ["active", "completed", "pending"]
364 adapter: "status_code_mapper"
365parity_tests: 47
366parity_passing: 47
367```
368
369---
370
371## Phase 8: Testing Strategy
372
373### Migration Testing Pyramid
374
375```
376 / Smoke Tests \ ← Whole system alive?
377 / Parity Tests \ ← Same behavior old vs new?
378 / Integration Tests \ ← Services work together?
379 / Contract Tests \ ← API contracts honored?
380 / Performance Tests \ ← Not slower than before?
381 / Data Validation Tests \ ← Data migrated correctly?
382 / Unit Tests \ ← New code works?
383```
384
385### Parity Testing Framework
386
387For EVERY migrated feature:
388
389```yaml
390feature: ""
391test_type: "api_parity | ui_parity | data_parity"
392method: "shadow traffic | replay | parallel run"
393sample_size: 0
394match_rate: "0%" # Target: 99.9%+
395mismatches_investigated: 0
396mismatches_accepted: 0 # Known intentional differences
397mismatches_bugs: 0
398sign_off: false
399```
400
401**Shadow traffic** — copy production requests to new system, compare responses (don't serve new responses to users yet).
402
403### Performance Regression Rules
404
405- P95 latency must not increase >10% vs legacy
406- Throughput must meet or exceed legacy under same load
407- Database query count must not increase per request
408- Memory usage must not increase >20%
409- If ANY metric regresses → investigate before proceeding
410
411---
412
413## Phase 9: Team & Process
414
415### Modernization Team Structure
416
417| Role | Responsibility | When Needed |
418|------|---------------|-------------|
419| **Modernization Lead** | Strategy, sequencing, blockers | Full-time |
420| **Legacy Expert** | Knows where the bodies are buried | Part-time, on-call |
421| **New Platform Engineer** | Builds target architecture | Full-time |
422| **Data Engineer** | Migration, sync, validation | Phase-dependent |
423| **QA/Test Engineer** | Parity testing, automation | Full-time |
424| **DevOps/Platform** | CI/CD, infrastructure | Part-time |
425| **Product Owner** | Business priority, acceptance | Part-time |
426
427### Knowledge Mining from Legacy
428
429The most dangerous part of modernization is **losing undocumented business rules**.
430
4311. **Code archaeology** — git blame, find oldest unchanged code, understand why
4322. **Interview stakeholders** — "What would break if we changed X?"
4333. **Production log analysis** — what edge cases actually occur?
4344. **Error handling review** — each catch block is a documented business rule
4355. **Test suite review** — tests describe expected behavior
4366. **Configuration review** — magic numbers, feature flags, overrides
437
438### Communication Plan
439
440| Audience | Frequency | Content |
441|----------|-----------|---------|
442| Executive sponsor | Bi-weekly | Progress, risks, budget, timeline |
443| Engineering team | Weekly | Sprint goals, technical decisions, blockers |
444| Dependent teams | Monthly | Upcoming changes, migration dates, API changes |
445| End users | Per migration | What's changing, when, how it affects them |
446
447---
448
449## Phase 10: Risk Management
450
451### Top 10 Modernization Risks (Pre-Built)
452
453| # | Risk | Likelihood | Impact | Mitigation |
454|---|------|-----------|--------|------------|
455| 1 | Undocumented business rules lost | High | Critical | Code archaeology + stakeholder interviews + parity tests |
456| 2 | Timeline underestimation | Very High | High | 2x initial estimate, phase-gated checkpoints |
457| 3 | Data migration corruption | Medium | Critical | Checksums, parallel runs, rollback plans |
458| 4 | Feature parity gaps | High | High | Shadow traffic testing, user acceptance testing |
459| 5 | Team knowledge loss (people leave) | Medium | High | Document everything, pair programming, knowledge sharing |
460| 6 | Legacy system changes during migration | High | Medium | Feature freeze or dual-write contract |
461| 7 | Performance regression | Medium | High | Load testing at every phase, performance budgets |
462| 8 | Scope creep (improve while migrating) | Very High | Medium | Strict "migrate, don't improve" rule for Phase 1 |
463| 9 | Integration failures | Medium | High | Contract testing, circuit breakers, fallback routing |
464| 10 | Stakeholder fatigue | High | Medium | Quick wins early, visible progress dashboard |
465
466### Kill Criteria
467
468Stop the modernization if:
469- Budget exceeds 2x initial estimate with <50% complete
470- Key business rules can't be verified after migration
471- Team attrition >30% during project
472- Legacy system stability degrades due to migration work
473- Business context changes (M&A, pivot, sunset)
474
475**If kill criteria triggered:** Stabilize what's done, document learnings, reassess in 6 months.
476
477---
478
479## Phase 11: Patterns & Playbooks
480
481### Language/Framework Migration Patterns
482
483**Java → Modern Java (8→17+)**
484- Records, sealed classes, pattern matching
485- Virtual threads (Project Loom) for thread-per-request
486- Migrate build: Maven→Gradle or update Maven plugins
487- Spring Boot 2→3: javax→jakarta namespace
488
489**Python 2→3**
490- Use `2to3` tool for automated conversion
491- Fix: print(), division, unicode, dict methods
492- Upgrade dependencies (check py3 compat)
493
494**jQuery→React/Vue**
495- Extract components from page sections
496- State management replaces DOM manipulation
497- Event handlers become component methods
498- Ajax calls become API service layer
499
500**Monolith→Microservices**
501- Strangler fig (see Phase 3)
502- Start with read models (reporting, search)
503- Extract stateless services first
504- Shared database → database-per-service last
505
506**On-Prem→Cloud**
507- Rehost first (lift & shift)
508- Then replatform (managed services)
509- Then re-architect (cloud-native patterns)
510- Never skip steps — each proves value
511
512### COBOL/Mainframe Modernization
513
5141. **API wrapping** — expose CICS/IMS transactions as REST APIs
5152. **Screen scraping** — automate 3270 terminal interactions
5163. **Gradual extraction** — one transaction at a time
5174. **Data replication** — DB2/VSAM → PostgreSQL/cloud DB
5185. **Rule extraction** — COBOL paragraphs → business rule engine
5196. **Never rewrite all at once** — decades of business logic = decades of edge cases
520
521### Microservices Anti-Patterns to Avoid
522
523| Anti-Pattern | Symptom | Fix |
524|-------------|---------|-----|
525| Distributed monolith | Services must deploy together | Identify and break coupling |
526| Shared database | Multiple services write same tables | Database-per-service |
527| Synchronous chains | A calls B calls C calls D | Async events, choreography |
528| Nano-services | Hundreds of tiny services | Merge related services |
529| Shared libraries for business logic | Library update breaks consumers | Duplicate code > shared coupling |
530| No API versioning | Breaking changes cascade | Semantic versioning, deprecation policy |
531
532---
533
534## Phase 12: Metrics & Reporting
535
536### Modernization Health Dashboard
537
538```yaml
539project: ""
540assessment_date: ""
541overall_health: "green | yellow | red"
542
543progress:
544 modules_total: 0
545 modules_migrated: 0
546 modules_in_progress: 0
547 percent_complete: "0%"
548
549velocity:
550 modules_per_sprint: 0
551 estimated_completion: ""
552 on_track: true
553
554quality:
555 parity_test_pass_rate: "0%"
556 production_incidents_from_migration: 0
557 rollbacks: 0
558
559risk:
560 open_risks: 0
561 p0_risks: 0
562 blocked_items: 0
563
564cost:
565 budget_total: "$0"
566 budget_spent: "$0"
567 budget_remaining: "$0"
568 burn_rate_monthly: "$0"
569```
570
571### 100-Point Modernization Quality Rubric
572
573| Dimension | Weight | Score (0-10) | Weighted |
574|-----------|--------|-------------|----------|
575| Strategy clarity | 15% | | |
576| Risk management | 15% | | |
577| Testing rigor | 15% | | |
578| Data integrity | 15% | | |
579| Architecture quality | 10% | | |
580| Team capability | 10% | | |
581| Stakeholder alignment | 10% | | |
582| Documentation | 10% | | |
583| **Total** | **100%** | | **/100** |
584
585**90-100**: Exemplary — reference project
586**70-89**: Strong — minor improvements
587**50-69**: Adequate — address gaps
588**Below 50**: At risk — pause and reassess
589
590### Weekly Status Template
591
592```markdown
593## Modernization Status — Week of [DATE]
594
595### Progress
596- Modules migrated this week: [N]
597- Total migrated: [N]/[TOTAL] ([X]%)
598- On track for [TARGET DATE]: [Yes/No]
599
600### Completed
601- [What shipped this week]
602
603### In Progress
604- [What's being worked on]
605
606### Blockers
607- [What's stuck and what's needed]
608
609### Risks
610- [New or changed risks]
611
612### Next Week
613- [Plan for next sprint]
614```
615
616---
617
618## Edge Cases
619
620### "We need to modernize but can't stop adding features"
621- **Strangler fig** — modernize around new features
622- Feature freeze on legacy module ONLY when that module is being migrated
623- New features build in new stack from day 1
624
625### "We don't know what the system does"
626- Start with observability: instrument logging, tracing, metrics
627- Run for 2-4 weeks to understand actual usage patterns
628- Code coverage analysis shows what code is actually executed
629- Interview longest-tenured team members
630
631### "Multiple systems need modernizing simultaneously"
632- Sequence by dependency order — downstream first
633- Shared services (auth, data) get modernized once, reused
634- Never parallelize more than 2 modernization streams
635
636### "The original developers are gone"
637- Treat code as the documentation
638- Invest 2-4 weeks in code archaeology before any migration work
639- Pair new developers with business stakeholders
640- Write tests for existing behavior before changing anything
641
642### "We're being acquired / merging systems"
643- Map overlapping functionality first
644- Pick "winner" system per domain — don't merge codebases
645- API integration layer between systems
646- 18-month realistic timeline for full consolidation
647
648### "Compliance requires the old system"
649- Maintain compliance evidence chain during migration
650- Dual-audit period with both systems
651- Get compliance team involved in migration planning from Day 1
652- Document control mapping: old control → new implementation
653
654---
655
656## Natural Language Commands
657
658| Command | Action |
659|---------|--------|
660| "Assess this system for modernization" | Run full Technical Debt Inventory |
661| "Which modernization strategy should we use?" | Walk through Strategy Decision Tree |
662| "Plan a strangler fig migration" | Generate Strangler Facade YAML + sequence |
663| "Decompose this monolith" | Domain discovery + Bounded Context mapping |
664| "Migrate this database" | Data Quality Gates + migration strategy |
665| "Check cloud readiness" | Run Cloud Readiness Assessment |
666| "Create a migration testing plan" | Build Testing Pyramid with parity tests |
667| "What are the risks?" | Generate Top 10 risk register |
668| "How do we migrate from [X] to [Y]?" | Pattern-specific playbook |
669| "Status update for modernization" | Generate Weekly Status Template |
670| "Score this modernization project" | Run 100-Point Quality Rubric |
671| "Should we kill this modernization?" | Evaluate Kill Criteria |