Capacity Planning
When to activate
After solution architecture is approved and before infrastructure provisioning. Triggered by: known growth trajectory, customer SLA requirements, or periodic review (quarterly). Essential before year-end planning or major feature launches. Run whenever user projections or load profile change significantly.
When NOT to use
Not for existing systems in steady state (use monitoring and alerting instead). Not without load testing data (estimate conservatively if missing). Not for small POCs that will be rebuilt. Not if you have <3 months of production data (too early to extrapolate).
Planning Checklist
Current Load Profile
Growth Projections
Performance Requirements
Resource Consumption Model
Cost Model
Output Format
Capacity Roadmap
Current State (as of [date])
| Metric |
Value |
| Daily Active Users |
10,000 |
| Peak Throughput |
500 req/sec |
| Stored Data |
500 GB |
| Concurrent Connections |
5,000 |
| API Latency (p99) |
200 ms |
| Database Size |
250 GB |
Growth Projections (3-year horizon)
| Period |
Users |
Throughput (req/sec) |
Stored Data |
Infrastructure |
| Today (M0) |
10K |
500 |
500 GB |
4x app servers, 1x db |
| 6 months (M6) |
25K |
1,200 |
1.2 TB |
8x app servers, 1x db (scaled) |
| 1 year (M12) |
50K |
2,500 |
2.5 TB |
16x app servers, 2x db (replica) |
| 2 years (M24) |
150K |
7,500 |
7 TB |
32x app servers, 4x db (sharded) |
| 3 years (M36) |
300K |
15,000 |
15 TB |
64x app servers, 8x db (distributed) |
Bottleneck Analysis
At each milestone, what will become the constraint?
| Milestone |
Bottleneck |
Mitigation |
Cost Impact |
| M0→M6 |
Database connections |
Add connection pooling |
+$2K/mo |
| M6→M12 |
Compute (CPU on app servers) |
Add load balancer, scale out |
+$5K/mo |
| M12→M24 |
Database query performance |
Sharding, replication |
+$10K/mo |
| M24→M36 |
Network bandwidth |
Multi-region, CDN |
+$15K/mo |
Server Sizing (per component)
Application Servers
- Current: 4x [size] instances
- M6: 8x [size] instances
- M12: 16x [size] instances
Reasoning: Each server handles ~125 req/sec. At M12 (2,500 req/sec), need 20 servers (over-provisioned by 1.25x for redundancy).
Database
- Current: 1x PostgreSQL [size], 500 GB storage
- M6: 1x PostgreSQL (scaled up), 1.2 TB storage
- M12: 1x primary + 1x replica, 2.5 TB storage
- M24: Sharded (3 shards), 7 TB storage
Reasoning: Single server can handle ~5K req/sec if well-tuned. At M24 (7.5K req/sec), sharding is required to maintain <100ms latency.
Cache Layer (Redis or Memcached)
- Current: 1x Redis [size], 50 GB heap
- M12: 1x Redis (scaled up), 200 GB heap
- M24: Redis cluster (3 nodes), 500 GB total
CDN & Static Serving
- Current: CloudFront, 10 GB/month bandwidth
- M12: CloudFront, 50 GB/month bandwidth
- M24: Multi-region edge, 200 GB/month bandwidth
Cost Projections
| Period |
Compute |
Database |
Network |
Storage |
Monitoring |
Total/Month |
| M0 |
$4K |
$2K |
$500 |
$1K |
$500 |
$8K |
| M6 |
$8K |
$4K |
$1K |
$2K |
$1K |
$16K |
| M12 |
$16K |
$8K |
$2K |
$4K |
$2K |
$32K |
| M24 |
$32K |
$16K |
$5K |
$8K |
$4K |
$65K |
| M36 |
$64K |
$32K |
$10K |
$15K |
$8K |
$129K |
Cost per User (unit economics)
- M0: $0.80/user/month
- M12: $0.64/user/month
- M24: $0.43/user/month
(Cost decreases with scale due to spreading fixed costs and improving efficiency.)
Contingency Plans
If Growth Exceeds Projections (2x faster):
- Implement aggressive caching (database queries halved)
- Move non-critical services to async queues
- Fast-track sharding implementation
- Consider serverless for traffic spikes
If Growth Stalls:
- Defer sharding/multi-region (save ~$30K/mo)
- Right-size instances (move from large to medium)
- Consolidate underutilized services
- Reduce monitoring/alerting complexity
Optimization Opportunities
| Opportunity |
Current Cost |
Optimized Cost |
Savings |
Implementation |
| Database indexing |
$8K |
$6K |
$2K/mo |
Q3 |
| Query optimization |
$8K |
$6K |
$2K/mo |
Q3 |
| Connection pooling |
$2K |
$1K |
$1K/mo |
Q2 |
| Compression (API responses) |
$1K |
$0.5K |
$0.5K/mo |
Q2 |
| Aggressive caching |
$2K |
$1K |
$1K/mo |
Q4 |
Scaling Strategy by Component
Vertical Scaling (bigger servers)
- Pros: Simple, less operational overhead
- Cons: Limited ceiling, expensive, eventual downtime for upgrades
- Best for: Small teams, low traffic, early stage
Horizontal Scaling (more servers)
- Pros: Unlimited, resilient, cost-efficient at scale
- Cons: More complex, requires load balancing, stateless design
- Best for: High traffic, growth trajectory, distributed systems
Recommendation: Start vertical (M0–M6), transition to horizontal after M6.
Monitoring & Metrics (for continuous capacity management)
Track weekly:
- Actual vs. projected user growth
- Actual vs. projected load (req/sec, data volume)
- Resource utilization (CPU, memory, disk, network)
- Cost actuals vs. budget
Capacity review quarterly:
- Update growth projections based on new data
- Adjust scaling timeline if needed
- Identify and implement optimizations
- Report to finance on cost efficiency
1---2name: capacity-planning3description: Forecasts infrastructure and resource needs based on projected growth. Calculates server count, database sizing, network bandwidth, and cost projections. Outputs capacity roadmap with spending milestones.4---56# Capacity Planning78## When to activate910After solution architecture is approved and before infrastructure provisioning. Triggered by: known growth trajectory, customer SLA requirements, or periodic review (quarterly). Essential before year-end planning or major feature launches. Run whenever user projections or load profile change significantly.1112## When NOT to use1314Not for existing systems in steady state (use monitoring and alerting instead). Not without load testing data (estimate conservatively if missing). Not for small POCs that will be rebuilt. Not if you have <3 months of production data (too early to extrapolate).1516## Planning Checklist17181. **Current Load Profile**19 - [ ] Peak throughput (req/sec, transactions/sec)20 - [ ] Concurrent users (current and realistic peak)21 - [ ] Data volume (stored data, growth rate per month)22 - [ ] Data retention policy (how long kept)23 - [ ] Backup/archival needs24252. **Growth Projections**26 - [ ] User growth forecast (conservative, realistic, aggressive)27 - [ ] Data growth forecast (documents, transactions, logs)28 - [ ] Seasonality (peaks in certain months?)29 - [ ] Planned feature launches (will they increase load?)30 - [ ] Geographic expansion (multi-region needs?)31323. **Performance Requirements**33 - [ ] Latency targets (p50, p95, p99)34 - [ ] Throughput targets (concurrent users, req/sec)35 - [ ] Availability SLA (uptime %, downtime tolerance)36 - [ ] Data consistency requirements (strong vs. eventual)37384. **Resource Consumption Model**39 - [ ] CPU per request (profile via load test)40 - [ ] Memory per request (session size, cache needs)41 - [ ] Storage per record (database, logs, backups)42 - [ ] Network per request (payload size, bandwidth)43 - [ ] Database connections (pool size per req/sec)44455. **Cost Model**46 - [ ] Compute costs (servers, containers, serverless)47 - [ ] Database costs (managed service, licensing, backups)48 - [ ] Network costs (ingress/egress, CDN)49 - [ ] Storage costs (primary, backup, archive)50 - [ ] Operational costs (monitoring, support, licenses)5152## Output Format5354### Capacity Roadmap5556**Current State** (as of [date])5758| Metric | Value |59|---|---|60| Daily Active Users | 10,000 |61| Peak Throughput | 500 req/sec |62| Stored Data | 500 GB |63| Concurrent Connections | 5,000 |64| API Latency (p99) | 200 ms |65| Database Size | 250 GB |6667**Growth Projections** (3-year horizon)6869| Period | Users | Throughput (req/sec) | Stored Data | Infrastructure |70|---|---|---|---|---|71| **Today (M0)** | 10K | 500 | 500 GB | 4x app servers, 1x db |72| **6 months (M6)** | 25K | 1,200 | 1.2 TB | 8x app servers, 1x db (scaled) |73| **1 year (M12)** | 50K | 2,500 | 2.5 TB | 16x app servers, 2x db (replica) |74| **2 years (M24)** | 150K | 7,500 | 7 TB | 32x app servers, 4x db (sharded) |75| **3 years (M36)** | 300K | 15,000 | 15 TB | 64x app servers, 8x db (distributed) |7677**Bottleneck Analysis**7879At each milestone, what will become the constraint?8081| Milestone | Bottleneck | Mitigation | Cost Impact |82|---|---|---|---|83| **M0→M6** | Database connections | Add connection pooling | +$2K/mo |84| **M6→M12** | Compute (CPU on app servers) | Add load balancer, scale out | +$5K/mo |85| **M12→M24** | Database query performance | Sharding, replication | +$10K/mo |86| **M24→M36** | Network bandwidth | Multi-region, CDN | +$15K/mo |8788**Server Sizing** (per component)8990**Application Servers**91- Current: 4x [size] instances92- M6: 8x [size] instances93- M12: 16x [size] instances9495Reasoning: Each server handles ~125 req/sec. At M12 (2,500 req/sec), need 20 servers (over-provisioned by 1.25x for redundancy).9697**Database**98- Current: 1x PostgreSQL [size], 500 GB storage99- M6: 1x PostgreSQL (scaled up), 1.2 TB storage100- M12: 1x primary + 1x replica, 2.5 TB storage101- M24: Sharded (3 shards), 7 TB storage102103Reasoning: Single server can handle ~5K req/sec if well-tuned. At M24 (7.5K req/sec), sharding is required to maintain <100ms latency.104105**Cache Layer** (Redis or Memcached)106- Current: 1x Redis [size], 50 GB heap107- M12: 1x Redis (scaled up), 200 GB heap108- M24: Redis cluster (3 nodes), 500 GB total109110**CDN & Static Serving**111- Current: CloudFront, 10 GB/month bandwidth112- M12: CloudFront, 50 GB/month bandwidth113- M24: Multi-region edge, 200 GB/month bandwidth114115**Cost Projections**116117| Period | Compute | Database | Network | Storage | Monitoring | Total/Month |118|---|---|---|---|---|---|---|119| **M0** | $4K | $2K | $500 | $1K | $500 | $8K |120| **M6** | $8K | $4K | $1K | $2K | $1K | $16K |121| **M12** | $16K | $8K | $2K | $4K | $2K | $32K |122| **M24** | $32K | $16K | $5K | $8K | $4K | $65K |123| **M36** | $64K | $32K | $10K | $15K | $8K | $129K |124125**Cost per User** (unit economics)126- M0: $0.80/user/month127- M12: $0.64/user/month128- M24: $0.43/user/month129130(Cost decreases with scale due to spreading fixed costs and improving efficiency.)131132**Contingency Plans**133134**If Growth Exceeds Projections (2x faster):**135- Implement aggressive caching (database queries halved)136- Move non-critical services to async queues137- Fast-track sharding implementation138- Consider serverless for traffic spikes139140**If Growth Stalls:**141- Defer sharding/multi-region (save ~$30K/mo)142- Right-size instances (move from large to medium)143- Consolidate underutilized services144- Reduce monitoring/alerting complexity145146**Optimization Opportunities**147148| Opportunity | Current Cost | Optimized Cost | Savings | Implementation |149|---|---|---|---|---|150| Database indexing | $8K | $6K | $2K/mo | Q3 |151| Query optimization | $8K | $6K | $2K/mo | Q3 |152| Connection pooling | $2K | $1K | $1K/mo | Q2 |153| Compression (API responses) | $1K | $0.5K | $0.5K/mo | Q2 |154| Aggressive caching | $2K | $1K | $1K/mo | Q4 |155156**Scaling Strategy by Component**157158**Vertical Scaling** (bigger servers)159- Pros: Simple, less operational overhead160- Cons: Limited ceiling, expensive, eventual downtime for upgrades161- Best for: Small teams, low traffic, early stage162163**Horizontal Scaling** (more servers)164- Pros: Unlimited, resilient, cost-efficient at scale165- Cons: More complex, requires load balancing, stateless design166- Best for: High traffic, growth trajectory, distributed systems167168**Recommendation:** Start vertical (M0–M6), transition to horizontal after M6.169170**Monitoring & Metrics** (for continuous capacity management)171172Track weekly:173- Actual vs. projected user growth174- Actual vs. projected load (req/sec, data volume)175- Resource utilization (CPU, memory, disk, network)176- Cost actuals vs. budget177178Capacity review quarterly:179- Update growth projections based on new data180- Adjust scaling timeline if needed181- Identify and implement optimizations182- Report to finance on cost efficiency183184---