Spot Instance Strategy
Phase 1: Workload Assessment
- Evaluate workloads for spot suitability
- Fault tolerance (can handle interruptions)
- Checkpointing capability
- Flexible start/completion times
- Stateless or state externalized
- Can run on multiple instance types
- Quantify potential savings per workload
Spot Suitability Matrix
| Workload | Fault Tolerant | Checkpointable | Flexible Timing | Multi-Instance | Spot Fit |
|---|---|---|---|---|---|
| [ ] | [ ] | [ ] | [ ] | High/Med/Low |
Phase 2: Instance Diversification
- Identify multiple instance types per workload (minimum 6-10)
- Spread across multiple availability zones
- Analyze historical spot pricing and interruption rates
- Define instance type priority based on price and availability
- Configure capacity-optimized allocation strategy
Instance Pool Configuration
| Instance Type | vCPU | Memory | Spot Price (avg) | Interruption Rate | Priority |
|---|---|---|---|---|---|
| <5% / 5-15% / >15% |
Phase 3: Interruption Handling
- Implement graceful shutdown handlers (2-minute warning)
- Set up checkpointing for long-running jobs
- Configure automatic replacement of interrupted instances
- Design queue-based architectures for batch workloads
- Implement health checks and automatic failover
Interruption Response Plan
| Scenario | Detection | Response | Recovery Time | Data Loss Risk |
|---|---|---|---|---|
| Single instance interruption | ||||
| Multiple simultaneous interruptions | ||||
| Availability zone capacity event |
Phase 4: Hybrid Strategy Design
- Define baseline capacity on reserved/on-demand instances
- Configure auto-scaling with spot for burst capacity
- Set maximum spot percentage per workload
- Implement fallback to on-demand when spot unavailable
- Balance cost savings with availability requirements
Phase 5: Implementation
- Configure spot fleet or managed instance group
- Set up monitoring for spot utilization and savings
- Implement cost tracking per workload
- Test interruption handling with fault injection
- Deploy to production with gradual spot percentage increase
Phase 6: Optimization
- Monitor actual vs. projected savings weekly
- Adjust instance type mix based on interruption patterns
- Refine capacity allocation between spot, reserved, on-demand
- Review and update maximum bid prices
- Expand spot usage to newly identified workloads
Counter-Rationalizations
| Shortcut | Counter | Why |
|---|---|---|
| "We can skip some steps for this case" | Adapt the workflow steps, don't skip them | Skipped steps are where incidents and oversights originate |
| "The user seems to already know what to do" | Complete all workflow phases with the user | The workflow catches blind spots that experience alone misses |
| "This is a minor case, full process is overkill" | Scale the process down, don't turn it off | Minor cases become major when unstructured; the process scales, not disappears |
| "I'll fill in the details later" | Complete each section before moving on | Deferred details are forgotten; real-time capture is more accurate |
| "The template output isn't necessary" | Always produce the structured output format | Structured output enables comparison, audit trails, and handoff to other teams |
Output Format
- Workload Assessment: Spot suitability classification per workload
- Instance Pool Design: Diversified instance types per workload
- Architecture Diagram: Hybrid spot/on-demand/reserved design
- Interruption Runbook: Automated and manual response procedures
- Savings Dashboard: Weekly tracking of spot savings vs. on-demand
Action Items
- Assess all workloads for spot suitability
- Select and test diversified instance pools
- Implement interruption handling and checkpointing
- Deploy spot strategy to non-critical workloads first
- Monitor and expand to additional workloads
- Report monthly savings to stakeholders