CPU Profiling Guide
Phase 1: Symptom Assessment
- Characterize the CPU issue
- Sustained high CPU (> 80% continuously)
- CPU spikes correlated with specific operations
- High CPU at low traffic (idle CPU waste)
- CPU-bound latency (slow responses despite available I/O)
- One core saturated (single-threaded bottleneck)
- Collect baseline metrics
- CPU utilization per core and per process
- Thread/goroutine count
- Context switch rate
- System vs. user CPU time ratio
- Correlate CPU with application events (requests, jobs, deployments)
CPU Baseline
| Metric | Current | Normal Baseline | Status |
|---|---|---|---|
| Overall CPU | % | % | |
| User CPU | % | % | |
| System CPU | % | % | |
| Context switches/sec | |||
| Active threads | |||
| Load average |
Phase 2: Profiler Setup
- Select appropriate profiler for runtime
- JVM: async-profiler, JFR (Java Flight Recorder)
- Node.js: --prof, 0x, clinic.js
- Python: py-spy, cProfile, Pyroscope
- Go: pprof (built-in)
- .NET: dotTrace, PerfView
- Native: perf, Instruments (macOS)
- Configure profiler with minimal overhead (< 2% CPU impact)
- Choose sampling rate (typically 99Hz or 997Hz)
- Profile in production or production-like environment
- Capture profiles during problematic period
Phase 3: Flame Graph Analysis
- Generate flame graphs from profile data
- Analyze flame graph patterns
- Wide plateaus: functions consuming most CPU time
- Deep stacks: excessive call depth or recursion
- Repeated patterns: hot loops
- Framework vs. application code ratio
- Identify top CPU-consuming functions
Top CPU Consumers
| Function | Self Time % | Total Time % | Calls/sec | Category | Actionable |
|---|---|---|---|---|---|
| % | % | App/Framework/GC | Yes/No |
Phase 4: Hot Path Investigation
- Investigate top CPU consumers
- Inefficient algorithms (O(n^2) or worse)
- Unnecessary computation in hot paths
- Excessive object allocation triggering GC
- Serialization/deserialization overhead
- Regular expression backtracking
- Logging in tight loops
- Lock contention causing spin-waits
- Redundant computation (missing caching)
- Review code for identified hot functions
- Check if work can be avoided, cached, or batched
Root Cause Checklist
| Issue Type | Location | CPU Impact % | Fix Effort | Priority |
|---|---|---|---|---|
| Algorithmic inefficiency | % | |||
| Excessive GC pressure | % | |||
| Unnecessary computation | % | |||
| Missing cache | % | |||
| Lock contention | % |
Phase 5: Optimization Implementation
- Apply targeted optimizations
- Replace inefficient algorithms
- Add caching for repeated computations
- Reduce allocation rate in hot paths
- Move work off hot path (async, batch, lazy)
- Optimize data structures
- Reduce serialization overhead
- Fix lock contention (finer-grained locks, lock-free)
- Benchmark each optimization in isolation
- Verify no functionality regression
Phase 6: Validation
- Re-profile after optimizations
- Generate comparison flame graphs (before vs. after)
- Verify CPU utilization reduction
- Load test to confirm improvement under peak traffic
- Monitor for regressions over time
Optimization Results
| Optimization | Before (CPU %) | After (CPU %) | Reduction | Latency Impact |
|---|---|---|---|---|
| % | % | -% | -ms | |
| Total | % | % | -% | -ms |
Counter-Rationalizations
| Shortcut | Counter | Why |
|---|---|---|
| "We can skip some steps for this case" | Adapt the workflow steps, don't skip them | Skipped steps are where incidents and oversights originate |
| "The user seems to already know what to do" | Complete all workflow phases with the user | The workflow catches blind spots that experience alone misses |
| "This is a minor case, full process is overkill" | Scale the process down, don't turn it off | Minor cases become major when unstructured; the process scales, not disappears |
| "I'll fill in the details later" | Complete each section before moving on | Deferred details are forgotten; real-time capture is more accurate |
| "The template output isn't necessary" | Always produce the structured output format | Structured output enables comparison, audit trails, and handoff to other teams |
Output Format
- CPU Baseline: Metrics before optimization
- Flame Graphs: Before and after comparison
- Hot Path Analysis: Top CPU consumers with root causes
- Optimization Summary: Changes made with measured impact
- Monitoring Setup: Continuous profiling and CPU alerts
Action Items
- Collect CPU baseline and characterize the problem
- Set up profiler in production-like environment
- Capture and analyze flame graphs
- Identify and prioritize hot path optimizations
- Implement and benchmark optimizations
- Validate with re-profiling and load testing
- Set up continuous profiling for future regressions