Performance Profiling
Goal
Provide tools to analyze simulation performance, identify bottlenecks, and recommend optimization strategies for computational materials science simulations.
Requirements
- Python 3.10+
- No external dependencies (uses Python standard library only)
- Works on Linux, macOS, and Windows
Inputs to Gather
Before running profiling scripts, collect from the user:
| Input |
Description |
Example |
| Simulation log |
Log file with timing information |
simulation.log |
| Scaling data |
JSON with multi-run performance data |
scaling_data.json |
| Simulation parameters |
JSON with mesh, fields, solver config |
params.json |
| Available memory |
System memory in GB (optional) |
16.0 |
Decision Guidance
When to Use Each Script
Need to identify slow phases?
├── YES → Use timing_analyzer.py
│ └── Parse simulation logs for timing data
│
Need to understand parallel performance?
├── YES → Use scaling_analyzer.py
│ └── Analyze strong or weak scaling efficiency
│
Need to estimate memory requirements?
├── YES → Use memory_profiler.py
│ └── Estimate memory from problem parameters
│
Need optimization recommendations?
└── YES → Use bottleneck_detector.py
└── Combine analyses and get actionable advice
Choosing Analysis Thresholds
| Metric |
Good |
Acceptable |
Poor |
| Phase dominance |
<30% |
30-50% |
>50% |
| Parallel efficiency |
>0.80 |
0.70-0.80 |
<0.70 |
| Memory usage |
<60% |
60-80% |
>80% |
Script Outputs (JSON Fields)
| Script |
Key Outputs |
timing_analyzer.py |
timing_data.phases, timing_data.slowest_phase, timing_data.total_time |
scaling_analyzer.py |
scaling_analysis.results, scaling_analysis.efficiency_threshold_processors |
memory_profiler.py |
memory_profile.total_memory_gb, memory_profile.per_process_gb, memory_profile.warnings |
bottleneck_detector.py |
bottlenecks, recommendations |
Workflow
Complete Profiling Workflow
- Analyze timing from simulation logs
- Analyze scaling from multi-run data (if available)
- Profile memory from simulation parameters
- Detect bottlenecks and get recommendations
- Implement optimizations based on recommendations
- Re-profile to verify improvements
Quick Profiling (Timing Only)
- Run timing analyzer on simulation log
- Identify dominant phases (>50% of runtime)
- Apply targeted optimizations to dominant phases
CLI Examples
Timing Analysis
# Basic timing analysis
python3 scripts/timing_analyzer.py \
--log simulation.log \
--json
# Custom timing pattern
python3 scripts/timing_analyzer.py \
--log simulation.log \
--pattern 'Step\s+(\w+)\s+took\s+([\d.]+)s' \
--json
Scaling Analysis
# Strong scaling (fixed problem size)
python3 scripts/scaling_analyzer.py \
--data scaling_data.json \
--type strong \
--json
# Weak scaling (constant work per processor)
python3 scripts/scaling_analyzer.py \
--data scaling_data.json \
--type weak \
--json
Memory Profiling
# Estimate memory requirements
python3 scripts/memory_profiler.py \
--params simulation_params.json \
--available-gb 16.0 \
--json
Bottleneck Detection
# Detect bottlenecks from timing only
python3 scripts/bottleneck_detector.py \
--timing timing_results.json \
--json
# Comprehensive analysis with all inputs
python3 scripts/bottleneck_detector.py \
--timing timing_results.json \
--scaling scaling_results.json \
--memory memory_results.json \
--json
Conversational Workflow Example
User: My simulation is taking too long. Can you help me identify what's slow?
Agent workflow:
- Ask for simulation log file
- Run timing analyzer:
python3 scripts/timing_analyzer.py --log simulation.log --json
- Interpret results:
- If solver dominates (>50%): Recommend preconditioner tuning
- If assembly dominates: Recommend caching or vectorization
- If I/O dominates: Recommend reducing output frequency
- If user has multi-run data, analyze scaling:
python3 scripts/scaling_analyzer.py --data scaling.json --type strong --json
- Generate comprehensive recommendations:
python3 scripts/bottleneck_detector.py --timing timing.json --scaling scaling.json --json
Interpretation Guidance
Timing Analysis
| Scenario |
Meaning |
Action |
| Solver >70% |
Solver-dominated |
Tune preconditioner, check tolerance |
| Assembly >50% |
Assembly-dominated |
Cache matrices, vectorize, parallelize |
| I/O >30% |
I/O-dominated |
Reduce frequency, use parallel I/O |
| Balanced (<30% each) |
Well-balanced |
Look for algorithmic improvements |
Scaling Analysis
| Efficiency |
Meaning |
Action |
| >0.80 |
Excellent scaling |
Continue scaling up |
| 0.70-0.80 |
Good scaling |
Monitor at larger scales |
| 0.50-0.70 |
Poor scaling |
Investigate communication/load balance |
| <0.50 |
Very poor scaling |
Reduce processor count or redesign |
Memory Profile
| Usage |
Meaning |
Action |
| <60% available |
Safe |
No action needed |
| 60-80% available |
Moderate |
Monitor, consider optimization |
| >80% available |
High |
Reduce resolution or increase processors |
| >100% available |
Exceeds capacity |
Must reduce problem size |
Error Handling
| Error |
Cause |
Resolution |
Log file not found |
Invalid path |
Verify log file path |
No timing data found |
Pattern mismatch |
Provide custom pattern with --pattern |
At least 2 runs required |
Insufficient data |
Provide more scaling runs |
Missing required parameters |
Incomplete params |
Add mesh and fields to params file |
Optimization Strategies by Bottleneck Type
Solver Bottlenecks
- Use algebraic multigrid (AMG) preconditioner
- Tighten solver tolerance if over-solving
- Consider direct solver for small problems
- Profile matrix assembly vs solve time
Assembly Bottlenecks
- Cache element matrices if geometry is static
- Use vectorized assembly routines
- Consider matrix-free methods
- Parallelize assembly with coloring
I/O Bottlenecks
- Reduce output frequency
- Use parallel I/O (HDF5, MPI-IO)
- Write to fast scratch storage
- Compress output data
Scaling Bottlenecks
- Investigate communication overhead
- Check for load imbalance
- Reduce synchronization points
- Use asynchronous communication
- Consider hybrid MPI+OpenMP
Memory Bottlenecks
- Reduce mesh resolution
- Use iterative solver (lower memory than direct)
- Enable out-of-core computation
- Increase number of processors
- Use single precision where appropriate
Security
Input Validation
- User-supplied
--pattern regex values are validated for length (500 chars max) and rejected if they contain constructs prone to catastrophic backtracking (ReDoS)
- Scaling data entries are validated for finite time values, integer processor counts, and bounded run count (10,000 max)
available_gb is validated as a positive finite number; mesh dimensions and field parameters are validated as positive integers
--type (scaling type) is validated against a fixed allowlist (strong, weak)
- All loaded JSON files must have an object (dict) as root element
File Access
timing_analyzer.py reads a single log file specified by --log; log files are capped at 500 MB and rejected before parsing
scaling_analyzer.py, memory_profiler.py, and bottleneck_detector.py read JSON files capped at 100 MB
- Phase names extracted from log files are truncated to 200 characters and stripped of control characters to prevent prompt-injection payloads from propagating into agent context
- No scripts write to the filesystem; all output goes to stdout
Tool Restrictions
- Read: Used to inspect script source, references, simulation logs, and result files
- Write: Used to save profiling reports or optimization recommendations; writes are scoped to the user's working directory
- Grep/Glob: Used to locate log files, result files, and search references
- The skill's
allowed-tools excludes Bash to prevent the agent from executing arbitrary commands when processing untrusted simulation logs or result files
Safety Measures
- No
eval(), exec(), or dynamic code generation
- All subprocess calls use explicit argument lists (no
shell=True)
- Reduced tool surface (no Bash) limits the agent to read/write operations only
- Phase names and diagnostic strings are sanitized before inclusion in output to prevent injection
Limitations
- Log parsing: Depends on pattern matching; may miss unusual formats
- Scaling analysis: Requires at least 2 runs for meaningful results
- Memory estimation: Approximate; actual usage may vary
- Recommendations: General guidance; may need domain-specific tuning
References
references/profiling_guide.md - Profiling concepts and interpretation
references/optimization_strategies.md - Detailed optimization approaches
Version History
- v1.0.0 (2025-01-22): Initial release with 4 profiling scripts
1---2name: performance-profiling3description: Identify computational bottlenecks, analyze parallel scaling, estimate memory requirements, and generate optimization recommendations for materials simulations — parse timing logs to find dominant phases (solver, assembly, I/O), evaluate strong and weak scaling efficiency, profile memory from mesh and field parameters, and detect bottlenecks with actionable fix suggestions. Use when a simulation is running slower than expected, investigating MPI scaling efficiency, planning HPC resource allocation, deciding whether to tune the preconditioner or reduce I/O frequency, or estimating if a problem fits in available RAM, even if the user only says "my simulation is too slow" or "how many nodes do I need."4---56# Performance Profiling78## Goal910Provide tools to analyze simulation performance, identify bottlenecks, and recommend optimization strategies for computational materials science simulations.1112## Requirements1314- Python 3.10+15- No external dependencies (uses Python standard library only)16- Works on Linux, macOS, and Windows1718## Inputs to Gather1920Before running profiling scripts, collect from the user:2122| Input | Description | Example |23|-------|-------------|---------|24| Simulation log | Log file with timing information | `simulation.log` |25| Scaling data | JSON with multi-run performance data | `scaling_data.json` |26| Simulation parameters | JSON with mesh, fields, solver config | `params.json` |27| Available memory | System memory in GB (optional) | `16.0` |2829## Decision Guidance3031### When to Use Each Script3233```34Need to identify slow phases?35├── YES → Use timing_analyzer.py36│ └── Parse simulation logs for timing data37│38Need to understand parallel performance?39├── YES → Use scaling_analyzer.py40│ └── Analyze strong or weak scaling efficiency41│42Need to estimate memory requirements?43├── YES → Use memory_profiler.py44│ └── Estimate memory from problem parameters45│46Need optimization recommendations?47└── YES → Use bottleneck_detector.py48 └── Combine analyses and get actionable advice49```5051### Choosing Analysis Thresholds5253| Metric | Good | Acceptable | Poor |54|--------|------|------------|------|55| Phase dominance | <30% | 30-50% | >50% |56| Parallel efficiency | >0.80 | 0.70-0.80 | <0.70 |57| Memory usage | <60% | 60-80% | >80% |5859## Script Outputs (JSON Fields)6061| Script | Key Outputs |62|--------|-------------|63| `timing_analyzer.py` | `timing_data.phases`, `timing_data.slowest_phase`, `timing_data.total_time` |64| `scaling_analyzer.py` | `scaling_analysis.results`, `scaling_analysis.efficiency_threshold_processors` |65| `memory_profiler.py` | `memory_profile.total_memory_gb`, `memory_profile.per_process_gb`, `memory_profile.warnings` |66| `bottleneck_detector.py` | `bottlenecks`, `recommendations` |6768## Workflow6970### Complete Profiling Workflow71721. **Analyze timing** from simulation logs732. **Analyze scaling** from multi-run data (if available)743. **Profile memory** from simulation parameters754. **Detect bottlenecks** and get recommendations765. **Implement optimizations** based on recommendations776. **Re-profile** to verify improvements7879### Quick Profiling (Timing Only)80811. **Run timing analyzer** on simulation log822. **Identify dominant phases** (>50% of runtime)833. **Apply targeted optimizations** to dominant phases8485## CLI Examples8687### Timing Analysis8889```bash90# Basic timing analysis91python3 scripts/timing_analyzer.py \92 --log simulation.log \93 --json9495# Custom timing pattern96python3 scripts/timing_analyzer.py \97 --log simulation.log \98 --pattern 'Step\s+(\w+)\s+took\s+([\d.]+)s' \99 --json100```101102### Scaling Analysis103104```bash105# Strong scaling (fixed problem size)106python3 scripts/scaling_analyzer.py \107 --data scaling_data.json \108 --type strong \109 --json110111# Weak scaling (constant work per processor)112python3 scripts/scaling_analyzer.py \113 --data scaling_data.json \114 --type weak \115 --json116```117118### Memory Profiling119120```bash121# Estimate memory requirements122python3 scripts/memory_profiler.py \123 --params simulation_params.json \124 --available-gb 16.0 \125 --json126```127128### Bottleneck Detection129130```bash131# Detect bottlenecks from timing only132python3 scripts/bottleneck_detector.py \133 --timing timing_results.json \134 --json135136# Comprehensive analysis with all inputs137python3 scripts/bottleneck_detector.py \138 --timing timing_results.json \139 --scaling scaling_results.json \140 --memory memory_results.json \141 --json142```143144## Conversational Workflow Example145146**User**: My simulation is taking too long. Can you help me identify what's slow?147148**Agent workflow**:1491. Ask for simulation log file1502. Run timing analyzer:151 ```bash152 python3 scripts/timing_analyzer.py --log simulation.log --json153 ```1543. Interpret results:155 - If solver dominates (>50%): Recommend preconditioner tuning156 - If assembly dominates: Recommend caching or vectorization157 - If I/O dominates: Recommend reducing output frequency1584. If user has multi-run data, analyze scaling:159 ```bash160 python3 scripts/scaling_analyzer.py --data scaling.json --type strong --json161 ```1625. Generate comprehensive recommendations:163 ```bash164 python3 scripts/bottleneck_detector.py --timing timing.json --scaling scaling.json --json165 ```166167## Interpretation Guidance168169### Timing Analysis170171| Scenario | Meaning | Action |172|----------|---------|--------|173| Solver >70% | Solver-dominated | Tune preconditioner, check tolerance |174| Assembly >50% | Assembly-dominated | Cache matrices, vectorize, parallelize |175| I/O >30% | I/O-dominated | Reduce frequency, use parallel I/O |176| Balanced (<30% each) | Well-balanced | Look for algorithmic improvements |177178### Scaling Analysis179180| Efficiency | Meaning | Action |181|------------|---------|--------|182| >0.80 | Excellent scaling | Continue scaling up |183| 0.70-0.80 | Good scaling | Monitor at larger scales |184| 0.50-0.70 | Poor scaling | Investigate communication/load balance |185| <0.50 | Very poor scaling | Reduce processor count or redesign |186187### Memory Profile188189| Usage | Meaning | Action |190|-------|---------|--------|191| <60% available | Safe | No action needed |192| 60-80% available | Moderate | Monitor, consider optimization |193| >80% available | High | Reduce resolution or increase processors |194| >100% available | Exceeds capacity | Must reduce problem size |195196## Error Handling197198| Error | Cause | Resolution |199|-------|-------|------------|200| `Log file not found` | Invalid path | Verify log file path |201| `No timing data found` | Pattern mismatch | Provide custom pattern with --pattern |202| `At least 2 runs required` | Insufficient data | Provide more scaling runs |203| `Missing required parameters` | Incomplete params | Add mesh and fields to params file |204205## Optimization Strategies by Bottleneck Type206207### Solver Bottlenecks208- Use algebraic multigrid (AMG) preconditioner209- Tighten solver tolerance if over-solving210- Consider direct solver for small problems211- Profile matrix assembly vs solve time212213### Assembly Bottlenecks214- Cache element matrices if geometry is static215- Use vectorized assembly routines216- Consider matrix-free methods217- Parallelize assembly with coloring218219### I/O Bottlenecks220- Reduce output frequency221- Use parallel I/O (HDF5, MPI-IO)222- Write to fast scratch storage223- Compress output data224225### Scaling Bottlenecks226- Investigate communication overhead227- Check for load imbalance228- Reduce synchronization points229- Use asynchronous communication230- Consider hybrid MPI+OpenMP231232### Memory Bottlenecks233- Reduce mesh resolution234- Use iterative solver (lower memory than direct)235- Enable out-of-core computation236- Increase number of processors237- Use single precision where appropriate238239## Security240241### Input Validation242- User-supplied `--pattern` regex values are validated for length (500 chars max) and rejected if they contain constructs prone to catastrophic backtracking (ReDoS)243- Scaling data entries are validated for finite time values, integer processor counts, and bounded run count (10,000 max)244- `available_gb` is validated as a positive finite number; mesh dimensions and field parameters are validated as positive integers245- `--type` (scaling type) is validated against a fixed allowlist (`strong`, `weak`)246- All loaded JSON files must have an object (dict) as root element247248### File Access249- `timing_analyzer.py` reads a single log file specified by `--log`; log files are capped at 500 MB and rejected before parsing250- `scaling_analyzer.py`, `memory_profiler.py`, and `bottleneck_detector.py` read JSON files capped at 100 MB251- Phase names extracted from log files are truncated to 200 characters and stripped of control characters to prevent prompt-injection payloads from propagating into agent context252- No scripts write to the filesystem; all output goes to stdout253254### Tool Restrictions255- **Read**: Used to inspect script source, references, simulation logs, and result files256- **Write**: Used to save profiling reports or optimization recommendations; writes are scoped to the user's working directory257- **Grep/Glob**: Used to locate log files, result files, and search references258- The skill's `allowed-tools` excludes `Bash` to prevent the agent from executing arbitrary commands when processing untrusted simulation logs or result files259260### Safety Measures261- No `eval()`, `exec()`, or dynamic code generation262- All subprocess calls use explicit argument lists (no `shell=True`)263- Reduced tool surface (no Bash) limits the agent to read/write operations only264- Phase names and diagnostic strings are sanitized before inclusion in output to prevent injection265266## Limitations267268- **Log parsing**: Depends on pattern matching; may miss unusual formats269- **Scaling analysis**: Requires at least 2 runs for meaningful results270- **Memory estimation**: Approximate; actual usage may vary271- **Recommendations**: General guidance; may need domain-specific tuning272273## References274275- `references/profiling_guide.md` - Profiling concepts and interpretation276- `references/optimization_strategies.md` - Detailed optimization approaches277278## Version History279280- **v1.0.0** (2025-01-22): Initial release with 4 profiling scripts