Performance Benchmarking
Overview
This skill defines how to systematically measure and track system performance. Without benchmarks, performance degradation goes unnoticed until it impacts users. Regular benchmarking ensures changes improve rather than degrade performance.
When to Use
- Before launching any user-facing feature or service
- After making changes that could affect performance (algorithms, data access, queries)
- When users report slowness or degraded responsiveness
- Before committing to a technology or approach with performance implications
- As part of regular system health monitoring
- Don't use when: Performance is clearly not a concern (static content, one-time operations)
Core Procedures
Step 1: Define Metrics
Select relevant metrics for what's being measured:
- Response Time: p50, p90, p95, p99 latencies
- Throughput: Requests/operations per second
- Resource Usage: CPU, memory, disk I/O, network, database connections
- Error Rate: Percentage of failed operations
- Scalability: How metrics change as load increases
Step 2: Establish Baseline
- Define test scenario (load, data set, environment)
- Run measurement under normal conditions
- Record results as baseline
- Ensure environment is stable and representative of production
- Document test conditions for reproducibility
Step 3: Execute Benchmark
- Run test under defined conditions (same as baseline)
- Run multiple iterations (minimum 5) for statistical significance
- Record all metrics and environmental conditions
- Save raw results for analysis
Step 4: Compare and Analyze
- Did performance improve, degrade, or stay flat vs baseline?
- Are any metrics outside acceptable thresholds?
- If performance changed, identify WHAT changed (not just THAT it changed)
- For degradation: is the regression acceptable for the improvement gained?
Step 5: Set Thresholds
Define acceptable ranges for each metric:
PERFORMANCE THRESHOLD
=====================
Metric: [response time, throughput, etc.]
Baseline: [typical value]
Warning: [degradation level that triggers investigation]
Critical: [degradation level that triggers rollback]
Target: [improvement goal]
Quality Checklist
Error Handling
- Error: Performance degrades significantly after change
Response: Do NOT deploy, diagnose root cause, fix or revert
- Error: Results are inconsistent between runs
Response: Investigate environmental factors, isolate variable, stabilize test conditions
- Error: No baseline data available
Response: Measure current state as baseline, cannot assess change impact until next measurement
Cross-Team Integration
Related Skills: testing-strategy, edge-case-analysis, incident-response, continuous-improvement
Used By: Performance testing engineers, developers, infrastructure teams, load testers, profiler agents
1---2name: performance-benchmarking3description: Use when evaluating, measuring, or comparing the performance of systems, functions, or services. This skill provides a framework for establishing baselines, measuring performance, and validating that changes meet performance requirements.4---56# Performance Benchmarking78## Overview9This skill defines how to systematically measure and track system performance. Without benchmarks, performance degradation goes unnoticed until it impacts users. Regular benchmarking ensures changes improve rather than degrade performance.1011## When to Use12- Before launching any user-facing feature or service13- After making changes that could affect performance (algorithms, data access, queries)14- When users report slowness or degraded responsiveness15- Before committing to a technology or approach with performance implications16- As part of regular system health monitoring17- **Don't use when:** Performance is clearly not a concern (static content, one-time operations)1819## Core Procedures2021### Step 1: Define Metrics22Select relevant metrics for what's being measured:23- **Response Time:** p50, p90, p95, p99 latencies24- **Throughput:** Requests/operations per second25- **Resource Usage:** CPU, memory, disk I/O, network, database connections26- **Error Rate:** Percentage of failed operations27- **Scalability:** How metrics change as load increases2829### Step 2: Establish Baseline301. Define test scenario (load, data set, environment)312. Run measurement under normal conditions323. Record results as baseline334. Ensure environment is stable and representative of production345. Document test conditions for reproducibility3536### Step 3: Execute Benchmark371. Run test under defined conditions (same as baseline)382. Run multiple iterations (minimum 5) for statistical significance393. Record all metrics and environmental conditions404. Save raw results for analysis4142### Step 4: Compare and Analyze43- Did performance improve, degrade, or stay flat vs baseline?44- Are any metrics outside acceptable thresholds?45- If performance changed, identify WHAT changed (not just THAT it changed)46- For degradation: is the regression acceptable for the improvement gained?4748### Step 5: Set Thresholds49Define acceptable ranges for each metric:50```51PERFORMANCE THRESHOLD52=====================53Metric: [response time, throughput, etc.]54Baseline: [typical value]55Warning: [degradation level that triggers investigation]56Critical: [degradation level that triggers rollback]57Target: [improvement goal]58```5960## Quality Checklist61- [ ] Relevant metrics selected for system type62- [ ] Baseline established with documented conditions63- [ ] Tests run multiple times for statistical reliability64- [ ] Results compared against baseline65- [ ] Thresholds defined for ongoing monitoring66- [ ] Raw results saved for trend analysis6768## Error Handling69- **Error:** Performance degrades significantly after change70 **Response:** Do NOT deploy, diagnose root cause, fix or revert71- **Error:** Results are inconsistent between runs72 **Response:** Investigate environmental factors, isolate variable, stabilize test conditions73- **Error:** No baseline data available74 **Response:** Measure current state as baseline, cannot assess change impact until next measurement7576## Cross-Team Integration77**Related Skills:** testing-strategy, edge-case-analysis, incident-response, continuous-improvement78**Used By:** Performance testing engineers, developers, infrastructure teams, load testers, profiler agents