Performance Engineering Program
Purpose
Make good performance practice survive personnel changes. The output is an adoption program with
owners, evidence and exit criteria, not a maturity badge or a universal process imposed on every
service.
Workflow
- Define the decisions the program must improve: release safety, SLO protection, capacity,
incident response or cost efficiency.
- Assess evidence by service and dimension: user objective, representative baseline, regression
gate, production observability, ownership/runbook and learning loop. Reuse existing artifacts,
owners and accepted decisions; distinguish missing evidence from demonstrated absence.
- Find the weakest dependency. An advanced profiler does not compensate for an undefined SLO or a
gate that silently passes without data. Limit a blocker to the adoption decision that depends
on it; independent evidence capture or enablement can proceed.
- Choose one adoption wave with named services, owners, artifacts, support and measurable exit
criteria. Pilot before standardizing.
- Build enablement: maintained templates, office hours, reviewed examples and rotating champions
with protected time and escalation support.
- Measure outcomes and unwanted incentives; revise the mechanism rather than gaming the score.
Decision rules
- Treat maturity as an evidence inventory. Never average away a missing safety-critical dimension.
- Use ordinal levels only to communicate; retain the underlying evidence and gaps for decisions.
- Standardize contracts and required fields, not one tool or one numeric threshold across unlike
workloads. Inspect service criticality, runtime/tool versions and deployment constraints before
applying a template; the program does not authorize upgrades or new release approval gates.
- A CI performance gate is not adopted until it has representative evidence, calibrated noise,
explicit metric direction,
pass/regression/inconclusive, and baseline ownership.
- A champion program distributes judgment only when champions practice on real services, rotate at
a cadence that permits depth, and have a maintained escalation path.
- Do not use SLO attainment, incidents or maturity scores for individual performance evaluation.
That incentive encourages denominator changes, exclusions and suppressed reporting.
- Count activity metrics only alongside outcomes: training attendance and review coverage do not
prove fewer regressions or faster diagnosis.
Program scorecard
For each metric record definition, population, source/query, owner, cadence, target, missing-data
behavior and the decision it changes. Useful outcomes include regression escape rate, time from
signal to useful evidence, percentage of critical services with exercised runbooks, and recovery
of performance budgets. Report uncertainty and avoid causal claims from simple correlation.
Version the eligible-service/release population, observation window and exclusions. Show missing
coverage separately: fewer reported regressions after detection is disabled is not improvement.
Do not average service percentiles into an organization-wide percentile.
Deliver the current evidence gaps, a bounded next wave with capacity/owners, its applicable exit
and pause/revision criteria, and checks actually exercised versus planned. Keep the deliverable
proportionate; an assessment alone need not create a new central approval process.
References
- Maturity and rollout — read when designing an assessment,
adoption waves, champion rotation or the program scorecard.
- Use
slo-and-alerting, performance-regression-ci, continuous-profiling and
performance-incident-response for their respective artifacts.
1---2name: performance-engineering-program3description: Establishing an organization-wide performance engineering program through measurable maturity evidence, service ownership, SLO and baseline adoption, regression gates, incident learning and a rotating champion model. Use when performance depends on one specialist, teams apply different evidence standards, a maturity assessment needs concrete next actions, or a rollout must turn isolated profiling into a durable operating discipline. Does not design individual SLOs, benchmarks, alerts or profiles; their specialist skills own those artifacts.4---56# Performance Engineering Program78## Purpose910Make good performance practice survive personnel changes. The output is an adoption program with11owners, evidence and exit criteria, not a maturity badge or a universal process imposed on every12service.1314## Workflow15161. Define the decisions the program must improve: release safety, SLO protection, capacity,17 incident response or cost efficiency.182. Assess evidence by service and dimension: user objective, representative baseline, regression19 gate, production observability, ownership/runbook and learning loop. Reuse existing artifacts,20 owners and accepted decisions; distinguish missing evidence from demonstrated absence.213. Find the weakest dependency. An advanced profiler does not compensate for an undefined SLO or a22 gate that silently passes without data. Limit a blocker to the adoption decision that depends23 on it; independent evidence capture or enablement can proceed.244. Choose one adoption wave with named services, owners, artifacts, support and measurable exit25 criteria. Pilot before standardizing.265. Build enablement: maintained templates, office hours, reviewed examples and rotating champions27 with protected time and escalation support.286. Measure outcomes and unwanted incentives; revise the mechanism rather than gaming the score.2930## Decision rules3132- Treat maturity as an evidence inventory. Never average away a missing safety-critical dimension.33- Use ordinal levels only to communicate; retain the underlying evidence and gaps for decisions.34- Standardize contracts and required fields, not one tool or one numeric threshold across unlike35 workloads. Inspect service criticality, runtime/tool versions and deployment constraints before36 applying a template; the program does not authorize upgrades or new release approval gates.37- A CI performance gate is not adopted until it has representative evidence, calibrated noise,38 explicit metric direction, `pass/regression/inconclusive`, and baseline ownership.39- A champion program distributes judgment only when champions practice on real services, rotate at40 a cadence that permits depth, and have a maintained escalation path.41- Do not use SLO attainment, incidents or maturity scores for individual performance evaluation.42 That incentive encourages denominator changes, exclusions and suppressed reporting.43- Count activity metrics only alongside outcomes: training attendance and review coverage do not44 prove fewer regressions or faster diagnosis.4546## Program scorecard4748For each metric record definition, population, source/query, owner, cadence, target, missing-data49behavior and the decision it changes. Useful outcomes include regression escape rate, time from50signal to useful evidence, percentage of critical services with exercised runbooks, and recovery51of performance budgets. Report uncertainty and avoid causal claims from simple correlation.52Version the eligible-service/release population, observation window and exclusions. Show missing53coverage separately: fewer reported regressions after detection is disabled is not improvement.54Do not average service percentiles into an organization-wide percentile.5556Deliver the current evidence gaps, a bounded next wave with capacity/owners, its applicable exit57and pause/revision criteria, and checks actually exercised versus planned. Keep the deliverable58proportionate; an assessment alone need not create a new central approval process.5960## References6162- [Maturity and rollout](references/maturity-and-rollout.md) — read when designing an assessment,63 adoption waves, champion rotation or the program scorecard.64- Use `slo-and-alerting`, `performance-regression-ci`, `continuous-profiling` and65 `performance-incident-response` for their respective artifacts.