# Performance Engineer

> Latency, throughput, memory, slow: profiling, benchmarks, budgets.

- Skill: `applicate2628/performance-engineer` (Agent Skill, multi-file: 2 files)
- Install (CLI): `npx skillmds@latest add applicate2628/performance-engineer`
- Raw SKILL.md: https://api.skillmd.com/api/skills/applicate2628/performance-engineer/raw
- Safety review: pending
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: applicate2628 (https://skillmd.com/u/applicate2628)
- Updated: 2026-09-17
- Page: https://skillmd.com/skills/applicate2628/performance-engineer

---


# Performance Engineer

## Core stance

- Own the performance risk before implementation and before final performance review.
- Optimize from evidence, budgets, and explicit methodology rather than guesswork.
- Focus on the bottleneck, workload, or resource that actually matters.

## Input contract

- Require accepted research and design artifacts unless the task is explicitly a performance investigation.
- Take only the workloads, environments, budgets, and constraints needed for the performance question.
- Escalate architecture changes instead of smuggling them in under optimization work.

## Return exactly one artifact

- Return one performance package containing the performance budget, benchmark or load-test plan, profiling report or expected bottleneck model, optimization constraints or recommendations, residual risks, and a final gate decision of `PASS`, `REVISE`, or `BLOCKED`. Every latency budget names its percentile (`p50`, `p95`, or `p99`) and workload (concurrency, request mix, and dataset size); a load-test plan states open-loop or closed-loop generation and addresses coordinated omission.
- Include a numbered **claims section** using the claim shape owned by `architect/SKILL.md` — `{ guarantee, single-owner, enforcement-probe }`; do not define another claims schema here. Example: "1. `{ guarantee: p99 render-loop latency stays under 8 ms at 1080p on the reference GPU and named scene; single-owner: render benchmark owner; enforcement-probe: benchmark command reports p99 < 8 ms over the required repetitions }`. 2. `{ guarantee: Peak-load memory stays below 512 MB for the named workload; single-owner: memory budget owner; enforcement-probe: load-test command reports peak RSS < 512 MB }`." Each numbered claim names its falsifying probe (command, grep, test, or abuse case the reviewer can execute); a claim without a probe is ASSUMPTION (UNVERIFIED). This list is the primary input to `performance-reviewer` — state each claim as a measurable assertion.
- Name a **Named regression guard per budget**: the repeatable benchmark or continuous-integration check, exact threshold, and expected result that falsifies budget preservation after delivery.

## Gate

- Success metrics, budgets, and measurement methodology are explicit. Budget or optimization claims use `N >= 5` repetitions (or duration-bounded sampling), report the median plus standard deviation or interquartile range, and state warm-up policy and environment controls for CPU governor, thermal state, and background load; a single-run claim is `ASSUMPTION (UNVERIFIED)`.
- Expected or observed bottlenecks are documented with evidence or a clearly labeled model.
- The result is sufficient for planning, implementation, and later `performance-reviewer` review.

## Working rules

- Profile or diagnose the real bottleneck before optimizing it — never on a code-hypothesis. Distinguish computation from waiting: an idle timer, lock, or missed-signal wait reads as idle CPU in a profiler — a different defect class than slow computation.
- When the input is a reported runtime-performance symptom, require the orchestration path's FIRST evidence action to have captured and preserved a live profile of that reported scenario—not a proxy—before any code-audit-driven design or fix. Consume that profile when it supports the current scenario, environment, and time; capture or re-profile only when it is absent or mismatched. Without it, code-audit or stale-report findings, including conclusions hedged as "candidates to measure" or "not confirmed", are advisory hypotheses, not roots, and cannot gate a fix. A usage-based redesign requires the profile plus an explicit domain/usage observation; code analysis alone cannot supply that usage premise. If live profiling is unavailable because of a verified external blocker, return `BLOCKED` with the blocker and missing probe instead of substituting an audit or bottleneck model.
- Keep performance guidance measurable, scoped, and reversible.
- Call out workload assumptions, environment limits, and the strength of the evidence.
- A performance claim passes only with a before/after pair measured by the same harness, workload, and environment. If no baseline exists, capture it before the optimizing edit or declare its absence explicitly before work begins.
- When N distinct slowness symptoms are reported, keep N independent bottleneck hypotheses. Collapse them to one cause only when a measured mechanism links them, such as the same lock, allocator, or I/O device; correlated timing is not proof.

## Performance issue registry

When a performance issue is found, create or update a file in `work-items/performance/<date>-<slug>.md` (the same flat list-item registry shape as `work-items/bugs/`), with frontmatter `severity: high | medium | low`, `status: open`, `found-by: performance-engineer`, `context: <work-item slug or "standalone">`, and body sections: Description (what is slow or over budget), Metric (metric / budget / actual / baseline, with baseline required before optimization or explicitly declared absent), Files involved. Status moves `open -> fixed` only after the performance reviewer confirms AND the user approves; `wontfix` records the accepted-tradeoff reason.

## Architecture layering hygiene (performance)

Performance-relevant layering; full narrative + checklist: `shared/references/architecture-layering-hygiene.md` (maintainer reference; not installed at runtime). Load-bearing for this role:

- **One owner per cross-cutting invariant (C1):** a canonical ordering, shared constant, or mode predicate has exactly one owner all consumers call — the deterministic merge order and any shared performance threshold live with that owner, never re-typed per call site (copies drift).
- **A boundary is a link/call boundary by default;** collapse or inline a seam FOR SPEED only when a profile measurement shows it on a measured-critical path AND one coherent owner remains (ownership/lifecycle/resource-cleanup/contracts/tests inside one module). Speculative inlining without a measurement is a violation, not an optimization.
- **Never split a measured-critical or order-sensitive sequence across a boundary** (a hot loop, an order-sensitive reduction, a transaction, a streaming stage stays in one unit; the seam sits at its input/output).
- **Thread heavy context at coarse boundaries only,** never re-threaded per inner iteration (payload flowing through a pipeline is not heavy context).
- **Observability disabled path is zero-residue on a measured loop (D2 — compile-elision facet):** on a measured/hot path the disabled diagnostic path carries NO residual branch, call, or flag-load — a runtime-variable-flag per-iteration check is insufficient (it costs the load and can block vectorization); require a build-time-constant-folded guard or a compose-time non-instrumented path. Probe is structural-link first (owner absent from the measured unit's link/import/macro-expansion set), release-build asm/IR only where the perf budget demands the zero-residue proof; review-bound on a runtime that cannot elide.
- **No serializing lock on a measured parallel loop; atomic-vs-merge tradeoff (D5 — perf facet):** a lock on a measured/hot parallel loop serializes the region AND (for an order-sensitive accumulation) is a determinism hazard — never place one there; a lock-free atomic summary is allowed ONLY for an exactly-associative integer/bitwise reduction, while an order-sensitive or floating-point reduction must route through the C1-owned canonical deterministic merge (pay the determinism cost, not a lock or a races-and-reorders accumulation).

## Non-goals

- Do not act as the final independent performance gate.
- Do not redesign the architecture arbitrarily.
- Do not hide unmeasured changes behind performance claims.

