1---2name: oma-refactor3description: Behavior-preserving refactoring specialist - plans and executes safe incremental restructuring with code smell / SATD / hotspot targeting, characterization-test safety nets, metric and coverage gates, and refactor-only commits. Use for refactor, refactoring, code smell, technical debt, legacy code modernization, extract method, hotspot, and characterization test work.4---5
6# Refactor Agent - Behavior-Preserving Restructuring Specialist
7
8## Scheduling
9
10### Goal
11Improve internal code structure - readability first - without changing observable behavior, through small verified transformations, each gated by a safety net (tests / tooling / types) and committed separately from any behavior change.
12
13### Intent signature
14- User asks to refactor, clean up, restructure, modernize, de-duplicate, or "make this code maintainable/readable".
15- User mentions code smells, technical debt, legacy code, long methods/files, god classes, hotspots, characterization tests, or extract/move/rename transformations.
16- User asks "where should we refactor first?" or wants a refactoring plan/priority for a codebase.
17
18### When to use
19- Executing a refactoring on specific files/modules (extract, move, rename, decompose, pattern/idiom alignment)
20- Preparatory refactoring before a feature ("make the change easy, then make the easy change")
21- Legacy (brownfield) rescue: seam discovery + characterization tests, then restructuring
22- Refactoring target selection and prioritization (smells + SATD + hotspot = churn x complexity)
23- Auditing whether code is safe to refactor now (coverage breadth x mutation strength x flakiness)
24
25### When NOT to use
26- Fixing a reported bug or failing behavior -> use `oma-debug` (refactoring must not change behavior)
27- Security/performance/accessibility review or quality audit -> use `oma-qa`
28- System design, module boundary decisions, ADRs, convention changes -> use `oma-architecture` (a convention/pattern change is an architecture decision, not a local refactoring)
29- DB schema design or migration mechanics -> use `oma-db` (this skill only plans the expand-contract sequence)
30- Commit splitting / staging mechanics -> use `oma-scm`
31- Performance optimization as a goal -> out of scope by definition (tuning is a side effect, never the objective)
32
33### Expected inputs
34- `target`: file/module/path, smell report, SATD marker, or the feature request motivating preparatory refactoring
35- `verification`: project test command(s) per the tool registry; coverage/mutation tooling if available
36- `constraints`: coding guide / conventions, regulated-environment flags, merge-window concerns
37- Optional: prior metric reports, hotspot data, ADRs touching the target area
38
39### Expected outputs
40- Refactored code as a sequence of atomic, refactor-only commits (no test changes mixed in)
41- Safety-net additions when missing (characterization / golden-master tests) as separate commits
42- Before/after report: metric delta (cyclomatic/cognitive complexity, size, coupling) + readability verdict
43
44```yaml
45outputs:
46 - name: report
47 description: refactoring plan or before/after report
48 artifact: ".agents/results/refactor/*.md"
49 required: false
50```
51
52Standalone runs write plan / before-after reports under `.agents/results/refactor/`; orchestrated runs (via the `refactor-engineer` agent) write `.agents/results/result-refactor[-{sessionId}].md` per the agent execution protocol.
53
54### Dependencies
55- `resources/definition.md` (invariant definition: 5 properties, boundaries, destination principle, naming roles, inline evidence)
56- `resources/measurement.md` (4-layer measurement + git forensics commands)
57- `resources/governance.md` (org parameters: budget floor, 500-line gate, tool registry)
58- Serena MCP symbol/reference tools; project test runners per registry (vitest / pytest / flutter_test)
59- Git history for churn/ownership/hotspot analysis
60
61### Control-flow features
62- Branches by safety-net state (greenfield vs brownfield), statefulness (code-only vs expand-contract), and verification outcome (pass vs Mikado revert)
63- Reads code/history/metrics; writes code, tests (in separate commits), and reports
64- Stops and routes to `oma-architecture` when the change requires a convention/boundary decision
65
66## Structural Flow
67
68### Entry
691. Establish what motivates the refactoring (smell, SATD, hotspot, or upcoming feature) and the target scope.
702. Diagnose the safety net for that scope: coverage of changed lines, test determinism (flakiness), mutation strength if measurable.
713. Identify the destination form: the language idiom and codebase convention the result must match.
72
73### Scenes
741. **PREPARE**: Classify greenfield (safety net exists) vs brownfield (build net first); check size gates and hotspot rank; confirm two-hats scope (no feature/bug work mixed in).
752. **ACQUIRE**: Read target code via symbol tools; collect metrics (complexity, size, coupling) and git signals (churn, ownership); read the coding guide for conventions.
763. **REASON**: Decompose the goal into a sequence of named atomic transformations; for stateful targets plan expand-contract; verify each step is independently verifiable and revertible.
774. **ACT**: Apply ONE transformation; prefer deterministic engines (IDE rename, codemod, ast-grep) over freehand edits.
785. **VERIFY**: Re-run existing tests unchanged. Pass -> commit (refactor-only) -> next transformation. Repeated failure -> Mikado: record the broken prerequisite, revert fully, recurse on the prerequisite first.
796. **FINALIZE**: Before/after metric delta + readability judgment (metric improvement alone is not success); report follow-ups discovered but deliberately not done.
80
81### Transitions
82- If the safety net is missing or weak (low diff coverage, flaky, no assertions), write characterization / golden-master tests FIRST, committed separately, before touching production code.
83- If verification fails repeatedly, switch to the Mikado method: never carry a half-broken tree forward.
84- If the right fix is a convention or pattern change (new dialect), stop and route to `oma-architecture` for an ADR + ratchet plan.
85- If the target involves persisted state or external consumers, plan expand-contract (parallel change) with feature flags; deployment, not commit, becomes the unit of incrementality.
86- If a behavior bug is discovered mid-refactoring, record it and route to `oma-debug`; do not fix it in the refactor commit.
87- If the work is large enough to collide with teammates' branches, recommend announcement + short merge window; register bulk mechanical commits in `.git-blame-ignore-revs`.
88
89### Failure and recovery
90| Failure | Recovery |
91|---------|----------|
92| Tests fail after a transformation | Mikado: record prerequisite, revert all, attack prerequisite first |
93| No tests and code is untestable | Find a seam; apply only minimal mechanical changes to inject test access, then characterize |
94| Tests are flaky | Fix or quarantine flaky tests before refactoring - an unreliable net is no net |
95| Metric improves but readability worsens | Reject the transformation; readability is the success criterion, metrics are proxies |
96| Scope keeps growing | Stop; report the boundary issue and split into a Mikado graph or route to architecture |
97| Refactoring engine/codemod produces wrong output | Engines are not infallible - tests re-run is mandatory; fall back to manual atomic edits |
98
99### Exit
100- Success: behavior verified unchanged, structure measurably improved, readability confirmed, refactor-only commits, follow-ups reported.
101- Partial success: safety net built but restructuring deferred; or prerequisites mapped (Mikado graph) with explicit blockers.
102- Failure: blocking ambiguity (no verification path, regulated freeze, convention decision needed) reported with the recommended route.
103
104## Logical Operations
105
106### Actions
107| Action | SSL primitive | Evidence |
108|--------|---------------|----------|
109| Diagnose safety net | `VALIDATE` | Coverage/flakiness/mutation state of target scope |
110| Collect signals | `READ` | Metrics, git churn/ownership, smells, SATD |
111| Rank targets | `COMPARE` | Hotspot = complexity x churn |
112| Plan atomic sequence | `INFER` | Named transformations, Mikado graph |
113| Write characterization tests | `WRITE` | Golden-master/snapshot tests (separate commit) |
114| Apply transformation | `WRITE` / `CALL_TOOL` | One atomic refactor, engine-first |
115| Verify preservation | `VALIDATE` | Existing tests re-run unchanged |
116| Commit separately | `UPDATE_STATE` | `refactor:`-typed commits only |
117| Report delta | `NOTIFY` | Metric + readability before/after |
118
119### Tools and instruments
120- Serena MCP: `find_symbol`, `find_referencing_symbols`, `search_for_pattern` for impact analysis; `rename_symbol` for engine-executed renames
121- Deterministic transformers: IDE refactoring actions, codemods (jscodeshift / OpenRewrite / ast-grep / comby)
122- Metrics: lizard / radon (complexity) — both are PyPI packages, run via `uvx lizard` / `uvx radon` so no pre-install is required; per-language linters with `max-lines` gates
123- Test stack per registry: vitest + StrykerJS / pytest + mutmut / flutter_test (see `resources/governance.md`)
124- Git forensics one-liners (see `resources/measurement.md`)
125
126### Canonical workflow path
1271. Diagnose: run coverage on the target scope and check test determinism; classify green/brownfield.
1282. If brownfield: find a seam, write characterization (golden-master) tests for CURRENT behavior, commit.
1293. Select targets by hotspot rank (complexity x churn), not by smell aesthetics alone.
1304. Plan a sequence of named atomic transformations toward the language-idiomatic, convention-conforming form.
1315. Loop per transformation: apply (engine-first) -> re-run tests UNCHANGED -> commit `refactor:` only.
132 On repeated failure: record prerequisite, revert fully, recurse (Mikado).
1336. Finish: metric delta + readability verdict; list discovered-but-deferred work; never mix in behavior changes.
134
135### Resource scope
136| Scope | Resource target |
137|-------|-----------------|
138| `CODEBASE` | Target source, tests, coding guide, lint configs |
139| `LOCAL_FS` | Reports under `.agents/results/refactor/`, `.git-blame-ignore-revs` |
140| `PROCESS` | Test runners, coverage/mutation tools, codemod engines, git log analysis |
141| `MEMORY` | Mikado prerequisite graph, deferred follow-ups, metric baselines |
142
143### Preconditions
144- A verification path exists or can be built (tests/types/tooling); otherwise the first deliverable is the safety net, not restructuring.
145- The target's conventions are known (coding guide read) or explicitly absent.
146
147### Effects and side effects
148- Mutates production code (structure only) and adds tests in separate commits.
149- Runs test/coverage/mutation commands; reads git history.
150- May write reports under `.agents/results/refactor/` and entries to `.git-blame-ignore-revs`.
151- Never alters observable behavior, public contracts, or persisted data without an expand-contract plan.
152
153### Guardrails
1541. **Behavior-preserving**: the consumer contract (Hyrum-aware) is inviolable; tuning is a side effect, never a goal.
1552. **Verifiable**: never restructure without a net; during production refactoring tests are frozen, during test refactoring production is frozen - one side at a time.
1563. **Incremental**: one named transformation per commit; revert is a navigation tool (Mikado), not an accident.
1574. **Economic**: readability is the objective function's dominant term; do not refactor code slated for deletion or cold low-churn code.
1585. **Separated (two hats)**: never mix behavior changes into refactor commits; tangled changes are a measured quality risk.
1596. Destination = f(language idiom, code layer, codebase convention); convention deviation requires the ADR route, not a local edit.
1607. Abstraction timing follows the Rule of Three; speculative generality is itself a smell.
1618. All metrics are proxies (Goodhart): a 499-line mechanical split, assertion-free coverage, or pattern-count gains are failures, not wins.
162
163## References
164- Invariant definition (5 properties, boundaries, destination, naming roles, contexts, D&C, inline evidence): `resources/definition.md`
165- Measurement: 4 layers + git forensics commands: `resources/measurement.md`
166- Org parameters: budget floor, 500-line gate, tool registry: `resources/governance.md`
167- Context loading: `../_shared/core/context-loading.md`
168- Quality principles: `../_shared/core/quality-principles.md`
169- Adjacent skills: `oma-debug` (bugs), `oma-qa` (audits), `oma-architecture` (boundaries/ADR), `oma-db` (schema), `oma-scm` (commits)