Production Upgrade
Run a risk-adjusted, evidence-bound modernization from discovery through a
pre-release maintainer checkpoint. The workflow is complete without a specific
model, subagent implementation, or task tracker, but it must honor stronger
project requirements when they exist.
Overview
The quality bar comes from the Databricks rebuild: understand real operator
pain, decide architecture before implementation, move load-bearing logic into
deterministic code, preserve compatibility intentionally, and prove safety with
negative and adversarial evidence. Match the depth to risk rather than matching
another project's document count.
When to use
Use for a legacy artifact, broad rewrite, breaking migration, unsafe integration,
or production-readiness claim. Do not trigger for a small typo, isolated bug fix,
routine dependency bump, or read-only status question unless the user explicitly
requests the full upgrade workflow.
Prerequisites
- Target repository or artifact and its project instructions.
- Authority to research and prepare local changes. Publication, deployment,
merge, destructive cleanup, and external messaging remain separate approvals.
- Current primary sources for externally defined contracts.
The pre-authorized Read, Write, and Edit capabilities apply only to local,
scoped implementation files. Network, shell, task-tracker, subagent, install,
and publication capabilities remain behind host and project approval.
Orchestration
Use five focused roles. When the host supports isolated subagents, dispatch the
role packets under references/roles and have the coordinator
reconcile their evidence. Otherwise execute the same packets sequentially in the
main context. A same-identity or inline review is self-review, never independent.
| Role |
Mutation authority |
Output |
| Researcher |
None |
Source ledger, pain catalog, explicit gaps |
| Architect |
None |
Scope synthesis, decision record, migration plan |
| Implementation engineer |
Local scoped writes only |
Minimal implementation and focused tests |
| Verification engineer |
None |
Reproduced commands, results, and hashes |
| Security adversary |
None |
Threat-driven findings and exploit attempts |
Instructions
1. Recover context and establish authority
- Read repository instructions, architecture owners, generated-file rules,
current status, worktrees, active reviews, and existing task state.
- If the repository uses Beads, follow
references/beads-workflow.md: prime, search,
create or reuse, claim before mutation, and attach receipts. Project policy
can make Beads mandatory.
- Preserve dirty work and contributor authorship. Isolate broad changes in a
branch or worktree when available.
- Record the exact initial revision and the actions currently authorized.
2. Research before designing
- Audit the existing artifact, every active and retired capability, consumers,
package identities, installation paths, and known defects.
- Research current primary sources for product, API, protocol, security, and
runtime contracts. Add community or issue evidence for real operator pain
when accessible and appropriate.
- Build a pain catalog: symptom, trigger, root cause, blast radius, current
workaround, evidence, and the right agent primitive. Record source gaps
rather than filling them with assumptions.
- Compare at least one relevant production benchmark for methodology, then
explain where narrower or deeper treatment is justified by risk.
3. Decide scope and safety
- Consolidate capabilities around distinct operator outcomes, not quotas.
- Write an architecture decision covering adopted, modified, and rejected
alternatives; authority boundaries; compatibility; migration; rollback; and
explicit non-goals.
- Threat-model inputs, outputs, credentials, network destinations, file paths,
dependencies, retries, mutations, reviewers, evidence, and publication.
- Make offline or read-only behavior the default. Unknown contracts, statuses,
fields, destinations, or permissions fail closed.
- Put deterministic classification, arithmetic, validation, transformation,
and policy decisions in reviewed scripts. Use model reasoning for synthesis
and ambiguity, not for load-bearing calculations.
4. Plan and implement
- Define measurable acceptance gates before editing. Include structure,
behavior, security, migration, provenance, and release evidence.
- Implement the smallest complete design. Keep the portable core independent
of host adapters and avoid infrastructure that does not add verified
capability.
- Preserve IDs or provide a machine-readable migration map. Breaking behavior
requires an explicit major-version decision and user-facing migration path.
- Never generate plaintext secrets, remote-pipe installers, unbounded retries,
silent destructive actions, fabricated provider responses, or blanket
scanner waivers.
5. Validate proportionately
Run the narrowest focused tests first, then repository-required gates.
Cover positive, negative, edge, adversarial, failure, and rollback paths.
Deliberately broken variants must fail the same gate when certification is
claimed.
Validate generated projections, packaging file lists, installation from a
disposable path, and removal or rollback where those surfaces changed.
Reproduce every material automated-review finding independently. Reviewer
silence or billing failure is unavailable evidence, not approval.
Record evidence using
references/evidence-contract.md and audit
it without executing recorded commands:
python3 scripts/audit_evidence.py upgrade-evidence.json --root <repository>
6. Stop at the approval boundary
Report the exact candidate revision, changed surfaces, test results, unresolved
risks, reviewer status, migration impact, rollback, and publication state.
Do not commit, push, open or update a PR, merge, tag, publish, deploy, delete, or
message externally unless that action is authorized by the user and project
policy. High-risk release approval is bound to the exact revision; a changed
revision requires renewed approval.
Output
Return a concise executive status plus links to the research, decision, threat
model, migration map, tests, and evidence manifest. Use these claim levels:
BLOCKED: a required safety or authority boundary failed.
CANDIDATE: implementation and local evidence exist; independent review or
approval remains.
REVIEWED: independent review is bound to the exact revision; release is not
yet authorized.
RELEASE-READY: all required gates and exact-revision authorization exist.
Error handling
- Missing Beads when project policy requires it: stop before mutation.
- Missing subagents: execute role packets inline and label review self-review.
- Missing current primary source: constrain or remove the affected capability.
- Conflicting authorities: stop and resolve the conflict at the named owner.
- Failed test or unknown reviewer finding: remain
BLOCKED or CANDIDATE;
never average it into a score.
- Dirty unrelated work: preserve and isolate; do not reset or overwrite it.
Examples
- A narrow API pack may need fewer documents than Databricks but still requires
an official contract audit, threat model, migration map, adversarial tests,
exact-revision evidence, and explicit research gaps.
- An MCP server with destructive methods requires stronger input, authorization,
rollback, and live-boundary evidence than an offline read-only skill.
- A model-neutral skill can be manually used by any capable model, while named
native support remains limited to harnesses with registry-backed receipts.
Resources
- Beads workflow
- Evidence contract
- Runtime portability
- Specialist role packets
- Pain catalog template
- Decision record template
1---2name: production-upgrade3description: Upgrade an existing skill, plugin, agent, MCP integration, or agent-system package to a security-first production standard using pain research, architecture decisions, migration planning, deterministic implementation, adversarial tests, independent review, and revision-bound evidence. Use when modernizing a legacy capability or asking for Databricks-level diligence. Trigger with "production upgrade", "modernize this skill", "bring this pack to production quality", or "audit and rebuild this plugin".4license: MIT5---6
7# Production Upgrade
8
9Run a risk-adjusted, evidence-bound modernization from discovery through a
10pre-release maintainer checkpoint. The workflow is complete without a specific
11model, subagent implementation, or task tracker, but it must honor stronger
12project requirements when they exist.
13
14## Overview
15
16The quality bar comes from the Databricks rebuild: understand real operator
17pain, decide architecture before implementation, move load-bearing logic into
18deterministic code, preserve compatibility intentionally, and prove safety with
19negative and adversarial evidence. Match the depth to risk rather than matching
20another project's document count.
21
22## When to use
23
24Use for a legacy artifact, broad rewrite, breaking migration, unsafe integration,
25or production-readiness claim. Do not trigger for a small typo, isolated bug fix,
26routine dependency bump, or read-only status question unless the user explicitly
27requests the full upgrade workflow.
28
29## Prerequisites
30
31- Target repository or artifact and its project instructions.
32- Authority to research and prepare local changes. Publication, deployment,
33 merge, destructive cleanup, and external messaging remain separate approvals.
34- Current primary sources for externally defined contracts.
35
36The pre-authorized `Read`, `Write`, and `Edit` capabilities apply only to local,
37scoped implementation files. Network, shell, task-tracker, subagent, install,
38and publication capabilities remain behind host and project approval.
39
40## Orchestration
41
42Use five focused roles. When the host supports isolated subagents, dispatch the
43role packets under [references/roles](references/roles/) and have the coordinator
44reconcile their evidence. Otherwise execute the same packets sequentially in the
45main context. A same-identity or inline review is self-review, never independent.
46
47| Role | Mutation authority | Output |
48| ----------------------- | ------------------------ | ------------------------------------------------ |
49| Researcher | None | Source ledger, pain catalog, explicit gaps |
50| Architect | None | Scope synthesis, decision record, migration plan |
51| Implementation engineer | Local scoped writes only | Minimal implementation and focused tests |
52| Verification engineer | None | Reproduced commands, results, and hashes |
53| Security adversary | None | Threat-driven findings and exploit attempts |
54
55## Instructions
56
57### 1. Recover context and establish authority
58
591. Read repository instructions, architecture owners, generated-file rules,
60 current status, worktrees, active reviews, and existing task state.
612. If the repository uses Beads, follow
62 [references/beads-workflow.md](references/beads-workflow.md): prime, search,
63 create or reuse, claim before mutation, and attach receipts. Project policy
64 can make Beads mandatory.
653. Preserve dirty work and contributor authorship. Isolate broad changes in a
66 branch or worktree when available.
674. Record the exact initial revision and the actions currently authorized.
68
69### 2. Research before designing
70
711. Audit the existing artifact, every active and retired capability, consumers,
72 package identities, installation paths, and known defects.
732. Research current primary sources for product, API, protocol, security, and
74 runtime contracts. Add community or issue evidence for real operator pain
75 when accessible and appropriate.
763. Build a pain catalog: symptom, trigger, root cause, blast radius, current
77 workaround, evidence, and the right agent primitive. Record source gaps
78 rather than filling them with assumptions.
794. Compare at least one relevant production benchmark for methodology, then
80 explain where narrower or deeper treatment is justified by risk.
81
82### 3. Decide scope and safety
83
841. Consolidate capabilities around distinct operator outcomes, not quotas.
852. Write an architecture decision covering adopted, modified, and rejected
86 alternatives; authority boundaries; compatibility; migration; rollback; and
87 explicit non-goals.
883. Threat-model inputs, outputs, credentials, network destinations, file paths,
89 dependencies, retries, mutations, reviewers, evidence, and publication.
904. Make offline or read-only behavior the default. Unknown contracts, statuses,
91 fields, destinations, or permissions fail closed.
925. Put deterministic classification, arithmetic, validation, transformation,
93 and policy decisions in reviewed scripts. Use model reasoning for synthesis
94 and ambiguity, not for load-bearing calculations.
95
96### 4. Plan and implement
97
981. Define measurable acceptance gates before editing. Include structure,
99 behavior, security, migration, provenance, and release evidence.
1002. Implement the smallest complete design. Keep the portable core independent
101 of host adapters and avoid infrastructure that does not add verified
102 capability.
1033. Preserve IDs or provide a machine-readable migration map. Breaking behavior
104 requires an explicit major-version decision and user-facing migration path.
1054. Never generate plaintext secrets, remote-pipe installers, unbounded retries,
106 silent destructive actions, fabricated provider responses, or blanket
107 scanner waivers.
108
109### 5. Validate proportionately
110
1111. Run the narrowest focused tests first, then repository-required gates.
1122. Cover positive, negative, edge, adversarial, failure, and rollback paths.
113 Deliberately broken variants must fail the same gate when certification is
114 claimed.
1153. Validate generated projections, packaging file lists, installation from a
116 disposable path, and removal or rollback where those surfaces changed.
1174. Reproduce every material automated-review finding independently. Reviewer
118 silence or billing failure is unavailable evidence, not approval.
1195. Record evidence using
120 [references/evidence-contract.md](references/evidence-contract.md) and audit
121 it without executing recorded commands:
122
123 ```bash
124 python3 scripts/audit_evidence.py upgrade-evidence.json --root <repository>
125 ```
126
127### 6. Stop at the approval boundary
128
129Report the exact candidate revision, changed surfaces, test results, unresolved
130risks, reviewer status, migration impact, rollback, and publication state.
131Do not commit, push, open or update a PR, merge, tag, publish, deploy, delete, or
132message externally unless that action is authorized by the user and project
133policy. High-risk release approval is bound to the exact revision; a changed
134revision requires renewed approval.
135
136## Output
137
138Return a concise executive status plus links to the research, decision, threat
139model, migration map, tests, and evidence manifest. Use these claim levels:
140
141- `BLOCKED`: a required safety or authority boundary failed.
142- `CANDIDATE`: implementation and local evidence exist; independent review or
143 approval remains.
144- `REVIEWED`: independent review is bound to the exact revision; release is not
145 yet authorized.
146- `RELEASE-READY`: all required gates and exact-revision authorization exist.
147
148## Error handling
149
150- Missing Beads when project policy requires it: stop before mutation.
151- Missing subagents: execute role packets inline and label review self-review.
152- Missing current primary source: constrain or remove the affected capability.
153- Conflicting authorities: stop and resolve the conflict at the named owner.
154- Failed test or unknown reviewer finding: remain `BLOCKED` or `CANDIDATE`;
155 never average it into a score.
156- Dirty unrelated work: preserve and isolate; do not reset or overwrite it.
157
158## Examples
159
160- A narrow API pack may need fewer documents than Databricks but still requires
161 an official contract audit, threat model, migration map, adversarial tests,
162 exact-revision evidence, and explicit research gaps.
163- An MCP server with destructive methods requires stronger input, authorization,
164 rollback, and live-boundary evidence than an offline read-only skill.
165- A model-neutral skill can be manually used by any capable model, while named
166 native support remains limited to harnesses with registry-backed receipts.
167
168## Resources
169
170- [Beads workflow](references/beads-workflow.md)
171- [Evidence contract](references/evidence-contract.md)
172- [Runtime portability](references/runtime-portability.md)
173- [Specialist role packets](references/roles/)
174- [Pain catalog template](templates/pain-catalog.md)
175- [Decision record template](templates/decision-record.md)