Codex compatibility note:
- Invoke repository skills with
$skill-name in Codex; this mirrored copy rewrites legacy Claude /skill-name references.
- Task tracker mandate: BEFORE executing any workflow or skill step, create/update task tracking for all steps and keep it synchronized as progress changes.
- User-question prompts mean to ask the user directly in Codex.
- Ignore Claude-specific mode-switch instructions when they appear.
- Strict execution contract: when a user explicitly invokes a skill, execute that skill protocol as written.
- Subagent authorization: when a skill is user-invoked or AI-detected and its protocol requires subagents, that skill activation authorizes use of the required
spawn_agent subagent(s) for that task.
- Do not skip, reorder, or merge protocol steps unless the user explicitly approves the deviation first.
- For workflow skills, execute each listed child-skill step explicitly and report step-by-step evidence.
- If a required step/tool cannot run in this environment, stop and ask the user before adapting.
Codex Project-Reference Loading (No Hooks)
Codex uses static project-reference loading instead of runtime-injected project docs.
When coding, planning, debugging, testing, or reviewing, open project docs explicitly using this routing.
Always read:
docs/project-config.json (project-specific paths, commands, modules, and workflow/test settings)
docs/project-reference/docs-index-reference.md (routes to the full docs/project-reference/* catalog)
docs/project-reference/lessons.md (always-on guardrails and anti-patterns)
Missing/stale context route: If docs/project-config.json, the docs index, lessons.md, CLAUDE.md, AGENTS.md, or any task-required reference doc is missing or stale, auto-run $project-init or the narrow setup route ($project-config, $docs-init, $scan-all, $scan --target=<key>, $claude-md-init) before ordinary project-specific work. If Codex mirrors or AGENTS.md are missing/stale, ask the user to run $sync-codex; do not auto-run it.
Situation-based docs:
- Project structure/architecture/tech-stack/deployment/setup (any layer — backend, frontend, or infra):
project-structure-reference.md
- Backend/CQRS/API/domain/entity changes:
backend-patterns-reference.md, domain-entities-reference.md
- Frontend/UI/styling/design-system:
frontend-patterns-reference.md, scss-styling-guide.md, design-system/README.md
- Spec authoring,
docs/specs/ pathing, or TC format: feature-spec-reference.md, spec-system-reference.md, spec-principles.md
- Behavior/public-contract changes or spec-test-code sync:
workflow-spec-test-code-cycle-reference.md plus the spec docs above
- Derived spec indexes/ERDs/reimplementation guides:
spec-system-reference.md and source Feature Specs under docs/specs/
- Integration test implementation/review:
integration-test-reference.md
- E2E test implementation/review:
e2e-test-reference.md
- Code review/audit work:
code-review-rules.md plus domain docs above based on changed files
Do not read all docs blindly. Start from docs-index-reference.md, then open only relevant files for the task.
[BLOCKING] Execute skill steps in declared order. NEVER skip, reorder, or merge steps without explicit user approval.
[BLOCKING] Before each step or sub-skill call, update task tracking: set in_progress when step starts, set completed when step ends.
[BLOCKING] Every completed/skipped step MUST include brief evidence or explicit skip reason.
[BLOCKING] If Task tools are unavailable, create and maintain an equivalent step-by-step plan tracker with the same status transitions.
Quick Summary
Goal: Ensure service/API changes are production-ready for observability, reliability, data integrity, and database performance — scoring each of these dimensions on service-layer and API changes.
Summary:
- Main steps (in order): (1) Resolve scope — args else
git diff --name-only uncommitted; backend service/API files only, skip frontend/tests/docs/config-only. (2) Score 12 criteria 0-2 across the 4 dimensions (/24). (3) Extended SRE Readiness gate — 8 pass/fail deploy-time + operate-time items; an unaccepted CRITICAL/HIGH fail blocks PASS regardless of the /24 score. Gating, NOT scored — does not change the /24 math. (4) Map score + gate → verdict. (5) Structural Impact Analysis — graph gate (blast-radius, tests_for, downstream trace) when graph.db exists. (6) Validated Fix + Full Re-Review loop on any finding. (7) Emit the SRE Review Results report — file:line evidence per score and per gate item. Execute in order; NEVER skip/merge a step — why: untracked steps get silently merged and gaps reach production.
- Score 12 criteria 0-2 across four dimensions (Observability/8, Reliability/8, Data Integrity/4, DB Performance/4) for a /24 PASS (19-24) / NEEDS WORK (13-18) / NOT READY (0-12) verdict — every score needs
file:line evidence or it is 0.
- The DB Performance Protocol is MANDATORY and non-advisory: ALL list queries must paginate (no unbounded GetAll/ToList) and ALL filter fields, foreign keys, and sort columns must have matching indexes.
- VERDICT is advisory only; the graph gate, validated-fix full re-review, and DB Performance Protocol are NEVER skippable regardless of change size — and when batched (≥10 files), re-score all 12 criteria holistically from combined cross-batch evidence, never by averaging per-batch scores.
- After applying any fix, validate findings first, then rerun the FULL review (fresh sub-agent with zero prior-round memory); a clean pass ENDS the loop.
When to use: After implementing backend service or API changes, before committing. Frontend-only changes exempt.
Why: Working code that can't be debugged, monitored, or rolled back is technical debt in disguise.
Deployment context: Read docs/project-config.json → infrastructure section:
containerization → check Dockerfiles, docker-compose
orchestration → check K8s manifests, Helm charts
cicd.tool → check pipeline configs
Your Mission
Review Mindset (NON-NEGOTIABLE)
Be skeptical. Every claim needs traced proof, confidence >80%.
- NEVER accept operational readiness at face value — verify by reading implementations
- Every score MUST have
file:line evidence — unprovable score = 0
- Question: "Is this really handled?" → trace error/retry/timeout path to confirm
- Challenge: "Are ALL failure modes covered?" → check behavior when dependencies fail
- Verify: "Can we debug this in production?" → check logging, correlation, metrics
Scope Resolution
- Arguments specify files/directories → review those
- Else → review uncommitted changes (
git diff --name-only)
- Focus: backend source files under service root (per the project's structure reference /
docs/project-config.json), API controllers, service classes
- Skip: frontend files, test files, documentation, config-only changes
Production Readiness Scoring
Score each criterion 0-2: 0 = not addressed, 1 = partially, 2 = fully.
MANDATORY when batched (≥10 files, SYNC:systematic-review-batching active): score the 12 criteria holistically across the FULL cross-batch scope, NOT by merging or averaging per-batch scores. Several criteria are cross-file — e.g. "all query filter fields have indexes" can have the query in one batch, the migration in another; a per-batch score sees only its ≤8 files and false-flags 0 when the satisfying file lives in a different batch. The synthesis/reduce tier MUST therefore RE-SCORE each of the 12 criteria from combined cross-batch evidence (batch agents surface evidence per criterion; reducer assigns the score). If holistic re-score is infeasible, do NOT batch production-readiness-review — fall back to whole-scope serial scoring.
Observability (max 8)
Think: If this service errors at 3am, can on-call engineer diagnose root cause from logs alone — without reproducing?
| # |
Criterion |
What to Check |
| 1 |
Structured Logging |
External API calls and critical operations log errors with context (request ID, user, parameters) |
| 2 |
Error Context |
Exceptions include enough context to diagnose without reproducing (entity IDs, operation type, input summary) |
| 3 |
Metrics Awareness |
Operations >100ms consider tracking duration. New endpoints consider latency monitoring |
| 4 |
Correlation |
Cross-service calls include or propagate correlation IDs for distributed tracing |
Reliability (max 8)
Think: If the downstream dependency is down or slow, does this service degrade gracefully or cascade-fail?
| # |
Criterion |
What to Check |
| 5 |
Retry Strategy |
Transient failures (HTTP, DB timeouts) have retry logic or documented reason for not retrying |
| 6 |
Timeout Configuration |
HTTP clients and external calls have explicit timeout (not relying on defaults) |
| 7 |
Error Handling |
Errors handled gracefully — no swallowed exceptions, no generic catch-all without logging |
| 8 |
Fallback Behavior |
Critical paths define behavior when dependencies fail (degraded mode, cached response, user-facing error) |
Data Integrity (max 4)
Think: If database wiped and reseeded from scratch, does system still reach a valid state?
| # |
Criterion |
What to Check |
| 9 |
Seed vs Migration |
Seed data (default records, system config) lives in startup data seeders, NOT in one-time migration executors |
| 10 |
Seeder Idempotency |
Data seeders use check-then-create pattern (query before insert) — safe for repeated runs on any environment |
Decision test: "If the database is reset, does this data still need to exist?" Yes → must be in seeder. No → migration acceptable.
Database Performance (max 4)
Think: At 10x current data volume, do these queries still complete in <1s?
Database Performance Protocol (MANDATORY):
- Paging Required — ALL list/collection queries use pagination. NEVER load all records into memory. Verify: no unbounded
GetAll(), ToList(), or Find() without Skip/Take or cursor-based paging.
- Index Required — ALL query filter fields, foreign keys, and sort columns have database indexes configured. Verify: entity expressions match index field order, database collections have index management methods, migrations include indexes for WHERE/JOIN/ORDER BY columns.
| # |
Criterion |
What to Check |
| 11 |
Pagination |
List/collection queries use pagination (Skip/Take, cursor). No unbounded GetAll/ToList loading all records into memory |
| 12 |
Database Indexes |
Query filter fields, foreign keys, and sort columns have matching database indexes. Migrations include index creation |
Spec-Loop Discipline for changed core logic (MANDATORY — gates the verdict, not a scored criterion):
- Mutation bar, not coverage % — for changed service/API core logic the bar is the MUTATION-SCORE gate: a surviving mutant on a changed line is a release blocker (it proves an invariant the tests do not assert), NEVER a line-coverage-% question. A green coverage number over un-asserted behavior does not clear this gate.
- Dual feedback — every production-readiness finding that changes behavior feeds BOTH the spec (NAME the contract/invariant in Section 8) AND a guarding test; a code-only fix is INCOMPLETE. A surviving mutant → add the killing test AND record the invariant it protects in the spec.
Extended SRE Readiness Gate (step-by-step, pass/fail — gating, NOT scored)
Runs as main step 3, after scoring, before verdict mapping. Deploy-time and operate-time SRE aspects the 12-criteria /24 model does NOT score. Check each item step by step; record pass / partial / fail with file:line evidence or explicit N/A — reason. Gate does not change /24 math — it overlays it: an unaccepted CRITICAL/HIGH fail blocks a PASS verdict regardless of score (per Severity Rubric — CRITICAL/HIGH must be resolved or owner-accepted before PASS). Read deployment context from docs/project-config.json → infrastructure (referenced above) to decide which items are N/A (e.g. no orchestration → readiness/liveness probes N/A with stated reason).
| # |
Gate Item |
What to Check |
Status |
Evidence |
| G1 |
Rollout & Rollback |
Deploy is staged/canary-able; a documented, fast rollback path exists (feature flag, versioned + reversible migration). No irreversible one-way change without a stated recovery plan. |
pass/partial/fail |
file:line or N/A — reason |
| G2 |
Health Checks |
Readiness + liveness endpoints/probes exist and reflect real dependency health (not an always-200 stub). |
pass/partial/fail |
... |
| G3 |
Alerting & Runbook |
New failure modes have an actionable alert (signal, not noise) and a runbook / escalation note. |
pass/partial/fail |
... |
| G4 |
SLO / Error-Budget |
Change respects an SLO or names the latency/availability target it affects; no silent new failure mode against the budget. |
pass/partial/fail |
... |
| G5 |
Capacity & Resource Limits |
Load ceilings, resource limits, autoscaling/back-pressure considered; no unbounded fan-out or unbounded in-memory growth. |
pass/partial/fail |
... |
| G6 |
Config & Secrets |
Required config present in all envs and fails fast if missing; no secrets committed in the diff. |
pass/partial/fail |
... |
| G7 |
Graceful Shutdown/Startup |
In-flight work drains on shutdown; startup waits for / degrades gracefully on unready dependencies. |
pass/partial/fail |
... |
| G8 |
Concurrency & Idempotency |
Operations are safe under retry / at-least-once delivery; no race on shared state; idempotency keys where needed. |
pass/partial/fail |
... |
Gate verdict: {n}/8 pass. Any CRITICAL/HIGH fail not explicitly owner-accepted ⇒ overall verdict cannot be PASS even at a 19-24 score.
Technique Applicability (advisory — NON-SCORING, NON-GATING)
Invoke SYNC:scale-technique-gate: derive the system's scale tier from evidence (users/RPS, SLO, data volume, tenancy, topology — cite file:line/config/infra + confidence), then emit the Technique Applicability Matrix (technique | tier-warranted? | present? | verdict | advice | evidence) across the 10 concern groups. Surface warranted-but-missing reliability/scale techniques (rate limiting, backups, DR, failover, graceful degradation) as advice; flag OVER-ENGINEERED techniques the tier does not warrant.
Advisory only — this matrix does NOT add a gate item, does NOT change the {n}/8 gate result, the /24 score, or the verdict. A MISSING-WARRANTED technique is guidance to consider at this tier, NOT a gate fail. N/A-by-scale for small systems is expected, never a failure. Full catalog → .claude/docs/scale-technique-catalog.md.
Scoring
| Score |
Verdict |
Recommendation |
| 19-24 |
PASS |
Production-ready. Proceed to commit. |
| 13-18 |
NEEDS WORK |
Address gaps before deploying to production. OK for dev/staging. |
| 0-12 |
NOT READY |
Significant operational gaps. Review Operational Readiness rules in code-review-rules.md. |
Run python .claude/scripts/code_graph connections <file> --json on service boundary files for cross-service impact.
Structural Impact Analysis (MANDATORY when graph.db exists)
python .claude/scripts/code_graph graph-blast-radius --json → blast radius >20 nodes = high-risk deployment
python .claude/scripts/code_graph query tests_for <function_name> --json → verify test coverage on changed functions
python .claude/scripts/code_graph trace <service-file> --direction downstream --json → verify all downstream event handlers, bus consumers, cross-service calls have error handling
Why-Review Findings Validation Gate (MANDATORY when findings exist)
Purpose: Adversarial validation of own findings BEFORE any fix. Catches over-flagged criteria, false positives, and severity/score inflation at the source rather than letting them drive fixes or ship downstream.
Trigger: Any finding produced (any severity). Skip ONLY when the verdict is unconditional PASS with literally zero findings.
Protocol:
- Read own finalized report from
plans/reports/{skill}-{date}-{slug}.md
- Invoke
$why-review --validate-findings plans/reports/{skill}-{date}-{slug}.md — verify each finding has file:line proof, steel-man each rejected interpretation, and stress-test every severity/score classification (each finding must clear why-review's finding-survival bar to be kept)
- Read the CLEAN / HAS-ISSUES verdict returned by why-review
- If why-review demotes/removes any finding: UPDATE own report with revised severities, remove false positives, and add a
## Why-Review Validation Notes section citing what changed and why
- If why-review confirms all findings: append a
## Why-Review Validation line stating "All N findings re-validated against actual code; no severity changes."
Skip conditions (record explicit reason if skipping): unconditional PASS with zero findings; why-review is itself the active context (avoid recursion).
Why this exists: SRE sub-agent reports inherit confirmation bias — the orchestrator absorbs severity claims as ground truth. Validate findings BEFORE the fix so no fix is ever driven by an inflated or false finding; this gate feeds the "Validated Fix + Full Re-Review" loop below.
Validated Fix + Full Re-Review (MANDATORY when fixes are applied)
When a review pass finds issues, validate findings before any fix. Do NOT spawn a fresh sub-agent only to re-review the same finding set before validation/fix. After validated SRE fixes applied, rerun the full SRE review. If that restarted review uses a sub-agent, spawn it with ZERO prior-round memory. A clean review pass ENDS the review.
When a fresh sub-agent is part of the restarted review, spawn via canonical template in SYNC:review-protocol-injection:
agent_type: code-reviewer
- Task:
"SRE production readiness review after validated fixes — score all 12 criteria (0-2) for {files reviewed in the current full scope}"
- Review mode:
"Fresh full re-review after validated fixes. Zero memory of prior rounds. Re-read ALL target files from scratch."
- Reference Docs:
docs/project-reference/code-review-rules.md
- Target Files: same files from Scope Resolution
- Integrate sub-agent report findings — DO NOT filter or override
Fresh re-review focus (what prior rounds typically miss):
- Operational concerns spanning multiple services
- Subtle reliability gaps (retry, circuit breakers, timeout handling)
- Missing observability (structured logging, correlation IDs, metrics)
- Data-integrity edge cases under concurrent load
Final verdict = every review pass that actually ran, combined.
Output Format
## SRE Review Results
**Scope:** {files reviewed}
**Date:** {date}
**Score:** {X}/24
**Verdict:** PASS / NEEDS WORK / NOT READY
### Observability ({X}/8)
| # | Criterion | Score | Evidence |
| --- | ------------------ | ----- | -------------------------- |
| 1 | Structured Logging | 0/1/2 | {file:line or "not found"} |
| 2 | Error Context | 0/1/2 | ... |
| 3 | Metrics Awareness | 0/1/2 | ... |
| 4 | Correlation | 0/1/2 | ... |
### Reliability ({X}/8)
| # | Criterion | Score | Evidence |
| --- | ----------------- | ----- | -------- |
| 5 | Retry Strategy | 0/1/2 | ... |
| 6 | Timeout Config | 0/1/2 | ... |
| 7 | Error Handling | 0/1/2 | ... |
| 8 | Fallback Behavior | 0/1/2 | ... |
### Data Integrity ({X}/4)
| # | Criterion | Score | Evidence |
| --- | ------------------ | ----- | -------- |
| 9 | Seed vs Migration | 0/1/2 | ... |
| 10 | Seeder Idempotency | 0/1/2 | ... |
### Database Performance ({X}/4)
| # | Criterion | Score | Evidence |
| --- | ---------------- | ----- | -------- |
| 11 | Pagination | 0/1/2 | ... |
| 12 | Database Indexes | 0/1/2 | ... |
### Extended SRE Readiness ({n}/8 gate — pass/fail, does not change /24)
| # | Gate Item | Status | Evidence |
| --- | -------------------------- | ----------------- | ----------------- |
| G1 | Rollout & Rollback | pass/partial/fail | `file:line` / N/A |
| G2 | Health Checks | pass/partial/fail | ... |
| G3 | Alerting & Runbook | pass/partial/fail | ... |
| G4 | SLO / Error-Budget | pass/partial/fail | ... |
| G5 | Capacity & Resource Limits | pass/partial/fail | ... |
| G6 | Config & Secrets | pass/partial/fail | ... |
| G7 | Graceful Shutdown/Startup | pass/partial/fail | ... |
| G8 | Concurrency & Idempotency | pass/partial/fail | ... |
_Any unaccepted CRITICAL/HIGH `fail` above blocks a PASS verdict regardless of the /24 score._
### Gaps to Address
- {specific actionable item}
### Recommendation
{Proceed / Address gaps first}
Important Notes
- Advisory (final VERDICT only) — score/verdict inform team but don't block commits; MANDATORY process steps (graph gate, validated-fix full re-review, Database Performance Protocol) are NEVER advisory
- Evidence-based — cite
file:line for every score; unprovable score = 0
- Proportional — small bug fixes need less rigor than new endpoints (applies to VERDICT interpretation, NOT to skipping MANDATORY steps)
- Extended SRE Readiness gate is pass/fail, NOT scored — does not change
/24 math; but an unaccepted CRITICAL/HIGH gate fail blocks a PASS verdict (Severity Rubric). Use docs/project-config.json → infrastructure to mark items N/A with stated reason
- Check framework patterns — background-job base handlers, base-controller error handling
Workflow Recommendation
MANDATORY — NO EXCEPTIONS: If NOT already in workflow, use ask the user directly to ask user:
- Activate
workflow-feature workflow (Recommended) — scout → investigate → plan → feature-implement → review → production-readiness-review → test → docs
- Execute
$production-readiness-review directly — run standalone
Next Steps
MANDATORY — NO EXCEPTIONS — after completing, use ask the user directly:
- "$watzup (Recommended)" — wrap up + check doc staleness
- "$test" — run tests before wrapping up
- "Skip, continue manually" — user decides
Combined audit: For a whole-project architecture + compliance + production-readiness audit in one pass, run $architecture-review-full (or $start-workflow workflow-architecture-audit) — fans out this skill, architecture-review, architecture-scalability-review as parallel sub-agents and synthesizes one consolidated report.
[IMPORTANT] Use task tracking to break ALL work into small tasks BEFORE starting. For simple tasks, AI MUST ask user whether to skip.
docs/project-reference/domain-entities-reference.md — Domain entity catalog, relationships, cross-service sync (read when task involves business entities/models)
Critical Purpose: Ensure quality — no flaws, no bugs, no missing updates, no stale content. Verify code AND documentation.
External Memory: Complex/lengthy work → write intermediate findings + final results to plans/reports/ — prevents context loss, serves as deliverable.
Evidence Gate: MANDATORY — every claim, finding, recommendation requires file:line proof or traced evidence with confidence percentage (>80% to act, <80% verify first).
Graph-Assisted Investigation — MANDATORY when .code-graph/graph.db exists.
HARD-GATE: MUST ATTENTION run at least ONE graph command on key files before concluding any investigation.
Pattern: Grep finds files → trace --direction both reveals full system flow → Grep verifies details
| Task |
Minimum Graph Action |
| Investigation/Scout |
trace --direction both on 2-3 entry files |
| Fix/Debug |
callers_of on buggy function + tests_for |
| Feature/Enhancement |
connections on files to be modified |
| Code Review |
tests_for on changed functions |
| Blast Radius |
trace --direction downstream |
CLI: python .claude/scripts/code_graph {command} --json. Use --node-mode file first (10-30x less noise), then --node-mode function for detail.
Sub-Agent Return Contract — When this skill spawns a sub-agent, the sub-agent MUST return ONLY this structure. Main agent reads only this summary — NEVER requests full sub-agent output inline.
## Sub-Agent Result: [skill-name]
Status: ✅ PASS | ⚠️ PARTIAL | ❌ FAIL
Confidence: [0-100]%
### Findings (Critical/High only — max 10 bullets)
- [severity] [file:line] [finding]
### Actions Taken
- [file changed] [what changed]
### Blockers (if any)
- [blocker description]
Full report: plans/reports/[skill-name]-[date]-[slug].md
Main agent reads Full report file ONLY when: (a) resolving a specific blocker, or (b) building a fix plan.
Sub-agent writes full report incrementally (per SYNC:incremental-persistence) — not held in memory.
Context budget — the return payload is a SUMMARY, not a transcript: ≤10 finding bullets, no raw file contents / full diffs / verbatim logs inline, no re-pasted source. Everything beyond the summary lives in the Full report on disk. A sub-agent that would exceed the summary shape MUST write the detail to its report and return only the pointer — the orchestrator's context is the scarce resource the whole map-reduce protects.
Nested Task Expansion Contract — For workflow-step invocation, the [Workflow] ... row is only a parent container; the child skill still creates visible phase tasks.
- Call the current task list first. If a matching active parent workflow row exists, set
nested=true and record parentTaskId; otherwise run standalone.
- Create one task per declared phase before phase work. When nested, prefix subjects
[N.M] $skill-name — phase.
- When nested, link the parent with
TaskUpdate(parentTaskId, addBlockedBy: [childIds]).
- Orchestrators must pre-expand a child skill's phase list and link the workflow row before invoking that child skill or sub-agent.
- Mark exactly one child
in_progress before work and completed immediately after evidence is written.
- Complete the parent only after all child tasks are completed or explicitly cancelled with reason.
Blocked until: the current task list done, child phases created, parent linked when nested, first child marked in_progress.
Project Reference Docs Gate — Run after task-tracking bootstrap and before target/source file reads, grep, edits, or analysis. Project docs override generic framework assumptions.
- Identify scope: file types, domain area, and operation.
- Read
docs/project-config.json first — the project's machine-readable map. It is the single source of truth for THIS repo (modules/paths, framework + search keywords, test/E2E/integration run-commands, design system, architecture rules, workflow patterns); ground exact paths, run-commands, and conventions on it before investigating, planning, or coding — never assume framework defaults (CLAUDE.md + reference docs are derived from it). If it — or the docs index, lessons.md, CLAUDE.md, AGENTS.md, or any required reference doc — is missing or stale, auto-run $project-init or the narrow route ($project-config, $docs-init, $scan-all, $scan --target=<key>, $claude-md-init) first; if Codex mirrors or AGENTS.md are stale, ask the user to run $sync-codex (never auto-run it).
- Required docs by trigger: always
docs/project-reference/lessons.md; doc lookup docs-index-reference.md; review code-review-rules.md; backend/CQRS/API backend-patterns-reference.md; domain/entity domain-entities-reference.md; frontend/UI frontend-patterns-reference.md; styles/design scss-styling-guide.md + design-system/design-system-canonical.md; integration tests integration-test-reference.md; E2E e2e-test-reference.md; feature docs/specs feature-spec-reference.md + spec-system-reference.md + spec-principles.md; behavior/public-contract/spec-test-code sync workflow-spec-test-code-cycle-reference.md; derived spec index/ERD/reimplementation guides spec-system-reference.md + source Feature Specs under docs/specs/; architecture/new area project-structure-reference.md.
- Read every required doc, then before target work state:
Reference docs read: ... | Not applicable: ....
Ready when: scope evaluated, docs/project-config.json consulted, required docs checked/read or setup route completed, lessons.md confirmed, citation emitted.
Task Tracking & External Report Persistence — Bootstrap this before execution; then run project-reference doc prefetch before target/source work.
- Create a small task breakdown before target file reads, grep, edits, or analysis. On context loss, inspect the current task list first.
- Mark one task
in_progress before work and completed immediately after evidence; never batch transitions.
- For plan/review work, create
plans/reports/{skill}-{YYMMDD}-{HHmm}-{slug}.md before first finding.
- Append findings after each file/section/decision and synthesize from the report file at the end.
- Final output cites
Full report: plans/reports/{filename}.
Blocked until: task breakdown exists, report path declared for plan/review work, first finding persisted before the next finding.
Critical Thinking Mindset — Apply critical thinking, sequential thinking. Every claim needs traced proof, confidence >80% to act.
Anti-hallucination: Never present guess as fact — cite sources for every claim, admit uncertainty freely, self-check output for errors, cross-reference independently, stay skeptical of own confidence — certainty without evidence root of all hallucination.
Evidence-Based Reasoning — Speculation is FORBIDDEN. Every claim needs proof.
- Cite
file:line, grep results, or framework docs for EVERY claim
- Declare confidence: >80% act freely, 60-80% verify first, <60% DO NOT recommend
- Cross-service validation required for architectural changes
- "I don't have enough evidence" is valid and expected output
BLOCKED until: - [ ] Evidence file path (file:line) - [ ] Grep search performed - [ ] 3+ similar patterns found - [ ] Confidence level stated
Forbidden without proof: "obviously", "I think", "should be", "probably", "this is because"
If incomplete → output: "Insufficient evidence. Verified: [...]. Not verified: [...]."
Validated-Finding Fix + Full Re-Review Loop — Re-review is triggered by a validated finding fix cycle, not by a round number. Review purpose: review → validate findings → fix validated findings → full re-review until a complete review pass clears the round's exit bar (see Severity floor below). A clean review ENDS the loop — no further rounds required.
aka Self-Review Convergence Loop. The name is historical — there is NO 2-round cap; "double-round-trip" only means a validated-finding fix cycle forces at least one fresh re-review. It runs until a clean pass, bounded by the 3-round ceiling below.
Round cap — 3 rounds MAX (a ceiling, NEVER a target). A clean pass ENDS the loop immediately at ANY round — round 1 included; the cap never obliges you to keep spinning. Hitting round 3 with blocking findings still open (severity floor applied) → STOP and escalate by asking the user directly with the still-open findings listed; NEVER emit a silent "good enough" PASS on cap exhaustion, and NEVER let the cap substitute for the clean-review requirement. The 2-repeated-no-progress blocker rule stays an EARLIER exit — escalate at whichever trips first.
Severity floor — from round 3, LOW stops blocking. The exit bar tightens by round, so the loop converges on consequence instead of spinning on polish:
Define one predicate everywhere: blocking_findings(round, findings) returns all validated findings in rounds 1–2 and only validated CRITICAL/HIGH/MEDIUM findings in round 3+. A binary gate (test-green, security must-fix, required artifact) is exempt only when its owning invariant explicitly says so.
| Round |
Exit bar — loop ENDS when the fresh full review has… |
Must be fixed to continue |
| 1-2 |
zero validated findings at ANY severity |
CRITICAL · HIGH · MEDIUM · LOW |
| 3+ |
zero validated CRITICAL / HIGH / MEDIUM findings — LOW-only is a PASS |
CRITICAL · HIGH · MEDIUM only |
From round 3 onward LOW findings are NOT required to be fixed: a round whose validated findings are ALL LOW ENDS the loop immediately — do not open another round for them. Severity tiers are SYNC:severity-rubric (CRITICAL block-merge · HIGH must-fix · MEDIUM should-fix · LOW nice-to-fix); rounds 1-2 are unchanged, so an easy LOW still gets fixed early when it is cheap.
Severity-floor rules:
- Never silently drop a deferred LOW. Every unfixed LOW is listed in the final report under
## Deferred LOW Findings (severity floor, round ≥3) with file, line, and description, so the owner can schedule it. Dropping it from the report is a protocol violation, not a clean pass.
- Never re-tier a finding to trigger the exit. Downgrading a real CRITICAL/HIGH/MEDIUM to LOW so the loop can end is a FALSE PASS. Severity is set by consequence per
SYNC:severity-rubric before the round bar is applied — never after, and never with the exit in view. — why: a floor that can be reached by relabeling is not a floor.
- The floor bounds the loop, not the standard. It ends iteration; it never authorizes shipping a known CRITICAL/HIGH/MEDIUM, and it never lowers the finding-survival bar that admits a finding in the first place.
- The floor never applies to a hard gate. Test-green gates (a suite must actually pass), security must-fix gates, and any gate whose criterion is binary rather than severity-rated are unaffected — a failing test is a failure, not a LOW finding.
Universal scope (any new output/judgment): any newly produced output or judgment gets ≥1 self-review; any new judgment gets ≥1 $why-review --validate-findings pass; anything flagged to re-check is re-checked ≥1 time — before that output is treated as final. This loop is the default convergence contract for ANY work-producing skill, not review skills only.
Routing invariant (author-facing): a skill that validates findings MUST route them through $why-review --validate-findings (the terminal validator) — NEVER fork an inline finding-validation. Routing through why-review is what makes the finding-survival bar and this loop apply; the verify-review-validate-coverage sensor enforces this exact route mechanically.
Round 1: Main-session review. Read target files, build understanding, note issues. Output findings + verdict (PASS / FAIL).
Decision after Round 1:
- No issues found (PASS, zero findings) → review ENDS. Do NOT spawn a fresh sub-agent for confirmation.
blocking_findings(round, findings) is non-empty → run the active review skill's findings-validation gate first; for review skills the default gate is $why-review --validate-findings <report-path>. Fix only validated findings, then restart the full review protocol from the beginning with a fresh task breakdown.
Fresh full re-review after every fix cycle: Re-run the whole review protocol over the current full target. When sub-agents are part of that protoc
…(truncated)
1---2name: production-readiness-review3description: [Code Quality] Use when reviewing service-layer and API changes for production readiness.4---56> Codex compatibility note:7>8> - Invoke repository skills with `$skill-name` in Codex; this mirrored copy rewrites legacy Claude `/skill-name` references.9> - Task tracker mandate: BEFORE executing any workflow or skill step, create/update task tracking for all steps and keep it synchronized as progress changes.10> - User-question prompts mean to ask the user directly in Codex.11> - Ignore Claude-specific mode-switch instructions when they appear.12> - Strict execution contract: when a user explicitly invokes a skill, execute that skill protocol as written.13> - Subagent authorization: when a skill is user-invoked or AI-detected and its protocol requires subagents, that skill activation authorizes use of the required `spawn_agent` subagent(s) for that task.14> - Do not skip, reorder, or merge protocol steps unless the user explicitly approves the deviation first.15> - For workflow skills, execute each listed child-skill step explicitly and report step-by-step evidence.16> - If a required step/tool cannot run in this environment, stop and ask the user before adapting.1718<!-- CODEX:PROJECT-REFERENCE-LOADING:START -->1920## Codex Project-Reference Loading (No Hooks)2122Codex uses static project-reference loading instead of runtime-injected project docs.23When coding, planning, debugging, testing, or reviewing, open project docs explicitly using this routing.2425**Always read:**2627- `docs/project-config.json` (project-specific paths, commands, modules, and workflow/test settings)28- `docs/project-reference/docs-index-reference.md` (routes to the full `docs/project-reference/*` catalog)29- `docs/project-reference/lessons.md` (always-on guardrails and anti-patterns)3031**Missing/stale context route:** If `docs/project-config.json`, the docs index, `lessons.md`, `CLAUDE.md`, `AGENTS.md`, or any task-required reference doc is missing or stale, auto-run `$project-init` or the narrow setup route (`$project-config`, `$docs-init`, `$scan-all`, `$scan --target=<key>`, `$claude-md-init`) before ordinary project-specific work. If Codex mirrors or `AGENTS.md` are missing/stale, ask the user to run `$sync-codex`; do not auto-run it.3233**Situation-based docs:**3435- Project structure/architecture/tech-stack/deployment/setup (any layer — backend, frontend, or infra): `project-structure-reference.md`36- Backend/CQRS/API/domain/entity changes: `backend-patterns-reference.md`, `domain-entities-reference.md`37- Frontend/UI/styling/design-system: `frontend-patterns-reference.md`, `scss-styling-guide.md`, `design-system/README.md`38- Spec authoring, `docs/specs/` pathing, or TC format: `feature-spec-reference.md`, `spec-system-reference.md`, `spec-principles.md`39- Behavior/public-contract changes or spec-test-code sync: `workflow-spec-test-code-cycle-reference.md` plus the spec docs above40- Derived spec indexes/ERDs/reimplementation guides: `spec-system-reference.md` and source Feature Specs under `docs/specs/`41- Integration test implementation/review: `integration-test-reference.md`42- E2E test implementation/review: `e2e-test-reference.md`43- Code review/audit work: `code-review-rules.md` plus domain docs above based on changed files4445Do not read all docs blindly. Start from `docs-index-reference.md`, then open only relevant files for the task.4647<!-- CODEX:PROJECT-REFERENCE-LOADING:END -->4849<!-- PROMPT-ENHANCE:STEP-TASK-ANCHOR:START -->5051> **[BLOCKING]** Execute skill steps in declared order. NEVER skip, reorder, or merge steps without explicit user approval.52> **[BLOCKING]** Before each step or sub-skill call, update task tracking: set `in_progress` when step starts, set `completed` when step ends.53> **[BLOCKING]** Every completed/skipped step MUST include brief evidence or explicit skip reason.54> **[BLOCKING]** If Task tools are unavailable, create and maintain an equivalent step-by-step plan tracker with the same status transitions.5556<!-- PROMPT-ENHANCE:STEP-TASK-ANCHOR:END -->5758## Quick Summary5960**Goal:** Ensure service/API changes are production-ready for observability, reliability, data integrity, and database performance — scoring each of these dimensions on service-layer and API changes.6162**Summary:**6364- **Main steps (in order):** (1) **Resolve scope** — args else `git diff --name-only` uncommitted; backend service/API files only, skip frontend/tests/docs/config-only. (2) **Score 12 criteria 0-2** across the 4 dimensions (/24). (3) **Extended SRE Readiness gate** — 8 pass/fail deploy-time + operate-time items; an unaccepted CRITICAL/HIGH fail blocks PASS regardless of the /24 score. Gating, NOT scored — does not change the /24 math. (4) **Map score + gate → verdict**. (5) **Structural Impact Analysis** — graph gate (blast-radius, `tests_for`, downstream trace) when `graph.db` exists. (6) **Validated Fix + Full Re-Review** loop on any finding. (7) **Emit the SRE Review Results report** — `file:line` evidence per score and per gate item. Execute in order; NEVER skip/merge a step — why: untracked steps get silently merged and gaps reach production.65- Score 12 criteria 0-2 across four dimensions (Observability/8, Reliability/8, Data Integrity/4, DB Performance/4) for a /24 PASS (19-24) / NEEDS WORK (13-18) / NOT READY (0-12) verdict — every score needs `file:line` evidence or it is 0.66- The DB Performance Protocol is MANDATORY and non-advisory: ALL list queries must paginate (no unbounded GetAll/ToList) and ALL filter fields, foreign keys, and sort columns must have matching indexes.67- VERDICT is advisory only; the graph gate, validated-fix full re-review, and DB Performance Protocol are NEVER skippable regardless of change size — and when batched (≥10 files), re-score all 12 criteria holistically from combined cross-batch evidence, never by averaging per-batch scores.68- After applying any fix, validate findings first, then rerun the FULL review (fresh sub-agent with zero prior-round memory); a clean pass ENDS the loop.6970**When to use:** After implementing backend service or API changes, before committing. Frontend-only changes exempt.7172**Why:** Working code that can't be debugged, monitored, or rolled back is technical debt in disguise.7374**Deployment context:** Read `docs/project-config.json` → `infrastructure` section:7576- `containerization` → check Dockerfiles, docker-compose77- `orchestration` → check K8s manifests, Helm charts78- `cicd.tool` → check pipeline configs7980## Your Mission8182<task>83$ARGUMENTS84</task>8586## Review Mindset (NON-NEGOTIABLE)8788**Be skeptical. Every claim needs traced proof, confidence >80%.**8990- NEVER accept operational readiness at face value — verify by reading implementations91- Every score MUST have `file:line` evidence — unprovable score = 092- Question: "Is this really handled?" → trace error/retry/timeout path to confirm93- Challenge: "Are ALL failure modes covered?" → check behavior when dependencies fail94- Verify: "Can we debug this in production?" → check logging, correlation, metrics9596## Scope Resolution97981. Arguments specify files/directories → review those992. Else → review uncommitted changes (`git diff --name-only`)1003. Focus: backend source files under service root (per the project's structure reference / `docs/project-config.json`), API controllers, service classes1014. Skip: frontend files, test files, documentation, config-only changes102103## Production Readiness Scoring104105Score each criterion 0-2: **0** = not addressed, **1** = partially, **2** = fully.106107> **MANDATORY when batched (≥10 files, `SYNC:systematic-review-batching` active):** score the 12 criteria **holistically across the FULL cross-batch scope**, NOT by merging or averaging per-batch scores. Several criteria are cross-file — e.g. "all query filter fields have indexes" can have the query in one batch, the migration in another; a per-batch score sees only its ≤8 files and false-flags `0` when the satisfying file lives in a different batch. The synthesis/reduce tier MUST therefore **RE-SCORE each of the 12 criteria from combined cross-batch evidence** (batch agents surface evidence per criterion; reducer assigns the score). If holistic re-score is infeasible, do NOT batch production-readiness-review — fall back to whole-scope serial scoring.108109### Observability (max 8)110111> **Think:** If this service errors at 3am, can on-call engineer diagnose root cause from logs alone — without reproducing?112113| # | Criterion | What to Check |114| --- | ---------------------- | ------------------------------------------------------------------------------------------------------------- |115| 1 | **Structured Logging** | External API calls and critical operations log errors with context (request ID, user, parameters) |116| 2 | **Error Context** | Exceptions include enough context to diagnose without reproducing (entity IDs, operation type, input summary) |117| 3 | **Metrics Awareness** | Operations >100ms consider tracking duration. New endpoints consider latency monitoring |118| 4 | **Correlation** | Cross-service calls include or propagate correlation IDs for distributed tracing |119120### Reliability (max 8)121122> **Think:** If the downstream dependency is down or slow, does this service degrade gracefully or cascade-fail?123124| # | Criterion | What to Check |125| --- | ------------------------- | --------------------------------------------------------------------------------------------------------- |126| 5 | **Retry Strategy** | Transient failures (HTTP, DB timeouts) have retry logic or documented reason for not retrying |127| 6 | **Timeout Configuration** | HTTP clients and external calls have explicit timeout (not relying on defaults) |128| 7 | **Error Handling** | Errors handled gracefully — no swallowed exceptions, no generic catch-all without logging |129| 8 | **Fallback Behavior** | Critical paths define behavior when dependencies fail (degraded mode, cached response, user-facing error) |130131### Data Integrity (max 4)132133> **Think:** If database wiped and reseeded from scratch, does system still reach a valid state?134135| # | Criterion | What to Check |136| --- | ---------------------- | ------------------------------------------------------------------------------------------------------------- |137| 9 | **Seed vs Migration** | Seed data (default records, system config) lives in startup data seeders, NOT in one-time migration executors |138| 10 | **Seeder Idempotency** | Data seeders use check-then-create pattern (query before insert) — safe for repeated runs on any environment |139140**Decision test:** _"If the database is reset, does this data still need to exist?"_ Yes → must be in seeder. No → migration acceptable.141142### Database Performance (max 4)143144> **Think:** At 10x current data volume, do these queries still complete in <1s?145146> **Database Performance Protocol (MANDATORY):**147>148> 1. **Paging Required** — ALL list/collection queries use pagination. NEVER load all records into memory. Verify: no unbounded `GetAll()`, `ToList()`, or `Find()` without `Skip/Take` or cursor-based paging.149> 2. **Index Required** — ALL query filter fields, foreign keys, and sort columns have database indexes configured. Verify: entity expressions match index field order, database collections have index management methods, migrations include indexes for WHERE/JOIN/ORDER BY columns.150151| # | Criterion | What to Check |152| --- | -------------------- | ---------------------------------------------------------------------------------------------------------------------- |153| 11 | **Pagination** | List/collection queries use pagination (Skip/Take, cursor). No unbounded GetAll/ToList loading all records into memory |154| 12 | **Database Indexes** | Query filter fields, foreign keys, and sort columns have matching database indexes. Migrations include index creation |155156> **Spec-Loop Discipline for changed core logic (MANDATORY — gates the verdict, not a scored criterion):**157>158> 1. **Mutation bar, not coverage %** — for changed service/API core logic the bar is the **MUTATION-SCORE gate**: a surviving mutant on a changed line is a release blocker (it proves an invariant the tests do not assert), NEVER a line-coverage-% question. A green coverage number over un-asserted behavior does not clear this gate.159> 2. **Dual feedback** — every production-readiness finding that changes behavior feeds BOTH the spec (NAME the contract/invariant in Section 8) AND a guarding test; a code-only fix is INCOMPLETE. A surviving mutant → add the killing test AND record the invariant it protects in the spec.160161## Extended SRE Readiness Gate (step-by-step, pass/fail — gating, NOT scored)162163> **Runs as main step 3, after scoring, before verdict mapping.** Deploy-time and operate-time SRE aspects the 12-criteria `/24` model does NOT score. Check each item **step by step**; record `pass` / `partial` / `fail` with `file:line` evidence or explicit `N/A — reason`. Gate does **not** change `/24` math — it overlays it: **an unaccepted CRITICAL/HIGH `fail` blocks a PASS verdict regardless of score** (per Severity Rubric — CRITICAL/HIGH must be resolved or owner-accepted before PASS). Read deployment context from `docs/project-config.json → infrastructure` (referenced above) to decide which items are `N/A` (e.g. no orchestration → readiness/liveness probes `N/A` with stated reason).164165| # | Gate Item | What to Check | Status | Evidence |166| --- | ------------------------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------- | ----------------------------- |167| G1 | **Rollout & Rollback** | Deploy is staged/canary-able; a documented, fast rollback path exists (feature flag, versioned + reversible migration). No irreversible one-way change without a stated recovery plan. | pass/partial/fail | `file:line` or `N/A — reason` |168| G2 | **Health Checks** | Readiness + liveness endpoints/probes exist and reflect real dependency health (not an always-200 stub). | pass/partial/fail | ... |169| G3 | **Alerting & Runbook** | New failure modes have an actionable alert (signal, not noise) and a runbook / escalation note. | pass/partial/fail | ... |170| G4 | **SLO / Error-Budget** | Change respects an SLO or names the latency/availability target it affects; no silent new failure mode against the budget. | pass/partial/fail | ... |171| G5 | **Capacity & Resource Limits** | Load ceilings, resource limits, autoscaling/back-pressure considered; no unbounded fan-out or unbounded in-memory growth. | pass/partial/fail | ... |172| G6 | **Config & Secrets** | Required config present in all envs and fails fast if missing; no secrets committed in the diff. | pass/partial/fail | ... |173| G7 | **Graceful Shutdown/Startup** | In-flight work drains on shutdown; startup waits for / degrades gracefully on unready dependencies. | pass/partial/fail | ... |174| G8 | **Concurrency & Idempotency** | Operations are safe under retry / at-least-once delivery; no race on shared state; idempotency keys where needed. | pass/partial/fail | ... |175176**Gate verdict:** `{n}/8 pass`. Any CRITICAL/HIGH `fail` not explicitly owner-accepted ⇒ overall verdict cannot be PASS even at a 19-24 score.177178## Technique Applicability (advisory — NON-SCORING, NON-GATING)179180Invoke `SYNC:scale-technique-gate`: derive the system's scale tier from evidence (users/RPS, SLO, data volume, tenancy, topology — cite `file:line`/config/infra + confidence), then emit the **Technique Applicability Matrix** (`technique | tier-warranted? | present? | verdict | advice | evidence`) across the 10 concern groups. Surface warranted-but-missing reliability/scale techniques (rate limiting, backups, DR, failover, graceful degradation) as **advice**; flag `OVER-ENGINEERED` techniques the tier does not warrant.181182> **Advisory only — this matrix does NOT add a gate item, does NOT change the `{n}/8` gate result, the `/24` score, or the verdict.** A `MISSING-WARRANTED` technique is guidance to consider at this tier, NOT a gate `fail`. `N/A-by-scale` for small systems is expected, never a failure. Full catalog → `.claude/docs/scale-technique-catalog.md`.183184## Scoring185186| Score | Verdict | Recommendation |187| ----- | -------------- | ----------------------------------------------------------------------------------------- |188| 19-24 | **PASS** | Production-ready. Proceed to commit. |189| 13-18 | **NEEDS WORK** | Address gaps before deploying to production. OK for dev/staging. |190| 0-12 | **NOT READY** | Significant operational gaps. Review Operational Readiness rules in code-review-rules.md. |191192> Run `python .claude/scripts/code_graph connections <file> --json` on service boundary files for cross-service impact.193194## Structural Impact Analysis (MANDATORY when graph.db exists)195196- `python .claude/scripts/code_graph graph-blast-radius --json` → blast radius >20 nodes = high-risk deployment197- `python .claude/scripts/code_graph query tests_for <function_name> --json` → verify test coverage on changed functions198- `python .claude/scripts/code_graph trace <service-file> --direction downstream --json` → verify all downstream event handlers, bus consumers, cross-service calls have error handling199200## Why-Review Findings Validation Gate (MANDATORY when findings exist)201202> **Purpose:** Adversarial validation of own findings BEFORE any fix. Catches over-flagged criteria, false positives, and severity/score inflation at the source rather than letting them drive fixes or ship downstream.203204**Trigger:** Any finding produced (any severity). Skip ONLY when the verdict is unconditional PASS with literally zero findings.205206**Protocol:**2072081. Read own finalized report from `plans/reports/{skill}-{date}-{slug}.md`2092. Invoke `$why-review --validate-findings plans/reports/{skill}-{date}-{slug}.md` — verify each finding has `file:line` proof, steel-man each rejected interpretation, and stress-test every severity/score classification (each finding must clear why-review's finding-survival bar to be kept)2103. Read the CLEAN / HAS-ISSUES verdict returned by why-review2114. **If why-review demotes/removes any finding:** UPDATE own report with revised severities, remove false positives, and add a `## Why-Review Validation Notes` section citing what changed and why2125. **If why-review confirms all findings:** append a `## Why-Review Validation` line stating "All N findings re-validated against actual code; no severity changes."213214**Skip conditions (record explicit reason if skipping):** unconditional PASS with zero findings; why-review is itself the active context (avoid recursion).215216**Why this exists:** SRE sub-agent reports inherit confirmation bias — the orchestrator absorbs severity claims as ground truth. Validate findings BEFORE the fix so no fix is ever driven by an inflated or false finding; this gate feeds the "Validated Fix + Full Re-Review" loop below.217218## Validated Fix + Full Re-Review (MANDATORY when fixes are applied)219220When a review pass finds issues, validate findings before any fix. Do NOT spawn a fresh sub-agent only to re-review the same finding set before validation/fix. After validated SRE fixes applied, rerun the full SRE review. If that restarted review uses a sub-agent, spawn it with ZERO prior-round memory. A clean review pass ENDS the review.221222**When a fresh sub-agent is part of the restarted review, spawn via canonical template in `SYNC:review-protocol-injection`:**2232241. `agent_type`: `code-reviewer`2252. Task: `"SRE production readiness review after validated fixes — score all 12 criteria (0-2) for {files reviewed in the current full scope}"`2263. Review mode: `"Fresh full re-review after validated fixes. Zero memory of prior rounds. Re-read ALL target files from scratch."`2274. Reference Docs: `docs/project-reference/code-review-rules.md`2285. Target Files: same files from Scope Resolution2296. Integrate sub-agent report findings — DO NOT filter or override230231**Fresh re-review focus** (what prior rounds typically miss):232233- Operational concerns spanning multiple services234- Subtle reliability gaps (retry, circuit breakers, timeout handling)235- Missing observability (structured logging, correlation IDs, metrics)236- Data-integrity edge cases under concurrent load237238**Final verdict = every review pass that actually ran, combined.**239240## Output Format241242```markdown243## SRE Review Results244245**Scope:** {files reviewed}246**Date:** {date}247**Score:** {X}/24248**Verdict:** PASS / NEEDS WORK / NOT READY249250### Observability ({X}/8)251252| # | Criterion | Score | Evidence |253| --- | ------------------ | ----- | -------------------------- |254| 1 | Structured Logging | 0/1/2 | {file:line or "not found"} |255| 2 | Error Context | 0/1/2 | ... |256| 3 | Metrics Awareness | 0/1/2 | ... |257| 4 | Correlation | 0/1/2 | ... |258259### Reliability ({X}/8)260261| # | Criterion | Score | Evidence |262| --- | ----------------- | ----- | -------- |263| 5 | Retry Strategy | 0/1/2 | ... |264| 6 | Timeout Config | 0/1/2 | ... |265| 7 | Error Handling | 0/1/2 | ... |266| 8 | Fallback Behavior | 0/1/2 | ... |267268### Data Integrity ({X}/4)269270| # | Criterion | Score | Evidence |271| --- | ------------------ | ----- | -------- |272| 9 | Seed vs Migration | 0/1/2 | ... |273| 10 | Seeder Idempotency | 0/1/2 | ... |274275### Database Performance ({X}/4)276277| # | Criterion | Score | Evidence |278| --- | ---------------- | ----- | -------- |279| 11 | Pagination | 0/1/2 | ... |280| 12 | Database Indexes | 0/1/2 | ... |281282### Extended SRE Readiness ({n}/8 gate — pass/fail, does not change /24)283284| # | Gate Item | Status | Evidence |285| --- | -------------------------- | ----------------- | ----------------- |286| G1 | Rollout & Rollback | pass/partial/fail | `file:line` / N/A |287| G2 | Health Checks | pass/partial/fail | ... |288| G3 | Alerting & Runbook | pass/partial/fail | ... |289| G4 | SLO / Error-Budget | pass/partial/fail | ... |290| G5 | Capacity & Resource Limits | pass/partial/fail | ... |291| G6 | Config & Secrets | pass/partial/fail | ... |292| G7 | Graceful Shutdown/Startup | pass/partial/fail | ... |293| G8 | Concurrency & Idempotency | pass/partial/fail | ... |294295_Any unaccepted CRITICAL/HIGH `fail` above blocks a PASS verdict regardless of the /24 score._296297### Gaps to Address298299- {specific actionable item}300301### Recommendation302303{Proceed / Address gaps first}304```305306## Important Notes307308- Advisory (final VERDICT only) — score/verdict inform team but don't block commits; MANDATORY process steps (graph gate, validated-fix full re-review, Database Performance Protocol) are NEVER advisory309- Evidence-based — cite `file:line` for every score; unprovable score = 0310- Proportional — small bug fixes need less rigor than new endpoints (applies to VERDICT interpretation, NOT to skipping MANDATORY steps)311- Extended SRE Readiness gate is pass/fail, NOT scored — does not change `/24` math; but an unaccepted CRITICAL/HIGH gate `fail` blocks a PASS verdict (Severity Rubric). Use `docs/project-config.json → infrastructure` to mark items `N/A` with stated reason312- Check framework patterns — background-job base handlers, base-controller error handling313314---315316## Workflow Recommendation317318> **MANDATORY — NO EXCEPTIONS:** If NOT already in workflow, use ask the user directly to ask user:319>320> 1. **Activate `workflow-feature` workflow** (Recommended) — scout → investigate → plan → feature-implement → review → production-readiness-review → test → docs321> 2. **Execute `$production-readiness-review` directly** — run standalone322323---324325## Next Steps326327**MANDATORY — NO EXCEPTIONS** — after completing, use ask the user directly:328329- **"$watzup (Recommended)"** — wrap up + check doc staleness330- **"$test"** — run tests before wrapping up331- **"Skip, continue manually"** — user decides332333> **Combined audit:** For a whole-project architecture + compliance + production-readiness audit in one pass, run `$architecture-review-full` (or `$start-workflow workflow-architecture-audit`) — fans out this skill, `architecture-review`, `architecture-scalability-review` as parallel sub-agents and synthesizes one consolidated report.334335---336337> **[IMPORTANT]** Use task tracking to break ALL work into small tasks BEFORE starting. For simple tasks, AI MUST ask user whether to skip.338339- `docs/project-reference/domain-entities-reference.md` — Domain entity catalog, relationships, cross-service sync (read when task involves business entities/models)340341> **Critical Purpose:** Ensure quality — no flaws, no bugs, no missing updates, no stale content. Verify code AND documentation.342343> **External Memory:** Complex/lengthy work → write intermediate findings + final results to `plans/reports/` — prevents context loss, serves as deliverable.344345> **Evidence Gate:** MANDATORY — every claim, finding, recommendation requires `file:line` proof or traced evidence with confidence percentage (>80% to act, <80% verify first).346347<!-- SYNC:graph-assisted-investigation -->348349> **Graph-Assisted Investigation** — MANDATORY when `.code-graph/graph.db` exists.350>351> **HARD-GATE:** MUST ATTENTION run at least ONE graph command on key files before concluding any investigation.352>353> **Pattern:** Grep finds files → `trace --direction both` reveals full system flow → Grep verifies details354>355> | Task | Minimum Graph Action |356> | ------------------- | -------------------------------------------- |357> | Investigation/Scout | `trace --direction both` on 2-3 entry files |358> | Fix/Debug | `callers_of` on buggy function + `tests_for` |359> | Feature/Enhancement | `connections` on files to be modified |360> | Code Review | `tests_for` on changed functions |361> | Blast Radius | `trace --direction downstream` |362>363> **CLI:** `python .claude/scripts/code_graph {command} --json`. Use `--node-mode file` first (10-30x less noise), then `--node-mode function` for detail.364365<!-- /SYNC:graph-assisted-investigation -->366367<!-- SYNC:subagent-return-contract -->368369> **Sub-Agent Return Contract** — When this skill spawns a sub-agent, the sub-agent MUST return ONLY this structure. Main agent reads only this summary — NEVER requests full sub-agent output inline.370>371> ```markdown372> ## Sub-Agent Result: [skill-name]373>374> Status: ✅ PASS | ⚠️ PARTIAL | ❌ FAIL375> Confidence: [0-100]%376>377> ### Findings (Critical/High only — max 10 bullets)378>379> - [severity] [file:line] [finding]380>381> ### Actions Taken382>383> - [file changed] [what changed]384>385> ### Blockers (if any)386>387> - [blocker description]388>389> Full report: plans/reports/[skill-name]-[date]-[slug].md390> ```391>392> Main agent reads `Full report` file ONLY when: (a) resolving a specific blocker, or (b) building a fix plan.393> Sub-agent writes full report incrementally (per SYNC:incremental-persistence) — not held in memory.394>395> **Context budget** — the return payload is a SUMMARY, not a transcript: ≤10 finding bullets, no raw file contents / full diffs / verbatim logs inline, no re-pasted source. Everything beyond the summary lives in the `Full report` on disk. A sub-agent that would exceed the summary shape MUST write the detail to its report and return only the pointer — the orchestrator's context is the scarce resource the whole map-reduce protects.396397<!-- /SYNC:subagent-return-contract -->398399<!-- SYNC:nested-task-creation -->400401> **Nested Task Expansion Contract** — For workflow-step invocation, the `[Workflow] ...` row is only a parent container; the child skill still creates visible phase tasks.402>403> 1. Call the current task list first. If a matching active parent workflow row exists, set `nested=true` and record `parentTaskId`; otherwise run standalone.404> 2. Create one task per declared phase before phase work. When nested, prefix subjects `[N.M] $skill-name — phase`.405> 3. When nested, link the parent with `TaskUpdate(parentTaskId, addBlockedBy: [childIds])`.406> 4. Orchestrators must pre-expand a child skill's phase list and link the workflow row before invoking that child skill or sub-agent.407> 5. Mark exactly one child `in_progress` before work and `completed` immediately after evidence is written.408> 6. Complete the parent only after all child tasks are completed or explicitly cancelled with reason.409>410> **Blocked until:** the current task list done, child phases created, parent linked when nested, first child marked `in_progress`.411412<!-- /SYNC:nested-task-creation -->413414<!-- SYNC:project-reference-docs-guide -->415416> **Project Reference Docs Gate** — Run after task-tracking bootstrap and before target/source file reads, grep, edits, or analysis. Project docs override generic framework assumptions.417>418> 1. Identify scope: file types, domain area, and operation.419> 2. **Read `docs/project-config.json` first — the project's machine-readable map.** It is the single source of truth for THIS repo (modules/paths, framework + search keywords, test/E2E/integration run-commands, design system, architecture rules, workflow patterns); ground exact paths, run-commands, and conventions on it **before investigating, planning, or coding** — never assume framework defaults (`CLAUDE.md` + reference docs are derived from it). If it — or the docs index, `lessons.md`, `CLAUDE.md`, `AGENTS.md`, or any required reference doc — is missing or stale, auto-run `$project-init` or the narrow route (`$project-config`, `$docs-init`, `$scan-all`, `$scan --target=<key>`, `$claude-md-init`) first; if Codex mirrors or `AGENTS.md` are stale, ask the user to run `$sync-codex` (never auto-run it).420> 3. Required docs by trigger: always `docs/project-reference/lessons.md`; doc lookup `docs-index-reference.md`; review `code-review-rules.md`; backend/CQRS/API `backend-patterns-reference.md`; domain/entity `domain-entities-reference.md`; frontend/UI `frontend-patterns-reference.md`; styles/design `scss-styling-guide.md` + `design-system/design-system-canonical.md`; integration tests `integration-test-reference.md`; E2E `e2e-test-reference.md`; feature docs/specs `feature-spec-reference.md` + `spec-system-reference.md` + `spec-principles.md`; behavior/public-contract/spec-test-code sync `workflow-spec-test-code-cycle-reference.md`; derived spec index/ERD/reimplementation guides `spec-system-reference.md` + source Feature Specs under `docs/specs/`; architecture/new area `project-structure-reference.md`.421> 4. Read every required doc, then before target work state: `Reference docs read: ... | Not applicable: ...`.422>423> **Ready when:** scope evaluated, `docs/project-config.json` consulted, required docs checked/read or setup route completed, `lessons.md` confirmed, citation emitted.424425<!-- /SYNC:project-reference-docs-guide -->426427<!-- SYNC:task-tracking-external-report -->428429> **Task Tracking & External Report Persistence** — Bootstrap this before execution; then run project-reference doc prefetch before target/source work.430>431> 1. Create a small task breakdown before target file reads, grep, edits, or analysis. On context loss, inspect the current task list first.432> 2. Mark one task `in_progress` before work and `completed` immediately after evidence; never batch transitions.433> 3. For plan/review work, create `plans/reports/{skill}-{YYMMDD}-{HHmm}-{slug}.md` before first finding.434> 4. Append findings after each file/section/decision and synthesize from the report file at the end.435> 5. Final output cites `Full report: plans/reports/{filename}`.436>437> **Blocked until:** task breakdown exists, report path declared for plan/review work, first finding persisted before the next finding.438439<!-- /SYNC:task-tracking-external-report -->440441<!-- SYNC:critical-thinking-mindset -->442443> **Critical Thinking Mindset** — Apply critical thinking, sequential thinking. Every claim needs traced proof, confidence >80% to act.444> **Anti-hallucination:** Never present guess as fact — cite sources for every claim, admit uncertainty freely, self-check output for errors, cross-reference independently, stay skeptical of own confidence — certainty without evidence root of all hallucination.445446<!-- /SYNC:critical-thinking-mindset -->447448<!-- SYNC:evidence-based-reasoning -->449450> **Evidence-Based Reasoning** — Speculation is FORBIDDEN. Every claim needs proof.451>452> 1. Cite `file:line`, grep results, or framework docs for EVERY claim453> 2. Declare confidence: >80% act freely, 60-80% verify first, <60% DO NOT recommend454> 3. Cross-service validation required for architectural changes455> 4. "I don't have enough evidence" is valid and expected output456>457> **BLOCKED until:** `- [ ]` Evidence file path (`file:line`) `- [ ]` Grep search performed `- [ ]` 3+ similar patterns found `- [ ]` Confidence level stated458>459> **Forbidden without proof:** "obviously", "I think", "should be", "probably", "this is because"460> **If incomplete →** output: `"Insufficient evidence. Verified: [...]. Not verified: [...]."`461462<!-- /SYNC:evidence-based-reasoning -->463464<!-- SYNC:double-round-trip-review -->465466> **Validated-Finding Fix + Full Re-Review Loop** — Re-review is triggered by a validated finding fix cycle, not by a round number. Review purpose: `review → validate findings → fix validated findings → full re-review` until a complete review pass clears the round's exit bar (see **Severity floor** below). **A clean review ENDS the loop — no further rounds required.**467>468> _aka **Self-Review Convergence Loop**._ The name is historical — there is **NO 2-round cap**; "double-round-trip" only means a validated-finding fix cycle forces at least one fresh re-review. It runs until a clean pass, bounded by the **3-round ceiling** below.469>470> **Round cap — 3 rounds MAX (a ceiling, NEVER a target).** A clean pass ENDS the loop immediately at ANY round — round 1 included; the cap never obliges you to keep spinning. Hitting round 3 with blocking findings still open (severity floor applied) → **STOP and escalate by asking the user directly** with the still-open findings listed; NEVER emit a silent "good enough" PASS on cap exhaustion, and NEVER let the cap substitute for the clean-review requirement. The 2-repeated-no-progress blocker rule stays an EARLIER exit — escalate at whichever trips first.471>472> **Severity floor — from round 3, LOW stops blocking.** The exit bar tightens by round, so the loop converges on consequence instead of spinning on polish:473474> Define one predicate everywhere: `blocking_findings(round, findings)` returns all validated findings in rounds 1–2 and only validated CRITICAL/HIGH/MEDIUM findings in round 3+. A binary gate (test-green, security must-fix, required artifact) is exempt only when its owning invariant explicitly says so.475>476> | Round | Exit bar — loop ENDS when the fresh full review has… | Must be fixed to continue |477> | ----- | ------------------------------------------------------------------------- | ------------------------------ |478> | 1-2 | zero validated findings at ANY severity | CRITICAL · HIGH · MEDIUM · LOW |479> | 3+ | zero validated CRITICAL / HIGH / MEDIUM findings — **LOW-only is a PASS** | CRITICAL · HIGH · MEDIUM only |480>481> From round 3 onward LOW findings are **NOT required to be fixed**: a round whose validated findings are ALL LOW **ENDS the loop immediately** — do not open another round for them. Severity tiers are `SYNC:severity-rubric` (CRITICAL block-merge · HIGH must-fix · MEDIUM should-fix · LOW nice-to-fix); rounds 1-2 are unchanged, so an easy LOW still gets fixed early when it is cheap.482>483> **Severity-floor rules:**484>485> - **Never silently drop a deferred LOW.** Every unfixed LOW is listed in the final report under `## Deferred LOW Findings (severity floor, round ≥3)` with file, line, and description, so the owner can schedule it. Dropping it from the report is a protocol violation, not a clean pass.486> - **Never re-tier a finding to trigger the exit.** Downgrading a real CRITICAL/HIGH/MEDIUM to LOW so the loop can end is a FALSE PASS. Severity is set by consequence per `SYNC:severity-rubric` before the round bar is applied — never after, and never with the exit in view. — why: a floor that can be reached by relabeling is not a floor.487> - **The floor bounds the loop, not the standard.** It ends _iteration_; it never authorizes shipping a known CRITICAL/HIGH/MEDIUM, and it never lowers the finding-survival bar that admits a finding in the first place.488> - **The floor never applies to a hard gate.** Test-green gates (a suite must actually pass), security must-fix gates, and any gate whose criterion is binary rather than severity-rated are unaffected — a failing test is a failure, not a LOW finding.489>490> **Universal scope (any new output/judgment):** any newly produced output or judgment gets **≥1 self-review**; any **new judgment** gets **≥1 `$why-review --validate-findings` pass**; anything flagged to re-check is re-checked **≥1 time** — before that output is treated as final. This loop is the default convergence contract for ANY work-producing skill, not review skills only.491>492> **Routing invariant (author-facing):** a skill that validates findings MUST route them through `$why-review --validate-findings` (the terminal validator) — NEVER fork an inline finding-validation. Routing through why-review is what makes the finding-survival bar and this loop apply; the `verify-review-validate-coverage` sensor enforces this exact route mechanically.493>494> **Round 1:** Main-session review. Read target files, build understanding, note issues. Output findings + verdict (PASS / FAIL).495>496> **Decision after Round 1:**497>498> - **No issues found (PASS, zero findings)** → review ENDS. Do NOT spawn a fresh sub-agent for confirmation.499> - **`blocking_findings(round, findings)` is non-empty** → run the active review skill's findings-validation gate first; for review skills the default gate is `$why-review --validate-findings <report-path>`. Fix only validated findings, then restart the full review protocol from the beginning with a fresh task breakdown.500>501> **Fresh full re-review after every fix cycle:** Re-run the whole review protocol over the current full target. When sub-agents are part of that protoc502503…(truncated)