[BLOCKING] Execute skill steps in declared order. NEVER skip, reorder, or merge steps without explicit user approval. [BLOCKING] Before each step or sub-skill call, update task tracking: set
in_progresswhen step starts, setcompletedwhen step ends. [BLOCKING] Every completed/skipped step MUST include brief evidence or explicit skip reason. [BLOCKING] If Task tools are unavailable, create and maintain an equivalent step-by-step plan tracker with the same status transitions.
Quick Summary
Goal: Ensure technical correctness: receiving feedback with verification (not performative agreement), requesting targeted systematic reviews via code-reviewer subagent, enforcing verification gates before completion claims.
Routing boundary: If the user asks to review current changes, uncommitted work, staged/unstaged diffs, or a branch-to-branch diff, use
review-changesinstead.
MANDATORY MUST ATTENTION Before reviewing, search for project-specific reference docs:
Coding standards — search:
code-review-rules,coding-standards,style-guide,contributingArchitecture — search:patterns-reference,architecture,adrTest conventions — search:integration-test-reference,test-guide,test-conventionsDesign system — search:design-system,design-tokens,component-libraryRead found docs before reviewing. None found → rely on tech stack knowledge from file extensions/directory structure.
Workflow:
- Create Review Report — Init
plans/reports/code-review-{date}-{slug}.md - Phase 0: Blast Radius — Run graph analysis first if
.code-graph/graph.dbexists - Phase 0.3: Risk Detection — Detect dependency, migration, bus/event, API, security, config, and infra risks
- Phase 0.5: Plan Compliance — Verify changed files and tests against active plan when present
- Phase 0.7: Surface Detection — Classify files by language + directory semantics + change nature → route sub-agents
- Phase 1: File-by-File — Review each file, update report with correctness, convention, DRY, intent, test, and docs checks
- Phase 2: Holistic — Re-read accumulated report, assess overall approach, architecture, duplication, and cross-boundary behavior
- Phase 3: Final Result — Update report with overall assessment, critical issues, recommendations, docs staleness, and test gaps
- Round 2: Fresh Sub-Agent — After a fix cycle, run a fresh code-reviewer for cross-cutting concerns, convention drift, and edge cases
Key Rules:
- Report-Driven: Build report incrementally; re-read for big picture
- Detect First: Run graph blast radius when available, then classify change types and file surfaces before any review
- No Performative Agreement: Technical evaluation only ("You're right!" banned)
- Verification Gates: Evidence required before completion claims
- Review Current Diffs Elsewhere: Current changes, staged/unstaged diffs, and branch diffs belong to
review-changes - A clean Round 1 ENDS the review. Spawn a fresh sub-agent for Round 2 ONLY after a fix cycle.
Code Review
Three practices: receiving feedback with technical rigor, requesting systematic reviews via code-reviewer subagent, enforcing verification gates before completion claims.
Run
python .claude/scripts/code_graph query tests_for <function> --jsonon changed functions to flag coverage gaps.
Review Mindset (NON-NEGOTIABLE)
Skeptical. Every claim needs traced proof file:line. Confidence >80% to act.
- NEVER accept code correctness at face value — trace call paths
- NEVER include finding without
file:lineevidence (grep results, read confirmations) - ALWAYS question: "Does this actually work?" → trace it. "Is this all?" → grep cross-service
- ALWAYS verify side effects: check consumers + dependents before approving
First Principle — Easy to Change
The success metric of every coding decision is future change cost. DRY, SRP, abstraction, design patterns, naming, layering, tests — every technique exists to serve one goal: making the next change cheaper.
When evaluating code, a refactor, a test, or an abstraction, ask: does this make the next change cheaper or more expensive?
- Reject "best practices" that raise change cost (premature abstraction, speculative generality, leaky indirection, ceremony without payoff).
- Name the real enemies in findings: coupling, hidden state, duplicated knowledge, unclear intent, irreversible decisions exposed too early.
- A simpler design that is easy to change beats a sophisticated design that isn't.
Apply this lens before invoking any specific rule, pattern, or checklist below — if a downstream rule would raise change cost, this principle wins.
Core Principles (ENFORCE ALL)
| Principle | Rule |
|---|---|
| YAGNI | Flag code solving hypothetical problems (unused params, speculative interfaces) |
| KISS | Flag unnecessary complexity. "Is there a simpler way?" |
| DRY | Grep for similar/duplicate code. 3+ similar patterns → flag for extraction |
| Clean Code | Readable > clever. Names reveal intent. Functions do ONE thing. Nesting <=3. Methods <30 lines |
| Convention | MUST ATTENTION grep 3+ existing examples before flagging violations. Codebase convention wins over textbook |
| No Bugs | Trace logic paths. Verify edge cases (null, empty, boundary). Check error handling |
| Proof Required | Every claim backed by file:line evidence. Speculation is forbidden |
| Doc Staleness | Cross-ref changed files against related docs. Flag stale/missing updates |
Technical correctness over social comfort. Verify before implementing. Evidence before claims.
Graph-Enhanced Review (RECOMMENDED if graph.db exists)
python .claude/scripts/code_graph graph-blast-radius --json— prioritize files by impact (most dependents first)python .claude/scripts/code_graph query tests_for <function_name> --json— flag untested changed functionspython .claude/scripts/code_graph trace <file> --direction downstream --json— downstream impact (events, bus, cross-service)python .claude/scripts/code_graph trace <file> --direction both --json— full flow context for controllers/commands/handlers- Wide blast radius (>20 impacted nodes) = high-risk. Flag in report.
Review Approach (Report-Driven Two-Phase — CRITICAL)
MANDATORY FIRST: Create Todo Tasks
| Task | Status |
|---|---|
[Review] Create report file |
in_progress |
[Review Phase 0] Run graph blast-radius if available |
pending |
[Review Phase 0.3] Detect high-risk change types |
pending |
[Review Phase 0.5] Plan compliance check (skip if no active plan) |
pending |
[Review Phase 0.7] Detect categories + route sub-agents |
pending |
[Review Phase 1] File-by-file review + update report |
pending |
[Review Phase 2] Holistic assessment |
pending |
[Review Phase 3] Final findings, docs triage, and test sync findings |
pending |
[Review Round 2] Fresh sub-agent re-review after fix cycle |
pending |
[Review Final] Consolidate all rounds |
pending |
Step 0: Create Report File
Create plans/reports/code-review-{date}-{slug}.md with Scope, Files to Review sections.
Phase 0: Graph Blast Radius (FIRST WHEN AVAILABLE)
If .code-graph/graph.db exists, run graph impact analysis before reviewing:
python .claude/scripts/code_graph graph-blast-radius --jsonor the project equivalent- Record impacted files count, untested changed functions, and risk level in the report
- Prioritize high-impact files during Phase 1
If graph data is unavailable, record "Graph not available — skipping blast radius" and continue.
Phase 0.3: Detect High-Risk Change Types
Before file review, inspect the target diff or explicit file set for:
- Dependency upgrades — semver, breaking changes, advisories, peer compatibility
- Migrations or schema changes — rollback, lock/volume impact, zero-downtime deployment, idempotent backfill
- Bus events/messages — consumer existence, idempotency, retries, poison/dead-letter handling
- API contract changes — backward compatibility, caller alignment, auth, required response fields
- Security changes — enforcement coverage, privilege escalation, negative tests, duplicated permission strings
- Config/env changes — all environments covered, no secrets, fail-fast behavior, setup docs
- Infra changes — dev/prod parity, pinned versions, CI/CD permissions, reproducible builds
Create focused review tasks for every true signal and complete them before dimensional review.
Phase 0.5: Plan Compliance Check (CONDITIONAL)
If active plan context exists, verify scope, test evidence, and success criteria against the plan before file review; otherwise record the skip reason.
Phase 0.7: Detect Review Categories
Before any review — classify the changeset and route sub-agents:
| Signal in changed files | Route to |
|---|---|
| Auth/permission/token/encryption files | security-auditor |
| Query files, caching, batch processing | performance-optimizer |
| Source code (logic, handlers, services) | code-reviewer |
| Docs, plans, specs, markdown | general-purpose |
| Mixed changeset with security/perf files | Spawn specialized sub-agent first, then code-reviewer |
Phase 0.7: Derive Review Categories
Group changed files by: file language (extension), directory semantics (path), change nature (new entity, schema, config, UI, test).
For each category: name it, create sub-task, derive concerns using SYNC:category-review-thinking (first principles — NOT a fixed checklist).
Category list = Phase 1 work breakdown. Each category → own section in report.
Phase 1: File-by-File Review (Build Report)
For EACH file, immediately update report:
- File path, Change Summary, Purpose, Issues Found
- Convention check: Grep 3+ similar patterns — does new code follow existing convention?
- Correctness check: Trace logic — null, empty, boundary, error cases handled?
- DRY check: Grep for similar/duplicate code — does this logic exist elsewhere?
- Intention check: Does the change serve the stated purpose? Flag unrelated modifications
- Test check: Changed behavior has corresponding test/spec coverage or a documented gap
- Documentation check: Related docs/specs/READMEs still match the changed behavior
Phase 2: Holistic Review (Re-read Report)
After all files reviewed, re-read accumulated report:
- Technical Solution: Overall approach coherent as unified plan?
- Responsibility: Logic in LOWEST layer? Business logic not in controllers?
- Data ownership: Constants/config in model/entity, not controller/component?
- Duplication: Grep to verify — duplicated logic across changes?
- Architecture: Clean Architecture? Service boundaries respected?
- Plan Compliance: If active plan → check
## Plan Context: impl matches requirements, TCs have code evidence (not "TBD"), no requirement unaddressed - Design Patterns: Pattern opportunities (switch→Strategy)? Anti-patterns (God Object, Copy-Paste, Circular Dep)? DRY via base classes?
- Cross-Boundary Behavior: Callers/callees aligned? API/event contracts consistent? New wiring reachable?
- Test Sync: Business logic changes have corresponding tests or explicit user-facing gap
- Translation Sync: Multilingual UI text changes have translation updates or explicit risk acceptance
MUST ATTENTION CHECK — Clean Code: YAGNI (unused params, speculative interfaces)? KISS (simpler exists)? Methods >30 lines or nesting >3?
MUST ATTENTION CHECK — Correctness: Null/empty/boundary handled? Error paths caught? Async race conditions? Trace happy + error paths.
Documentation Staleness Check:
For each changed file — grep file name/module across docs/ and AI tooling dirs. Changed behavior → flag stale doc (specific section + what changed). Do NOT auto-fix — flag only.
Common staleness patterns: count/limit changed → docs embedding that number | API/contract changed → API usage docs | hook/skill added/removed → catalogs/README | schema changed → entity reference docs.
Phase 3: Final Review Result
Update report: Overall Assessment, Critical Issues, High Priority, Architecture Recommendations, Documentation Staleness, Positive Observations.
If documentation staleness is detected, recommend docs-update and list exact stale sections; do not silently pass stale docs.
Round 2+: Fresh Sub-Agent Re-Review (MANDATORY)
After Phase 3 (Round 1), spawn fresh code-reviewer sub-agent for Round 2 using canonical template from SYNC:review-protocol-injection:
- Copy Agent call shape from
SYNC:review-protocol-injectionverbatim - Embed full verbatim body of all 10 SYNC blocks:
SYNC:evidence-based-reasoning,SYNC:bug-detection,SYNC:design-patterns-quality,SYNC:complexity-prevention,SYNC:logic-and-intention-review,SYNC:test-spec-verification,SYNC:fix-layer-accountability,SYNC:rationalization-prevention,SYNC:graph-assisted-investigation,SYNC:understand-code-first - Task:
"Review the assigned code-review scope. Focus: cross-cutting concerns, interaction bugs, convention drift, missing pieces, subtle edge cases, logic errors, test spec gaps." - Target Files:
"use the explicit files, plan scope, or reviewer-provided target range" - Report:
plans/reports/code-review-round{N}-{date}.md
After sub-agent returns:
- Read report from
plans/reports/code-review-round{N}-{date}.md - Integrate findings as
## Round {N} Findings (Fresh Sub-Agent)— DO NOT filter or override - If FAIL: fix issues → spawn NEW Round N+1 fresh sub-agent (never reuse)
- Max 3 fresh rounds — escalate via
AskUserQuestionif still failing after 3 rounds
Clean Code Rules (MUST ATTENTION CHECK)
| # | Rule | Details |
|---|---|---|
| 1 | No Magic Values | All literals → named constants |
| 2 | Type Annotations | Explicit parameter and return types on all functions |
| 3 | Single Responsibility | One concern per method/class. Event handlers/consumers: one handler = one concern. NEVER bundle — platform swallows exceptions silently |
| 4 | DRY | No duplication; extract shared logic |
| 5 | Naming | Specific (employeeRecords not data), Verb+Noun methods, is/has/can/should booleans, no abbreviations |
| 6 | Performance | No O(n²) (use dictionary). Project in query (not load-all). ALWAYS paginate. Batch-by-IDs (not N+1) |
| 7 | Entity Indexes | Collections: index management methods. EF Core: composite indexes. Expression fields match index order. Text search → text indexes |
Data Lifecycle Rules (MUST ATTENTION CHECK)
Decision test: "Delete the DB and start fresh — does this data still need to exist?" Yes → Seeder/fixture. No → Migration.
| Type | Contains | NEVER contains |
|---|---|---|
| Seeder / Fixture | Default records, system config, reference data (idempotent — safe to run every startup) | Schema changes |
| Migration | Schema changes, column adds/removes, data transforms, index changes | Default records, permission seeds, system config |
Apply project's language/framework conventions. Principle universal — implementation project-specific.
Legacy Pattern Compliance
When reviewing files with legacy and modern patterns:
- Detect legacy signals — search
project-config.json,package.json, or equivalent for"legacy", version flags, feature annotations - Read what "legacy" means — grep 3+ legacy files to understand pattern constraints vs. modern files
- Derive compliance rules — what lifecycle/memory management differences exist between legacy/modern for this tech stack?
- Apply tech stack knowledge to flag anti-patterns
NEVER assume any specific framework's lifecycle. Derive from codebase evidence.
When to Use This Skill
| Practice | Triggers | MUST ATTENTION READ |
|---|---|---|
| Receiving Feedback | Review comments received, feedback unclear/questionable, conflicts with existing decisions | references/code-review-reception.md |
| Requesting Review | After each subagent task, major feature done, targeted review scope, after complex bug fix | references/requesting-code-review.md |
| Verification Gates | Before any completion claim, commit, push, or PR. ANY success/satisfaction statement | references/verification-before-completion.md |
Quick Decision Tree
SITUATION?
│
├─ Received feedback
│ ├─ Unclear items? → STOP, ask for clarification first
│ ├─ From human partner? → Understand, then implement
│ └─ From external reviewer? → Verify technically before implementing
│
├─ Completed work
│ ├─ Major feature/task? → Request code-reviewer subagent review
│ └─ Before merge? → Request code-reviewer subagent review
│
└─ About to claim status
├─ Have fresh verification? → State claim WITH evidence
└─ No fresh verification? → RUN verification command first
Receiving Feedback Protocol
Pattern: READ → UNDERSTAND → VERIFY → EVALUATE → RESPOND → IMPLEMENT
- NEVER use performative agreement ("You're right!", "Great point!", "Thanks for...")
- NEVER implement before verification
- MUST ATTENTION restate requirement, ask questions, or push back with technical reasoning
- MUST ATTENTION ask for clarification on ALL unclear items BEFORE starting
- MUST ATTENTION grep for usage before implementing suggested "proper" features (YAGNI check)
Source handling: Human partner → implement after understanding. External reviewer → verify technically, push back if wrong.
Full protocol: references/code-review-reception.md
Requesting Review Protocol
- Get git SHAs:
BASE_SHA=$(git rev-parse HEAD~1)andHEAD_SHA=$(git rev-parse HEAD) - Dispatch code-reviewer subagent with: WHAT_WAS_IMPLEMENTED, PLAN_OR_REQUIREMENTS, BASE_SHA, HEAD_SHA, DESCRIPTION
- Act on feedback: Critical → fix immediately. Important → fix before proceeding. Minor → note for later.
Full protocol: references/requesting-code-review.md
Verification Gates Protocol
Iron Law: NO COMPLETION CLAIMS WITHOUT FRESH VERIFICATION EVIDENCE
Gate: IDENTIFY command → RUN it → READ output → VERIFY it confirms claim → THEN claim. Skip any step = lying.
| Claim | Required Evidence |
|---|---|
| Tests pass | Test output shows 0 failures |
| Build succeeds | Build command exit 0 |
| Bug fixed | Original symptom test passes |
| Requirements met | Line-by-line checklist verified |
Red Flags — STOP: "should"/"probably"/"seems to", satisfaction before verification, committing without verification, trusting agent reports.
Full protocol: references/verification-before-completion.md
Related
code-simplifierdebug-investigaterefactoring
Systematic Review Protocol (10+ changed files)
For large changesets: categorize files by concern → fire parallel code-reviewer sub-agents per category → synchronize findings → holistic assessment. See review-changes/SKILL.md § "Systematic Review Protocol" for full 4-step protocol.
Workflow Recommendation
MANDATORY MUST ATTENTION — NO EXCEPTIONS: If NOT already in a workflow, use
AskUserQuestionto ask user:
- Activate
quality-auditworkflow (Recommended) — code-review → plan → code → review-changes → test- Execute
/code-reviewdirectly — run standalone
Architecture Boundary Check
For each changed file, verify no forbidden layer imports:
- Read rules from
docs/project-config.json→architectureRules.layerBoundaries - Determine layer — match file path against each rule's
pathsglob patterns - Scan imports — grep for
using(C#) orimport(TS) statements - Check violations — import path contains forbidden layer name → violation
- Exclude framework — skip files matching
architectureRules.excludePatterns - BLOCK on violation —
"BLOCKED: {layer} layer file {filePath} imports from {forbiddenLayer} ({importStatement})"
If architectureRules absent in project-config.json → skip silently.
Phase 4: Why-Review Self-Validation Gate (MANDATORY when findings exist)
Purpose: Adversarial validation of own findings BEFORE handoff. Catches over-flagged Highs, false positives, and severity inflation at the source rather than letting them propagate downstream.
Trigger: Any finding produced (Critical, High, Medium, OR Low). Skip ONLY when the report's verdict is unconditional PASS with literally zero findings.
Protocol:
- Read own finalized report from
plans/reports/{skill}-{date}-{slug}.md - Invoke
/why-reviewskill with arg:validate findings in plans/reports/{skill}-{date}-{slug}.md — verify each finding has file:line proof, steel-man each rejected interpretation, and stress-test severity classifications - Read why-review output from
plans/reports/why-review-{date}.md - If why-review demotes/removes any finding: UPDATE own finalized report with revised severities, remove false positives, and add a
## Why-Review Validation Notessection citing what changed and why - If why-review confirms all findings: Append
## Why-Review Validationline to own report stating "All N findings re-validated against actual code; no severity changes."
Skip conditions (record explicit reason if skipping):
- Verdict is unconditional PASS with zero findings → log "Skipped — no findings to validate"
- Why-review skill itself is the active context (avoid recursion)
Why this exists: AI sub-agent reports inherit confirmation bias — the orchestrator absorbs severity claims as ground truth. The 2026-05-09 review incident produced 5 Highs; adversarial validation demoted 3 of them. Codify this as standard practice.
Next Steps
MANDATORY MUST ATTENTION — NO EXCEPTIONS after completing, use AskUserQuestion:
- "/fix (Recommended)" — review found issues needing fixes
- "/watzup" — review clean, wrap up session
- "Skip, continue manually" — user decides
AI Agent Integrity Gate (NON-NEGOTIABLE)
Completion ≠ Correctness. Before reporting ANY work done:
- Grep every removed name. Extraction/rename/delete → grep confirms 0 dangling refs across ALL file types.
- Ask WHY before changing. Existing values intentional until proven otherwise.
- Verify ALL outputs. One build passing ≠ all builds passing.
- Evaluate pattern fit. Copying nearby code? Verify preconditions match — scope, lifetime, base class, constraints.
- New artifact = wired artifact. Created something? Prove it's registered, imported, reachable by all consumers.
[IMPORTANT] Use
TaskCreateto break ALL work into small tasks BEFORE starting — including tasks for each file read. This prevents context loss from long files. For simple tasks, AI MUST ATTENTION ask user whether to skip.
Critical Purpose: Ensure quality — no flaws, bugs, missing updates, stale content. Verify code AND documentation.
External Memory: Complex work → write findings incrementally to
plans/reports/— prevents context loss, serves as deliverable.
Evidence Gate: MANDATORY MUST ATTENTION — every claim, finding, recommendation requires
file:lineproof + confidence % (>80% act, <80% verify first).
OOP & DRY: MANDATORY MUST ATTENTION — flag patterns extractable to base class/generic/helper. Same-suffix/lifecycle/responsibility classes MUST ATTENTION share common base. Apply idiomatic abstraction (base class, mixin, trait, protocol) for project's language. Verify linting/analyzer configured.
Graph-Assisted Investigation — MANDATORY when
.code-graph/graph.dbexists.HARD-GATE: MUST ATTENTION run at least ONE graph command on key files before concluding any investigation.
Pattern: Grep finds files →
trace --direction bothreveals full system flow → Grep verifies details
Task Minimum Graph Action Investigation/Scout trace --direction bothon 2-3 entry filesFix/Debug callers_ofon buggy function +tests_forFeature/Enhancement connectionson files to be modifiedCode Review tests_foron changed functionsBlast Radius trace --direction downstreamCLI:
python .claude/scripts/code_graph {command} --json. Use--node-mode filefirst (10-30x less noise), then--node-mode functionfor detail.
Category Review Thinking — For each category of changed files, think from first principles. Do NOT use a fixed checklist — derive concerns based on the category's domain.
Step 1: Understand the category's role What is this category responsible for? What are its invariants? Who are its consumers (callers, dependents, downstream systems)?
Step 2: Read project conventions for this category Grep 3+ existing similar files in this category. What patterns do they follow? What base classes/interfaces/abstractions do they use?
Step 3: Derive concerns from first principles Given the category's role and invariants, what could go wrong? Start from universal concerns, then expand with category-specific knowledge:
- Correctness: Does the change do what it claims? Are contracts maintained?
- Contracts: Does the change preserve consumer-facing behavior?
- Security: What trust assumptions does this category make? Are they still valid?
- Performance: Does the change introduce O(n²), unbounded queries, or unnecessary I/O?
- Maintainability: Does the change follow existing patterns? Does it introduce hidden coupling?
- Tests: Is the changed behavior observable and testable?
- Documentation: Does the change invalidate any existing docs or specs?
These are starting points — your domain knowledge of the tech stack should expand this list. Do NOT limit yourself to what's listed above.
Step 4: Create sub-tasks and execute with file:line evidence Convert derived concerns into concrete review tasks. Each task must produce
file:lineevidence. No findings without proof.Examples of categories (illustrative — NOT exhaustive):
- Logic/domain files (business rules, handlers, services)
- Data/schema files (migrations, models, ORM definitions)
- API/contract files (controllers, routes, serializers, proto definitions)
- Configuration/environment files (env vars, feature flags, secrets)
- Infrastructure files (Dockerfiles, CI pipelines, manifests)
- UI/style files (components, templates, stylesheets)
- Test files (unit, integration, e2e)
- Documentation files (markdown, specs, ADRs)
- Security artifacts (auth middleware, permission definitions, crypto)
- Tooling/build files (build configs, linting rules, dependency manifests)
Sub-Agent Return Contract — When this skill spawns a sub-agent, the sub-agent MUST return ONLY this structure. Main agent reads only this summary — NEVER requests full sub-agent output inline.
## Sub-Agent Result: [skill-name] Status: ✅ PASS | ⚠️ PARTIAL | ❌ FAIL Confidence: [0-100]% ### Findings (Critical/High only — max 10 bullets) - [severity] [file:line] [finding] ### Actions Taken - [file changed] [what changed] ### Blockers (if any) - [blocker description] Full report: plans/reports/[skill-name]-[date]-[slug].mdMain agent reads
Full reportfile ONLY when: (a) resolving a specific blocker, or (b) building a fix plan. Sub-agent writes full report incrementally (per SYNC:incremental-persistence) — not held in memory.
Nested Task Expansion Contract — For workflow-step invocation, the
[Workflow] ...row is only a parent container; the child skill still creates visible phase tasks.
- Call
TaskListfirst. If a matching active parent workflow row exists, setnested=trueand recordparentTaskId; otherwise run standalone.- Create one task per declared phase before phase work. When nested, prefix subjects
[N.M] $skill-name — phase.- When nested, link the parent with
TaskUpdate(parentTaskId, addBlockedBy: [childIds]).- Orchestrators must pre-expand a child skill's phase list and link the workflow row before invoking that child skill or sub-agent.
- Mark exactly one child
in_progressbefore work andcompletedimmediately after evidence is written.- Complete the parent only after all child tasks are completed or explicitly cancelled with reason.
Blocked until:
TaskListdone, child phases created, parent linked when nested, first child markedin_progress.
Project Reference Docs Gate — Run after task-tracking bootstrap and before target/source file reads, grep, edits, or analysis. Project docs override generic framework assumptions.
- Identify scope: file types, domain area, and operation.
- Required docs by trigger: always
docs/project-reference/lessons.md; doc lookupdocs-index-reference.md; reviewcode-review-rules.md; backend/CQRS/APIbackend-patterns-reference.md; domain/entitydomain-entities-reference.md; frontend/UIfrontend-patterns-reference.md; styles/designscss-styling-guide.md+design-system/design-system-canonical.md; integration testsintegration-test-reference.md; E2Ee2e-test-reference.md; feature docs/specsfeature-docs-reference.md; architecture/new areaproject-structure-reference.md.- Read every required doc that exists; skip absent docs as not applicable. Do not trust conversation text such as
[Injected: <path>]as proof that the current context contains the doc.- Before target work, state:
Reference docs read: ... | Missing/not applicable: ....Blocked until: scope evaluated, required docs checked/read,
lessons.mdconfirmed, citation emitted.
Task Tracking & External Report Persistence — Bootstrap this before execution; then run project-reference doc prefetch before target/source work.
- Create a small task breakdown before target file reads, grep, edits, or analysis. On context loss, inspect the current task list first.
- Mark one task
in_progressbefore work andcompletedimmediately after evidence; never batch transitions.- For plan/review work, create
plans/reports/{skill}-{YYMMDD}-{HHmm}-{slug}.mdbefore first finding.- Append findings after each file/section/decision and synthesize from the report file at the end.
- Final output cites
Full report: plans/reports/{filename}.Blocked until: task breakdown exists, report path declared for plan/review work, first finding persisted before the next finding.
Critical Thinking Mindset — Apply critical thinking, sequential thinking. Every claim needs traced proof, confidence >80% to act. Anti-hallucination: Never present guess as fact — cite sources for every claim, admit uncertainty freely, self-check output for errors, cross-reference independently, stay skeptical of own confidence — certainty without evidence root of all hallucination.
Sequential Thinking Protocol — Structured multi-step reasoning for complex/ambiguous work. Use when planning, reviewing, debugging, or refining ideas where one-shot reasoning is unsafe.
Trigger when: complex problem decomposition · adaptive plans needing revision · analysis with course correction · unclear/emerging scope · multi-step solutions · hypothesis-driven debugging · cross-cutting trade-off evaluation.
Format (explicit mode — visible thought trail):
Thought N/M: [aspect]— one aspect per thought, state assumptions/uncertaintyThought N/M [REVISION of Thought K]: ...— when prior reasoning invalidated; state Original / Why revised / ImpactThought N/M [BRANCH A from Thought K]: ...— explore alternative; converge with decision rationaleThought N/M [HYPOTHESIS]: ...then[VERIFICATION]: ...— test before actingThought N/N [FINAL]— only when verified, all critical aspects addressed, confidence >80%Mandatory closers: Confidence % stated · Assumptions listed · Open questions surfaced · Next action concrete.
Stop conditions: confidence <80% on any critical decision → escalate via AskUserQuestion · ≥3 revisions on same thought → re-frame the problem · branch count >3 → split into sub-task.
Implicit mode: apply methodology internally without visible markers when adding markers would clutter the response (routine work where reasoning aids accuracy).
Deep-dive: see
/sequential-thinkingskill (.claude/skills/sequential-thinking/SKILL.md) for worked examples (api-design, debug, architecture), advanced techniques (spiral refinement, hypothesis testing, convergence), and meta-strategies (uncertainty handling, revision cascades).
Evidence-Based Reasoning — Speculation is FORBIDDEN. Every claim needs proof.
- Cite
file:line, grep results, or framework docs for EVERY claim- Declare confidence: >80% act freely, 60-80% verify first, <60% DO NOT recommend
- Cross-service validation required for architectural changes
- "I don't have enough evidence" is valid and expected output
BLOCKED until:
- [ ]Evidence file path (file:line)- [ ]Grep search performed- [ ]3+ similar patterns found- [ ]Confidence level statedForbidden without proof: "obviously", "I think", "should be", "probably", "this is because" If incomplete → output:
"Insufficient evidence. Verified: [...]. Not verified: [...]."
Design Patterns Quality — Priority checks for every code change:
- DRY via OOP: Identify classes/modules with the same purpose, naming pattern, or lifecycle. Apply your knowledge of the project's language/framework to determine the idiomatic abstraction (base class, mixin, trait, protocol, decorator). 3+ similar patterns → extract to shared abstraction.
- Right Responsibility: Logic in LOWEST layer (Entity > Domain Service > Application Service > Controller). Never business logic in controllers.
- SOLID: Single responsibility (one reason to change). Open-closed (extend, don't modify). Liskov (subtypes substitutable). Interface segregation (small interfaces). Dependency inversion (depend on abstractions).
- After extraction/move/rename: Grep ENTIRE scope for dangling references. Zero tolerance.
- YAGNI gate: NEVER recommend patterns unless 3+ occurrences exist. Don't extract for hypothetical future use.
Anti-patterns to flag: God Object, Copy-Paste inheritance, Circular Dependency, Leaky Abstraction.
Serial Attention for Design Quality — DO NOT scan all quality concerns simultaneously. Split attention misses violations that focused passes catch.
- Identify applicable dimensions — Based on the code's language, domain, and patterns, determine which quality dimensions apply: DRY, SOLID principles (SRP/OCP/LSP/ISP/DIP), OOP idioms, cohesion/coupling, GRASP, Law of Demeter, CQRS invariants, etc. Your list is NOT fixed — derive from what the code actually does.
- One focused pass per dimension — Dedicate single-focus attention to EACH dimension in sequence. Do NOT mix concerns across passes.
- Threshold: 3+ similar patterns = MANDATORY extraction — Not optional suggestion. Flag as mandatory structural fix requiring action.
- 2+ violations of same kind = structural finding — Report as "pattern problem" needing architectural resolution, not a list of individual instances.
Complexity Prevention (Ousterhout) — MANDATORY. Measure code by cost of change: one business change should map to one code change. Flag ALL of the following in review:
- Change amplification — small business change forces edits in >3 places → structural flaw. Count edit sites for a plausible future change (add variant, add field, add authorization). >3 = reject.
- Cognitive load — reader must hold too much context to safely modify. Flag deep inheritance, long parameter lists, boolean traps, implicit ordering dependencies.
- Cross-cutting duplication at entry points — logging, error handling, val
…(truncated)