Codex compatibility note:
- Invoke repository skills with
$skill-name in Codex; this mirrored copy rewrites legacy Claude /skill-name references.
- Task tracker mandate: BEFORE executing any workflow or skill step, create/update task tracking for all steps and keep it synchronized as progress changes.
- User-question prompts mean to ask the user directly in Codex.
- Ignore Claude-specific mode-switch instructions when they appear.
- Strict execution contract: when a user explicitly invokes a skill, execute that skill protocol as written.
- Subagent authorization: when a skill is user-invoked or AI-detected and its protocol requires subagents, that skill activation authorizes use of the required
spawn_agent subagent(s) for that task.
- Do not skip, reorder, or merge protocol steps unless the user explicitly approves the deviation first.
- For workflow skills, execute each listed child-skill step explicitly and report step-by-step evidence.
- If a required step/tool cannot run in this environment, stop and ask the user before adapting.
Codex Project-Reference Loading (No Hooks)
Codex uses static project-reference loading instead of runtime-injected project docs.
When coding, planning, debugging, testing, or reviewing, open project docs explicitly using this routing.
Always read:
docs/project-config.json (project-specific paths, commands, modules, and workflow/test settings)
docs/project-reference/docs-index-reference.md (routes to the full docs/project-reference/* catalog)
docs/project-reference/lessons.md (always-on guardrails and anti-patterns)
Missing/stale context route: If docs/project-config.json, the docs index, lessons.md, CLAUDE.md, AGENTS.md, or any task-required reference doc is missing or stale, auto-run $project-init or the narrow setup route ($project-config, $docs-init, $scan-all, $scan --target=<key>, $claude-md-init) before ordinary project-specific work. If Codex mirrors or AGENTS.md are missing/stale, ask the user to run $sync-codex; do not auto-run it.
Situation-based docs:
- Project structure/architecture/tech-stack/deployment/setup (any layer — backend, frontend, or infra):
project-structure-reference.md
- Backend/CQRS/API/domain/entity changes:
backend-patterns-reference.md, domain-entities-reference.md
- Frontend/UI/styling/design-system:
frontend-patterns-reference.md, scss-styling-guide.md, design-system/README.md
- Spec authoring,
docs/specs/ pathing, or TC format: feature-spec-reference.md, spec-system-reference.md, spec-principles.md
- Behavior/public-contract changes or spec-test-code sync:
workflow-spec-test-code-cycle-reference.md plus the spec docs above
- Derived spec indexes/ERDs/reimplementation guides:
spec-system-reference.md and source Feature Specs under docs/specs/
- Integration test implementation/review:
integration-test-reference.md
- E2E test implementation/review:
e2e-test-reference.md
- Code review/audit work:
code-review-rules.md plus domain docs above based on changed files
Do not read all docs blindly. Start from docs-index-reference.md, then open only relevant files for the task.
[BLOCKING] Execute skill steps in declared order. NEVER skip, reorder, or merge steps without explicit user approval.
[BLOCKING] Before each step or sub-skill call, update task tracking: set in_progress when step starts, set completed when step ends.
[BLOCKING] Every completed/skipped step MUST include brief evidence or explicit skip reason.
[BLOCKING] If Task tools are unavailable, create and maintain an equivalent step-by-step plan tracker with the same status transitions.
Quick Summary
Portability: docs/specs/ is the fixed Feature Spec root. docs/templates/detailed-feature-spec-template.md remains the default template unless workflowPatterns.featureDocTemplate points to another template.
[IMPORTANT] task tracking — Break ALL work into small tasks BEFORE starting. For simple tasks, ask user whether to skip.
Goal: Own the entire Feature Spec lifecycle in one skill — author/maintain tech-free 8-section business Feature Specs (code evidence carried only in Section 8 test-case anchors, never prose), generate Section 8 test specifications, and reconcile those TCs with executing test code — producing a tech-free, AI-implementable Feature Spec whose Section 8 TC registry stays the single source of truth, traceable to test code, so any team can rebuild the feature on any stack from the spec alone. The mode you run determines which references/ body drives work; the shared §8 contract, M1-M7 mandates, and quality philosophy below apply every mode.
Summary:
- Purpose: one skill owns the whole Feature Spec lifecycle across 7 modes —
draft | init | update | audit | amend | tests | sync — producing/maintaining a tech-free 8-section business Feature Spec whose §8 TC registry is the single source of truth, traceable to executing test code, so any team can rebuild the feature on any stack from the spec alone.
- Main steps (every run): (1) resolve mode FIRST — explicit
[mode=<x>] wins, else infer from request + repo state, ambiguous → ask the user directly before any mutating mode; (2) read the matching references/{author,tests,sync}.md body — NEVER run a mode from memory; (3) task tracking-break the work (one task per file read) before starting; (4) execute the mode's procedure/gates from its body; (5) cross-service check before concluding.
- §1-7 prose STRICTLY tech-free — no framework/product/language/persistence/messaging/auth names (banned tokens →
spec-principles.md §3.2); technical identifiers live ONLY in evidence carriers, frontmatter, and mermaid blocks. — why: M1/M5 require rebuild-on-any-stack from prose alone.
- Section 8 is the canonical TC registry for business TCs (
TC-{FEATURE}-{NNN}) — every TC carries verifiable [Source: namespace/service/id] evidence (sole exception mode=draft → Evidence: TBD + provisional flag, upgraded to a real anchor on the first code-sourced run); NEVER overwrite existing TCs during update — tests owns generation, sync reconciles drift.
- Honor the M1-M7 mandates (
sdd-artifact-contract.md) + canonical TC format (shared/tc-format.md) — any M1-M5 or M7 violation FAILS the artifact (M6 binds the REVIEWER, not the artifact); INDEX.md/ERD are DERIVED — flag refresh need in update, NEVER trigger $spec-index here. — why: separation of concerns keeps the canonical spec the only source of truth.
Renamed: formerly /feature-spec (and earlier /feature-docs); the former /spec-tests skill is now folded in as mode=tests / mode=sync. Those names no longer resolve as slash commands — use $spec with the matching mode.
Modes (resolve mode FIRST — BLOCKING)
| Mode |
Use when… |
Body |
draft |
Author a provisional spec from an idea/requirement/prompt — no code yet (TDD-first, §8 Evidence: TBD, provisional marker) |
references/author.md |
init |
No docs/specs/{Bucket}/ exists — author a full 8-section spec from source |
references/author.md |
update |
Docs exist + code changed — section-impact-mapped updates |
references/author.md |
audit |
--audit flag or user asks — staleness report per section (never mutates docs) |
references/author.md |
amend |
[mode=amend] from the bugfix workflow — minimal regression-scoped §3/§4/§8 touch only |
references/author.md |
tests |
Generate or update Section 8 TC-{FEATURE}-{NNN} test specifications |
references/tests.md |
sync |
Reconcile §8 TCs ↔ executing test code (forward/reverse/harvest, orphan, staleness); harvest captures a SPEC-SILENT invariant into §4/§5/§8 |
references/sync.md |
Mode resolution (do this before any work):
- Parse the mode from the invocation: explicit
[mode=<x>] arg wins; else infer from request + repo state ("from idea/requirements/prompt", "draft spec", "no code yet" → draft; no docs/specs/{Bucket}/ AND code exists to source from → init; docs exist + diff → update; "audit/stale" → audit; bugfix caller → amend; "write/update test specs", "TCs" → tests; "sync tests", "reconcile §8 with tests" → sync). draft vs init: both author a new spec, but draft sources from idea/requirement text (no code → Evidence: TBD, provisional) while init sources from existing code (real [Source:] evidence). "No docs" alone does NOT imply init — check whether code exists to source from. draft never auto-overwrites existing §8 TCs.
- If ambiguous, present the detected mode by asking the user directly before proceeding — NEVER auto-start a mutating mode.
- Read the matching
references/ body — it is the single source of truth for that mode's procedure, gates, and output contract. Do not run a mode from memory.
Key Rules (all modes):
- [BLOCKING] Read
docs/project-reference/spec-principles.md — repo-local prose/evidence rules (§3 prose scope + §3.2 banned prose-token list). For the AI-implementability criteria + tech-agnostic mandates, read .claude/skills/shared/sdd-artifact-contract.md ("AI-Implementability Gate" + mandates M1-M7) — those are the canonical authority, not the local stub.
- [BLOCKING] EVERY test case MUST carry verifiable code evidence as a
[Source: namespace/service/id] abstract anchor in its Section 8 hidden carrier — physical file:line lives only in the provenance sidecar.
Exception (mode=draft): an idea-sourced spec has no code yet — its §8 TCs carry Evidence: TBD (reference-only) and the spec is flagged provisional (provisional: true frontmatter + a "DRAFT — unverified until code lands" header banner). The first update/init run against real code MUST upgrade every TBD to a real [Source:] anchor and clear the provisional flag. This mirrors existing TDD-first handling — it relaxes evidence ONLY for draft, never for code-sourced modes.
- [BLOCKING] Section 8 is the canonical TC registry for business TCs — §8 business TCs are the source of truth; test code implements them. The
tests mode owns generation; sync mode reconciles drift; the author modes (draft/init) populate §8 at authoring time (draft with Evidence: TBD, init with real [Source:]) and MUST NOT overwrite existing TCs during UPDATE.
- Authored docs MUST match the master template's 8 tech-free sections (Overview, Glossary, User Stories & AC, Business Rules, Domain Model, Process Flows & Interaction Surface, Permissions & Roles, Test Specifications) + YAML frontmatter — zero technical terms in prose, size caps enforced.
- [BLOCKING] Canonical TC format authority:
.claude/skills/shared/tc-format.md (GWT template, Evidence carrier, decade-numbering, Preservation Tests). M1-M7 mandates: .claude/skills/shared/sdd-artifact-contract.md — any M1-M5 or M7 violation FAILS the artifact (M6 binds the REVIEWER, not the artifact).
docs/project-reference/feature-spec-reference.md — project-specific Feature Spec patterns (read directly when relevant). docs/project-reference/domain-entities-reference.md — domain entity catalog, relationships, cross-service sync.
8-Section Feature Spec Rules (canonical reference)
Canonical home for the Feature Spec rules; applies to any edit under the Feature Spec docs root.
Format: Tech-free 8-section Feature Spec. Activate the $spec skill before editing.
Read first: docs/project-reference/feature-spec-reference.md, docs/project-reference/spec-system-reference.md, and docs/project-reference/spec-principles.md. For behavior/public-contract changes, also read docs/project-reference/workflow-spec-test-code-cycle-reference.md.
8 sections (exact order): 1. Overview · 2. Glossary · 3. User Stories & Acceptance Criteria · 4. Business Rules · 5. Domain Model · 6. Process Flows & Interaction Surface · 7. Permissions & Roles · 8. Test Specifications. No technical sections (Commands/Events/API/Cross-Service/Performance/Troubleshooting) — code is the technical source of truth.
§6 carries a tech-agnostic interaction surface (views/nav/observable states/per-story click-paths) per the SYNC:ui-intent-layer block this skill carries; backend-only specs state the skip reason explicitly. This does NOT contradict "No technical sections" — the interaction surface is tech-agnostic INTENT (UX roles, information, states, flows), not a technical "UI Pages" section; M1-clean keeps it free of framework/route/CSS/component-class names.
Mandatory:
- §1-7 prose is STRICTLY tech-free — no framework/product/language/persistence/messaging/auth names (banned tokens →
spec-principles.md §3.2). Technical identifiers live ONLY in evidence carriers.
- Section 5 (Domain Model): Mermaid ERD +
[Source: component/{service}/{id}] abstract anchor per entity (cannot be omitted)
- Section 4 (Business Rules):
[Source: rule/{service}/{id}] abstract anchor per rule group
- Section 8 (Test Specifications): canonical business TC source — TC-{FEATURE}-{NNN} IDs, each carrying a hidden
[Source: namespace/service/id] carrier + a CoveredBy: field. Legacy IntegrationTest: fields are accepted only as migration input.
Rules:
- TC IDs live in Section 8 only — never authored in
docs/specs/ directly
- Section 8 authored via
$spec [mode=tests]; $spec [mode=init] populates it only during initial authoring
- No line-count cap applies to Feature Specs. Split the capability only when TCs>40 or distinct module-level capabilities emerge.
M1-M5 + M7 Compliance (BLOCKING — applies to every authored spec and every TC)
See .claude/skills/shared/sdd-artifact-contract.md → "AI-SDD Mandates (M1-M7)" for the full BLOCKING criteria. In brief: M1 tech-agnostic prose; M2 no source code in prose; M3 logical-IDs-first traceability with a SEPARATE [Source:] carrier; M4 AI-implementability (one interpretation, named success/failure); M5 rebuild-from-scratch on any stack from §1-8 prose alone. Tech terms are allowed ONLY inside evidence carriers (**Evidence**, CoveredBy, legacy IntegrationTest, [Source:]), YAML frontmatter, and ```mermaid ``` blocks.
M7 — Business-visibility. Apply the demo test to the case BODY: "what would a stakeholder SEE change?" — no answer → FAIL as TECHNICAL-ONLY. A When that is an invocation (a handler runs, a consumer receives, a job fires, data syncs) or a Then asserting schema/type/nullability/call-count FAILS. Judge the BODY, never the title or ID.
M1 governs vocabulary; M7 governs subject matter. A technical case in impeccably tech-free prose satisfies M1 while violating M7.
M6 is absent from this list by design: it binds the REVIEWER (a review that passes an M1-M5/M7 violation is itself defective), never the artifact. The artifact-facing set reads M1-M5 and M7.
Derived-Index Delegation
This skill owns the canonical Feature Spec (§1-8) and its §8 TC registry. The bucket INDEX.md and cross-capability ERD are derived artifacts regenerated by $spec-index FROM these specs — never a source of truth, and never authored here. In update mode, flag "derived spec artifact refresh may be required" but do NOT trigger $spec-index directly (separation of concerns).
Workflow Recommendation
MANDATORY IMPORTANT MUST ATTENTION — NO EXCEPTIONS: If you are NOT already in a workflow, you MUST ATTENTION use ask the user directly to ask the user. Do NOT judge task complexity or decide this is "simple enough to skip" — the user decides whether to use a workflow, not you:
- Activate
workflow-feature workflow (Recommended) — spec-driven with tests by default: scout → investigate → spec-discovery → domain-analysis → why-review → spec → spec-clarify → plan → plan-review → plan-validate → why-review → spec [mode=tests] → why-review → artifact-review --type=spec-tests → plan → plan-review → feature-implement → domain-entities-review → spec [mode=tests] → why-review → artifact-review --type=spec-tests → spec [mode=sync] → integration-test → integration-test-review → integration-test-verify → workflow-review-changes → production-readiness-review → security-review → changelog → test → docs-update → workflow-end → watzup
- Execute
$spec directly — run this skill standalone in the resolved mode
Next Steps
[BLOCKING] After completing, use ask the user directly to present options. Do NOT skip — user decides:
- "$spec [mode=tests] (Recommended)" — Generate/update Section 8 test specs for the documented features (if you just authored/updated a spec)
- "$spec [mode=sync]" — Reconcile Section 8 TCs ↔ executing test code
- "$artifact-review --type=spec-tests" — Audit TC coverage + GIVEN/WHEN/THEN quality
- "Skip, continue manually" — user decides
Related Skills
| Skill |
Relationship |
When to Call |
$spec-index |
Derived consumer — assembles a regenerable navigation index/ERD FROM these Feature Specs (never a source of truth) |
AFTER specs exist — (re)generate the bucket INDEX.md / cross-capability ERD over the canonical specs |
$artifact-review --type=spec-tests |
Reviewer — audits TC coverage in Section 8 |
After spec [mode=tests], to validate TC completeness and GIVEN/WHEN/THEN quality |
$integration-test |
End consumer — generates test code from TCs in Section 8 |
After spec [mode=tests], to produce actual integration test files |
$docs-update |
Orchestrator — calls this skill as Phase 2 |
Run $docs-update for full chain sync; it calls $spec internally |
$changes-review |
Trigger — detects feature doc staleness |
Calls $docs-update when a business doc is stale relative to code changes |
[IMPORTANT] Use task tracking to break ALL work into small tasks BEFORE starting — including tasks for each file read. This prevents context loss from long files. For simple tasks, AI MUST ATTENTION ask user whether to skip.
[BLOCKING] Capture a tech-agnostic UI/UX intent layer in every UI-bearing spec — a reader must be able to visualize how the feature works without naming any technology. When the feature has a user interface, the spec MUST ATTENTION carry an interaction-surface section so the application — not just its API — can be rebuilt on any stack:
- View Inventory — list each view/screen by its UX ROLE and purpose (e.g. "list of items", "item editor", "confirmation step") and what information it presents. Describe by role, never by an implementation name.
- Navigation Map — how a user moves between views: entry points, transitions, and exits. Trace how this surface connects to neighboring features already in the system.
- Key observable UI States — the distinct states a user can observe per view (empty, loading, populated, error, success, permission-denied, etc.) — described as what the user perceives, not how it is rendered.
- Per-story interaction flow — for each user story, the step-by-step click/action path from intent to outcome, cross-referenced to the logical IDs the spec already owns (
US-/OP-/BR-).
- Couple to the companion design artifact — keep deep visual fidelity (layout, tokens, pixel detail) OUT of the spec; it lives in the linked
design-spec/mockup. Record that companion's path in the spec frontmatter so the spec stays the navigable hub.
M1-clean (NON-NEGOTIABLE): the prose names ZERO frameworks, routes/URLs, CSS, or component-class names — only roles, information, states, and flows. Technology detail belongs in the companion design artifact, never here.
Skip ONLY when the feature is backend-only (no UI) — state that reason explicitly in the section.
Cross-Service Check — Microservices/event-driven: MANDATORY before concluding investigation, plan, spec, or feature doc. Missing downstream consumer = silent regression.
| Boundary |
Grep terms |
| Event producers |
Publish, Dispatch, Send, emit, EventBus, outbox, IntegrationEvent |
| Event consumers |
Consumer, EventHandler, Subscribe, @EventListener, inbox |
| Sagas/orchestration |
Saga, ProcessManager, Choreography, Workflow, Orchestrator |
| Sync service calls |
HTTP/gRPC calls to/from other services |
| Shared contracts |
OpenAPI spec, proto, shared DTO — flag breaking changes |
| Data ownership |
Other service reads/writes same table/collection → Shared-DB anti-pattern |
Per touchpoint: owner service · message name · consumers · risk (NONE / ADDITIVE / BREAKING).
BLOCKED until: Producers scanned · Consumers scanned · Sagas checked · Contracts reviewed · Breaking-change risk flagged
Evidence-Based Reasoning — Speculation is FORBIDDEN. Every claim needs proof.
- Cite
file:line, grep results, or framework docs for EVERY claim
- Declare confidence: >80% act freely, 60-80% verify first, <60% DO NOT recommend
- Cross-service validation required for architectural changes
- "I don't have enough evidence" is valid and expected output
BLOCKED until: - [ ] Evidence file path (file:line) - [ ] Grep search performed - [ ] 3+ similar patterns found - [ ] Confidence level stated
Forbidden without proof: "obviously", "I think", "should be", "probably", "this is because"
If incomplete → output: "Insufficient evidence. Verified: [...]. Not verified: [...]."
Critical Thinking Mindset — Apply critical thinking, sequential thinking. Every claim needs traced proof, confidence >80% to act.
Anti-hallucination: Never present guess as fact — cite sources for every claim, admit uncertainty freely, self-check output for errors, cross-reference independently, stay skeptical of own confidence — certainty without evidence root of all hallucination.
Spec ↔ Tests ↔ Code Triangulation — The unit of review is the WHOLE PACKAGE (spec + tests + code), not the diff alone. Load all three faces together and reason mutual-consistency FIRST, before any isolated per-file check.
- Locate all three faces for the changed behavior: the governing Feature Spec section(s) (§3 ACs / §4 BRs / §8 TCs), the tests that guard it, and the production code. A missing face is a finding (SPEC-GAP / TEST-GAP / DEAD-SPEC).
- Triangulate pairwise — classify which face is wrong on every disagreement:
- code vs spec → CODE-EXTRA / SPEC-STALE / CODE-WRONG (a [HARD] §4 rule or §5 invariant with no enforcing path is CODE-WRONG).
- tests vs spec → TEST-GAP / SPEC-SILENT.
- tests vs code → TEST-GAP / WEAK-TEST (a test that survives a deliberately broken invariant).
- Capture hidden rules — an invariant the code enforces but the spec never states (SPEC-SILENT) is surfaced as a finding, added into §3/§4/§8, and guarded with a test: the enrichment loop, never a silent pass.
- Re-review after enrichment — when triangulation adds spec content or a test, re-review the package against the enriched spec; converge only when a full pass surfaces no new disagreement.
NEVER mark PASS while any face disagrees without a logged finding. The diff is the entry point; the package is the unit of judgment.
Spec drift adjudication (code-wrong vs spec-stale). Whenever changed behavior diverges from a canonical Feature Spec (business rule, acceptance criterion, flow, state transition, or §8 TC under docs/specs/), you MUST NOT silently pick a side. Adjudicate per shared/sdd-artifact-contract.md → Drift Gates:
- Detect — compare the change against the spec's documented intent. No divergence → record
Spec in sync and move on.
- Classify the divergence:
- CODE-WRONG — the spec correctly states intended behavior and the change violates it → BLOCKING finding; fix the code/test against intended behavior (write/adjust a regression TC first).
- SPEC-STALE — the change is the new intended behavior and the spec now documents the old/wrong behavior → update the spec FIRST via
$spec [mode=update], then sync $spec [mode=tests] + $spec [mode=sync].
- AMBIGUOUS — intended behavior is unclear → ask the user directly (or the canonical spec owner) before editing either side.
- SPEC-SILENT — the code correctly enforces an invariant/behavior that NO canonical spec artifact (§3 AC, §4 BR, §5 invariant, §8 TC) states → not drift but an UNWRITTEN rule discovered by review. ENRICH the spec via the Invariant Harvest pass (
$spec [mode=sync] direction=harvest → spec/references/sync.md): prove it is always-true (≥2 enforcement points or a rejecting guard), express it as a universally-quantified property, then add the rule to §4 (or §3/§5) AND a §8 TC via $spec [update] + $spec [mode=tests] and add the guarding test. A discovered invariant left only in code (or only in tests) is INCOMPLETE — this is the highest-value capture (the rule nobody wrote down).
- Never normalize drift just because code/tests are green — green can encode the drift itself. Reconcile to canonical intent, never to whichever side currently passes.
A behavior-changing review/implementation that leaves a spec divergence unadjudicated is INCOMPLETE; an unwritten-but-enforced invariant left uncaptured (no §4/§8 entry) is equally INCOMPLETE.
AI Mistake Prevention — Failure modes to avoid on every task:
Re-read files after context changes. Context compaction, resume, or long-running work can make memory stale; verify current files before acting.
Verify generated content against source evidence. AI hallucinates APIs, names, claims, and document facts. Check the relevant source before documenting or referencing.
Check downstream references before deleting or renaming. Removing an artifact can stale docs, generated mirrors, configs, and callers; map references first.
Trace the full impact chain after edits. Changing a definition can miss derived outputs and consumers. Follow the affected chain before declaring done.
Verify ALL affected outputs, not just the first. One green check is not all green checks; validate every output surface the change can affect.
Assume existing values are intentional — ask WHY before changing OR flagging one as a defect. Before changing or reporting a constant, limit, flag, cutoff, wording, or pattern, read nearby context and history, the CALLER's ordering, and 2+ sibling call sites of the same convention. A doc stating WHAT without WHY is missing rationale, not proof of a missing guard.
Surface ambiguity before acting — don't pick silently. Multiple valid interpretations require an explicit question or stated assumption with risk.
Assert the outcome your system owns, not the intermediate state your infrastructure owns. When verifying async work, assert the final business state — never the delivery/retry bookkeeping held in shared infrastructure that any co-running process can write. Such a check passes when run alone and flakes the moment anything else shares that infrastructure.
Keep shared guidance role-relevant. Universal guidance must help every receiving skill or agent; code-specific obligations belong only in code-specific protocols.
IMPORTANT MUST ATTENTION cite file:line evidence for every claim. Confidence >80% to act, <60% = do NOT recommend.
MUST ATTENTION apply critical + sequential thinking — every claim needs appropriate traced evidence (file:line for repo/code claims; source URL or artifact section for research, product, content, and docs claims); confidence >80% to act, <60% DO NOT recommend. Anti-hallucination: never present guess as fact, admit uncertainty freely, cross-reference independently, stay skeptical of own confidence.
MUST ATTENTION apply AI mistake prevention — verify generated content against evidence, trace downstream references before deleting or renaming, verify all affected outputs, re-read files after context loss, and surface ambiguity before acting.
- MANDATORY For UI-bearing specs, author/maintain the tech-agnostic interaction-surface layer (View Inventory + Navigation Map + observable UI States + per-story
US-/OP-/BR--traced flow); keep deep visual fidelity in the linked design-spec/mockup recorded in frontmatter; name ZERO frameworks/routes/CSS/component classes; skip ONLY for backend-only features with a stated reason.
Project Protocol Overlay — Before executing this skill, resolve any PROJECT overlay rules layered onto it: match this skill's name against the Target column of the project's skill-protocol index (docs/project-reference/skill-protocols-reference.md by default; a referenceDocs entry in docs/project-config.json overrides the path), taking the most specific matching tier ONLY — exact name > glob > *. That precedence orders overlays against EACH OTHER, never against this skill. Read ONLY the matched bodies, resolved as <protocols-dir>/<Name>.md; a row's Body link is display text, never a read path. A matched body that is missing or malformed is REPORTED and skipped — never reconstructed from the index Description. No index, or no match -> proceed with no overlay, silently. Full contract: .claude/skills/project-skill-protocol/references/registry.md.
Overlays are ADDITIVE ONLY: they ADD rules on top of this skill's own protocol and NEVER replace, override, disable, or reinterpret a rule it already states — removing every overlay must return this skill to exactly its documented behavior. An overlay is a BRIEF, not an authority escalation: it can NEVER waive a workflow gate, git discipline, a review gate, or a user-confirmation gate. A genuine overlay-vs-skill conflict, or two equally-specific overlays that directly contradict -> surface both to the user; NEVER resolve silently.
MUST ATTENTION resolve project protocol overlays for this skill BEFORE executing — most specific matching tier only (exact > glob > *, which ranks overlays against each other, NEVER against this skill), read only matched bodies at <protocols-dir>/<Name>.md; a missing or malformed body is reported, never reconstructed. Overlays are ADDITIVE ONLY (they never replace this skill's own rules) and are a brief, NEVER an authority escalation; an equal-specificity contradiction goes to the user.
Closing Reminders
- IMPORTANT MUST ATTENTION Goal: Produce a tech-free, AI-implementable Feature Spec whose Section 8 TC registry stays the single source of truth, traceable to executing test code — so any team can rebuild the feature on any stack from the spec alone
Protocols in force (concise digest of the SYNC/shared blocks this skill carries — MUST ATTENTION honor each canonical body):
Cross-Service Check: ALWAYS scan producers, consumers, sagas, contracts before concluding; missing consumer = silent regression.
Evidence: cite file:line for every claim; confidence >80% to act, <60% NEVER recommend.
Critical Thinking: apply critical + sequential thinking; NEVER present a guess as fact.
Spec↔Tests↔Code Triangulation: the unit of judgment is the WHOLE PACKAGE (spec §3/§4/§8 + tests + code) — reason mutual-consistency first; a disagreeing or missing face is a logged finding, NEVER a silent pass.
Spec Drift Adjudication: on behavior divergence from a canonical spec, classify CODE-WRONG / SPEC-STALE / AMBIGUOUS / SPEC-SILENT and harvest unwritten invariants into §4/§8 + a guarding test — NEVER normalize drift to whichever side is green.
AI Mistake Prevention: verify generated content against evidence, trace downstream references, verify all affected outputs, re-read after context loss, surface ambiguity.
IMPORTANT MUST ATTENTION [BLOCKING] Resolve the mode FIRST and read its references/{author,tests,sync}.md body — NEVER run draft/init/update/audit/amend/tests/sync from memory; ambiguous → ask the user directly before any mutating mode — why: each mode's gates + output contract live in its body, not in this entry skill
IMPORTANT MUST ATTENTION [BLOCKING] EVERY test case MUST carry verifiable code evidence as a [Source: namespace/service/id] abstract anchor in its Section 8 hidden carrier — physical file:line → provenance sidecar only; sole exception mode=draft (Evidence: TBD + provisional flag, upgraded to real anchor on first code-sourced run) — why: a TC without evidence is unverifiable and silently rots
IMPORTANT MUST ATTENTION [BLOCKING] Section 8 is the canonical TC registry for business TCs — existing TCs MUST NOT be overwritten during update; tests mode owns generation, sync mode reconciles drift — why: test code implements §8, so overwriting it orphans real tests
IMPORTANT MUST ATTENTION [BLOCKING] §1-7 prose is STRICTLY tech-free — no framework/product/language/persistence/messaging/auth names (banned tokens → spec-principles.md §3.2); technical identifiers live ONLY in evidence carriers, frontmatter, and mermaid blocks — why: M1/M5 require rebuild-from-scratch on any stack
IMPORTANT MUST ATTENTION [BLOCKING] Honor the M1-M7 mandates (.claude/skills/shared/sdd-artifact-contract.md) + canonical TC format (.claude/skills/shared/tc-format.md) — any M1-M5 or M7 violation FAILS the artifact (M6 binds the REVIEWER, not the artifact); M7: judge the case BODY by the demo test — an invocation-shaped When is TECHNICAL-ONLY however tech-free its wording; no line-count cap applies, split only for TC volume or distinct capabilities
IMPORTANT MUST ATTENTION INDEX.md/ERD are DERIVED — flag refresh need in update, NEVER trigger $spec-index here — why: separation of concerns keeps the canonical spec the only source of truth
IMPORTANT MUST ATTENTION evidence gate — cite file:line/grep for every claim, confidence >80% to act, <60% do NOT recommend; verify AI-generated TC/source anchors against ACTUAL code (grep to confirm) before authoring — why: hallucinated [Source:] anchors break traceability
IMPORTANT MUST ATTENTION cross-service check before concluding any spec/§8 work — scan producers, consumers, sagas, contracts; per touchpoint owner · message · risk (NONE/ADDITIVE/BREAKING) — why: a missing downstream consumer is a silent regression
IMPORTANT MUST ATTENTION [BLOCKING] Break work into small task tracking tasks BEFORE starting (one per file read) + a final review task; on context loss the current task list first, never duplicate — why: long spec files exhaust context and lose un-tracked progress
IMPORTANT MUST ATTENTION Search codebase for 3+ similar patterns and read existing spec siblings before authoring new content — match local conventions over generic defaults
Anti-Rationalization:
| Evasion |
Rebuttal |
"Mode is obvious, skip the references/ body" |
The body owns gates + output contract — running from memory drifts. Read it every time. |
| "This TC's source is clear, skip the anchor" |
No [Source:] carrier (or Evidence: TBD for non-draft) = unverifiable TC. Add the anchor. |
"update — just regenerate Section 8" |
§8 is canonical; integration tests implement it. NEVER overwrite — sync reconciles drift. |
| "One tech name in prose is harmless" |
One banned token fails M1 and breaks rebuild-on-any-stack. Move it to an evidence carrier. |
| "Small spec, skip task tracking" |
Skip depth, NEVER skip tracking — context loss wipes un-tracked progress. |
"Index looks stale, I'll just run $spec-index" |
Not this skill's job — flag the refresh need; derived artifacts regenerate separately. |
[TASK-PLANNING] MUST ATTENTION analyze task scope and break into small todo tasks/sub-tasks via task tracking before acting.
IMPORTANT MUST ATTENTION Resolve the mode FIRST + read its references/ body · EVERY §8 TC carries a [Source:] evidence anchor (except mode=draft) · §1-7 prose STRICTLY tech-free — the three rules this skill must never skip.
Hookless Prompt Protocol Mirror (Auto-Synced)
Source: .claude/.ck.json + .claude/skills/shared/sync-inline-versions.md (:full blocks) + .claude/scripts/lib/hookless-prompt-protocol.cjs
[WORKFLOW-EXECUTION-PROTOCOL] [BLOCKING] Workflow Execution Protoco
…(truncated)
1---2name: spec3description: [Documentation] Use to author, audit, amend, or test-spec a business Feature Spec. The single spec skill — modes draft|init|update|audit|amend create/maintain the tech-free 8-section Feature Spec; draft authors a provisional spec from an idea/requirement (no code yet, Evidence: TBD); tests generates Section 8 TC-{FEATURE}-{NNN} test specifications; sync reconciles §8 TCs ↔ executing test code. Per-mode procedure lives in references/{author,tests,sync}.md.4---5
6> Codex compatibility note:
7>
8> - Invoke repository skills with `$skill-name` in Codex; this mirrored copy rewrites legacy Claude `/skill-name` references.
9> - Task tracker mandate: BEFORE executing any workflow or skill step, create/update task tracking for all steps and keep it synchronized as progress changes.
10> - User-question prompts mean to ask the user directly in Codex.
11> - Ignore Claude-specific mode-switch instructions when they appear.
12> - Strict execution contract: when a user explicitly invokes a skill, execute that skill protocol as written.
13> - Subagent authorization: when a skill is user-invoked or AI-detected and its protocol requires subagents, that skill activation authorizes use of the required `spawn_agent` subagent(s) for that task.
14> - Do not skip, reorder, or merge protocol steps unless the user explicitly approves the deviation first.
15> - For workflow skills, execute each listed child-skill step explicitly and report step-by-step evidence.
16> - If a required step/tool cannot run in this environment, stop and ask the user before adapting.
17
18<!-- CODEX:PROJECT-REFERENCE-LOADING:START -->
19
20## Codex Project-Reference Loading (No Hooks)
21
22Codex uses static project-reference loading instead of runtime-injected project docs.
23When coding, planning, debugging, testing, or reviewing, open project docs explicitly using this routing.
24
25**Always read:**
26
27- `docs/project-config.json` (project-specific paths, commands, modules, and workflow/test settings)
28- `docs/project-reference/docs-index-reference.md` (routes to the full `docs/project-reference/*` catalog)
29- `docs/project-reference/lessons.md` (always-on guardrails and anti-patterns)
30
31**Missing/stale context route:** If `docs/project-config.json`, the docs index, `lessons.md`, `CLAUDE.md`, `AGENTS.md`, or any task-required reference doc is missing or stale, auto-run `$project-init` or the narrow setup route (`$project-config`, `$docs-init`, `$scan-all`, `$scan --target=<key>`, `$claude-md-init`) before ordinary project-specific work. If Codex mirrors or `AGENTS.md` are missing/stale, ask the user to run `$sync-codex`; do not auto-run it.
32
33**Situation-based docs:**
34
35- Project structure/architecture/tech-stack/deployment/setup (any layer — backend, frontend, or infra): `project-structure-reference.md`
36- Backend/CQRS/API/domain/entity changes: `backend-patterns-reference.md`, `domain-entities-reference.md`
37- Frontend/UI/styling/design-system: `frontend-patterns-reference.md`, `scss-styling-guide.md`, `design-system/README.md`
38- Spec authoring, `docs/specs/` pathing, or TC format: `feature-spec-reference.md`, `spec-system-reference.md`, `spec-principles.md`
39- Behavior/public-contract changes or spec-test-code sync: `workflow-spec-test-code-cycle-reference.md` plus the spec docs above
40- Derived spec indexes/ERDs/reimplementation guides: `spec-system-reference.md` and source Feature Specs under `docs/specs/`
41- Integration test implementation/review: `integration-test-reference.md`
42- E2E test implementation/review: `e2e-test-reference.md`
43- Code review/audit work: `code-review-rules.md` plus domain docs above based on changed files
44
45Do not read all docs blindly. Start from `docs-index-reference.md`, then open only relevant files for the task.
46
47<!-- CODEX:PROJECT-REFERENCE-LOADING:END -->
48
49<!-- PROMPT-ENHANCE:STEP-TASK-ANCHOR:START -->
50
51> **[BLOCKING]** Execute skill steps in declared order. NEVER skip, reorder, or merge steps without explicit user approval.
52> **[BLOCKING]** Before each step or sub-skill call, update task tracking: set `in_progress` when step starts, set `completed` when step ends.
53> **[BLOCKING]** Every completed/skipped step MUST include brief evidence or explicit skip reason.
54> **[BLOCKING]** If Task tools are unavailable, create and maintain an equivalent step-by-step plan tracker with the same status transitions.
55
56<!-- PROMPT-ENHANCE:STEP-TASK-ANCHOR:END -->
57
58## Quick Summary
59
60> **Portability:** `docs/specs/` is the fixed Feature Spec root. `docs/templates/detailed-feature-spec-template.md` remains the default template unless `workflowPatterns.featureDocTemplate` points to another template.
61
62**[IMPORTANT] task tracking** — Break ALL work into small tasks BEFORE starting. For simple tasks, ask user whether to skip.
63
64**Goal:** Own the entire Feature Spec lifecycle in one skill — author/maintain tech-free 8-section business Feature Specs (code evidence carried only in Section 8 test-case anchors, never prose), generate Section 8 test specifications, and reconcile those TCs with executing test code — producing a tech-free, AI-implementable Feature Spec whose Section 8 TC registry stays the single source of truth, traceable to test code, so any team can rebuild the feature on any stack from the spec alone. The mode you run determines which `references/` body drives work; the shared §8 contract, M1-M7 mandates, and quality philosophy below apply every mode.
65
66**Summary:**
67
68- **Purpose:** one skill owns the whole Feature Spec lifecycle across 7 modes — `draft | init | update | audit | amend | tests | sync` — producing/maintaining a tech-free 8-section business Feature Spec whose §8 TC registry is the single source of truth, traceable to executing test code, so any team can rebuild the feature on any stack from the spec alone.
69- **Main steps (every run):** (1) resolve mode FIRST — explicit `[mode=<x>]` wins, else infer from request + repo state, ambiguous → ask the user directly before any mutating mode; (2) read the matching `references/{author,tests,sync}.md` body — NEVER run a mode from memory; (3) task tracking-break the work (one task per file read) before starting; (4) execute the mode's procedure/gates from its body; (5) cross-service check before concluding.
70- **§1-7 prose STRICTLY tech-free** — no framework/product/language/persistence/messaging/auth names (banned tokens → `spec-principles.md` §3.2); technical identifiers live ONLY in evidence carriers, frontmatter, and mermaid blocks. — why: M1/M5 require rebuild-on-any-stack from prose alone.
71- **Section 8 is the canonical TC registry for business TCs** (`TC-{FEATURE}-{NNN}`) — every TC carries verifiable `[Source: namespace/service/id]` evidence (sole exception `mode=draft` → `Evidence: TBD` + provisional flag, upgraded to a real anchor on the first code-sourced run); NEVER overwrite existing TCs during `update` — `tests` owns generation, `sync` reconciles drift.
72- **Honor the M1-M7 mandates** (`sdd-artifact-contract.md`) + canonical TC format (`shared/tc-format.md`) — any **M1-M5 or M7** violation FAILS the artifact (M6 binds the REVIEWER, not the artifact); `INDEX.md`/ERD are DERIVED — flag refresh need in `update`, NEVER trigger `$spec-index` here. — why: separation of concerns keeps the canonical spec the only source of truth.
73
74> **Renamed:** formerly `/feature-spec` (and earlier `/feature-docs`); the former `/spec-tests` skill is now folded in as `mode=tests` / `mode=sync`. Those names no longer resolve as slash commands — use `$spec` with the matching mode.
75
76### Modes (resolve mode FIRST — BLOCKING)
77
78| Mode | Use when… | Body |
79| -------- | --------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------- |
80| `draft` | Author a provisional spec from an idea/requirement/prompt — **no code yet** (TDD-first, §8 `Evidence: TBD`, provisional marker) | `references/author.md` |
81| `init` | No `docs/specs/{Bucket}/` exists — author a full 8-section spec from source | `references/author.md` |
82| `update` | Docs exist + code changed — section-impact-mapped updates | `references/author.md` |
83| `audit` | `--audit` flag or user asks — staleness report per section (never mutates docs) | `references/author.md` |
84| `amend` | `[mode=amend]` from the bugfix workflow — minimal regression-scoped §3/§4/§8 touch only | `references/author.md` |
85| `tests` | Generate or update Section 8 `TC-{FEATURE}-{NNN}` test specifications | `references/tests.md` |
86| `sync` | Reconcile §8 TCs ↔ executing test code (forward/reverse/harvest, orphan, staleness); `harvest` captures a SPEC-SILENT invariant into §4/§5/§8 | `references/sync.md` |
87
88**Mode resolution (do this before any work):**
89
901. Parse the mode from the invocation: explicit `[mode=<x>]` arg wins; else infer from request + repo state ("from idea/requirements/prompt", "draft spec", "no code yet" → `draft`; no `docs/specs/{Bucket}/` AND code exists to source from → `init`; docs exist + diff → `update`; "audit/stale" → `audit`; bugfix caller → `amend`; "write/update test specs", "TCs" → `tests`; "sync tests", "reconcile §8 with tests" → `sync`). **`draft` vs `init`:** both author a new spec, but `draft` sources from idea/requirement text (no code → `Evidence: TBD`, provisional) while `init` sources from existing code (real `[Source:]` evidence). "No docs" alone does NOT imply `init` — check whether code exists to source from. `draft` never auto-overwrites existing §8 TCs.
912. If ambiguous, present the detected mode by asking the user directly before proceeding — NEVER auto-start a mutating mode.
923. **Read the matching `references/` body** — it is the single source of truth for that mode's procedure, gates, and output contract. Do not run a mode from memory.
93
94**Key Rules (all modes):**
95
96- **[BLOCKING]** Read `docs/project-reference/spec-principles.md` — repo-local prose/evidence rules (§3 prose scope + §3.2 banned prose-token list). For the AI-implementability criteria + tech-agnostic mandates, read `.claude/skills/shared/sdd-artifact-contract.md` ("AI-Implementability Gate" + mandates M1-M7) — those are the canonical authority, not the local stub.
97- **[BLOCKING]** EVERY test case MUST carry verifiable code evidence as a `[Source: namespace/service/id]` abstract anchor in its Section 8 hidden carrier — physical `file:line` lives only in the provenance sidecar.
98 > **Exception (`mode=draft`):** an idea-sourced spec has no code yet — its §8 TCs carry `Evidence: TBD` (reference-only) and the spec is flagged provisional (`provisional: true` frontmatter + a "DRAFT — unverified until code lands" header banner). The first `update`/`init` run against real code MUST upgrade every `TBD` to a real `[Source:]` anchor and clear the provisional flag. This mirrors existing TDD-first handling — it relaxes evidence ONLY for draft, never for code-sourced modes.
99- **[BLOCKING]** Section 8 is the canonical TC registry for business TCs — §8 business TCs are the source of truth; test code implements them. The `tests` mode owns generation; `sync` mode reconciles drift; the author modes (`draft`/`init`) populate §8 at authoring time (`draft` with `Evidence: TBD`, `init` with real `[Source:]`) and MUST NOT overwrite existing TCs during UPDATE.
100- Authored docs MUST match the master template's **8 tech-free sections** (Overview, Glossary, User Stories & AC, Business Rules, Domain Model, Process Flows & Interaction Surface, Permissions & Roles, Test Specifications) + YAML frontmatter — zero technical terms in prose, size caps enforced.
101- **[BLOCKING] Canonical TC format authority:** `.claude/skills/shared/tc-format.md` (GWT template, Evidence carrier, decade-numbering, Preservation Tests). **M1-M7 mandates:** `.claude/skills/shared/sdd-artifact-contract.md` — any **M1-M5 or M7** violation FAILS the artifact (M6 binds the REVIEWER, not the artifact).
102
103> `docs/project-reference/feature-spec-reference.md` — project-specific Feature Spec patterns (read directly when relevant). `docs/project-reference/domain-entities-reference.md` — domain entity catalog, relationships, cross-service sync.
104
105### 8-Section Feature Spec Rules (canonical reference)
106
107> Canonical home for the Feature Spec rules; applies to any edit under the Feature Spec docs root.
108
109**Format:** Tech-free 8-section Feature Spec. Activate the `$spec` skill before editing.
110
111**Read first:** `docs/project-reference/feature-spec-reference.md`, `docs/project-reference/spec-system-reference.md`, and `docs/project-reference/spec-principles.md`. For behavior/public-contract changes, also read `docs/project-reference/workflow-spec-test-code-cycle-reference.md`.
112
113**8 sections (exact order):** 1. Overview · 2. Glossary · 3. User Stories & Acceptance Criteria · 4. Business Rules · 5. Domain Model · 6. Process Flows & Interaction Surface · 7. Permissions & Roles · 8. Test Specifications. No technical sections (Commands/Events/API/Cross-Service/Performance/Troubleshooting) — code is the technical source of truth.
114
115> §6 carries a tech-agnostic interaction surface (views/nav/observable states/per-story click-paths) per the `SYNC:ui-intent-layer` block this skill carries; backend-only specs state the skip reason explicitly. This does NOT contradict "No technical sections" — the interaction surface is tech-agnostic INTENT (UX roles, information, states, flows), not a technical "UI Pages" section; M1-clean keeps it free of framework/route/CSS/component-class names.
116
117**Mandatory:**
118
119- §1-7 prose is STRICTLY tech-free — no framework/product/language/persistence/messaging/auth names (banned tokens → `spec-principles.md` §3.2). Technical identifiers live ONLY in evidence carriers.
120- Section 5 (Domain Model): Mermaid ERD + `[Source: component/{service}/{id}]` abstract anchor per entity (cannot be omitted)
121- Section 4 (Business Rules): `[Source: rule/{service}/{id}]` abstract anchor per rule group
122- Section 8 (Test Specifications): canonical business TC source — TC-{FEATURE}-{NNN} IDs, each carrying a hidden `[Source: namespace/service/id]` carrier + a `CoveredBy:` field. Legacy `IntegrationTest:` fields are accepted only as migration input.
123
124**Rules:**
125
126- TC IDs live in Section 8 only — never authored in `docs/specs/` directly
127- Section 8 authored via `$spec [mode=tests]`; `$spec [mode=init]` populates it only during initial authoring
128- No line-count cap applies to Feature Specs. Split the capability only when TCs>40 or distinct module-level capabilities emerge.
129
130### M1-M5 + M7 Compliance (BLOCKING — applies to every authored spec and every TC)
131
132See `.claude/skills/shared/sdd-artifact-contract.md` → "AI-SDD Mandates (M1-M7)" for the full BLOCKING criteria. In brief: **M1** tech-agnostic prose; **M2** no source code in prose; **M3** logical-IDs-first traceability with a SEPARATE `[Source:]` carrier; **M4** AI-implementability (one interpretation, named success/failure); **M5** rebuild-from-scratch on any stack from §1-8 prose alone. Tech terms are allowed ONLY inside evidence carriers (`**Evidence**`, `CoveredBy`, legacy `IntegrationTest`, `[Source:]`), YAML frontmatter, and ` ```mermaid ``` ` blocks.
133
134**M7 — Business-visibility.** Apply the demo test to the case BODY: _"what would a stakeholder SEE change?"_ — no answer → FAIL as TECHNICAL-ONLY. A `When` that is an invocation (a handler runs, a consumer receives, a job fires, data syncs) or a `Then` asserting schema/type/nullability/call-count FAILS. Judge the BODY, never the title or ID.
135
136> **M1 governs vocabulary; M7 governs subject matter.** A technical case in impeccably tech-free prose satisfies M1 while violating M7.
137>
138> M6 is absent from this list by design: it binds the REVIEWER (a review that passes an M1-M5/M7 violation is itself defective), never the artifact. The artifact-facing set reads **M1-M5 and M7**.
139
140---
141
142## Derived-Index Delegation
143
144This skill owns the **canonical** Feature Spec (§1-8) and its §8 TC registry. The bucket `INDEX.md` and cross-capability ERD are **derived** artifacts regenerated by `$spec-index` FROM these specs — never a source of truth, and never authored here. In `update` mode, flag "derived spec artifact refresh may be required" but do NOT trigger `$spec-index` directly (separation of concerns).
145
146---
147
148## Workflow Recommendation
149
150> **MANDATORY IMPORTANT MUST ATTENTION — NO EXCEPTIONS:** If you are NOT already in a workflow, you MUST ATTENTION use ask the user directly to ask the user. Do NOT judge task complexity or decide this is "simple enough to skip" — the user decides whether to use a workflow, not you:
151>
152> 1. **Activate `workflow-feature` workflow** (Recommended) — spec-driven with tests by default: scout → investigate → spec-discovery → domain-analysis → why-review → spec → spec-clarify → plan → plan-review → plan-validate → why-review → spec [mode=tests] → why-review → artifact-review --type=spec-tests → plan → plan-review → feature-implement → domain-entities-review → spec [mode=tests] → why-review → artifact-review --type=spec-tests → spec [mode=sync] → integration-test → integration-test-review → integration-test-verify → workflow-review-changes → production-readiness-review → security-review → changelog → test → docs-update → workflow-end → watzup
153> 2. **Execute `$spec` directly** — run this skill standalone in the resolved mode
154
155---
156
157## Next Steps
158
159**[BLOCKING]** After completing, use ask the user directly to present options. Do NOT skip — user decides:
160
161- **"$spec [mode=tests] (Recommended)"** — Generate/update Section 8 test specs for the documented features (if you just authored/updated a spec)
162- **"$spec [mode=sync]"** — Reconcile Section 8 TCs ↔ executing test code
163- **"$artifact-review --type=spec-tests"** — Audit TC coverage + GIVEN/WHEN/THEN quality
164- **"Skip, continue manually"** — user decides
165
166---
167
168## Related Skills
169
170| Skill | Relationship | When to Call |
171| ------------------------------------ | ---------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------ |
172| `$spec-index` | **Derived consumer** — assembles a regenerable navigation index/ERD FROM these Feature Specs (never a source of truth) | AFTER specs exist — (re)generate the bucket `INDEX.md` / cross-capability ERD over the canonical specs |
173| `$artifact-review --type=spec-tests` | **Reviewer** — audits TC coverage in Section 8 | After `spec [mode=tests]`, to validate TC completeness and GIVEN/WHEN/THEN quality |
174| `$integration-test` | **End consumer** — generates test code from TCs in Section 8 | After `spec [mode=tests]`, to produce actual integration test files |
175| `$docs-update` | **Orchestrator** — calls this skill as Phase 2 | Run `$docs-update` for full chain sync; it calls `$spec` internally |
176| `$changes-review` | **Trigger** — detects feature doc staleness | Calls `$docs-update` when a business doc is stale relative to code changes |
177
178---
179
180> **[IMPORTANT]** Use task tracking to break ALL work into small tasks BEFORE starting — including tasks for each file read. This prevents context loss from long files. For simple tasks, AI MUST ATTENTION ask user whether to skip.
181
182<!-- SYNC:ui-intent-layer -->
183
184> **[BLOCKING] Capture a tech-agnostic UI/UX intent layer in every UI-bearing spec — a reader must be able to visualize how the feature works without naming any technology.** When the feature has a user interface, the spec MUST ATTENTION carry an interaction-surface section so the application — not just its API — can be rebuilt on any stack:
185>
186> 1. **View Inventory** — list each view/screen by its UX ROLE and purpose (e.g. "list of items", "item editor", "confirmation step") and what information it presents. Describe by role, never by an implementation name.
187> 2. **Navigation Map** — how a user moves between views: entry points, transitions, and exits. Trace how this surface connects to neighboring features already in the system.
188> 3. **Key observable UI States** — the distinct states a user can observe per view (empty, loading, populated, error, success, permission-denied, etc.) — described as what the user perceives, not how it is rendered.
189> 4. **Per-story interaction flow** — for each user story, the step-by-step click/action path from intent to outcome, cross-referenced to the logical IDs the spec already owns (`US-`/`OP-`/`BR-`).
190> 5. **Couple to the companion design artifact** — keep deep visual fidelity (layout, tokens, pixel detail) OUT of the spec; it lives in the linked `design-spec`/mockup. Record that companion's path in the spec frontmatter so the spec stays the navigable hub.
191>
192> **M1-clean (NON-NEGOTIABLE):** the prose names ZERO frameworks, routes/URLs, CSS, or component-class names — only roles, information, states, and flows. Technology detail belongs in the companion design artifact, never here.
193>
194> **Skip ONLY** when the feature is backend-only (no UI) — state that reason explicitly in the section.
195
196<!-- /SYNC:ui-intent-layer -->
197
198<!-- SYNC:cross-service-check -->
199
200> **Cross-Service Check** — Microservices/event-driven: MANDATORY before concluding investigation, plan, spec, or feature doc. Missing downstream consumer = silent regression.
201>
202> | Boundary | Grep terms |
203> | ------------------- | ------------------------------------------------------------------------------- |
204> | Event producers | `Publish`, `Dispatch`, `Send`, `emit`, `EventBus`, `outbox`, `IntegrationEvent` |
205> | Event consumers | `Consumer`, `EventHandler`, `Subscribe`, `@EventListener`, `inbox` |
206> | Sagas/orchestration | `Saga`, `ProcessManager`, `Choreography`, `Workflow`, `Orchestrator` |
207> | Sync service calls | HTTP/gRPC calls to/from other services |
208> | Shared contracts | OpenAPI spec, proto, shared DTO — flag breaking changes |
209> | Data ownership | Other service reads/writes same table/collection → Shared-DB anti-pattern |
210>
211> **Per touchpoint:** owner service · message name · consumers · risk (NONE / ADDITIVE / BREAKING).
212>
213> **BLOCKED until:** Producers scanned · Consumers scanned · Sagas checked · Contracts reviewed · Breaking-change risk flagged
214
215<!-- /SYNC:cross-service-check -->
216
217<!-- SYNC:evidence-based-reasoning -->
218
219> **Evidence-Based Reasoning** — Speculation is FORBIDDEN. Every claim needs proof.
220>
221> 1. Cite `file:line`, grep results, or framework docs for EVERY claim
222> 2. Declare confidence: >80% act freely, 60-80% verify first, <60% DO NOT recommend
223> 3. Cross-service validation required for architectural changes
224> 4. "I don't have enough evidence" is valid and expected output
225>
226> **BLOCKED until:** `- [ ]` Evidence file path (`file:line`) `- [ ]` Grep search performed `- [ ]` 3+ similar patterns found `- [ ]` Confidence level stated
227>
228> **Forbidden without proof:** "obviously", "I think", "should be", "probably", "this is because"
229> **If incomplete →** output: `"Insufficient evidence. Verified: [...]. Not verified: [...]."`
230
231<!-- /SYNC:evidence-based-reasoning -->
232
233<!-- SYNC:critical-thinking-mindset -->
234
235> **Critical Thinking Mindset** — Apply critical thinking, sequential thinking. Every claim needs traced proof, confidence >80% to act.
236> **Anti-hallucination:** Never present guess as fact — cite sources for every claim, admit uncertainty freely, self-check output for errors, cross-reference independently, stay skeptical of own confidence — certainty without evidence root of all hallucination.
237
238<!-- /SYNC:critical-thinking-mindset -->
239
240<!-- SYNC:spec-tests-code-triangulation -->
241
242> **Spec ↔ Tests ↔ Code Triangulation** — The unit of review is the WHOLE PACKAGE (spec + tests + code), not the diff alone. Load all three faces together and reason mutual-consistency FIRST, before any isolated per-file check.
243>
244> 1. **Locate all three faces** for the changed behavior: the governing Feature Spec section(s) (§3 ACs / §4 BRs / §8 TCs), the tests that guard it, and the production code. A missing face is a finding (SPEC-GAP / TEST-GAP / DEAD-SPEC).
245> 2. **Triangulate pairwise** — classify which face is wrong on every disagreement:
246> - code vs spec → CODE-EXTRA / SPEC-STALE / CODE-WRONG (a [HARD] §4 rule or §5 invariant with no enforcing path is CODE-WRONG).
247> - tests vs spec → TEST-GAP / SPEC-SILENT.
248> - tests vs code → TEST-GAP / WEAK-TEST (a test that survives a deliberately broken invariant).
249> 3. **Capture hidden rules** — an invariant the code enforces but the spec never states (SPEC-SILENT) is surfaced as a finding, added into §3/§4/§8, and guarded with a test: the enrichment loop, never a silent pass.
250> 4. **Re-review after enrichment** — when triangulation adds spec content or a test, re-review the package against the enriched spec; converge only when a full pass surfaces no new disagreement.
251>
252> NEVER mark PASS while any face disagrees without a logged finding. The diff is the entry point; the package is the unit of judgment.
253
254<!-- /SYNC:spec-tests-code-triangulation -->
255
256<!-- SYNC:spec-drift-adjudication -->
257
258> **Spec drift adjudication (code-wrong vs spec-stale).** Whenever changed behavior diverges from a canonical Feature Spec (business rule, acceptance criterion, flow, state transition, or §8 TC under `docs/specs/`), you MUST NOT silently pick a side. Adjudicate per `shared/sdd-artifact-contract.md` → **Drift Gates**:
259>
260> 1. **Detect** — compare the change against the spec's documented intent. No divergence → record `Spec in sync` and move on.
261> 2. **Classify** the divergence:
262> - **CODE-WRONG** — the spec correctly states intended behavior and the change violates it → BLOCKING finding; fix the code/test against intended behavior (write/adjust a regression TC first).
263> - **SPEC-STALE** — the change is the new intended behavior and the spec now documents the old/wrong behavior → update the spec FIRST via `$spec [mode=update]`, then sync `$spec [mode=tests]` + `$spec [mode=sync]`.
264> - **AMBIGUOUS** — intended behavior is unclear → ask the user directly (or the canonical spec owner) before editing either side.
265> - **SPEC-SILENT** — the code correctly enforces an invariant/behavior that NO canonical spec artifact (§3 AC, §4 BR, §5 invariant, §8 TC) states → not drift but an UNWRITTEN rule discovered by review. ENRICH the spec via the **Invariant Harvest** pass (`$spec [mode=sync] direction=harvest` → `spec/references/sync.md`): prove it is always-true (≥2 enforcement points or a rejecting guard), express it as a universally-quantified property, then add the rule to §4 (or §3/§5) AND a §8 TC via `$spec [update]` + `$spec [mode=tests]` and add the guarding test. A discovered invariant left only in code (or only in tests) is INCOMPLETE — this is the highest-value capture (the rule nobody wrote down).
266> 3. **Never normalize drift just because code/tests are green** — green can encode the drift itself. Reconcile to canonical intent, never to whichever side currently passes.
267>
268> A behavior-changing review/implementation that leaves a spec divergence unadjudicated is INCOMPLETE; an unwritten-but-enforced invariant left uncaptured (no §4/§8 entry) is equally INCOMPLETE.
269
270<!-- /SYNC:spec-drift-adjudication -->
271
272<!-- SYNC:ai-mistake-prevention -->
273
274> **AI Mistake Prevention** — Failure modes to avoid on every task:
275>
276> **Re-read files after context changes.** Context compaction, resume, or long-running work can make memory stale; verify current files before acting.
277> **Verify generated content against source evidence.** AI hallucinates APIs, names, claims, and document facts. Check the relevant source before documenting or referencing.
278> **Check downstream references before deleting or renaming.** Removing an artifact can stale docs, generated mirrors, configs, and callers; map references first.
279> **Trace the full impact chain after edits.** Changing a definition can miss derived outputs and consumers. Follow the affected chain before declaring done.
280> **Verify ALL affected outputs, not just the first.** One green check is not all green checks; validate every output surface the change can affect.
281> **Assume existing values are intentional — ask WHY before changing OR flagging one as a defect.** Before changing or reporting a constant, limit, flag, cutoff, wording, or pattern, read nearby context and history, the CALLER's ordering, and 2+ sibling call sites of the same convention. A doc stating WHAT without WHY is missing rationale, not proof of a missing guard.
282> **Surface ambiguity before acting — don't pick silently.** Multiple valid interpretations require an explicit question or stated assumption with risk.
283> **Assert the outcome your system owns, not the intermediate state your infrastructure owns.** When verifying async work, assert the final business state — never the delivery/retry bookkeeping held in shared infrastructure that any co-running process can write. Such a check passes when run alone and flakes the moment anything else shares that infrastructure.
284> **Keep shared guidance role-relevant.** Universal guidance must help every receiving skill or agent; code-specific obligations belong only in code-specific protocols.
285
286<!-- /SYNC:ai-mistake-prevention -->
287
288<!-- SYNC:evidence-based-reasoning:reminder -->
289
290**IMPORTANT MUST ATTENTION** cite `file:line` evidence for every claim. Confidence >80% to act, <60% = do NOT recommend.
291
292<!-- /SYNC:evidence-based-reasoning:reminder -->
293
294<!-- SYNC:critical-thinking-mindset:reminder -->
295
296**MUST ATTENTION** apply critical + sequential thinking — every claim needs appropriate traced evidence (`file:line` for repo/code claims; source URL or artifact section for research, product, content, and docs claims); confidence >80% to act, <60% DO NOT recommend. Anti-hallucination: never present guess as fact, admit uncertainty freely, cross-reference independently, stay skeptical of own confidence.
297
298<!-- /SYNC:critical-thinking-mindset:reminder -->
299
300<!-- SYNC:ai-mistake-prevention:reminder -->
301
302**MUST ATTENTION** apply AI mistake prevention — verify generated content against evidence, trace downstream references before deleting or renaming, verify all affected outputs, re-read files after context loss, and surface ambiguity before acting.
303
304<!-- /SYNC:ai-mistake-prevention:reminder -->
305
306<!-- SYNC:ui-intent-layer:reminder -->
307
308- **MANDATORY** For UI-bearing specs, author/maintain the tech-agnostic interaction-surface layer (View Inventory + Navigation Map + observable UI States + per-story `US-/OP-/BR-`-traced flow); keep deep visual fidelity in the linked `design-spec`/mockup recorded in frontmatter; name ZERO frameworks/routes/CSS/component classes; skip ONLY for backend-only features with a stated reason.
309
310<!-- /SYNC:ui-intent-layer:reminder -->
311
312<!-- SYNC:project-protocol-overlay -->
313
314> **Project Protocol Overlay** — Before executing this skill, resolve any PROJECT overlay rules layered onto it: match this skill's name against the `Target` column of the project's skill-protocol index (`docs/project-reference/skill-protocols-reference.md` by default; a `referenceDocs` entry in `docs/project-config.json` overrides the path), taking the most specific matching tier ONLY — exact name > glob > `*`. **That precedence orders overlays against EACH OTHER, never against this skill.** Read ONLY the matched bodies, resolved as `<protocols-dir>/<Name>.md`; a row's Body link is display text, never a read path. A matched body that is missing or malformed is REPORTED and skipped — never reconstructed from the index Description. No index, or no match -> proceed with no overlay, silently. Full contract: `.claude/skills/project-skill-protocol/references/registry.md`.
315>
316> Overlays are **ADDITIVE ONLY**: they ADD rules on top of this skill's own protocol and NEVER replace, override, disable, or reinterpret a rule it already states — removing every overlay must return this skill to exactly its documented behavior. An overlay is a BRIEF, not an authority escalation: it can NEVER waive a workflow gate, git discipline, a review gate, or a user-confirmation gate. A genuine overlay-vs-skill conflict, or two equally-specific overlays that directly contradict -> surface both to the user; NEVER resolve silently.
317
318<!-- /SYNC:project-protocol-overlay -->
319
320<!-- SYNC:project-protocol-overlay:reminder -->
321
322**MUST ATTENTION** resolve project protocol overlays for this skill BEFORE executing — most specific matching tier only (exact > glob > `*`, which ranks overlays against each other, NEVER against this skill), read only matched bodies at `<protocols-dir>/<Name>.md`; a missing or malformed body is reported, never reconstructed. Overlays are ADDITIVE ONLY (they never replace this skill's own rules) and are a brief, NEVER an authority escalation; an equal-specificity contradiction goes to the user.
323
324<!-- /SYNC:project-protocol-overlay:reminder -->
325
326## Closing Reminders
327
328- **IMPORTANT MUST ATTENTION Goal:** Produce a tech-free, AI-implementable Feature Spec whose Section 8 TC registry stays the single source of truth, traceable to executing test code — so any team can rebuild the feature on any stack from the spec alone
329
330**Protocols in force (concise digest of the SYNC/shared blocks this skill carries — MUST ATTENTION honor each canonical body):**
331
332- **Cross-Service Check:** ALWAYS scan producers, consumers, sagas, contracts before concluding; missing consumer = silent regression.
333- **Evidence:** cite `file:line` for every claim; confidence >80% to act, <60% NEVER recommend.
334- **Critical Thinking:** apply critical + sequential thinking; NEVER present a guess as fact.
335- **Spec↔Tests↔Code Triangulation:** the unit of judgment is the WHOLE PACKAGE (spec §3/§4/§8 + tests + code) — reason mutual-consistency first; a disagreeing or missing face is a logged finding, NEVER a silent pass.
336- **Spec Drift Adjudication:** on behavior divergence from a canonical spec, classify CODE-WRONG / SPEC-STALE / AMBIGUOUS / SPEC-SILENT and harvest unwritten invariants into §4/§8 + a guarding test — NEVER normalize drift to whichever side is green.
337- **AI Mistake Prevention:** verify generated content against evidence, trace downstream references, verify all affected outputs, re-read after context loss, surface ambiguity.
338
339- **IMPORTANT MUST ATTENTION [BLOCKING]** Resolve the mode FIRST and read its `references/{author,tests,sync}.md` body — NEVER run `draft`/`init`/`update`/`audit`/`amend`/`tests`/`sync` from memory; ambiguous → ask the user directly before any mutating mode — why: each mode's gates + output contract live in its body, not in this entry skill
340- **IMPORTANT MUST ATTENTION [BLOCKING]** EVERY test case MUST carry verifiable code evidence as a `[Source: namespace/service/id]` abstract anchor in its Section 8 hidden carrier — physical `file:line` → provenance sidecar only; sole exception `mode=draft` (`Evidence: TBD` + provisional flag, upgraded to real anchor on first code-sourced run) — why: a TC without evidence is unverifiable and silently rots
341- **IMPORTANT MUST ATTENTION [BLOCKING]** Section 8 is the canonical TC registry for business TCs — existing TCs MUST NOT be overwritten during `update`; `tests` mode owns generation, `sync` mode reconciles drift — why: test code implements §8, so overwriting it orphans real tests
342- **IMPORTANT MUST ATTENTION [BLOCKING]** §1-7 prose is STRICTLY tech-free — no framework/product/language/persistence/messaging/auth names (banned tokens → `spec-principles.md` §3.2); technical identifiers live ONLY in evidence carriers, frontmatter, and mermaid blocks — why: M1/M5 require rebuild-from-scratch on any stack
343- **IMPORTANT MUST ATTENTION [BLOCKING]** Honor the M1-M7 mandates (`.claude/skills/shared/sdd-artifact-contract.md`) + canonical TC format (`.claude/skills/shared/tc-format.md`) — any **M1-M5 or M7** violation FAILS the artifact (M6 binds the REVIEWER, not the artifact); M7: judge the case BODY by the demo test — an invocation-shaped `When` is TECHNICAL-ONLY however tech-free its wording; no line-count cap applies, split only for TC volume or distinct capabilities
344- **IMPORTANT MUST ATTENTION** `INDEX.md`/ERD are DERIVED — flag refresh need in `update`, NEVER trigger `$spec-index` here — why: separation of concerns keeps the canonical spec the only source of truth
345- **IMPORTANT MUST ATTENTION** evidence gate — cite `file:line`/grep for every claim, confidence >80% to act, <60% do NOT recommend; verify AI-generated TC/source anchors against ACTUAL code (grep to confirm) before authoring — why: hallucinated `[Source:]` anchors break traceability
346- **IMPORTANT MUST ATTENTION** cross-service check before concluding any spec/§8 work — scan producers, consumers, sagas, contracts; per touchpoint owner · message · risk (NONE/ADDITIVE/BREAKING) — why: a missing downstream consumer is a silent regression
347- **IMPORTANT MUST ATTENTION [BLOCKING]** Break work into small task tracking tasks BEFORE starting (one per file read) + a final review task; on context loss the current task list first, never duplicate — why: long spec files exhaust context and lose un-tracked progress
348- **IMPORTANT MUST ATTENTION** Search codebase for 3+ similar patterns and read existing spec siblings before authoring new content — match local conventions over generic defaults
349
350**Anti-Rationalization:**
351
352| Evasion | Rebuttal |
353| ------------------------------------------------ | -------------------------------------------------------------------------------------------- |
354| "Mode is obvious, skip the `references/` body" | The body owns gates + output contract — running from memory drifts. Read it every time. |
355| "This TC's source is clear, skip the anchor" | No `[Source:]` carrier (or `Evidence: TBD` for non-draft) = unverifiable TC. Add the anchor. |
356| "`update` — just regenerate Section 8" | §8 is canonical; integration tests implement it. NEVER overwrite — `sync` reconciles drift. |
357| "One tech name in prose is harmless" | One banned token fails M1 and breaks rebuild-on-any-stack. Move it to an evidence carrier. |
358| "Small spec, skip task tracking" | Skip depth, NEVER skip tracking — context loss wipes un-tracked progress. |
359| "Index looks stale, I'll just run `$spec-index`" | Not this skill's job — flag the refresh need; derived artifacts regenerate separately. |
360
361**[TASK-PLANNING]** MUST ATTENTION analyze task scope and break into small todo tasks/sub-tasks via task tracking before acting.
362
363**IMPORTANT MUST ATTENTION** Resolve the mode FIRST + read its `references/` body · EVERY §8 TC carries a `[Source:]` evidence anchor (except `mode=draft`) · §1-7 prose STRICTLY tech-free — the three rules this skill must never skip.
364
365---
366
367<!-- CODEX:SYNC-PROMPT-PROTOCOLS:START -->
368
369## Hookless Prompt Protocol Mirror (Auto-Synced)
370
371Source: `.claude/.ck.json` + `.claude/skills/shared/sync-inline-versions.md` (`:full` blocks) + `.claude/scripts/lib/hookless-prompt-protocol.cjs`
372
373## [WORKFLOW-EXECUTION-PROTOCOL] [BLOCKING] Workflow Execution Protoco
374
375…(truncated)