Codex compatibility note:
- Invoke repository skills with
$skill-name in Codex; this mirrored copy rewrites legacy Claude /skill-name references.
- Task tracker mandate: BEFORE executing any workflow or skill step, create/update task tracking for all steps and keep it synchronized as progress changes.
- User-question prompts mean to ask the user directly in Codex.
- Ignore Claude-specific mode-switch instructions when they appear.
- Strict execution contract: when a user explicitly invokes a skill, execute that skill protocol as written.
- Subagent authorization: when a skill is user-invoked or AI-detected and its protocol requires subagents, that skill activation authorizes use of the required
spawn_agent subagent(s) for that task.
- Do not skip, reorder, or merge protocol steps unless the user explicitly approves the deviation first.
- For workflow skills, execute each listed child-skill step explicitly and report step-by-step evidence.
- If a required step/tool cannot run in this environment, stop and ask the user before adapting.
Codex Project-Reference Loading (No Hooks)
Codex uses static project-reference loading instead of runtime-injected project docs.
When coding, planning, debugging, testing, or reviewing, open project docs explicitly using this routing.
Always read:
docs/project-config.json (project-specific paths, commands, modules, and workflow/test settings)
docs/project-reference/docs-index-reference.md (routes to the full docs/project-reference/* catalog)
docs/project-reference/lessons.md (always-on guardrails and anti-patterns)
Missing/stale context route: If docs/project-config.json, the docs index, lessons.md, CLAUDE.md, AGENTS.md, or any task-required reference doc is missing or stale, auto-run $project-init or the narrow setup route ($project-config, $docs-init, $scan-all, $scan --target=<key>, $claude-md-init) before ordinary project-specific work. If Codex mirrors or AGENTS.md are missing/stale, ask the user to run $sync-codex; do not auto-run it.
Situation-based docs:
- Project structure/architecture/tech-stack/deployment/setup (any layer — backend, frontend, or infra):
project-structure-reference.md
- Backend/CQRS/API/domain/entity changes:
backend-patterns-reference.md, domain-entities-reference.md
- Frontend/UI/styling/design-system:
frontend-patterns-reference.md, scss-styling-guide.md, design-system/README.md
- Spec authoring,
docs/specs/ pathing, or TC format: feature-spec-reference.md, spec-system-reference.md, spec-principles.md
- Behavior/public-contract changes or spec-test-code sync:
workflow-spec-test-code-cycle-reference.md plus the spec docs above
- Derived spec indexes/ERDs/reimplementation guides:
spec-system-reference.md and source Feature Specs under docs/specs/
- Integration test implementation/review:
integration-test-reference.md
- E2E test implementation/review:
e2e-test-reference.md
- Code review/audit work:
code-review-rules.md plus domain docs above based on changed files
Do not read all docs blindly. Start from docs-index-reference.md, then open only relevant files for the task.
[BLOCKING] Execute skill steps in declared order. NEVER skip, reorder, or merge steps without explicit user approval.
[BLOCKING] Before each step or sub-skill call, update task tracking: set in_progress when step starts, set completed when step ends.
[BLOCKING] Every completed/skipped step MUST include brief evidence or explicit skip reason.
[BLOCKING] If Task tools are unavailable, create and maintain an equivalent step-by-step plan tracker with the same status transitions.
Quick Summary
Goal: Deliver user-confirmed, per-layer tech stack — each choice backed by 3+ researched options, weighted scoring, cited evidence, confidence % — by acting as solution architect: derive technical requirements from business analysis, research current market, produce detailed comparison report, so team commits to stack fit for scale, budget, skills, timeline, NOT familiarity.
Summary:
- Purpose: act as solution architect — derive technical requirements from business analysis, research current market, produce per-layer comparison report so team commits to stack fit for scale/budget/skills/timeline, NOT familiarity.
- All 7 main steps (run in order): (1) Load Business Context → (2) Derive Technical Requirements + user-confirm → (3) Research Per Layer (WebSearch 3+ options each) → (4) Deep Comparison Matrix → (5) Weighted Score & Ranking → (6) Generate Report → (7) User Validation Interview.
- Requirements BEFORE research: load prior business/domain/PBI artifacts (Step 1), map business signals → technical requirements (Step 2), gate on user confirmation (ask the user directly) before any WebSearch (Step 3).
- Evaluate every stack layer (backend, frontend, database, messaging, infra, auth) independently — minimum 3 WebSearched options per layer, each with cited evidence (URL, benchmark, case study), NEVER familiarity (Steps 3-4).
- Score with weighted 8-criteria matrix (High=3x / Medium=2x / Low=1x), rank each layer with confidence %; capped <=200-line report →
{plan-dir}/research/tech-stack-comparison.md (Steps 5-6).
- End-of-skill user validation interview (5-8 questions) mandatory, NEVER skipped — only confirmed decisions written to
phase-02-tech-stack.md as status: confirmed (Step 7).
Workflow:
- Load Business Context — Read business evaluation, domain model, refined PBI artifacts
- Derive Technical Requirements — Map business needs to technical constraints
- Research Per Layer — WebSearch top 3 options for each stack responsibility
- Deep Compare — Pros/cons matrix, benchmarks, community health, team fit
- Score & Rank — Weighted scoring across 8 criteria
- Generate Report — Structured comparison report with recommendation
- User Validation — Present findings, ask 5-8 questions, confirm choices
Key Rules:
- MANDATORY IMPORTANT MUST ATTENTION research minimum 3 options per stack layer
- MANDATORY IMPORTANT MUST ATTENTION include confidence % with evidence for every recommendation
- MANDATORY IMPORTANT MUST ATTENTION run user validation interview at end (NEVER skip)
- All claims must cite sources (URL, benchmark, case study)
- Recommend on benchmarked evidence (URL, benchmark, case study); NEVER on familiarity alone
Be skeptical. Apply critical thinking, sequential thinking. Every claim needs traced proof, confidence percentages (Idea should be more than 80%).
Step 1: Load Business Context
Read artifacts from prior workflow steps (search plans/, team-artifacts/):
- Business evaluation report (viability, scale, constraints)
- Domain model / ERD (complexity, entity count, relationships)
- Refined PBI (acceptance criteria, scope)
- Discovery interview notes (team skills, budget, timeline)
Extract and summarize:
| Signal |
Value |
Source |
| Expected users |
... |
discovery interview |
| Domain complexity |
Low/Med/High |
domain model |
| Team skills |
... |
discovery interview |
| Budget constraint |
... |
business evaluation |
| Timeline |
... |
business evaluation |
| Compliance needs |
... |
business evaluation |
| Real-time needs |
Yes/No |
refined PBI |
| Integration complexity |
Low/Med/High |
domain model |
Step 2: Derive Technical Requirements
Map business signals to technical requirements:
| Business Signal |
Technical Requirement |
Priority |
| High user scale |
Horizontal scaling, connection pooling |
Must |
| Complex domain |
Strong type system, ORM with migrations |
Must |
| Real-time features |
WebSocket/SSE support, event-driven arch |
Must |
| Small team |
Low learning curve, good DX, batteries-included |
Should |
| Tight budget |
Open-source, low hosting cost |
Should |
| Compliance |
Audit trail, encryption, auth framework |
Must |
MANDATORY IMPORTANT MUST ATTENTION validate derived requirements with user by asking the user directly before proceeding to research.
Step 3: Research Per Stack Layer
For EACH layer, research top 3 options via WebSearch (minimum 5 queries total):
Stack Layers to Evaluate
| Layer |
Example Options |
Research Focus |
| Backend Framework |
Candidate backend runtimes/frameworks |
Performance, type safety, ecosystem |
| Frontend Framework |
Candidate frontend frameworks |
DX, ecosystem, hiring, enterprise fit |
| Database |
Candidate database engines/stores |
Scale, query complexity, cost |
| Messaging/Events |
Candidate messaging/event systems |
Throughput, reliability, complexity |
| Infrastructure |
Docker+K8s, Serverless, PaaS |
Cost, ops overhead, scaling |
| Auth |
Keycloak, Auth0, custom |
Cost, compliance, flexibility |
WebSearch Queries (minimum 5 per layer)
"{option_A} vs {option_B} {current_year} comparison"
"{option} enterprise production case studies"
"{option} community size github stars"
"{option} performance benchmarks {use_case}"
"{option} security track record vulnerabilities"
Step 4: Deep Comparison Matrix
For EACH stack layer, produce comparison table:
| Criteria |
Option A |
Option B |
Option C |
Weight |
| Team Fit |
score + rationale |
... |
... |
High |
| Scalability |
score + rationale |
... |
... |
High |
| Time-to-Market |
score + rationale |
... |
... |
High |
| Ecosystem/Libs |
score + rationale |
... |
... |
Medium |
| Hiring Market |
score + rationale |
... |
... |
Medium |
| Cost (hosting) |
score + rationale |
... |
... |
Medium |
| Learning Curve |
score + rationale |
... |
... |
Medium |
| Community Health |
score + rationale |
... |
... |
Low |
Scoring: 1-5 scale. Weight: High=3x, Medium=2x, Low=1x.
Per-Option Detail Block
For each option, document:
### {Layer}: {Option Name}
**Pros:**
- {Pro 1} — {evidence/source}
- {Pro 2} — {evidence/source}
- {Pro 3} — {evidence/source}
**Cons:**
- {Con 1} — {evidence/source}
- {Con 2} — {evidence/source}
**Best suited when:** {conditions}
**Not suitable when:** {conditions}
**Production examples:** {2-3 real companies using this}
Step 5: Weighted Score & Ranking
Calculate weighted total per option per layer. Present ranking:
### {Layer} Ranking
1. **{Option A}** — Score: {X}/100 — Confidence: {Y}%
2. **{Option B}** — Score: {X}/100 — Confidence: {Y}%
3. **{Option C}** — Score: {X}/100 — Confidence: {Y}%
**Recommendation:** {Option A}
**Why:** {2-3 sentence rationale linking to team skills, scale, and constraints}
Step 6: Generate Report
Write report to {plan-dir}/research/tech-stack-comparison.md with:
- Executive summary (recommended full stack in 5 lines)
- Technical requirements table (from Step 2)
- Per-layer comparison matrices (from Step 4)
- Per-layer rankings with recommendations (from Step 5)
- Combined recommended stack diagram
- Risk assessment for recommended stack
- Alternative stack (second-best combo) for comparison
- Unresolved questions
Report must be <=200 lines. Use tables over prose.
Step 7: User Validation Interview
MANDATORY IMPORTANT MUST ATTENTION present findings and ask 5-8 questions by asking the user directly:
Required Questions
- Per-layer recommendation confirmation — "For {layer}, I recommend {option}. Agree?"
- Options: Agree (Recommended) | Prefer {option B} | Need more research
- Risk tolerance — "The recommended stack has {risk}. Acceptable?"
- Team readiness — "Team needs to learn {X}. Training plan needed?"
- Budget alignment — "Estimated infra cost: ${X}/month. Within budget?"
- Timeline fit — "This stack enables MVP in {X} months. Acceptable?"
Optional Deep-Dive Questions (pick 2-3 based on context)
- "Should we consider {emerging tech} for {layer}?"
- "Any compliance requirements I haven't captured?"
- "Preference for managed services vs self-hosted?"
- "Monorepo or polyrepo for this team size?"
After user confirms, update report with final decisions, mark status: confirmed.
Output
{plan-dir}/research/tech-stack-comparison.md # Full comparison report
{plan-dir}/phase-02-tech-stack.md # Final confirmed tech stack decisions
MANDATORY IMPORTANT MUST ATTENTION break work into small todo tasks using task tracking BEFORE starting.
MANDATORY IMPORTANT MUST ATTENTION validate EVERY recommendation with user by asking the user directly — NEVER auto-decide.
MANDATORY IMPORTANT MUST ATTENTION include confidence % and evidence citations for all claims.
MANDATORY IMPORTANT MUST ATTENTION add a final review todo task to verify work quality.
Next Steps
MANDATORY IMPORTANT MUST ATTENTION — NO EXCEPTIONS after completing this skill, you MUST ATTENTION use ask the user directly to present these options. Do NOT skip because task seems "simple"/"obvious" — the user decides:
- "$architecture-design (Recommended)" — Design solution architecture with chosen tech stack
- "$plan" — If architecture already decided
- "Skip, continue manually" — user decides
Council escalation (always-offer, second prompt)
After the existing ## Next Steps prompt above resolves, present a second, independent ask the user directly call:
- "Skip council — proceed with chosen stack (Recommended)" — Continue with the selected tech stack as-is.
- "Escalate to $llm-council" — Run 11 sub-agent council. Best applied when 2+ stacks score within 15% on the comparison matrix or you have unfamiliar/strategic dependencies. Cheaper alternatives:
$why-review, $plan-validate.
Scenario Stress & Resilience Evaluation — CONDITIONAL, evidence-gated, business-criticality-aware. The top-down companion to SYNC:scale-technique-gate: instead of "is technique X present?", put the system UNDER concrete failure/load scenarios and judge whether it SURVIVES, SELF-HEALS, and whether its BUSINESS needs it to. ADVICE-ONLY: emit the Scenario Stress Matrix as guidance; NEVER mutate any score, verdict band, or gate pass/fail.
- Reuse the scale tier derived by
SYNC:scale-technique-gate (or derive it identically from evidence); also derive business-criticality B0–B3 from specs/SLA/product docs + the domain, cite file:line + confidence. B0 best-effort · B1 important · B2 business-critical · B3 mission-critical/regulated. Unknown → state the assumption, do NOT default to B3/T3. Criticality-signal floor (both-directions safety): regulated / PII / financial / health data, money movement, auth/identity, or legal-compliance scope raises B to at least B2 even absent SLA/SLO docs; anti-over-engineering lowers hardening ONLY when NO such signal is present. B (blast if it fails) and T (scale of load/data) are independent — a low-traffic payroll run is low-T, high-B.
- Select in-scope scenarios — only those the system's
B/T combination warrants (a B0 internal PoC skips region-loss/DR entirely; a B3/T0 regulated service still needs backups + DR by BUSINESS, not scale).
- Walk each in-scope scenario: simulate the stimulus → trace the break path → name the failure signature → answer the self-heal/recovery question (auto-recover? MTTR? manual runbook?) → name the trade-off it forces. Families: traffic spike · sustained growth · data-volume growth · write/ingest burst · dependency down/slow · instance/node loss · zone/region loss · data loss/corruption · poison-message/retry-storm · cascading failure/backpressure · cold-start/deploy-blip · clock-skew/duplicate-delivery.
- Assign one verdict per scenario:
WITHSTANDS · DEGRADES-GRACEFULLY · FAILS-HARD (→ advise only) · N/A-by-business (not warranted → skip, not a gap) · OVER-HARDENED (resilience beyond business need → advise AGAINST, cite carrying cost).
- Anti-over-engineering guard (first-class): a lean system whose business does not need HA/DR is a PASS;
OVER-HARDENED flags resilience the business does not warrant. This guard is symmetric with the criticality-signal floor above — never under-harden a B2+ system just because its traffic is low.
- Output — Scenario Stress Matrix:
scenario | in-scope (B/T)? | verdict | self-heal | trade-off | evidence (file:line/config/infra). Full catalog + Business×Scale in-scope baseline + verdict/tier tables → .claude/docs/scenario-stress-catalog.md. ADVISORY-ONLY: NEVER mutate any /20, /24, verdict band, or gate pass/fail. Drift-guard: scenarios/verdicts/business-tiers are AUTHORITATIVE in the catalog — update it FIRST, then re-run .claude/scripts/inject_scenario_stress_gate.py. Scale tier stays single-sourced in scale-technique-catalog.md.
BLOCKED until: - [ ] scale tier + business-criticality (with criticality-signal floor) derived from evidence - [ ] in-scope scenarios selected - [ ] matrix emitted - [ ] over-hardening guard applied - [ ] advisory-only (no score/verdict mutation) confirmed
IMPORTANT MUST ATTENTION scale-technique gate: derive the scale tier from evidence FIRST (T0 internal · T1 <10k · T2 10k–1M · T3 millions+), then judge each warranted technique PRESENT/MISSING-WARRANTED/N/A-by-scale/OVER-ENGINEERED. Advise on warranted-but-missing gaps AND advise AGAINST unwarranted heavyweight techniques (anti-over-engineering). ADVICE-ONLY — emit the Technique Applicability Matrix as guidance; NEVER mutate any score, verdict band, or gate pass/fail. Full catalog → .claude/docs/scale-technique-catalog.md (authoritative for tier thresholds & per-technique warranting tiers — on any change update the catalog FIRST, then re-run inject_scale_technique_gate.py).
IMPORTANT MUST ATTENTION scenario-stress gate: reuse the scale tier T0–T3 AND derive business-criticality B0–B3 from evidence first — apply the criticality-signal floor (regulated/PII/financial/health data · money movement · auth/identity · legal-compliance → at least B2 even absent SLA docs; do NOT default to B3). Select only the scenarios the B/T combination warrants, then walk each (simulate → trace → failure signature → self-heal/MTTR → trade-off) and assign WITHSTANDS/DEGRADES-GRACEFULLY/FAILS-HARD/N/A-by-business/OVER-HARDENED. Anti-over-engineering is first-class (a lean system that needs no HA/DR is a PASS) AND symmetric (never under-harden a B2+ system for low traffic). ADVICE-ONLY — emit the Scenario Stress Matrix as guidance; NEVER mutate any score, verdict band, or gate pass/fail. Full catalog → .claude/docs/scenario-stress-catalog.md (authoritative for scenarios/verdicts/business-tiers — on any change update the catalog FIRST, then re-run inject_scenario_stress_gate.py; scale tier stays single-sourced in scale-technique-catalog.md).
Prompt-Enhance Closing Anchors
- IMPORTANT MUST ATTENTION follow declared step order for this skill; NEVER skip, reorder, or merge steps without explicit user approval
- IMPORTANT MUST ATTENTION for every step/sub-skill call: set
in_progress before execution, set completed after execution
- IMPORTANT MUST ATTENTION every skipped step MUST include explicit reason; every completed step MUST include concise evidence
- IMPORTANT MUST ATTENTION if Task tools unavailable, maintain an equivalent step-by-step plan tracker with synchronized statuses
Project Protocol Overlay — Before executing this skill, resolve any PROJECT overlay rules layered onto it: match this skill's name against the Target column of the project's skill-protocol index (docs/project-reference/skill-protocols-reference.md by default; a referenceDocs entry in docs/project-config.json overrides the path), taking the most specific matching tier ONLY — exact name > glob > *. That precedence orders overlays against EACH OTHER, never against this skill. Read ONLY the matched bodies, resolved as <protocols-dir>/<Name>.md; a row's Body link is display text, never a read path. A matched body that is missing or malformed is REPORTED and skipped — never reconstructed from the index Description. No index, or no match -> proceed with no overlay, silently. Full contract: .claude/skills/project-skill-protocol/references/registry.md.
Overlays are ADDITIVE ONLY: they ADD rules on top of this skill's own protocol and NEVER replace, override, disable, or reinterpret a rule it already states — removing every overlay must return this skill to exactly its documented behavior. An overlay is a BRIEF, not an authority escalation: it can NEVER waive a workflow gate, git discipline, a review gate, or a user-confirmation gate. A genuine overlay-vs-skill conflict, or two equally-specific overlays that directly contradict -> surface both to the user; NEVER resolve silently.
MUST ATTENTION resolve project protocol overlays for this skill BEFORE executing — most specific matching tier only (exact > glob > *, which ranks overlays against each other, NEVER against this skill), read only matched bodies at <protocols-dir>/<Name>.md; a missing or malformed body is reported, never reconstructed. Overlays are ADDITIVE ONLY (they never replace this skill's own rules) and are a brief, NEVER an authority escalation; an equal-specificity contradiction goes to the user.
Closing Reminders
IMPORTANT MUST ATTENTION Goal: deliver user-confirmed, per-layer tech stack — each choice backed by 3+ researched options, weighted 8-criteria scoring, cited evidence, confidence % — so team commits to a stack fit for scale, budget, skills, timeline, NOT familiarity.
IMPORTANT MUST ATTENTION — run ALL 7 steps in declared order, none skipped: (1) Load Business Context → (2) Derive Technical Requirements (+ ask the user directly confirm) → (3) Research Per Layer (WebSearch 3+ options each) → (4) Deep Comparison Matrix → (5) Weighted Score & Ranking (confidence %) → (6) Generate Report (<=200 lines) → (7) User Validation Interview (5-8 questions, write status: confirmed) — why: AI keeps collapsing this into "just pick a stack" and dropping requirements-derivation, scoring, and the confirmation gate that make the choice defensible.
Protocols in force (concise digest of the SYNC/shared blocks this skill carries):
- Critical Thinking: MUST ATTENTION apply critical + sequential thinking; traced proof, confidence >80% to act, NEVER guess as fact.
- AI Mistake Prevention: verify generated content against evidence, trace downstream references, verify all affected outputs, re-read after context loss, surface ambiguity.
IMPORTANT MUST ATTENTION research minimum 3 WebSearched options per stack layer (backend, frontend, database, messaging, infra, auth); every recommendation carries confidence % + cited evidence (URL, benchmark, case study) — NEVER recommend on familiarity alone — why: familiarity bias commits the team to the wrong stack that surfaces only at scale.
IMPORTANT MUST ATTENTION gate on user by asking the user directly at EVERY decision point — confirm derived requirements before research (Step 2), confirm each layer recommendation in the end interview (Step 7) — NEVER auto-decide — why: the team owns the stack, not the AI.
MANDATORY IMPORTANT MUST ATTENTION break work into small todo tasks using task tracking BEFORE starting; mark one in_progress, completed immediately after evidence; add a final review todo.
Scalability & Production-Readiness Technique Gate — CONDITIONAL, evidence-gated, scale-tiered. Judge which system-design techniques a system warrants at its scale — flag warranted-but-missing gaps AND advise AGAINST unwarranted heavyweight ones. ADVICE-ONLY: emit the matrix as guidance; NEVER mutate any score, verdict band, or gate pass/fail.
- Derive the scale tier FIRST — from evidence, never assumed. Read users/RPS, SLO/latency targets, data volume, tenancy, topology from config/infra/specs; cite
file:line + confidence. Tiers: T0 internal/single-instance · T1 small SaaS (<10k users) · T2 high-scale (10k–1M) · T3 massive/multi-region (millions+). Unknown tier → state assumption, do NOT default to T3.
- Judge each concern group only at/above its warranting tier (member techniques → owning review skill for depth):
- Traffic & Edge — Rate Limiting, Load Balancing, Reverse Proxy, API Gateway, CDN, Edge Caching, WAF, DDoS (T1+; CDN/WAF T2+) → security-review owns WAF/DDoS
- Caching & Data Access — Caching, Cache Invalidation, DB Indexing, Query Optimization, N+1, Connection Pooling (T1+) → performance-review owns depth
- Data Scaling & Consistency — Read Replicas, Sharding, Partitioning, Replication, CAP, Eventual Consistency, Locks, Leader Election (T2+; sharding/multi-region T3) → performance-review
- Async & Messaging — Message Queues, Pub/Sub, Event-Driven, Saga, DLQ, Distributed Transactions, Backpressure, Webhooks, WebSockets/SSE (T2+)
- Resilience — Circuit Breakers, Timeouts, Retries, Backoff, Idempotency, Health Checks, Liveness/Readiness, Failover, Graceful Degradation (T1+) → production-readiness-review
- Scaling & Compute — Autoscaling, Horizontal/Vertical Scaling, Serverless Limits, Cold Starts, Cron Jobs, Thread Safety, GC/Memory Leaks (T1+; autoscaling T2+)
- Deployment & Release — CI/CD, Docker, Kubernetes, Blue-Green/Canary/Rolling, Rollbacks, Feature Flags, IaC/Terraform/Helm, Build Caching (CI/CD T0+; K8s/canary T2+)
- Observability — Monitoring, Logging, Distributed Tracing, Metrics, Alerting, SLOs/SLIs, Error Budgets (T1+; tracing/error-budgets T2+) → production-readiness-review
- Security & Compliance — Secrets Management, IAM, OAuth, JWT Rotation, TLS, Encryption at Rest/Transit, CORS, CSRF, SQLi, XSS, SSRF (T0+) → security-review owns
- DR & Infra — Backups, Disaster Recovery, Multi-Region, Chaos Engineering, Schema Versioning, DB Migrations, Cost Optimization (backups T1+; DR/multi-region/chaos T3) → production-readiness-review
- Assign one of 4 verdicts per warranted technique:
PRESENT · MISSING-WARRANTED (→ advise only — guidance, NOT a score/gate lever) · N/A-by-scale (below warranting tier) · OVER-ENGINEERED (present but unwarranted at this tier → advise AGAINST).
- Anti-over-engineering guard (first-class): do NOT recommend K8s, sharding, multi-region, service mesh, event sourcing, or distributed transactions below their warranting tier. A correctly-lean small system is a PASS, never a gap.
- Output — Technique Applicability Matrix:
technique | tier-warranted? | present? | verdict | advice | evidence (file:line/config/infra). Full grouped catalog + per-tier baseline → .claude/docs/scale-technique-catalog.md. Hosting reviews surface this matrix WITHOUT changing any /20, /24, verdict band, or PASS/FAIL (per user decision 2026-07-06). Drift-guard: tier thresholds & per-technique warranting tiers are AUTHORITATIVE in .claude/docs/scale-technique-catalog.md — the inline tier summary above is a condensed pointer; on any tier/technique change, update the catalog FIRST, then re-run .claude/scripts/inject_scale_technique_gate.py to re-propagate this block.
BLOCKED until: - [ ] tier derived from evidence (not assumed) - [ ] matrix emitted - [ ] over-engineering guard applied - [ ] advisory-only (no score/verdict mutation) confirmed
MUST ATTENTION apply critical + sequential thinking — every claim needs appropriate traced evidence (file:line for repo/code claims; source URL or artifact section for research, product, content, and docs claims); confidence >80% to act, <60% DO NOT recommend. Anti-hallucination: never present guess as fact, admit uncertainty freely, cross-reference independently, stay skeptical of own confidence.
MUST ATTENTION apply AI mistake prevention — verify generated content against evidence, trace downstream references before deleting or renaming, verify all affected outputs, re-read files after context loss, and surface ambiguity before acting.
[TASK-PLANNING] Before acting, analyze task scope and systematically break it into small todo tasks and sub-tasks using task tracking.
[IMPORTANT] Analyze how big the task is and break it into many small todo tasks systematically before starting — this is very important.
Critical Thinking Mindset — Apply critical thinking, sequential thinking. Every claim needs traced proof, confidence >80% to act.
Anti-hallucination: Never present guess as fact — cite sources for every claim, admit uncertainty freely, self-check output for errors, cross-reference independently, stay skeptical of own confidence — certainty without evidence root of all hallucination.
AI Mistake Prevention — Failure modes to avoid on every task:
Re-read files after context changes. Context compaction, resume, or long-running work can make memory stale; verify current files before acting.
Verify generated content against source evidence. AI hallucinates APIs, names, claims, and document facts. Check the relevant source before documenting or referencing.
Check downstream references before deleting or renaming. Removing an artifact can stale docs, generated mirrors, configs, and callers; map references first.
Trace the full impact chain after edits. Changing a definition can miss derived outputs and consumers. Follow the affected chain before declaring done.
Verify ALL affected outputs, not just the first. One green check is not all green checks; validate every output surface the change can affect.
Assume existing values are intentional — ask WHY before changing OR flagging one as a defect. Before changing or reporting a constant, limit, flag, cutoff, wording, or pattern, read nearby context and history, the CALLER's ordering, and 2+ sibling call sites of the same convention. A doc stating WHAT without WHY is missing rationale, not proof of a missing guard.
Surface ambiguity before acting — don't pick silently. Multiple valid interpretations require an explicit question or stated assumption with risk.
Assert the outcome your system owns, not the intermediate state your infrastructure owns. When verifying async work, assert the final business state — never the delivery/retry bookkeeping held in shared infrastructure that any co-running process can write. Such a check passes when run alone and flakes the moment anything else shares that infrastructure.
Keep shared guidance role-relevant. Universal guidance must help every receiving skill or agent; code-specific obligations belong only in code-specific protocols.
IMPORTANT MUST ATTENTION requirements come BEFORE research — load prior business/domain/PBI artifacts (Step 1), map business signals to technical requirements (Step 2), user-confirm them, THEN WebSearch (Step 3) — NEVER research before requirements are derived and confirmed — why: researching first picks tech then back-fits the problem, the reverse of architecture.
IMPORTANT MUST ATTENTION score every layer with the weighted 8-criteria matrix (High=3x / Medium=2x / Low=1x), rank with confidence %, cap the {plan-dir}/research/tech-stack-comparison.md report at <=200 lines using tables over prose — why: an unscored or unbounded report hides the trade-off the decision turns on.
**IMPORTANT MUST ATTENTION** only user-confirmed decisions get written to phase-02-tech-stack.md as status: confirmed — the end interview (5-8 ask the user directly questions) is mandatory and NEVER skipped even when the choice seems "obvious" — why: an unconfirmed stack is a guess the team will pay for.
**IMPORTANT MUST ATTENTION** every claim, finding, and recommendation requires file:line/URL proof or traced evidence + confidence % (>80% act, 60-80% verify first, <60% DO NOT recommend) — NEVER present a guess as fact — why: a stack chosen on speculation fails silently until production.
IMPORTANT MUST ATTENTION evaluate fit before copying a reference stack from another project — verify the new context shares the same scale, budget, team skills, compliance, and timeline constraints — why: the closest example rarely matches preconditions, and a mismatched copy compiles but fails the real requirements.
Anti-Rationalization:
| Evasion |
Rebuttal |
| "Stack is obvious — skip the research" |
3+ WebSearched options per layer with cited evidence anyway — familiarity is not evidence. |
| "I already know this is the best framework" |
Show the weighted 8-criteria score + confidence %. No matrix = no recommendation. |
| "Skip the user interview, the choice is clear" |
The end interview is MANDATORY — only status: confirmed decisions get written. |
| "Just research the stack, requirements are fine" |
Derive + user-confirm technical requirements FIRST (Steps 1-2), then research. |
| "One source is enough for this layer" |
Cite URL + benchmark + case study; a single anecdote is not benchmarked evidence. |
External Memory: For research/analysis work, write intermediate findings and final results to a report file in plans/reports/ — prevents context loss and serves as deliverable.
Evidence Gate: MANDATORY IMPORTANT MUST ATTENTION — every claim, finding, recommendation requires file:line/URL proof or traced evidence with confidence percentage (>80% to act, <80% must verify first).
Hookless Prompt Protocol Mirror (Auto-Synced)
Source: .claude/.ck.json + .claude/skills/shared/sync-inline-versions.md (:full blocks) + .claude/scripts/lib/hookless-prompt-protocol.cjs
[WORKFLOW-EXECUTION-PROTOCOL] [BLOCKING] Workflow Execution Protocol — MANDATORY IMPORTANT MUST CRITICAL. Do not skip for any reason.
Generic portability boundary: Reusable skills and protocol text stay project-neutral; project-specific conventions are discovered from docs/project-config.json and docs/project-reference/. Apply shared AI-SDD from shared/sdd-artifact-contract.md. Read docs/project-config.json and docs/project-reference/docs-index-reference.md, then open the project reference docs named there. For spec, test-case, behavior-change, public-contract, or docs/specs/ work, route through the local spec docs named by the docs index: feature-spec-reference.md, spec-system-reference.md, spec-principles.md, and workflow-spec-test-code-cycle-reference.md when specs/tests/code must stay synchronized. If either file or a required reference doc is missing or stale, auto-run $project-init (or the narrow lower-level route such as $project-config, $docs-init, $scan-all, or $scan --target=<key>) before ordinary project-specific work. Any supported AI tool may execute when this shared context and local docs are available.
- DETECT: If the prompt starts with an explicit slash skill/workflow command, execute it directly. Otherwise match the prompt against the workflow catalog and skill list.
- ANALYZE: Choose the best option: execute directly, invoke a skill, activate a standard workflow, or compose a custom step combination.
- AUTO-SELECT: Pick the best option yourself. Do not ask the user to choose between direct execution, skill, standard workflow, or custom workflow.
- ACTIVATE: For a selected workflow, call
$start-workflow <workflowId>; for a selected skill, invoke that skill; for a custom workflow, sequence custom steps directly; for direct execution, proceed with the task.
- CREATE TASKS: task tracking for ALL workflow/skill/custom steps before execution when the selected path has multiple steps.
- PARALLELIZE: Before executing the task list, tag each task
PAR (independent inputs + write set disjoint from every other PAR task) or SEQ (name the blocking dependency), group PAR tasks into waves, declare the wave plan, and spawn each wave's sub-agents in ONE message — all-return barrier per wave, fan-out one level deep unless a sub-agent's own definition authorizes further fan-out. Sequential-by-default is a defect when tasks are independent; do not parallelize shared write targets, output-consuming tasks, trivial single-file work, ordering a skill or workflow explicitly fixes, or user-approval gates.
- EXECUTE: Advance per the Workflow Step Advancement & Parallel Phases rule in your context instructions — model-driven; a sub-agent completion advances a step identically to an inline call; a parallel-phase group is an all-return barrier (advance only after ALL members return, never serialize it)
Shared AI-SDD Protocol Markers
Source: .claude/skills/shared/sync-inline-versions.md
SYNC:ai-sdd-artifact-contract
AI-SDD Artifact Contract — Shared spec-driven development rules stay portable and source-owned.
- Keep reusable AI-SDD principles in
.claude; put repository-specific paths, commands, owners, products, and formats in project config/reference docs.
- Preserve cycle:
spec -> plan -> tasks -> implement -> verify -> update spec/docs.
- Trace every requirement or invariant through decision, task, TC/test, source evidence, and docs/spec update.
- Treat code-to-spec extraction as reference-only until accepted by the canonical spec owner.
- Any supported AI tool may plan, implement, review, or verify with synced context; using multiple tools is optional.
- Update
.claude source first, then sync generated mirrors; do not manually edit .agents, .codex, or
…(truncated)
1---2name: tech-stack-research3description: [Architecture] Use when you need to research, analyze, and compare tech stack options as a solution architect.4---5
6> Codex compatibility note:
7>
8> - Invoke repository skills with `$skill-name` in Codex; this mirrored copy rewrites legacy Claude `/skill-name` references.
9> - Task tracker mandate: BEFORE executing any workflow or skill step, create/update task tracking for all steps and keep it synchronized as progress changes.
10> - User-question prompts mean to ask the user directly in Codex.
11> - Ignore Claude-specific mode-switch instructions when they appear.
12> - Strict execution contract: when a user explicitly invokes a skill, execute that skill protocol as written.
13> - Subagent authorization: when a skill is user-invoked or AI-detected and its protocol requires subagents, that skill activation authorizes use of the required `spawn_agent` subagent(s) for that task.
14> - Do not skip, reorder, or merge protocol steps unless the user explicitly approves the deviation first.
15> - For workflow skills, execute each listed child-skill step explicitly and report step-by-step evidence.
16> - If a required step/tool cannot run in this environment, stop and ask the user before adapting.
17
18<!-- CODEX:PROJECT-REFERENCE-LOADING:START -->
19
20## Codex Project-Reference Loading (No Hooks)
21
22Codex uses static project-reference loading instead of runtime-injected project docs.
23When coding, planning, debugging, testing, or reviewing, open project docs explicitly using this routing.
24
25**Always read:**
26
27- `docs/project-config.json` (project-specific paths, commands, modules, and workflow/test settings)
28- `docs/project-reference/docs-index-reference.md` (routes to the full `docs/project-reference/*` catalog)
29- `docs/project-reference/lessons.md` (always-on guardrails and anti-patterns)
30
31**Missing/stale context route:** If `docs/project-config.json`, the docs index, `lessons.md`, `CLAUDE.md`, `AGENTS.md`, or any task-required reference doc is missing or stale, auto-run `$project-init` or the narrow setup route (`$project-config`, `$docs-init`, `$scan-all`, `$scan --target=<key>`, `$claude-md-init`) before ordinary project-specific work. If Codex mirrors or `AGENTS.md` are missing/stale, ask the user to run `$sync-codex`; do not auto-run it.
32
33**Situation-based docs:**
34
35- Project structure/architecture/tech-stack/deployment/setup (any layer — backend, frontend, or infra): `project-structure-reference.md`
36- Backend/CQRS/API/domain/entity changes: `backend-patterns-reference.md`, `domain-entities-reference.md`
37- Frontend/UI/styling/design-system: `frontend-patterns-reference.md`, `scss-styling-guide.md`, `design-system/README.md`
38- Spec authoring, `docs/specs/` pathing, or TC format: `feature-spec-reference.md`, `spec-system-reference.md`, `spec-principles.md`
39- Behavior/public-contract changes or spec-test-code sync: `workflow-spec-test-code-cycle-reference.md` plus the spec docs above
40- Derived spec indexes/ERDs/reimplementation guides: `spec-system-reference.md` and source Feature Specs under `docs/specs/`
41- Integration test implementation/review: `integration-test-reference.md`
42- E2E test implementation/review: `e2e-test-reference.md`
43- Code review/audit work: `code-review-rules.md` plus domain docs above based on changed files
44
45Do not read all docs blindly. Start from `docs-index-reference.md`, then open only relevant files for the task.
46
47<!-- CODEX:PROJECT-REFERENCE-LOADING:END -->
48
49<!-- PROMPT-ENHANCE:STEP-TASK-ANCHOR:START -->
50
51> **[BLOCKING]** Execute skill steps in declared order. NEVER skip, reorder, or merge steps without explicit user approval.
52> **[BLOCKING]** Before each step or sub-skill call, update task tracking: set `in_progress` when step starts, set `completed` when step ends.
53> **[BLOCKING]** Every completed/skipped step MUST include brief evidence or explicit skip reason.
54> **[BLOCKING]** If Task tools are unavailable, create and maintain an equivalent step-by-step plan tracker with the same status transitions.
55
56<!-- PROMPT-ENHANCE:STEP-TASK-ANCHOR:END -->
57
58## Quick Summary
59
60**Goal:** Deliver user-confirmed, per-layer tech stack — each choice backed by 3+ researched options, weighted scoring, cited evidence, confidence % — by acting as solution architect: derive technical requirements from business analysis, research current market, produce detailed comparison report, so team commits to stack fit for scale, budget, skills, timeline, NOT familiarity.
61
62**Summary:**
63
64- **Purpose:** act as solution architect — derive technical requirements from business analysis, research current market, produce per-layer comparison report so team commits to stack fit for scale/budget/skills/timeline, NOT familiarity.
65- **All 7 main steps (run in order):** (1) Load Business Context → (2) Derive Technical Requirements + user-confirm → (3) Research Per Layer (WebSearch 3+ options each) → (4) Deep Comparison Matrix → (5) Weighted Score & Ranking → (6) Generate Report → (7) User Validation Interview.
66- Requirements BEFORE research: load prior business/domain/PBI artifacts (Step 1), map business signals → technical requirements (Step 2), gate on user confirmation (ask the user directly) before any WebSearch (Step 3).
67- Evaluate every stack layer (backend, frontend, database, messaging, infra, auth) independently — minimum 3 WebSearched options per layer, each with cited evidence (URL, benchmark, case study), NEVER familiarity (Steps 3-4).
68- Score with weighted 8-criteria matrix (High=3x / Medium=2x / Low=1x), rank each layer with confidence %; capped <=200-line report → `{plan-dir}/research/tech-stack-comparison.md` (Steps 5-6).
69- End-of-skill user validation interview (5-8 questions) mandatory, NEVER skipped — only confirmed decisions written to `phase-02-tech-stack.md` as `status: confirmed` (Step 7).
70
71**Workflow:**
72
731. **Load Business Context** — Read business evaluation, domain model, refined PBI artifacts
742. **Derive Technical Requirements** — Map business needs to technical constraints
753. **Research Per Layer** — WebSearch top 3 options for each stack responsibility
764. **Deep Compare** — Pros/cons matrix, benchmarks, community health, team fit
775. **Score & Rank** — Weighted scoring across 8 criteria
786. **Generate Report** — Structured comparison report with recommendation
797. **User Validation** — Present findings, ask 5-8 questions, confirm choices
80
81**Key Rules:**
82
83- **MANDATORY IMPORTANT MUST ATTENTION** research minimum 3 options per stack layer
84- **MANDATORY IMPORTANT MUST ATTENTION** include confidence % with evidence for every recommendation
85- **MANDATORY IMPORTANT MUST ATTENTION** run user validation interview at end (NEVER skip)
86- All claims must cite sources (URL, benchmark, case study)
87- Recommend on benchmarked evidence (URL, benchmark, case study); NEVER on familiarity alone
88
89**Be skeptical. Apply critical thinking, sequential thinking. Every claim needs traced proof, confidence percentages (Idea should be more than 80%).**
90
91## Step 1: Load Business Context
92
93Read artifacts from prior workflow steps (search `plans/`, `team-artifacts/`):
94
95- Business evaluation report (viability, scale, constraints)
96- Domain model / ERD (complexity, entity count, relationships)
97- Refined PBI (acceptance criteria, scope)
98- Discovery interview notes (team skills, budget, timeline)
99
100Extract and summarize:
101
102| Signal | Value | Source |
103| ---------------------- | ------------ | ------------------- |
104| Expected users | ... | discovery interview |
105| Domain complexity | Low/Med/High | domain model |
106| Team skills | ... | discovery interview |
107| Budget constraint | ... | business evaluation |
108| Timeline | ... | business evaluation |
109| Compliance needs | ... | business evaluation |
110| Real-time needs | Yes/No | refined PBI |
111| Integration complexity | Low/Med/High | domain model |
112
113## Step 2: Derive Technical Requirements
114
115Map business signals to technical requirements:
116
117| Business Signal | Technical Requirement | Priority |
118| ------------------ | ----------------------------------------------- | -------- |
119| High user scale | Horizontal scaling, connection pooling | Must |
120| Complex domain | Strong type system, ORM with migrations | Must |
121| Real-time features | WebSocket/SSE support, event-driven arch | Must |
122| Small team | Low learning curve, good DX, batteries-included | Should |
123| Tight budget | Open-source, low hosting cost | Should |
124| Compliance | Audit trail, encryption, auth framework | Must |
125
126**MANDATORY IMPORTANT MUST ATTENTION** validate derived requirements with user by asking the user directly before proceeding to research.
127
128## Step 3: Research Per Stack Layer
129
130For EACH layer, research top 3 options via WebSearch (minimum 5 queries total):
131
132### Stack Layers to Evaluate
133
134| Layer | Example Options | Research Focus |
135| ---------------------- | ------------------------------------- | ------------------------------------- |
136| **Backend Framework** | Candidate backend runtimes/frameworks | Performance, type safety, ecosystem |
137| **Frontend Framework** | Candidate frontend frameworks | DX, ecosystem, hiring, enterprise fit |
138| **Database** | Candidate database engines/stores | Scale, query complexity, cost |
139| **Messaging/Events** | Candidate messaging/event systems | Throughput, reliability, complexity |
140| **Infrastructure** | Docker+K8s, Serverless, PaaS | Cost, ops overhead, scaling |
141| **Auth** | Keycloak, Auth0, custom | Cost, compliance, flexibility |
142
143### WebSearch Queries (minimum 5 per layer)
144
145```
146"{option_A} vs {option_B} {current_year} comparison"
147"{option} enterprise production case studies"
148"{option} community size github stars"
149"{option} performance benchmarks {use_case}"
150"{option} security track record vulnerabilities"
151```
152
153## Step 4: Deep Comparison Matrix
154
155For EACH stack layer, produce comparison table:
156
157| Criteria | Option A | Option B | Option C | Weight |
158| -------------------- | ----------------- | -------- | -------- | ------ |
159| **Team Fit** | score + rationale | ... | ... | High |
160| **Scalability** | score + rationale | ... | ... | High |
161| **Time-to-Market** | score + rationale | ... | ... | High |
162| **Ecosystem/Libs** | score + rationale | ... | ... | Medium |
163| **Hiring Market** | score + rationale | ... | ... | Medium |
164| **Cost (hosting)** | score + rationale | ... | ... | Medium |
165| **Learning Curve** | score + rationale | ... | ... | Medium |
166| **Community Health** | score + rationale | ... | ... | Low |
167
168Scoring: 1-5 scale. Weight: High=3x, Medium=2x, Low=1x.
169
170### Per-Option Detail Block
171
172For each option, document:
173
174```markdown
175### {Layer}: {Option Name}
176
177**Pros:**
178
179- {Pro 1} — {evidence/source}
180- {Pro 2} — {evidence/source}
181- {Pro 3} — {evidence/source}
182
183**Cons:**
184
185- {Con 1} — {evidence/source}
186- {Con 2} — {evidence/source}
187
188**Best suited when:** {conditions}
189**Not suitable when:** {conditions}
190**Production examples:** {2-3 real companies using this}
191```
192
193## Step 5: Weighted Score & Ranking
194
195Calculate weighted total per option per layer. Present ranking:
196
197```markdown
198### {Layer} Ranking
199
2001. **{Option A}** — Score: {X}/100 — Confidence: {Y}%
2012. **{Option B}** — Score: {X}/100 — Confidence: {Y}%
2023. **{Option C}** — Score: {X}/100 — Confidence: {Y}%
203
204**Recommendation:** {Option A}
205**Why:** {2-3 sentence rationale linking to team skills, scale, and constraints}
206```
207
208## Step 6: Generate Report
209
210Write report to `{plan-dir}/research/tech-stack-comparison.md` with:
211
2121. Executive summary (recommended full stack in 5 lines)
2132. Technical requirements table (from Step 2)
2143. Per-layer comparison matrices (from Step 4)
2154. Per-layer rankings with recommendations (from Step 5)
2165. Combined recommended stack diagram
2176. Risk assessment for recommended stack
2187. Alternative stack (second-best combo) for comparison
2198. Unresolved questions
220
221Report must be **<=200 lines**. Use tables over prose.
222
223## Step 7: User Validation Interview
224
225**MANDATORY IMPORTANT MUST ATTENTION** present findings and ask 5-8 questions by asking the user directly:
226
227### Required Questions
228
2291. **Per-layer recommendation confirmation** — "For {layer}, I recommend {option}. Agree?"
230 - Options: Agree (Recommended) | Prefer {option B} | Need more research
2312. **Risk tolerance** — "The recommended stack has {risk}. Acceptable?"
2323. **Team readiness** — "Team needs to learn {X}. Training plan needed?"
2334. **Budget alignment** — "Estimated infra cost: ${X}/month. Within budget?"
2345. **Timeline fit** — "This stack enables MVP in {X} months. Acceptable?"
235
236### Optional Deep-Dive Questions (pick 2-3 based on context)
237
238- "Should we consider {emerging tech} for {layer}?"
239- "Any compliance requirements I haven't captured?"
240- "Preference for managed services vs self-hosted?"
241- "Monorepo or polyrepo for this team size?"
242
243After user confirms, update report with final decisions, mark `status: confirmed`.
244
245## Output
246
247```
248{plan-dir}/research/tech-stack-comparison.md # Full comparison report
249{plan-dir}/phase-02-tech-stack.md # Final confirmed tech stack decisions
250```
251
252---
253
254**MANDATORY IMPORTANT MUST ATTENTION** break work into small todo tasks using task tracking BEFORE starting.
255**MANDATORY IMPORTANT MUST ATTENTION** validate EVERY recommendation with user by asking the user directly — NEVER auto-decide.
256**MANDATORY IMPORTANT MUST ATTENTION** include confidence % and evidence citations for all claims.
257**MANDATORY IMPORTANT MUST ATTENTION** add a final review todo task to verify work quality.
258
259---
260
261## Next Steps
262
263**MANDATORY IMPORTANT MUST ATTENTION — NO EXCEPTIONS** after completing this skill, you MUST ATTENTION use ask the user directly to present these options. Do NOT skip because task seems "simple"/"obvious" — the user decides:
264
265- **"$architecture-design (Recommended)"** — Design solution architecture with chosen tech stack
266- **"$plan"** — If architecture already decided
267- **"Skip, continue manually"** — user decides
268
269### Council escalation (always-offer, second prompt)
270
271After the existing `## Next Steps` prompt above resolves, present a **second**, independent ask the user directly call:
272
273- **"Skip council — proceed with chosen stack (Recommended)"** — Continue with the selected tech stack as-is.
274- **"Escalate to $llm-council"** — Run 11 sub-agent council. Best applied when 2+ stacks score within 15% on the comparison matrix or you have unfamiliar/strategic dependencies. Cheaper alternatives: `$why-review`, `$plan-validate`.
275
276<!-- SYNC:scenario-stress-eval -->
277
278> **Scenario Stress & Resilience Evaluation** — CONDITIONAL, evidence-gated, business-criticality-aware. The top-down companion to `SYNC:scale-technique-gate`: instead of _"is technique X present?"_, put the system UNDER concrete failure/load scenarios and judge whether it SURVIVES, SELF-HEALS, and whether its BUSINESS needs it to. **ADVICE-ONLY: emit the Scenario Stress Matrix as guidance; NEVER mutate any score, verdict band, or gate pass/fail.**
279>
280> 1. **Reuse the scale tier** derived by `SYNC:scale-technique-gate` (or derive it identically from evidence); **also derive business-criticality `B0`–`B3`** from specs/SLA/product docs + the domain, cite `file:line` + confidence. `B0` best-effort · `B1` important · `B2` business-critical · `B3` mission-critical/regulated. Unknown → state the assumption, do **NOT** default to `B3`/`T3`. **Criticality-signal floor (both-directions safety):** regulated / PII / financial / health data, money movement, auth/identity, or legal-compliance scope raises `B` to **at least `B2` even absent SLA/SLO docs**; anti-over-engineering lowers hardening ONLY when NO such signal is present. `B` (blast if it fails) and `T` (scale of load/data) are independent — a low-traffic payroll run is low-`T`, high-`B`.
281> 2. **Select in-scope scenarios** — only those the system's `B`/`T` combination warrants (a `B0` internal PoC skips region-loss/DR entirely; a `B3`/`T0` regulated service still needs backups + DR by BUSINESS, not scale).
282> 3. **Walk each in-scope scenario:** simulate the stimulus → trace the break path → name the failure signature → answer the self-heal/recovery question (auto-recover? MTTR? manual runbook?) → name the trade-off it forces. Families: traffic spike · sustained growth · data-volume growth · write/ingest burst · dependency down/slow · instance/node loss · zone/region loss · **data loss/corruption** · poison-message/retry-storm · cascading failure/backpressure · cold-start/deploy-blip · clock-skew/duplicate-delivery.
283> 4. **Assign one verdict per scenario:** `WITHSTANDS` · `DEGRADES-GRACEFULLY` · `FAILS-HARD` (→ **advise only**) · `N/A-by-business` (not warranted → skip, not a gap) · `OVER-HARDENED` (resilience beyond business need → **advise AGAINST**, cite carrying cost).
284> 5. **Anti-over-engineering guard (first-class):** a lean system whose business does not need HA/DR is a PASS; `OVER-HARDENED` flags resilience the business does not warrant. This guard is symmetric with the criticality-signal floor above — never under-harden a `B2`+ system just because its traffic is low.
285> 6. **Output — Scenario Stress Matrix:** `scenario | in-scope (B/T)? | verdict | self-heal | trade-off | evidence (file:line/config/infra)`. Full catalog + Business×Scale in-scope baseline + verdict/tier tables → `.claude/docs/scenario-stress-catalog.md`. **ADVISORY-ONLY: NEVER mutate any `/20`, `/24`, verdict band, or gate pass/fail. Drift-guard: scenarios/verdicts/business-tiers are AUTHORITATIVE in the catalog — update it FIRST, then re-run `.claude/scripts/inject_scenario_stress_gate.py`. Scale tier stays single-sourced in `scale-technique-catalog.md`.**
286>
287> **BLOCKED until:** `- [ ]` scale tier + business-criticality (with criticality-signal floor) derived from evidence `- [ ]` in-scope scenarios selected `- [ ]` matrix emitted `- [ ]` over-hardening guard applied `- [ ]` advisory-only (no score/verdict mutation) confirmed
288
289<!-- /SYNC:scenario-stress-eval -->
290
291<!-- SYNC:scale-technique-gate:reminder -->
292
293**IMPORTANT MUST ATTENTION** scale-technique gate: derive the scale tier from evidence FIRST (T0 internal · T1 <10k · T2 10k–1M · T3 millions+), then judge each warranted technique `PRESENT`/`MISSING-WARRANTED`/`N/A-by-scale`/`OVER-ENGINEERED`. Advise on warranted-but-missing gaps AND advise AGAINST unwarranted heavyweight techniques (anti-over-engineering). **ADVICE-ONLY — emit the Technique Applicability Matrix as guidance; NEVER mutate any score, verdict band, or gate pass/fail.** Full catalog → `.claude/docs/scale-technique-catalog.md` (authoritative for tier thresholds & per-technique warranting tiers — on any change update the catalog FIRST, then re-run `inject_scale_technique_gate.py`).
294
295<!-- /SYNC:scale-technique-gate:reminder -->
296
297<!-- SYNC:scenario-stress-eval:reminder -->
298
299**IMPORTANT MUST ATTENTION** scenario-stress gate: reuse the scale tier `T0`–`T3` AND derive business-criticality `B0`–`B3` from evidence first — apply the **criticality-signal floor** (regulated/PII/financial/health data · money movement · auth/identity · legal-compliance → at least `B2` even absent SLA docs; do NOT default to `B3`). Select only the scenarios the `B`/`T` combination warrants, then walk each (simulate → trace → failure signature → self-heal/MTTR → trade-off) and assign `WITHSTANDS`/`DEGRADES-GRACEFULLY`/`FAILS-HARD`/`N/A-by-business`/`OVER-HARDENED`. Anti-over-engineering is first-class (a lean system that needs no HA/DR is a PASS) AND symmetric (never under-harden a `B2`+ system for low traffic). **ADVICE-ONLY — emit the Scenario Stress Matrix as guidance; NEVER mutate any score, verdict band, or gate pass/fail.** Full catalog → `.claude/docs/scenario-stress-catalog.md` (authoritative for scenarios/verdicts/business-tiers — on any change update the catalog FIRST, then re-run `inject_scenario_stress_gate.py`; scale tier stays single-sourced in `scale-technique-catalog.md`).
300
301<!-- /SYNC:scenario-stress-eval:reminder -->
302
303<!-- PROMPT-ENHANCE:STEP-TASK-CLOSING:START -->
304
305## Prompt-Enhance Closing Anchors
306
307- **IMPORTANT MUST ATTENTION** follow declared step order for this skill; NEVER skip, reorder, or merge steps without explicit user approval
308- **IMPORTANT MUST ATTENTION** for every step/sub-skill call: set `in_progress` before execution, set `completed` after execution
309- **IMPORTANT MUST ATTENTION** every skipped step MUST include explicit reason; every completed step MUST include concise evidence
310- **IMPORTANT MUST ATTENTION** if Task tools unavailable, maintain an equivalent step-by-step plan tracker with synchronized statuses
311
312<!-- PROMPT-ENHANCE:STEP-TASK-CLOSING:END -->
313
314<!-- SYNC:project-protocol-overlay -->
315
316> **Project Protocol Overlay** — Before executing this skill, resolve any PROJECT overlay rules layered onto it: match this skill's name against the `Target` column of the project's skill-protocol index (`docs/project-reference/skill-protocols-reference.md` by default; a `referenceDocs` entry in `docs/project-config.json` overrides the path), taking the most specific matching tier ONLY — exact name > glob > `*`. **That precedence orders overlays against EACH OTHER, never against this skill.** Read ONLY the matched bodies, resolved as `<protocols-dir>/<Name>.md`; a row's Body link is display text, never a read path. A matched body that is missing or malformed is REPORTED and skipped — never reconstructed from the index Description. No index, or no match -> proceed with no overlay, silently. Full contract: `.claude/skills/project-skill-protocol/references/registry.md`.
317>
318> Overlays are **ADDITIVE ONLY**: they ADD rules on top of this skill's own protocol and NEVER replace, override, disable, or reinterpret a rule it already states — removing every overlay must return this skill to exactly its documented behavior. An overlay is a BRIEF, not an authority escalation: it can NEVER waive a workflow gate, git discipline, a review gate, or a user-confirmation gate. A genuine overlay-vs-skill conflict, or two equally-specific overlays that directly contradict -> surface both to the user; NEVER resolve silently.
319
320<!-- /SYNC:project-protocol-overlay -->
321
322<!-- SYNC:project-protocol-overlay:reminder -->
323
324**MUST ATTENTION** resolve project protocol overlays for this skill BEFORE executing — most specific matching tier only (exact > glob > `*`, which ranks overlays against each other, NEVER against this skill), read only matched bodies at `<protocols-dir>/<Name>.md`; a missing or malformed body is reported, never reconstructed. Overlays are ADDITIVE ONLY (they never replace this skill's own rules) and are a brief, NEVER an authority escalation; an equal-specificity contradiction goes to the user.
325
326<!-- /SYNC:project-protocol-overlay:reminder -->
327
328## Closing Reminders
329
330**IMPORTANT MUST ATTENTION Goal:** deliver user-confirmed, per-layer tech stack — each choice backed by 3+ researched options, weighted 8-criteria scoring, cited evidence, confidence % — so team commits to a stack fit for scale, budget, skills, timeline, NOT familiarity.
331
332**IMPORTANT MUST ATTENTION — run ALL 7 steps in declared order, none skipped:** (1) Load Business Context → (2) Derive Technical Requirements (+ ask the user directly confirm) → (3) Research Per Layer (WebSearch 3+ options each) → (4) Deep Comparison Matrix → (5) Weighted Score & Ranking (confidence %) → (6) Generate Report (<=200 lines) → (7) User Validation Interview (5-8 questions, write `status: confirmed`) — why: AI keeps collapsing this into "just pick a stack" and dropping requirements-derivation, scoring, and the confirmation gate that make the choice defensible.
333
334**Protocols in force (concise digest of the SYNC/shared blocks this skill carries):**
335
336- **Critical Thinking:** MUST ATTENTION apply critical + sequential thinking; traced proof, confidence >80% to act, NEVER guess as fact.
337- **AI Mistake Prevention:** verify generated content against evidence, trace downstream references, verify all affected outputs, re-read after context loss, surface ambiguity.
338
339**IMPORTANT MUST ATTENTION** research minimum 3 WebSearched options per stack layer (backend, frontend, database, messaging, infra, auth); every recommendation carries confidence % + cited evidence (URL, benchmark, case study) — NEVER recommend on familiarity alone — why: familiarity bias commits the team to the wrong stack that surfaces only at scale.
340**IMPORTANT MUST ATTENTION** gate on user by asking the user directly at EVERY decision point — confirm derived requirements before research (Step 2), confirm each layer recommendation in the end interview (Step 7) — NEVER auto-decide — why: the team owns the stack, not the AI.
341**MANDATORY IMPORTANT MUST ATTENTION** break work into small todo tasks using task tracking BEFORE starting; mark one `in_progress`, `completed` immediately after evidence; add a final review todo.
342
343<!-- SYNC:scale-technique-gate -->
344
345> **Scalability & Production-Readiness Technique Gate** — CONDITIONAL, evidence-gated, scale-tiered. Judge which system-design techniques a system _warrants_ at its scale — flag warranted-but-missing gaps AND advise AGAINST unwarranted heavyweight ones. **ADVICE-ONLY: emit the matrix as guidance; NEVER mutate any score, verdict band, or gate pass/fail.**
346>
347> 1. **Derive the scale tier FIRST — from evidence, never assumed.** Read users/RPS, SLO/latency targets, data volume, tenancy, topology from config/infra/specs; cite `file:line` + confidence. Tiers: `T0` internal/single-instance · `T1` small SaaS (<10k users) · `T2` high-scale (10k–1M) · `T3` massive/multi-region (millions+). Unknown tier → state assumption, do NOT default to T3.
348> 2. **Judge each concern group only at/above its warranting tier** (member techniques → owning review skill for depth):
349> - Traffic & Edge — Rate Limiting, Load Balancing, Reverse Proxy, API Gateway, CDN, Edge Caching, WAF, DDoS (T1+; CDN/WAF T2+) → security-review owns WAF/DDoS
350> - Caching & Data Access — Caching, Cache Invalidation, DB Indexing, Query Optimization, N+1, Connection Pooling (T1+) → performance-review owns depth
351> - Data Scaling & Consistency — Read Replicas, Sharding, Partitioning, Replication, CAP, Eventual Consistency, Locks, Leader Election (T2+; sharding/multi-region T3) → performance-review
352> - Async & Messaging — Message Queues, Pub/Sub, Event-Driven, Saga, DLQ, Distributed Transactions, Backpressure, Webhooks, WebSockets/SSE (T2+)
353> - Resilience — Circuit Breakers, Timeouts, Retries, Backoff, Idempotency, Health Checks, Liveness/Readiness, Failover, Graceful Degradation (T1+) → production-readiness-review
354> - Scaling & Compute — Autoscaling, Horizontal/Vertical Scaling, Serverless Limits, Cold Starts, Cron Jobs, Thread Safety, GC/Memory Leaks (T1+; autoscaling T2+)
355> - Deployment & Release — CI/CD, Docker, Kubernetes, Blue-Green/Canary/Rolling, Rollbacks, Feature Flags, IaC/Terraform/Helm, Build Caching (CI/CD T0+; K8s/canary T2+)
356> - Observability — Monitoring, Logging, Distributed Tracing, Metrics, Alerting, SLOs/SLIs, Error Budgets (T1+; tracing/error-budgets T2+) → production-readiness-review
357> - Security & Compliance — Secrets Management, IAM, OAuth, JWT Rotation, TLS, Encryption at Rest/Transit, CORS, CSRF, SQLi, XSS, SSRF (T0+) → security-review owns
358> - DR & Infra — Backups, Disaster Recovery, Multi-Region, Chaos Engineering, Schema Versioning, DB Migrations, Cost Optimization (backups T1+; DR/multi-region/chaos T3) → production-readiness-review
359> 3. **Assign one of 4 verdicts per warranted technique:** `PRESENT` · `MISSING-WARRANTED` (→ **advise only** — guidance, NOT a score/gate lever) · `N/A-by-scale` (below warranting tier) · `OVER-ENGINEERED` (present but unwarranted at this tier → advise AGAINST).
360> 4. **Anti-over-engineering guard (first-class):** do NOT recommend K8s, sharding, multi-region, service mesh, event sourcing, or distributed transactions below their warranting tier. A correctly-lean small system is a PASS, never a gap.
361> 5. **Output — Technique Applicability Matrix:** `technique | tier-warranted? | present? | verdict | advice | evidence (file:line/config/infra)`. Full grouped catalog + per-tier baseline → `.claude/docs/scale-technique-catalog.md`. Hosting reviews surface this matrix WITHOUT changing any `/20`, `/24`, verdict band, or PASS/FAIL (per user decision 2026-07-06). **Drift-guard: tier thresholds & per-technique warranting tiers are AUTHORITATIVE in `.claude/docs/scale-technique-catalog.md` — the inline tier summary above is a condensed pointer; on any tier/technique change, update the catalog FIRST, then re-run `.claude/scripts/inject_scale_technique_gate.py` to re-propagate this block.**
362>
363> **BLOCKED until:** `- [ ]` tier derived from evidence (not assumed) `- [ ]` matrix emitted `- [ ]` over-engineering guard applied `- [ ]` advisory-only (no score/verdict mutation) confirmed
364
365<!-- /SYNC:scale-technique-gate -->
366
367<!-- SYNC:critical-thinking-mindset:reminder -->
368
369**MUST ATTENTION** apply critical + sequential thinking — every claim needs appropriate traced evidence (`file:line` for repo/code claims; source URL or artifact section for research, product, content, and docs claims); confidence >80% to act, <60% DO NOT recommend. Anti-hallucination: never present guess as fact, admit uncertainty freely, cross-reference independently, stay skeptical of own confidence.
370
371<!-- /SYNC:critical-thinking-mindset:reminder -->
372<!-- SYNC:ai-mistake-prevention:reminder -->
373
374**MUST ATTENTION** apply AI mistake prevention — verify generated content against evidence, trace downstream references before deleting or renaming, verify all affected outputs, re-read files after context loss, and surface ambiguity before acting.
375
376<!-- /SYNC:ai-mistake-prevention:reminder -->
377
378**[TASK-PLANNING]** Before acting, analyze task scope and systematically break it into small todo tasks and sub-tasks using task tracking.
379
380> **[IMPORTANT]** Analyze how big the task is and break it into many small todo tasks systematically before starting — this is very important.
381
382<!-- SYNC:critical-thinking-mindset -->
383
384> **Critical Thinking Mindset** — Apply critical thinking, sequential thinking. Every claim needs traced proof, confidence >80% to act.
385> **Anti-hallucination:** Never present guess as fact — cite sources for every claim, admit uncertainty freely, self-check output for errors, cross-reference independently, stay skeptical of own confidence — certainty without evidence root of all hallucination.
386
387<!-- /SYNC:critical-thinking-mindset -->
388
389<!-- SYNC:ai-mistake-prevention -->
390
391> **AI Mistake Prevention** — Failure modes to avoid on every task:
392>
393> **Re-read files after context changes.** Context compaction, resume, or long-running work can make memory stale; verify current files before acting.
394> **Verify generated content against source evidence.** AI hallucinates APIs, names, claims, and document facts. Check the relevant source before documenting or referencing.
395> **Check downstream references before deleting or renaming.** Removing an artifact can stale docs, generated mirrors, configs, and callers; map references first.
396> **Trace the full impact chain after edits.** Changing a definition can miss derived outputs and consumers. Follow the affected chain before declaring done.
397> **Verify ALL affected outputs, not just the first.** One green check is not all green checks; validate every output surface the change can affect.
398> **Assume existing values are intentional — ask WHY before changing OR flagging one as a defect.** Before changing or reporting a constant, limit, flag, cutoff, wording, or pattern, read nearby context and history, the CALLER's ordering, and 2+ sibling call sites of the same convention. A doc stating WHAT without WHY is missing rationale, not proof of a missing guard.
399> **Surface ambiguity before acting — don't pick silently.** Multiple valid interpretations require an explicit question or stated assumption with risk.
400> **Assert the outcome your system owns, not the intermediate state your infrastructure owns.** When verifying async work, assert the final business state — never the delivery/retry bookkeeping held in shared infrastructure that any co-running process can write. Such a check passes when run alone and flakes the moment anything else shares that infrastructure.
401> **Keep shared guidance role-relevant.** Universal guidance must help every receiving skill or agent; code-specific obligations belong only in code-specific protocols.
402
403<!-- /SYNC:ai-mistake-prevention -->
404
405**IMPORTANT MUST ATTENTION** requirements come BEFORE research — load prior business/domain/PBI artifacts (Step 1), map business signals to technical requirements (Step 2), user-confirm them, THEN WebSearch (Step 3) — NEVER research before requirements are derived and confirmed — why: researching first picks tech then back-fits the problem, the reverse of architecture.
406**IMPORTANT MUST ATTENTION** score every layer with the weighted 8-criteria matrix (High=3x / Medium=2x / Low=1x), rank with confidence %, cap the `{plan-dir}/research/tech-stack-comparison.md` report at <=200 lines using tables over prose — why: an unscored or unbounded report hides the trade-off the decision turns on.
407**IMPORTANT MUST ATTENTION** only user-confirmed decisions get written to `phase-02-tech-stack.md` as `status: confirmed` — the end interview (5-8 ask the user directly questions) is mandatory and NEVER skipped even when the choice seems "obvious" — why: an unconfirmed stack is a guess the team will pay for.
408**IMPORTANT MUST ATTENTION** every claim, finding, and recommendation requires `file:line`/URL proof or traced evidence + confidence % (>80% act, 60-80% verify first, <60% DO NOT recommend) — NEVER present a guess as fact — why: a stack chosen on speculation fails silently until production.
409**IMPORTANT MUST ATTENTION** evaluate fit before copying a reference stack from another project — verify the new context shares the same scale, budget, team skills, compliance, and timeline constraints — why: the closest example rarely matches preconditions, and a mismatched copy compiles but fails the real requirements.
410
411**Anti-Rationalization:**
412
413| Evasion | Rebuttal |
414| ------------------------------------------------ | ------------------------------------------------------------------------------------------ |
415| "Stack is obvious — skip the research" | 3+ WebSearched options per layer with cited evidence anyway — familiarity is not evidence. |
416| "I already know this is the best framework" | Show the weighted 8-criteria score + confidence %. No matrix = no recommendation. |
417| "Skip the user interview, the choice is clear" | The end interview is MANDATORY — only `status: confirmed` decisions get written. |
418| "Just research the stack, requirements are fine" | Derive + user-confirm technical requirements FIRST (Steps 1-2), then research. |
419| "One source is enough for this layer" | Cite URL + benchmark + case study; a single anecdote is not benchmarked evidence. |
420
421> **External Memory:** For research/analysis work, write intermediate findings and final results to a report file in `plans/reports/` — prevents context loss and serves as deliverable.
422
423> **Evidence Gate:** MANDATORY IMPORTANT MUST ATTENTION — every claim, finding, recommendation requires `file:line`/URL proof or traced evidence with confidence percentage (>80% to act, <80% must verify first).
424
425<!-- CODEX:SYNC-PROMPT-PROTOCOLS:START -->
426
427## Hookless Prompt Protocol Mirror (Auto-Synced)
428
429Source: `.claude/.ck.json` + `.claude/skills/shared/sync-inline-versions.md` (`:full` blocks) + `.claude/scripts/lib/hookless-prompt-protocol.cjs`
430
431## [WORKFLOW-EXECUTION-PROTOCOL] [BLOCKING] Workflow Execution Protocol — MANDATORY IMPORTANT MUST CRITICAL. Do not skip for any reason.
432
433**Generic portability boundary:** Reusable skills and protocol text stay project-neutral; project-specific conventions are discovered from docs/project-config.json and docs/project-reference/. Apply shared AI-SDD from `shared/sdd-artifact-contract.md`. Read `docs/project-config.json` and `docs/project-reference/docs-index-reference.md`, then open the project reference docs named there. For spec, test-case, behavior-change, public-contract, or `docs/specs/` work, route through the local spec docs named by the docs index: `feature-spec-reference.md`, `spec-system-reference.md`, `spec-principles.md`, and `workflow-spec-test-code-cycle-reference.md` when specs/tests/code must stay synchronized. If either file or a required reference doc is missing or stale, auto-run `$project-init` (or the narrow lower-level route such as `$project-config`, `$docs-init`, `$scan-all`, or `$scan --target=<key>`) before ordinary project-specific work. Any supported AI tool may execute when this shared context and local docs are available.
434
4351. **DETECT:** If the prompt starts with an explicit slash skill/workflow command, execute it directly. Otherwise match the prompt against the workflow catalog and skill list.
4362. **ANALYZE:** Choose the best option: execute directly, invoke a skill, activate a standard workflow, or compose a custom step combination.
4373. **AUTO-SELECT:** Pick the best option yourself. Do not ask the user to choose between direct execution, skill, standard workflow, or custom workflow.
4384. **ACTIVATE:** For a selected workflow, call `$start-workflow <workflowId>`; for a selected skill, invoke that skill; for a custom workflow, sequence custom steps directly; for direct execution, proceed with the task.
4395. **CREATE TASKS:** task tracking for ALL workflow/skill/custom steps before execution when the selected path has multiple steps.
4406. **PARALLELIZE:** Before executing the task list, tag each task `PAR` (independent inputs + write set disjoint from every other `PAR` task) or `SEQ` (name the blocking dependency), group `PAR` tasks into waves, declare the wave plan, and spawn each wave's sub-agents in ONE message — all-return barrier per wave, fan-out one level deep unless a sub-agent's own definition authorizes further fan-out. Sequential-by-default is a defect when tasks are independent; do not parallelize shared write targets, output-consuming tasks, trivial single-file work, ordering a skill or workflow explicitly fixes, or user-approval gates.
4417. **EXECUTE:** Advance per the **Workflow Step Advancement & Parallel Phases** rule in your context instructions — model-driven; a sub-agent completion advances a step identically to an inline call; a parallel-phase group is an all-return barrier (advance only after ALL members return, never serialize it)
442
443## Shared AI-SDD Protocol Markers
444
445Source: `.claude/skills/shared/sync-inline-versions.md`
446
447## SYNC:ai-sdd-artifact-contract
448
449> **AI-SDD Artifact Contract** — Shared spec-driven development rules stay portable and source-owned.
450>
451> 1. Keep reusable AI-SDD principles in `.claude`; put repository-specific paths, commands, owners, products, and formats in project config/reference docs.
452> 2. Preserve cycle: `spec -> plan -> tasks -> implement -> verify -> update spec/docs`.
453> 3. Trace every requirement or invariant through decision, task, TC/test, source evidence, and docs/spec update.
454> 4. Treat code-to-spec extraction as reference-only until accepted by the canonical spec owner.
455> 5. Any supported AI tool may plan, implement, review, or verify with synced context; using multiple tools is optional.
456> 6. Update `.claude` source first, then sync generated mirrors; do not manually edit `.agents`, `.codex`, or
457
458…(truncated)