Codex compatibility note:
- Invoke repository skills with
$skill-name in Codex; this mirrored copy rewrites legacy Claude /skill-name references.
- Task tracker mandate: BEFORE executing any workflow or skill step, create/update task tracking for all steps and keep it synchronized as progress changes.
- User-question prompts mean to ask the user directly in Codex.
- Ignore Claude-specific mode-switch instructions when they appear.
- Strict execution contract: when a user explicitly invokes a skill, execute that skill protocol as written.
- Subagent authorization: when a skill is user-invoked or AI-detected and its protocol requires subagents, that skill activation authorizes use of the required
spawn_agent subagent(s) for that task.
- Do not skip, reorder, or merge protocol steps unless the user explicitly approves the deviation first.
- For workflow skills, execute each listed child-skill step explicitly and report step-by-step evidence.
- If a required step/tool cannot run in this environment, stop and ask the user before adapting.
Codex Project-Reference Loading (No Hooks)
Codex uses static project-reference loading instead of runtime-injected project docs.
When coding, planning, debugging, testing, or reviewing, open project docs explicitly using this routing.
Always read:
docs/project-config.json (project-specific paths, commands, modules, and workflow/test settings)
docs/project-reference/docs-index-reference.md (routes to the full docs/project-reference/* catalog)
docs/project-reference/lessons.md (always-on guardrails and anti-patterns)
Missing/stale context route: If docs/project-config.json, the docs index, lessons.md, CLAUDE.md, AGENTS.md, or any task-required reference doc is missing or stale, auto-run $project-init or the narrow setup route ($project-config, $docs-init, $scan-all, $scan --target=<key>, $claude-md-init) before ordinary project-specific work. If Codex mirrors or AGENTS.md are missing/stale, ask the user to run $sync-codex; do not auto-run it.
Situation-based docs:
- Project structure/architecture/tech-stack/deployment/setup (any layer — backend, frontend, or infra):
project-structure-reference.md
- Backend/CQRS/API/domain/entity changes:
backend-patterns-reference.md, domain-entities-reference.md
- Frontend/UI/styling/design-system:
frontend-patterns-reference.md, scss-styling-guide.md, design-system/README.md
- Spec authoring,
docs/specs/ pathing, or TC format: feature-spec-reference.md, spec-system-reference.md, spec-principles.md
- Behavior/public-contract changes or spec-test-code sync:
workflow-spec-test-code-cycle-reference.md plus the spec docs above
- Derived spec indexes/ERDs/reimplementation guides:
spec-system-reference.md and source Feature Specs under docs/specs/
- Integration test implementation/review:
integration-test-reference.md
- E2E test implementation/review:
e2e-test-reference.md
- Code review/audit work:
code-review-rules.md plus domain docs above based on changed files
Do not read all docs blindly. Start from docs-index-reference.md, then open only relevant files for the task.
[BLOCKING] Execute skill steps in declared order. NEVER skip, reorder, or merge steps without explicit user approval.
[BLOCKING] Before each step or sub-skill call, update task tracking: set in_progress when step starts, set completed when step ends.
[BLOCKING] Every completed/skipped step MUST include brief evidence or explicit skip reason.
[BLOCKING] If Task tools are unavailable, create and maintain an equivalent step-by-step plan tracker with the same status transitions.
Quick Summary
Goal: Stand up configurable, local-dev-only test-data seeders — enabled by default ONLY on local/dev — that auto-seed each feature's happy-path scenarios by calling the same public entry-point application commands a real user / QC tester would (NEVER direct DB writes), so the system self-tests its main cases like a QC engineer exercising them by hand; a configurable seed count repeats each scenario to BOTH cover the main cases AND enrich data volume (many simulated users) for performance testing + realistic first-time-init data, idempotent + restart-safe (never re-seeds already-seeded data; resumes from the last count toward target X, at 50% → continue until X), defaulting the count small when nothing is configured — and ALWAYS finding the project's existing seed-data convention FIRST.
Summary:
- Find the existing convention FIRST. Before designing anything, discover the project's seeder base class, env-gate key, count config key, and registration with
file:line evidence (Step 1) — match it exactly; never invent a parallel mechanism.
- Seeders orchestrate the real app pipeline like a real user: invoke the public entry-point application commands (which own validation, domain logic, and event side-effects) — never repo/DB inserts for domain entities, never duplicate command logic in the seeder.
- Dual purpose, one mechanism — a configurable count: repeat each happy-path scenario N times to (a) self-test the main cases (QC mimic) and (b) enrich data volume for many-users / performance / first-init realism. Read the count from config (never hardcode); default small when unset; zero → no-op.
- Four non-negotiable gates in order: (1) environment gate as the FIRST check (local-dev/enabled-config only), (2) count-before-seed idempotency (no re-seed when already seeded), (3) restart-safe loop from
existing_count to target_count (never 0 — resume the remainder after stop/restart), (4) scoped DI per iteration — a shared scope silently corrupts the DbContext/session.
- Always pre-read
docs/project-reference/seed-test-data-reference.md + project-config Data Seeders group, then close with a fresh zero-memory code-reviewer round; re-review fully only after a validated fix.
- Two modes — surface the flag: default Generate (implement / enhance / fix a seeder);
--mode=review = READ-ONLY convention audit grading a target (prompt → current changes → work-context) against EVERY universal rule + project conventions with file:line PASS/FAIL — routes confirmed fixes back to Generate, NEVER edits the seeder itself.
- Main steps to run (Generate, in order — do not skip): Phase 0 detect task type (new/enhance/fix) → Step 1 discover conventions (base class, env-gate key, count key, registration) → Step 1.5 verify dev-config keys exist → Step 2 feature scope + application commands → Step 3 find/create seeder → Step 4 implement (env-gate FIRST → config count → idempotency → restart-safe loop → scoped DI) → Step 5 validate every gate with
file:line → Step 7 --mode=review self-audit on the changed code → fresh code-reviewer round → $changes-review (final).
Workflow (Generate mode — default):
- Phase 0 — Detect seeder task type (new / enhance / fix)
- Step 1 — Discover project seeder patterns, env gate key, count key
- Step 2 — Analyze feature scope + application commands
- Step 3 — Find or create seeder file
- Step 4 — Implement using language-agnostic algorithm
- Step 5 — Validate against universal rules
- Self-Review — Re-run THIS skill in
--mode=review over the changed seeder code (convention gate)
- Review — Fresh sub-agent review round, then hand off to
$changes-review
Modes:
- Default (generate) — implement / enhance / fix seeders. Everything in the Generate-mode Protocol below applies. The generate-mode task plan MUST end by re-running this skill in
--mode=review (Step 7) BEFORE the $changes-review hand-off.
--mode=review (read-only convention audit) — review a target against EVERY universal seed-data rule AND the project-specific seeder conventions, with file:line evidence and a PASS/FAIL verdict. Makes NO code changes; reports findings and routes confirmed defects back to generate mode for the fix. See Mode: Review.
Key Rules:
- ALWAYS find the project's existing seed-data convention FIRST — read
docs/project-reference/seed-test-data-reference.md and docs/project-config.json (Data Seeders context group) before writing any seeder changes; match the discovered pattern, never invent a new one
- ENABLE seeding by default ONLY on a local/development environment — the environment gate is the FIRST check, NEVER production
- SEED like a real user / QC tester — call the public entry-point application commands; NEVER call repository/DB directly for domain data
- NEVER duplicate command logic — seeder orchestrates, commands own validation
- ALWAYS make the seed count configurable and read it from config (NEVER hardcode); default to a small number when nothing is configured; zero → no-op
- A configurable count serves BOTH goals — self-test the main happy-path cases AND enrich data volume (many simulated users) for performance testing and realistic first-time-init data
- GUARANTEE idempotency — check count before seeding; never re-seed already-seeded data on restart
- ALWAYS loop from
existing_count to target_count so stop/restart resumes the remainder (target X, at 50% → continue until X), never re-seeding from 0
- SEED only states the application could actually produce — application-level operations guarantee reachability by construction; a direct store write fabricating an otherwise-unreachable state MUST be commented with why it is legitimate, and seeded entities MUST carry plausible relative timing rather than one shared instant
Mode Routing (FIRST decision)
Before Phase 0, route on the invocation flag:
| Signal |
Mode |
Go to |
--mode=review flag, OR prompt asks to review/audit/check a seeder |
Review |
Mode: Review |
| Any other invocation (implement / enhance / fix a seeder) |
Generate (default) |
Phase 0 below |
MUST ATTENTION Generate mode OWNS the fix; Review mode is READ-ONLY and only reports. When Review mode finds a defect, it routes the fix back through Generate mode — it never edits the seeder itself.
Phase 0: Detect Seeder Task Type (Generate mode)
Before any other step, classify the request:
| Task Type |
Detection |
Action |
| New seeder |
No existing seeder for feature area |
Create following discovered base class pattern |
| Enhance existing |
Seeder exists, needs new scenarios |
Read existing seeder, add without breaking |
| Fix broken |
Seeder fails env gate / idempotency / DI scope |
Diagnose via Universal Rules, fix at root |
| Unknown |
Request ambiguous |
Ask user — NEVER assume |
rg "{Feature}Seeder|{Feature}SeedData|{Feature}TestData" {configured-source-roots} -l
Universal Seed Data Rules
Rule 0 — Convention First (priority before all others): ALWAYS discover and follow the project's EXISTING seed-data convention before doing anything — base class, env-gate key, count config key, registration, seeder marker. Match it with file:line evidence; never invent a parallel mechanism. If no convention exists, propose the smallest one that fits the project's stack.
- Environment Gate (local-dev default-only) — First check in seeder. Enabled by DEFAULT only on a local/development environment (or an explicit enable-config flag). NEVER seeds in production. The purpose is auto-setting-up feature test data on local/first-init, not a production data path.
- Command-Based (mimic a real user / QC tester) — Seeds by calling the same PUBLIC entry-point application commands a real user or QC engineer would invoke, via the full pipeline (validation + domain logic + events). This is automated happy-path self-testing. NEVER direct DB/repo writes for domain entities.
- No Duplicate Logic — Seeder provides realistic inputs. Commands own validation, domain logic, event side-effects.
- Idempotency (no re-seed when already seeded) — Check existing count → calculate remaining → seed only the difference. On restart with data already seeded, seed NOTHING. Running N times converges to the target, never duplicating.
- Count-Configurable (dual purpose, small default) — Read the seed count/times from the project config key (discovered Step 1); NEVER hardcode. Default to a small number when nothing is configured. The same count serves BOTH goals: repeating each scenario like a QC tester running it many times exercises the main cases AND enriches data volume (many simulated users) for performance testing and realistic first-time-init data. Zero → no-op.
- Restart-Safe (resume from last count) — Supports stop/start/restart any number of times: the loop runs from
existing_count to target_count, so if the target is X and only 50% of X is currently seeded, it continues until X is reached — never restarting from 0.
- Real-World Reachable State (seed only what the app could produce) — Every seeded entity MUST represent a state the application itself could have produced. This is the deeper reason Rule 2 exists: an application-level operation can only ever leave reachable state behind. Where a direct store write is genuinely unavoidable, it MUST carry a comment stating WHY that state is legitimate (bootstrapping legacy/migrated data, an externally-owned record, a deliberately corrupt fixture for repair testing) — unexplained, it is a defect, not a fixture. Seeded entities MUST also carry plausible relative timing: stagger creation/update/activity stamps across a realistic span instead of stamping every record with one shared instant. — why: a corpus the application could never produce makes every test over it prove nothing, and a corpus where everything happened in the same millisecond hides ordering defects and makes time-window, sort, and pagination behaviour untestable.
- Spec-Consistent (Spec-Loop Discipline — tailored) — Seeders are orchestration, NOT business logic, so property/metamorphic generation and the MUTATION-SCORE gate are N/A here — do not force them. Apply the dual-feedback half: every seeded scenario MUST stay consistent with the §5 invariants (commands own validation; a seeder that produces state violating an invariant is a bug, not a fixture). If a seeder encodes a domain rule — a required precondition, a status/relationship the scenario assumes, a business default — that rule belongs in the spec, not silently in the seeder: feed it into BOTH the spec (the rule) AND, where it is testable, the tests — never a seeder-only fix.
Protocol (Generate mode)
Generate-mode Task Plan (task tracking — required)
MUST ATTENTION task tracking ALL of these BEFORE the first edit. The plan ALWAYS ends with a --mode=review self-audit, and --mode=review ALWAYS precedes the $changes-review hand-off — changes-review stays the final step.
- Discover seeder patterns, env-gate key, count key (Step 1) —
file:line evidence.
- Verify dev config has env-gate + count keys (Step 1.5).
- Analyze feature scope + application commands (Step 2).
- Find or create the seeder file (Step 3).
- Implement using the language-agnostic algorithm (Step 4).
- Validate against the universal rules (Step 5) —
file:line for every gate.
- Self-review the changed seeder code by re-running THIS skill in
--mode=review (convention gate over the just-changed code — MUST be a task, not optional). Fix any FAIL through this generate flow, then re-review.
- Fresh zero-memory
code-reviewer round (Review Loop).
- Hand off to
$changes-review (final step — review all changes before commit).
- Analyze AI mistakes & lessons learned.
Step 1: Discover Seeder Patterns
Search for project seeder conventions:
# Search configured source roots using the repository's discovered seed-data naming conventions
rg "{configured-seeder-interface-or-base-patterns}|seeder|SeedData|DataSeed" {configured-source-roots} -l
Record with file:line evidence:
- Seeder base class / interface
- Seeder registration mechanism (DI, module, startup hook)
- Environment gate method/key name
- Count multiplier config key name
Step 1.5: Verify Dev Config Keys
Confirm dev config has both env gate key and count key. If absent, add following project's dev config convention. — why: missing keys silently disable the gate or count, producing no-op or unbounded seeding.
Step 2: Feature Scope Analysis
Identify before writing any code:
- Feature area — domain entity/aggregate being seeded
- Application commands —
rg "{Feature}.*Command|{configured-command-handler-patterns}" {configured-source-roots} -l
- Dependencies — data must exist (users, orgs, prerequisite records)
- Scenarios — 3–5 realistic variations (standard, boundary, multi-actor)
- Target count — clarify: 1 scenario or N repetitions per scenario
Step 3: Find or Create Seeder
rg "{Feature}TestSeeder|{Feature}SeedingHelper|{Feature}TestDataSeeder" {configured-source-roots} -l
- Exists → enhance with new scenarios, do NOT break existing ones
- Absent → create following discovered base class pattern
Step 4: Implement
Algorithm (language-agnostic):
seeder():
if not is_local_development_environment(): return # default-enabled on local/dev only, NEVER prod
if not seed_enabled_in_config(): return # explicit enable flag (default on for local)
target = config.get("SeedCount", SMALL_DEFAULT) # configurable; small default when unset
if target <= 0: return # zero → no-op
existing = count_by_seeder_marker() # how much is already seeded
if existing >= target: return # idempotent: already seeded → seed NOTHING
for i from existing to target: # restart-safe: resume the remainder (e.g. 50% → target)
call_application_command(build_scenario_input(i)) # public entry command, like a real user / QC tester
Seeder marker — stable predicate identifying seeded vs user data:
- Email prefix, created-by field, name prefix, or dedicated boolean flag
- MUST be deterministic across restarts
Step 5: Validate
MUST ATTENTION verify all before complete:
- MUST ATTENTION environment gate is FIRST check —
file:line evidence required
- MUST ATTENTION count-before-seed idempotency gate present —
file:line evidence
- MUST ATTENTION loop starts at
existing_count, not 0 — file:line evidence
- MUST ATTENTION only application-layer commands used for domain entities — NEVER repo/DB
- MUST ATTENTION no business logic or validation duplicated in seeder
- MUST ATTENTION seeder registered via project DI mechanism —
file:line evidence
- MUST ATTENTION count config key read correctly (zero → no-op, NEVER hardcoded)
- MUST ATTENTION scoped DI per iteration — shared scope = DbContext/session corruption
- MUST ATTENTION every seeded state is one the application could actually produce; any unavoidable direct store write carries a comment justifying WHY that state is legitimate —
file:line evidence
- MUST ATTENTION seeded entities carry plausible relative timing (staggered stamps), NEVER one shared instant
Sub-Agent Routing
| Task |
Sub-Agent |
When |
| Discover seeders + commands across large codebase |
general-purpose |
Steps 1-2 |
| Review seeder compliance |
code-reviewer |
Round 1 post-implementation |
| Seeder handles credentials/PII |
security-auditor |
Security-sensitive patterns |
| Seeder runs 1000+ records |
performance-optimizer |
Performance-intensive |
All sub-agent prompts MUST include:
Graph DB active. After grep finds key files, run:
python .claude/scripts/code_graph trace <file> --direction both --json
Pattern: grep → trace → grep verify.
Anti-Patterns
| Anti-Pattern |
Correct |
| Direct repo insert for domain entities |
Call application command |
| Seeder validates business rules |
Command owns validation; seeder provides valid inputs |
| No idempotency check |
Check count first; seed only remaining |
Hardcoded count (for i in 0..10) |
Read count from config key (discovered Step 1) |
| No environment gate |
Check project env gate key first |
| Shared DI scope across loop iterations |
Use project's scoped DI per iteration (prevents DbContext corruption) |
| Seeded state no application operation could produce |
Seed it through an application-level operation; if a direct write is unavoidable, comment WHY that state is legitimate |
| Every seeded record sharing one creation instant |
Stagger stamps across a realistic span so ordering / time-window behaviour stays testable |
| Batch-all-then-write sub-agent findings |
Persist findings per file; NEVER batch at end |
Review Loop
Round 1: After implementation, spawn fresh code-reviewer sub-agent with zero memory of implementation:
Review seeder at [file:path]. Verify with file:line evidence for each:
1. Environment gate is FIRST check
2. Idempotency: count-before-seed pattern present
3. Loop starts at existing_count not 0
4. Zero application-layer command bypasses (direct repo/DB = FAIL)
5. No hardcoded count — config key read
6. Scoped DI per iteration
Report: PASS or FAIL with file:line for each finding.
Fix loop: If FAIL → validate findings → fix validated findings → restart full review from first phase. When restarted review uses sub-agents, NEVER reuse them across rounds. If same blocker repeats across 2 full invocations with no progress, escalate to user.
NEVER fix unvalidated findings. Do not spawn a fresh sub-agent only to re-review known findings before validation/fix.
Mode: Review (seed-data convention audit)
Invoke with --mode=review. READ-ONLY audit of a seeder target against EVERY Universal Seed Data Rule AND the project-specific seeder conventions. Produces a per-principle PASS/FAIL with file:line evidence. This mode makes NO code changes — it reports findings and routes confirmed defects back to Generate mode for the fix.
R0 — Resolve the review target
Determine WHAT to review, in priority order:
- Explicit target in the user prompt — a named seeder file / class / feature area → review exactly that.
- Else → current changes —
git diff --name-only plus staged (git diff --cached --name-only) and untracked, filtered to seeder files using the discovered seeder naming (Step 1 / reference doc). Review every changed or added seeder.
- Else → current work-context result — the seeder(s) created or edited earlier in THIS session / work context.
If none resolve → ask the user which seeder to review. NEVER assume a target.
R1 — Read the conventions BEFORE reviewing (BLOCKING)
MUST ATTENTION read, in full, before forming ANY verdict:
docs/project-reference/seed-test-data-reference.md — project seeder locations, base class, env-gate key, count config key, DI/UoW scope strategy, Required Patterns, Verification Checklist.
docs/project-config.json → Data Seeders context group — configured source roots, naming conventions, run commands.
- The Universal Seed Data Rules (1–8) in this skill — the principles being graded.
- The target seeder file(s) themselves — re-read in full; NEVER review from memory.
- Step 1 discovery: confirm the project's ACTUAL seeder base class, env-gate key, and count key with
file:line — the review grades against THESE, not generic defaults.
If the reference doc is still a skeleton (TODO placeholders), say so explicitly, grade against discovered file:line conventions instead, and raise the missing/incomplete project reference as its own finding.
R2 — Review checklist (grade EVERY item: file:line evidence or FAIL)
Universal rules:
Project-specific conventions (from the reference doc):
R3 — Verdict
Per item: PASS / FAIL / N/A with file:line evidence and confidence (>80% required to assert a FAIL; <60% → "insufficient evidence", verify before grading). Overall verdict is PASS only if ZERO universal-rule FAILs.
- PASS → report the evidence table; if idempotency/count tests are absent, suggest
$integration-test.
- FAIL → list each violation with the responsible
file:line and the correct pattern (from the Anti-Patterns table / reference doc). Route the fix back through Generate mode (Phase 0 → "Fix broken"); after the fix lands, RE-RUN --mode=review over the changed code. NEVER edit the seeder inside review mode.
Review-mode task plan (task tracking — required)
- Resolve the review target (prompt → current changes → work-context).
- Read
seed-test-data-reference.md + project-config Data Seeders group + Universal Rules + the target file(s).
- Discover/confirm base class, env-gate key, count key with
file:line evidence.
- Grade every universal + project-specific checklist item (R2).
- Produce the PASS/FAIL verdict with per-item
file:line evidence (R3).
- If FAIL → hand confirmed defects to Generate mode and re-review after the fix; else report PASS + next-step suggestion.
- Analyze AI mistakes & lessons learned.
Workflow Recommendation
MUST ATTENTION — NOT IN WORKFLOW YET: Use ask the user directly:
- Activate
workflow-seed-test-data (Recommended) — scout → investigate → seed-test-data → changes-review → code-simplifier → docs-update
- Execute
$seed-test-data directly — run this skill standalone
Next Steps
MUST ATTENTION after completing (Generate mode): use ask the user directly — do NOT skip. Step 7 self-review (--mode=review) MUST have run on the changed code BEFORE these:
- "$workflow-review-changes (Recommended)" — final step: review all changes before commit (runs AFTER the
--mode=review convention self-audit)
- "$integration-test" — write tests verifying idempotency and count compliance
- "Skip, continue manually" — user decides
[IMPORTANT] task tracking for ALL tasks BEFORE starting. For simple tasks, ask user whether to skip.
Critical Thinking Mindset — Apply critical thinking, sequential thinking. Every claim needs traced proof, confidence >80% to act.
Anti-hallucination: Never present guess as fact — cite sources for every claim, admit uncertainty freely, self-check output for errors, cross-reference independently, stay skeptical of own confidence — certainty without evidence root of all hallucination.
Understand Code First — HARD-GATE: Do NOT write, plan, or fix until you READ existing code.
- Search 3+ similar patterns (
grep/glob) — cite file:line evidence
- Read existing files in target area — understand structure, base classes, conventions
- Run
python .claude/scripts/code_graph trace <file> --direction both --json when .code-graph/graph.db exists
- Map dependencies via
connections or callers_of — know what depends on your target
- Write investigation to
.ai/workspace/analysis/ for non-trivial tasks (3+ files)
- Re-read analysis file before implementing — never work from memory alone. — why: long context drifts from the file; the file is ground truth
- NEVER invent new patterns when existing ones work — match exactly or document deviation. — why: divergent patterns fragment the codebase and slow every future reader
BLOCKED until: - [ ] Read target files - [ ] Grep 3+ patterns - [ ] Graph trace (if graph.db exists) - [ ] Assumptions verified with evidence
Evidence-Based Reasoning — Speculation is FORBIDDEN. Every claim needs proof.
- Cite
file:line, grep results, or framework docs for EVERY claim
- Declare confidence: >80% act freely, 60-80% verify first, <60% DO NOT recommend
- Cross-service validation required for architectural changes
- "I don't have enough evidence" is valid and expected output
BLOCKED until: - [ ] Evidence file path (file:line) - [ ] Grep search performed - [ ] 3+ similar patterns found - [ ] Confidence level stated
Forbidden without proof: "obviously", "I think", "should be", "probably", "this is because"
If incomplete → output: "Insufficient evidence. Verified: [...]. Not verified: [...]."
AI Mistake Prevention — Failure modes to avoid on every task:
Re-read files after context changes. Context compaction, resume, or long-running work can make memory stale; verify current files before acting.
Verify generated content against source evidence. AI hallucinates APIs, names, claims, and document facts. Check the relevant source before documenting or referencing.
Check downstream references before deleting or renaming. Removing an artifact can stale docs, generated mirrors, configs, and callers; map references first.
Trace the full impact chain after edits. Changing a definition can miss derived outputs and consumers. Follow the affected chain before declaring done.
Verify ALL affected outputs, not just the first. One green check is not all green checks; validate every output surface the change can affect.
Assume existing values are intentional — ask WHY before changing OR flagging one as a defect. Before changing or reporting a constant, limit, flag, cutoff, wording, or pattern, read nearby context and history, the CALLER's ordering, and 2+ sibling call sites of the same convention. A doc stating WHAT without WHY is missing rationale, not proof of a missing guard.
Surface ambiguity before acting — don't pick silently. Multiple valid interpretations require an explicit question or stated assumption with risk.
Assert the outcome your system owns, not the intermediate state your infrastructure owns. When verifying async work, assert the final business state — never the delivery/retry bookkeeping held in shared infrastructure that any co-running process can write. Such a check passes when run alone and flakes the moment anything else shares that infrastructure.
Keep shared guidance role-relevant. Universal guidance must help every receiving skill or agent; code-specific obligations belong only in code-specific protocols.
Real-World Fidelity Gate — MANDATORY when authoring, reviewing, or repairing any integration / E2E / system test.
A test earns trust by reproducing a situation the system can actually meet in production. A scenario that could never occur in real life proves nothing when it passes, and wastes hours when it fails.
- Ask the fidelity question BEFORE writing the setup: "Can this sequence, timing, and data actually occur in production?" If no, the test is mis-specified — fix the SCENARIO, never the assertion.
- Model real pacing between actor steps. Two distinct actor actions that production separates by seconds, minutes, or hours MUST NOT be fired back-to-back in the same millisecond. Compressed pacing manufactures races the system was never designed to survive, then reports them as product defects.
- Wait on a real signal, never a blind sleep. Find an observable proving the prior step finished — a persisted state change, an audit/version stamp, a queue/worker idle marker, a completion event — and poll until it settles (unchanged across a short stability window). Use a fixed delay ONLY when no observable exists, and say so in a comment.
- Barriers belong in ARRANGE, never in ASSERT. Waiting for a precondition is fidelity. Widening an assertion's timeout, loosening a comparison, adding a retry around a failing assertion, or skipping the test is masking. NEVER do the latter to force green.
- Distinguish harness-amplified from real. Test topologies (shared infra, fan-out consumers, parallel suites, cold starts) can make a rare production race routine locally. Before filing a product defect, state whether the trigger exists in production and at what likelihood.
- Keep the protected invariant intact. Improving fidelity must NEVER reduce what the test protects. If a realistic scenario no longer exercises the rule, the rule needs a DIFFERENT realistic scenario — not a weaker assertion.
- Deliberate impossible-state tests are allowed, but MUST be labelled. Corruption-repair, migration, and fail-safe tests intentionally construct states production should never reach; comment WHY the state is reachable (upstream bug, partial write, legacy data), so they are never confused with unrealistic setups.
IMPORTANT MUST ATTENTION search 3+ existing patterns and read code BEFORE writing any seeder.
MUST ATTENTION cite file:line for every claim; declare confidence; "I don't have enough evidence" is valid output.
MUST ATTENTION apply critical + sequential thinking — every claim needs appropriate traced evidence (file:line for repo/code claims; source URL or artifact section for research, product, content, and docs claims); confidence >80% to act, <60% DO NOT recommend. Anti-hallucination: never present guess as fact, admit uncertainty freely, cross-reference independently, stay skeptical of own confidence.
MUST ATTENTION apply AI mistake prevention — verify generated content against evidence, trace downstream references before deleting or renaming, verify all affected outputs, re-read files after context loss, and surface ambiguity before acting.
Prompt-Enhance Closing Anchors
IMPORTANT MUST ATTENTION follow declared step order for this skill; NEVER skip, reorder, or merge steps without explicit user approval
IMPORTANT MUST ATTENTION for every step/sub-skill call: set in_progress before execution, set completed after execution
IMPORTANT MUST ATTENTION every skipped step MUST include explicit reason; every completed step MUST include concise evidence
IMPORTANT MUST ATTENTION if Task tools unavailable, maintain an equivalent step-by-step plan tracker with synchronized statuses
Parallel Sub-Agent Dispatch — Plan parallelism the moment a task breakdown exists, BEFORE executing it — running provably independent tasks sequentially wastes wall-clock. Applies to every multi-step job: workflow steps, planning, batch updates, investigation, research, scans, reviews, doc sync. Plan execution is metadata-gated, NEVER default-parallel — fan-out follows ONLY what the plan declares (PAR/SEQ tags + per-phase write set); an untagged plan runs sequentially — why: a derived write set cannot see cascade or generated writes.
- Tag every task
PAR or SEQ. PAR = inputs exclude every pending task's output AND write set disjoint from every other PAR. Else SEQ — MUST ATTENTION name the dependency forcing it.
- Group
PAR into waves. No edge between members. Two writers of one file NEVER share a wave. Read-only work (search, investigation, review, research) parallelizes freely.
- Declare before dispatch:
Parallel plan: wave 1 = [...] · wave 2 = [...] · SEQ = [...] (reason).
- Spawn each wave in ONE message — every
spawn_agent call in one response, NEVER dripped per turn. Route each task to its specialist (`.claude/skills/sh
…(truncated)
1---2name: seed-test-data3description: [Dev Data] Use when you need to implement or enhance test data seeders that simulate QC happy-path scenarios via application-layer commands. Flag: --mode=review reviews a target seeder (or the current changes / current work-context result) against every universal seed-data rule AND the project-specific seeder conventions — read-only, evidence-backed PASS/FAIL.4---5
6> Codex compatibility note:
7>
8> - Invoke repository skills with `$skill-name` in Codex; this mirrored copy rewrites legacy Claude `/skill-name` references.
9> - Task tracker mandate: BEFORE executing any workflow or skill step, create/update task tracking for all steps and keep it synchronized as progress changes.
10> - User-question prompts mean to ask the user directly in Codex.
11> - Ignore Claude-specific mode-switch instructions when they appear.
12> - Strict execution contract: when a user explicitly invokes a skill, execute that skill protocol as written.
13> - Subagent authorization: when a skill is user-invoked or AI-detected and its protocol requires subagents, that skill activation authorizes use of the required `spawn_agent` subagent(s) for that task.
14> - Do not skip, reorder, or merge protocol steps unless the user explicitly approves the deviation first.
15> - For workflow skills, execute each listed child-skill step explicitly and report step-by-step evidence.
16> - If a required step/tool cannot run in this environment, stop and ask the user before adapting.
17
18<!-- CODEX:PROJECT-REFERENCE-LOADING:START -->
19
20## Codex Project-Reference Loading (No Hooks)
21
22Codex uses static project-reference loading instead of runtime-injected project docs.
23When coding, planning, debugging, testing, or reviewing, open project docs explicitly using this routing.
24
25**Always read:**
26
27- `docs/project-config.json` (project-specific paths, commands, modules, and workflow/test settings)
28- `docs/project-reference/docs-index-reference.md` (routes to the full `docs/project-reference/*` catalog)
29- `docs/project-reference/lessons.md` (always-on guardrails and anti-patterns)
30
31**Missing/stale context route:** If `docs/project-config.json`, the docs index, `lessons.md`, `CLAUDE.md`, `AGENTS.md`, or any task-required reference doc is missing or stale, auto-run `$project-init` or the narrow setup route (`$project-config`, `$docs-init`, `$scan-all`, `$scan --target=<key>`, `$claude-md-init`) before ordinary project-specific work. If Codex mirrors or `AGENTS.md` are missing/stale, ask the user to run `$sync-codex`; do not auto-run it.
32
33**Situation-based docs:**
34
35- Project structure/architecture/tech-stack/deployment/setup (any layer — backend, frontend, or infra): `project-structure-reference.md`
36- Backend/CQRS/API/domain/entity changes: `backend-patterns-reference.md`, `domain-entities-reference.md`
37- Frontend/UI/styling/design-system: `frontend-patterns-reference.md`, `scss-styling-guide.md`, `design-system/README.md`
38- Spec authoring, `docs/specs/` pathing, or TC format: `feature-spec-reference.md`, `spec-system-reference.md`, `spec-principles.md`
39- Behavior/public-contract changes or spec-test-code sync: `workflow-spec-test-code-cycle-reference.md` plus the spec docs above
40- Derived spec indexes/ERDs/reimplementation guides: `spec-system-reference.md` and source Feature Specs under `docs/specs/`
41- Integration test implementation/review: `integration-test-reference.md`
42- E2E test implementation/review: `e2e-test-reference.md`
43- Code review/audit work: `code-review-rules.md` plus domain docs above based on changed files
44
45Do not read all docs blindly. Start from `docs-index-reference.md`, then open only relevant files for the task.
46
47<!-- CODEX:PROJECT-REFERENCE-LOADING:END -->
48
49<!-- PROMPT-ENHANCE:STEP-TASK-ANCHOR:START -->
50
51> **[BLOCKING]** Execute skill steps in declared order. NEVER skip, reorder, or merge steps without explicit user approval.
52> **[BLOCKING]** Before each step or sub-skill call, update task tracking: set `in_progress` when step starts, set `completed` when step ends.
53> **[BLOCKING]** Every completed/skipped step MUST include brief evidence or explicit skip reason.
54> **[BLOCKING]** If Task tools are unavailable, create and maintain an equivalent step-by-step plan tracker with the same status transitions.
55
56<!-- PROMPT-ENHANCE:STEP-TASK-ANCHOR:END -->
57
58## Quick Summary
59
60**Goal:** Stand up **configurable, local-dev-only** test-data seeders — **enabled by default ONLY on local/dev** — that auto-seed each feature's happy-path scenarios by calling the **same public entry-point application commands a real user / QC tester would** (NEVER direct DB writes), so the system **self-tests its main cases** like a QC engineer exercising them by hand; a **configurable seed count** repeats each scenario to BOTH **cover the main cases** AND **enrich data volume** (many simulated users) for **performance testing** + realistic **first-time-init** data, **idempotent + restart-safe** (never re-seeds already-seeded data; resumes from the last count toward target X, at 50% → continue until X), **defaulting the count small** when nothing is configured — and **ALWAYS finding the project's existing seed-data convention FIRST**.
61
62**Summary:**
63
64- **Find the existing convention FIRST.** Before designing anything, discover the project's seeder base class, env-gate key, count config key, and registration with `file:line` evidence (Step 1) — match it exactly; never invent a parallel mechanism.
65- Seeders orchestrate the real app pipeline like a real user: invoke the **public entry-point application commands** (which own validation, domain logic, and event side-effects) — never repo/DB inserts for domain entities, never duplicate command logic in the seeder.
66- **Dual purpose, one mechanism — a configurable count:** repeat each happy-path scenario N times to (a) self-test the main cases (QC mimic) and (b) enrich data volume for many-users / performance / first-init realism. Read the count from config (never hardcode); **default small** when unset; zero → no-op.
67- Four non-negotiable gates in order: (1) **environment gate** as the FIRST check (local-dev/enabled-config only), (2) **count-before-seed idempotency** (no re-seed when already seeded), (3) **restart-safe loop** from `existing_count` to `target_count` (never 0 — resume the remainder after stop/restart), (4) scoped DI per iteration — a shared scope silently corrupts the DbContext/session.
68- Always pre-read `docs/project-reference/seed-test-data-reference.md` + project-config `Data Seeders` group, then close with a fresh zero-memory `code-reviewer` round; re-review fully only after a validated fix.
69- **Two modes — surface the flag:** default **Generate** (implement / enhance / fix a seeder); **`--mode=review`** = READ-ONLY convention audit grading a target (prompt → current changes → work-context) against EVERY universal rule + project conventions with `file:line` PASS/FAIL — routes confirmed fixes back to Generate, NEVER edits the seeder itself.
70- **Main steps to run (Generate, in order — do not skip):** Phase 0 detect task type (new/enhance/fix) → Step 1 discover conventions (base class, env-gate key, count key, registration) → Step 1.5 verify dev-config keys exist → Step 2 feature scope + application commands → Step 3 find/create seeder → Step 4 implement (env-gate FIRST → config count → idempotency → restart-safe loop → scoped DI) → Step 5 validate every gate with `file:line` → Step 7 `--mode=review` self-audit on the changed code → fresh `code-reviewer` round → `$changes-review` (final).
71
72**Workflow (Generate mode — default):**
73
741. **Phase 0** — Detect seeder task type (new / enhance / fix)
752. **Step 1** — Discover project seeder patterns, env gate key, count key
763. **Step 2** — Analyze feature scope + application commands
774. **Step 3** — Find or create seeder file
785. **Step 4** — Implement using language-agnostic algorithm
796. **Step 5** — Validate against universal rules
807. **Self-Review** — Re-run THIS skill in `--mode=review` over the changed seeder code (convention gate)
818. **Review** — Fresh sub-agent review round, then hand off to `$changes-review`
82
83**Modes:**
84
85- **Default (generate)** — implement / enhance / fix seeders. Everything in the Generate-mode Protocol below applies. The generate-mode task plan MUST end by re-running this skill in `--mode=review` (Step 7) BEFORE the `$changes-review` hand-off.
86- **`--mode=review`** (read-only convention audit) — review a target against EVERY universal seed-data rule AND the project-specific seeder conventions, with `file:line` evidence and a PASS/FAIL verdict. Makes NO code changes; reports findings and routes confirmed defects back to generate mode for the fix. See [Mode: Review](#mode-review-seed-data-convention-audit).
87
88**Key Rules:**
89
90- ALWAYS find the project's existing seed-data convention FIRST — read `docs/project-reference/seed-test-data-reference.md` and `docs/project-config.json` (`Data Seeders` context group) before writing any seeder changes; match the discovered pattern, never invent a new one
91- ENABLE seeding by default ONLY on a local/development environment — the environment gate is the FIRST check, NEVER production
92- SEED like a real user / QC tester — call the public entry-point application commands; NEVER call repository/DB directly for domain data
93- NEVER duplicate command logic — seeder orchestrates, commands own validation
94- ALWAYS make the seed count configurable and read it from config (NEVER hardcode); **default to a small number when nothing is configured**; zero → no-op
95- A configurable count serves BOTH goals — self-test the main happy-path cases AND enrich data volume (many simulated users) for performance testing and realistic first-time-init data
96- GUARANTEE idempotency — check count before seeding; never re-seed already-seeded data on restart
97- ALWAYS loop from `existing_count` to `target_count` so stop/restart resumes the remainder (target X, at 50% → continue until X), never re-seeding from 0
98- SEED only states the application could actually produce — application-level operations guarantee reachability by construction; a direct store write fabricating an otherwise-unreachable state MUST be commented with why it is legitimate, and seeded entities MUST carry plausible relative timing rather than one shared instant
99
100## Mode Routing (FIRST decision)
101
102Before Phase 0, route on the invocation flag:
103
104| Signal | Mode | Go to |
105| ------------------------------------------------------------------- | ---------------------- | ------------------------------------------------------- |
106| `--mode=review` flag, OR prompt asks to review/audit/check a seeder | **Review** | [Mode: Review](#mode-review-seed-data-convention-audit) |
107| Any other invocation (implement / enhance / fix a seeder) | **Generate** (default) | Phase 0 below |
108
109> **MUST ATTENTION** Generate mode OWNS the fix; Review mode is READ-ONLY and only reports. When Review mode finds a defect, it routes the fix back through Generate mode — it never edits the seeder itself.
110
111## Phase 0: Detect Seeder Task Type _(Generate mode)_
112
113Before any other step, classify the request:
114
115| Task Type | Detection | Action |
116| ---------------- | ---------------------------------------------- | ---------------------------------------------- |
117| New seeder | No existing seeder for feature area | Create following discovered base class pattern |
118| Enhance existing | Seeder exists, needs new scenarios | Read existing seeder, add without breaking |
119| Fix broken | Seeder fails env gate / idempotency / DI scope | Diagnose via Universal Rules, fix at root |
120| Unknown | Request ambiguous | Ask user — NEVER assume |
121
122```bash
123rg "{Feature}Seeder|{Feature}SeedData|{Feature}TestData" {configured-source-roots} -l
124```
125
126## Universal Seed Data Rules
127
128> **Rule 0 — Convention First (priority before all others):** ALWAYS discover and follow the project's EXISTING seed-data convention before doing anything — base class, env-gate key, count config key, registration, seeder marker. Match it with `file:line` evidence; never invent a parallel mechanism. If no convention exists, propose the smallest one that fits the project's stack.
129
1301. **Environment Gate (local-dev default-only)** — First check in seeder. Enabled by DEFAULT only on a local/development environment (or an explicit enable-config flag). NEVER seeds in production. The purpose is auto-setting-up feature test data on local/first-init, not a production data path.
1312. **Command-Based (mimic a real user / QC tester)** — Seeds by calling the same PUBLIC entry-point application commands a real user or QC engineer would invoke, via the full pipeline (validation + domain logic + events). This is automated happy-path self-testing. NEVER direct DB/repo writes for domain entities.
1323. **No Duplicate Logic** — Seeder provides realistic inputs. Commands own validation, domain logic, event side-effects.
1334. **Idempotency (no re-seed when already seeded)** — Check existing count → calculate remaining → seed only the difference. On restart with data already seeded, seed NOTHING. Running N times converges to the target, never duplicating.
1345. **Count-Configurable (dual purpose, small default)** — Read the seed count/times from the project config key (discovered Step 1); NEVER hardcode. **Default to a small number when nothing is configured.** The same count serves BOTH goals: repeating each scenario like a QC tester running it many times exercises the main cases AND enriches data volume (many simulated users) for performance testing and realistic first-time-init data. Zero → no-op.
1356. **Restart-Safe (resume from last count)** — Supports stop/start/restart any number of times: the loop runs from `existing_count` to `target_count`, so if the target is X and only 50% of X is currently seeded, it continues until X is reached — never restarting from 0.
1367. **Real-World Reachable State (seed only what the app could produce)** — Every seeded entity MUST represent a state the application itself could have produced. This is the deeper reason Rule 2 exists: an application-level operation can only ever leave reachable state behind. Where a direct store write is genuinely unavoidable, it MUST carry a comment stating WHY that state is legitimate (bootstrapping legacy/migrated data, an externally-owned record, a deliberately corrupt fixture for repair testing) — unexplained, it is a defect, not a fixture. Seeded entities MUST also carry **plausible relative timing**: stagger creation/update/activity stamps across a realistic span instead of stamping every record with one shared instant. — why: a corpus the application could never produce makes every test over it prove nothing, and a corpus where everything happened in the same millisecond hides ordering defects and makes time-window, sort, and pagination behaviour untestable.
1378. **Spec-Consistent (Spec-Loop Discipline — tailored)** — Seeders are orchestration, NOT business logic, so property/metamorphic generation and the MUTATION-SCORE gate are **N/A here** — do not force them. Apply the dual-feedback half: every seeded scenario MUST stay consistent with the **§5 invariants** (commands own validation; a seeder that produces state violating an invariant is a bug, not a fixture). If a seeder encodes a **domain rule** — a required precondition, a status/relationship the scenario assumes, a business default — that rule belongs in the **spec**, not silently in the seeder: feed it into BOTH the spec (the rule) AND, where it is testable, the tests — never a seeder-only fix.
138
139## Protocol _(Generate mode)_
140
141### Generate-mode Task Plan (task tracking — required)
142
143> **MUST ATTENTION** task tracking ALL of these BEFORE the first edit. The plan ALWAYS ends with a `--mode=review` self-audit, and `--mode=review` ALWAYS precedes the `$changes-review` hand-off — changes-review stays the final step.
144
1451. Discover seeder patterns, env-gate key, count key (Step 1) — `file:line` evidence.
1462. Verify dev config has env-gate + count keys (Step 1.5).
1473. Analyze feature scope + application commands (Step 2).
1484. Find or create the seeder file (Step 3).
1495. Implement using the language-agnostic algorithm (Step 4).
1506. Validate against the universal rules (Step 5) — `file:line` for every gate.
1517. **Self-review the changed seeder code by re-running THIS skill in `--mode=review`** (convention gate over the just-changed code — MUST be a task, not optional). Fix any FAIL through this generate flow, then re-review.
1528. Fresh zero-memory `code-reviewer` round (Review Loop).
1539. Hand off to `$changes-review` (final step — review all changes before commit).
15410. Analyze AI mistakes & lessons learned.
155
156### Step 1: Discover Seeder Patterns
157
158Search for project seeder conventions:
159
160```bash
161# Search configured source roots using the repository's discovered seed-data naming conventions
162rg "{configured-seeder-interface-or-base-patterns}|seeder|SeedData|DataSeed" {configured-source-roots} -l
163```
164
165Record with `file:line` evidence:
166
167- Seeder base class / interface
168- Seeder registration mechanism (DI, module, startup hook)
169- Environment gate method/key name
170- Count multiplier config key name
171
172### Step 1.5: Verify Dev Config Keys
173
174Confirm dev config has both env gate key and count key. If absent, add following project's dev config convention. — why: missing keys silently disable the gate or count, producing no-op or unbounded seeding.
175
176### Step 2: Feature Scope Analysis
177
178Identify before writing any code:
179
1801. **Feature area** — domain entity/aggregate being seeded
1812. **Application commands** — `rg "{Feature}.*Command|{configured-command-handler-patterns}" {configured-source-roots} -l`
1823. **Dependencies** — data must exist (users, orgs, prerequisite records)
1834. **Scenarios** — 3–5 realistic variations (standard, boundary, multi-actor)
1845. **Target count** — clarify: 1 scenario or N repetitions per scenario
185
186### Step 3: Find or Create Seeder
187
188```bash
189rg "{Feature}TestSeeder|{Feature}SeedingHelper|{Feature}TestDataSeeder" {configured-source-roots} -l
190```
191
192- **Exists** → enhance with new scenarios, do NOT break existing ones
193- **Absent** → create following discovered base class pattern
194
195### Step 4: Implement
196
197**Algorithm (language-agnostic):**
198
199```
200seeder():
201 if not is_local_development_environment(): return # default-enabled on local/dev only, NEVER prod
202 if not seed_enabled_in_config(): return # explicit enable flag (default on for local)
203 target = config.get("SeedCount", SMALL_DEFAULT) # configurable; small default when unset
204 if target <= 0: return # zero → no-op
205 existing = count_by_seeder_marker() # how much is already seeded
206 if existing >= target: return # idempotent: already seeded → seed NOTHING
207 for i from existing to target: # restart-safe: resume the remainder (e.g. 50% → target)
208 call_application_command(build_scenario_input(i)) # public entry command, like a real user / QC tester
209```
210
211**Seeder marker** — stable predicate identifying seeded vs user data:
212
213- Email prefix, created-by field, name prefix, or dedicated boolean flag
214- MUST be deterministic across restarts
215
216### Step 5: Validate
217
218MUST ATTENTION verify all before complete:
219
220- MUST ATTENTION environment gate is FIRST check — `file:line` evidence required
221- MUST ATTENTION count-before-seed idempotency gate present — `file:line` evidence
222- MUST ATTENTION loop starts at `existing_count`, not 0 — `file:line` evidence
223- MUST ATTENTION only application-layer commands used for domain entities — NEVER repo/DB
224- MUST ATTENTION no business logic or validation duplicated in seeder
225- MUST ATTENTION seeder registered via project DI mechanism — `file:line` evidence
226- MUST ATTENTION count config key read correctly (zero → no-op, NEVER hardcoded)
227- MUST ATTENTION scoped DI per iteration — shared scope = DbContext/session corruption
228- MUST ATTENTION every seeded state is one the application could actually produce; any unavoidable direct store write carries a comment justifying WHY that state is legitimate — `file:line` evidence
229- MUST ATTENTION seeded entities carry plausible relative timing (staggered stamps), NEVER one shared instant
230
231## Sub-Agent Routing
232
233| Task | Sub-Agent | When |
234| ------------------------------------------------- | ----------------------- | --------------------------- |
235| Discover seeders + commands across large codebase | `general-purpose` | Steps 1-2 |
236| Review seeder compliance | `code-reviewer` | Round 1 post-implementation |
237| Seeder handles credentials/PII | `security-auditor` | Security-sensitive patterns |
238| Seeder runs 1000+ records | `performance-optimizer` | Performance-intensive |
239
240**All sub-agent prompts MUST include:**
241
242```
243Graph DB active. After grep finds key files, run:
244python .claude/scripts/code_graph trace <file> --direction both --json
245Pattern: grep → trace → grep verify.
246```
247
248## Anti-Patterns
249
250| Anti-Pattern | Correct |
251| --------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------- |
252| Direct repo insert for domain entities | Call application command |
253| Seeder validates business rules | Command owns validation; seeder provides valid inputs |
254| No idempotency check | Check count first; seed only remaining |
255| Hardcoded count (`for i in 0..10`) | Read count from config key (discovered Step 1) |
256| No environment gate | Check project env gate key first |
257| Shared DI scope across loop iterations | Use project's scoped DI per iteration (prevents DbContext corruption) |
258| Seeded state no application operation could produce | Seed it through an application-level operation; if a direct write is unavoidable, comment WHY that state is legitimate |
259| Every seeded record sharing one creation instant | Stagger stamps across a realistic span so ordering / time-window behaviour stays testable |
260| Batch-all-then-write sub-agent findings | Persist findings per file; NEVER batch at end |
261
262## Review Loop
263
264**Round 1:** After implementation, spawn fresh `code-reviewer` sub-agent with zero memory of implementation:
265
266```
267Review seeder at [file:path]. Verify with file:line evidence for each:
2681. Environment gate is FIRST check
2692. Idempotency: count-before-seed pattern present
2703. Loop starts at existing_count not 0
2714. Zero application-layer command bypasses (direct repo/DB = FAIL)
2725. No hardcoded count — config key read
2736. Scoped DI per iteration
274Report: PASS or FAIL with file:line for each finding.
275```
276
277**Fix loop:** If FAIL → validate findings → fix validated findings → restart full review from first phase. When restarted review uses sub-agents, NEVER reuse them across rounds. If same blocker repeats across 2 full invocations with no progress, escalate to user.
278NEVER fix unvalidated findings. Do not spawn a fresh sub-agent only to re-review known findings before validation/fix.
279
280---
281
282## Mode: Review (seed-data convention audit)
283
284> **Invoke with `--mode=review`.** READ-ONLY audit of a seeder target against EVERY [Universal Seed Data Rule](#universal-seed-data-rules) AND the project-specific seeder conventions. Produces a per-principle PASS/FAIL with `file:line` evidence. This mode makes **NO code changes** — it reports findings and routes confirmed defects back to Generate mode for the fix.
285
286### R0 — Resolve the review target
287
288Determine WHAT to review, in priority order:
289
2901. **Explicit target in the user prompt** — a named seeder file / class / feature area → review exactly that.
2912. **Else → current changes** — `git diff --name-only` plus staged (`git diff --cached --name-only`) and untracked, filtered to seeder files using the discovered seeder naming (Step 1 / reference doc). Review every changed or added seeder.
2923. **Else → current work-context result** — the seeder(s) created or edited earlier in THIS session / work context.
293
294If none resolve → ask the user which seeder to review. NEVER assume a target.
295
296### R1 — Read the conventions BEFORE reviewing (BLOCKING)
297
298MUST ATTENTION read, in full, before forming ANY verdict:
299
300- `docs/project-reference/seed-test-data-reference.md` — project seeder locations, base class, env-gate key, count config key, DI/UoW scope strategy, Required Patterns, Verification Checklist.
301- `docs/project-config.json` → `Data Seeders` context group — configured source roots, naming conventions, run commands.
302- The [Universal Seed Data Rules](#universal-seed-data-rules) (1–8) in this skill — the principles being graded.
303- The target seeder file(s) themselves — re-read in full; NEVER review from memory.
304- Step 1 discovery: confirm the project's ACTUAL seeder base class, env-gate key, and count key with `file:line` — the review grades against THESE, not generic defaults.
305
306> If the reference doc is still a skeleton (`TODO` placeholders), say so explicitly, grade against discovered `file:line` conventions instead, and raise the missing/incomplete project reference as its own finding.
307
308### R2 — Review checklist (grade EVERY item: `file:line` evidence or FAIL)
309
310**Universal rules:**
311
312- [ ] **Environment gate is the FIRST check** — dev/enabled-config only, NEVER production.
313- [ ] **Command-based** — domain entities created ONLY via application-layer commands; ZERO direct repo/DB writes.
314- [ ] **No duplicated logic** — seeder feeds realistic inputs; commands own validation / domain / event side-effects.
315- [ ] **Idempotency** — count-before-seed gate present; running N times converges to target (no duplicates).
316- [ ] **Count-configurable** — count read from the discovered config key; NEVER hardcoded (zero → no-op).
317- [ ] **Restart-safe loop** — loop starts at `existing_count`, NEVER 0.
318- [ ] **Scoped DI per iteration** — fresh scope per loop iteration; no shared DbContext/session.
319- [ ] **Real-world reachable state** — every seeded entity is a state the application itself could produce; any direct store write fabricating an otherwise-unreachable state carries a comment justifying WHY it is legitimate; seeded entities carry plausible relative timing, not one shared instant.
320- [ ] **Spec-consistency** — every seeded scenario satisfies the §5 invariants; any encoded domain rule (precondition / status / default) is reflected in the spec (and tests where testable), not seeder-only.
321
322**Project-specific conventions (from the reference doc):**
323
324- [ ] Seeder lives in the configured folder and extends the project's discovered base class / interface.
325- [ ] Registered via the project's documented DI / registration mechanism.
326- [ ] Env-gate key + count key match the documented keys AND exist in dev config.
327- [ ] Seeder marker (email/name prefix, created-by, dedicated flag) is deterministic across restarts.
328- [ ] Conforms to the reference doc's Required Patterns + Verification Checklist.
329
330### R3 — Verdict
331
332Per item: **PASS / FAIL / N/A** with `file:line` evidence and confidence (>80% required to assert a FAIL; <60% → "insufficient evidence", verify before grading). Overall verdict is **PASS only if ZERO universal-rule FAILs**.
333
334- **PASS** → report the evidence table; if idempotency/count tests are absent, suggest `$integration-test`.
335- **FAIL** → list each violation with the responsible `file:line` and the correct pattern (from the [Anti-Patterns](#anti-patterns) table / reference doc). Route the fix back through **Generate mode** (Phase 0 → "Fix broken"); after the fix lands, RE-RUN `--mode=review` over the changed code. NEVER edit the seeder inside review mode.
336
337### Review-mode task plan (task tracking — required)
338
3391. Resolve the review target (prompt → current changes → work-context).
3402. Read `seed-test-data-reference.md` + project-config `Data Seeders` group + Universal Rules + the target file(s).
3413. Discover/confirm base class, env-gate key, count key with `file:line` evidence.
3424. Grade every universal + project-specific checklist item (R2).
3435. Produce the PASS/FAIL verdict with per-item `file:line` evidence (R3).
3446. If FAIL → hand confirmed defects to Generate mode and re-review after the fix; else report PASS + next-step suggestion.
3457. Analyze AI mistakes & lessons learned.
346
347---
348
349## Workflow Recommendation
350
351> **MUST ATTENTION — NOT IN WORKFLOW YET:** Use ask the user directly:
352>
353> 1. **Activate `workflow-seed-test-data`** (Recommended) — scout → investigate → seed-test-data → changes-review → code-simplifier → docs-update
354> 2. **Execute `$seed-test-data` directly** — run this skill standalone
355
356---
357
358## Next Steps
359
360> **MUST ATTENTION** after completing (Generate mode): use ask the user directly — do NOT skip. Step 7 self-review (`--mode=review`) MUST have run on the changed code BEFORE these:
361
362- **"$workflow-review-changes (Recommended)"** — final step: review all changes before commit (runs AFTER the `--mode=review` convention self-audit)
363- **"$integration-test"** — write tests verifying idempotency and count compliance
364- **"Skip, continue manually"** — user decides
365
366---
367
368> **[IMPORTANT]** task tracking for ALL tasks BEFORE starting. For simple tasks, ask user whether to skip.
369
370<!-- SYNC:critical-thinking-mindset -->
371
372> **Critical Thinking Mindset** — Apply critical thinking, sequential thinking. Every claim needs traced proof, confidence >80% to act.
373> **Anti-hallucination:** Never present guess as fact — cite sources for every claim, admit uncertainty freely, self-check output for errors, cross-reference independently, stay skeptical of own confidence — certainty without evidence root of all hallucination.
374
375<!-- /SYNC:critical-thinking-mindset -->
376
377<!-- SYNC:understand-code-first -->
378
379> **Understand Code First** — HARD-GATE: Do NOT write, plan, or fix until you READ existing code.
380>
381> 1. Search 3+ similar patterns (`grep`/`glob`) — cite `file:line` evidence
382> 2. Read existing files in target area — understand structure, base classes, conventions
383> 3. Run `python .claude/scripts/code_graph trace <file> --direction both --json` when `.code-graph/graph.db` exists
384> 4. Map dependencies via `connections` or `callers_of` — know what depends on your target
385> 5. Write investigation to `.ai/workspace/analysis/` for non-trivial tasks (3+ files)
386> 6. Re-read analysis file before implementing — never work from memory alone. — why: long context drifts from the file; the file is ground truth
387> 7. NEVER invent new patterns when existing ones work — match exactly or document deviation. — why: divergent patterns fragment the codebase and slow every future reader
388>
389> **BLOCKED until:** `- [ ]` Read target files `- [ ]` Grep 3+ patterns `- [ ]` Graph trace (if graph.db exists) `- [ ]` Assumptions verified with evidence
390
391<!-- /SYNC:understand-code-first -->
392
393<!-- SYNC:evidence-based-reasoning -->
394
395> **Evidence-Based Reasoning** — Speculation is FORBIDDEN. Every claim needs proof.
396>
397> 1. Cite `file:line`, grep results, or framework docs for EVERY claim
398> 2. Declare confidence: >80% act freely, 60-80% verify first, <60% DO NOT recommend
399> 3. Cross-service validation required for architectural changes
400> 4. "I don't have enough evidence" is valid and expected output
401>
402> **BLOCKED until:** `- [ ]` Evidence file path (`file:line`) `- [ ]` Grep search performed `- [ ]` 3+ similar patterns found `- [ ]` Confidence level stated
403>
404> **Forbidden without proof:** "obviously", "I think", "should be", "probably", "this is because"
405> **If incomplete →** output: `"Insufficient evidence. Verified: [...]. Not verified: [...]."`
406
407<!-- /SYNC:evidence-based-reasoning -->
408
409<!-- SYNC:ai-mistake-prevention -->
410
411> **AI Mistake Prevention** — Failure modes to avoid on every task:
412>
413> **Re-read files after context changes.** Context compaction, resume, or long-running work can make memory stale; verify current files before acting.
414> **Verify generated content against source evidence.** AI hallucinates APIs, names, claims, and document facts. Check the relevant source before documenting or referencing.
415> **Check downstream references before deleting or renaming.** Removing an artifact can stale docs, generated mirrors, configs, and callers; map references first.
416> **Trace the full impact chain after edits.** Changing a definition can miss derived outputs and consumers. Follow the affected chain before declaring done.
417> **Verify ALL affected outputs, not just the first.** One green check is not all green checks; validate every output surface the change can affect.
418> **Assume existing values are intentional — ask WHY before changing OR flagging one as a defect.** Before changing or reporting a constant, limit, flag, cutoff, wording, or pattern, read nearby context and history, the CALLER's ordering, and 2+ sibling call sites of the same convention. A doc stating WHAT without WHY is missing rationale, not proof of a missing guard.
419> **Surface ambiguity before acting — don't pick silently.** Multiple valid interpretations require an explicit question or stated assumption with risk.
420> **Assert the outcome your system owns, not the intermediate state your infrastructure owns.** When verifying async work, assert the final business state — never the delivery/retry bookkeeping held in shared infrastructure that any co-running process can write. Such a check passes when run alone and flakes the moment anything else shares that infrastructure.
421> **Keep shared guidance role-relevant.** Universal guidance must help every receiving skill or agent; code-specific obligations belong only in code-specific protocols.
422
423<!-- /SYNC:ai-mistake-prevention -->
424
425<!-- SYNC:real-world-fidelity-testing -->
426
427> **Real-World Fidelity Gate** — MANDATORY when authoring, reviewing, or repairing any integration / E2E / system test.
428>
429> A test earns trust by reproducing a situation the system can actually meet in production. A scenario that could never occur in real life proves nothing when it passes, and wastes hours when it fails.
430>
431> 1. **Ask the fidelity question BEFORE writing the setup:** _"Can this sequence, timing, and data actually occur in production?"_ If no, the test is mis-specified — fix the SCENARIO, never the assertion.
432> 2. **Model real pacing between actor steps.** Two distinct actor actions that production separates by seconds, minutes, or hours MUST NOT be fired back-to-back in the same millisecond. Compressed pacing manufactures races the system was never designed to survive, then reports them as product defects.
433> 3. **Wait on a real signal, never a blind sleep.** Find an observable proving the prior step finished — a persisted state change, an audit/version stamp, a queue/worker idle marker, a completion event — and poll until it settles (unchanged across a short stability window). Use a fixed delay ONLY when no observable exists, and say so in a comment.
434> 4. **Barriers belong in ARRANGE, never in ASSERT.** Waiting for a precondition is fidelity. Widening an assertion's timeout, loosening a comparison, adding a retry around a failing assertion, or skipping the test is masking. NEVER do the latter to force green.
435> 5. **Distinguish harness-amplified from real.** Test topologies (shared infra, fan-out consumers, parallel suites, cold starts) can make a rare production race routine locally. Before filing a product defect, state whether the trigger exists in production and at what likelihood.
436> 6. **Keep the protected invariant intact.** Improving fidelity must NEVER reduce what the test protects. If a realistic scenario no longer exercises the rule, the rule needs a DIFFERENT realistic scenario — not a weaker assertion.
437> 7. **Deliberate impossible-state tests are allowed, but MUST be labelled.** Corruption-repair, migration, and fail-safe tests intentionally construct states production should never reach; comment WHY the state is reachable (upstream bug, partial write, legacy data), so they are never confused with unrealistic setups.
438
439<!-- /SYNC:real-world-fidelity-testing -->
440
441<!-- SYNC:understand-code-first:reminder -->
442
443**IMPORTANT MUST ATTENTION** search 3+ existing patterns and read code BEFORE writing any seeder.
444
445<!-- /SYNC:understand-code-first:reminder -->
446
447<!-- SYNC:evidence-based-reasoning:reminder -->
448
449**MUST ATTENTION** cite `file:line` for every claim; declare confidence; "I don't have enough evidence" is valid output.
450
451<!-- /SYNC:evidence-based-reasoning:reminder -->
452
453<!-- SYNC:critical-thinking-mindset:reminder -->
454
455**MUST ATTENTION** apply critical + sequential thinking — every claim needs appropriate traced evidence (`file:line` for repo/code claims; source URL or artifact section for research, product, content, and docs claims); confidence >80% to act, <60% DO NOT recommend. Anti-hallucination: never present guess as fact, admit uncertainty freely, cross-reference independently, stay skeptical of own confidence.
456
457<!-- /SYNC:critical-thinking-mindset:reminder -->
458
459<!-- SYNC:ai-mistake-prevention:reminder -->
460
461**MUST ATTENTION** apply AI mistake prevention — verify generated content against evidence, trace downstream references before deleting or renaming, verify all affected outputs, re-read files after context loss, and surface ambiguity before acting.
462
463<!-- /SYNC:ai-mistake-prevention:reminder -->
464
465<!-- PROMPT-ENHANCE:STEP-TASK-CLOSING:START -->
466
467## Prompt-Enhance Closing Anchors
468
469**IMPORTANT MUST ATTENTION** follow declared step order for this skill; NEVER skip, reorder, or merge steps without explicit user approval
470**IMPORTANT MUST ATTENTION** for every step/sub-skill call: set `in_progress` before execution, set `completed` after execution
471**IMPORTANT MUST ATTENTION** every skipped step MUST include explicit reason; every completed step MUST include concise evidence
472**IMPORTANT MUST ATTENTION** if Task tools unavailable, maintain an equivalent step-by-step plan tracker with synchronized statuses
473
474<!-- PROMPT-ENHANCE:STEP-TASK-CLOSING:END -->
475
476<!-- SYNC:parallel-subagent-dispatch -->
477
478> **Parallel Sub-Agent Dispatch** — Plan parallelism the moment a task breakdown exists, BEFORE executing it — running provably independent tasks sequentially wastes wall-clock. Applies to every multi-step job: workflow steps, planning, batch updates, investigation, research, scans, reviews, doc sync. **Plan execution is metadata-gated, NEVER default-parallel** — fan-out follows ONLY what the plan declares (`PAR`/`SEQ` tags + per-phase write set); an untagged plan runs sequentially — why: a derived write set cannot see cascade or generated writes.
479>
480> 1. **Tag every task `PAR` or `SEQ`.** `PAR` = inputs exclude every pending task's output AND write set disjoint from every other `PAR`. Else `SEQ` — MUST ATTENTION name the dependency forcing it.
481> 2. **Group `PAR` into waves.** No edge between members. Two writers of one file NEVER share a wave. Read-only work (search, investigation, review, research) parallelizes freely.
482> 3. **Declare before dispatch:** `Parallel plan: wave 1 = [...] · wave 2 = [...] · SEQ = [...] (reason)`.
483> 4. **Spawn each wave in ONE message** — every `spawn_agent` call in one response, NEVER dripped per turn. Route each task to its specialist (`.claude/skills/sh
484
485…(truncated)