Harness Planning
Implementation planning with atomic tasks, goal-backward must-haves, and complete executable instructions. Every task fits in one context window.
When to Use
- After a design spec is approved (output of harness-brainstorming) and implementation needs planning
- When starting a new feature or project needing structured task decomposition
- When
on_new_feature or on_project_init triggers fire and the work is non-trivial
- When resuming a stalled project that needs a fresh plan
- NOT for small tasks (under 15 minutes, single file — just do it)
- NOT for problem exploration (use harness-brainstorming)
- NOT when a plan exists and needs execution (use harness-execution)
Process
Iron Law
Every task in the plan must be completable in one context window (2-5 minutes). If a task is larger, split it.
A plan with vague tasks like "add validation" or "implement the service" is not a plan — it is a wish list. Every task must contain exact file paths, exact commands, and complete code snippets.
Rigor Levels
The rigorLevel is passed by autopilot (or set via --fast/--thorough flags). Default is standard.
| Phase |
fast |
standard (default) |
thorough |
| SCOPE |
No change. |
No change. |
No change. |
| KNOWLEDGE |
Skip entirely. |
Run detect; fix if gaps found. |
Run detect; fix if gaps found. |
| DECOMPOSE |
Skip skeleton. Full tasks directly after file map. |
Skeleton if tasks >= 8; full tasks if < 8. |
Always skeleton. Require approval before expanding. |
| SEQUENCE |
No change. |
No change. |
No change. |
| VALIDATE |
No change. |
No change. |
No change. |
The skeleton pass is the primary rigor lever. Fast mode goes straight to full detail. Thorough mode validates direction before investing tokens in expansion.
Argument Resolution
When invoked by autopilot (or with explicit arguments), resolve paths before starting:
- Session slug: If
session-slug argument provided, set {sessionDir} = .harness/sessions/<session-slug>/. Pass to gather_context({ session: "<session-slug>", include: ["state", "learnings", "handoff", "graph", "businessKnowledge", "sessions", "validation"] }). All handoff writes go to {sessionDir}/handoff.json.
- Spec path: If
spec-path argument provided, read spec from that path. Otherwise, discover from {sessionDir}/handoff.json (read upstream brainstorming output) or prompt the user.
- Rigor level: If
fast/thorough argument provided, use it. Otherwise default to standard.
When no arguments are provided (standalone invocation), discover spec from context or prompt. Global .harness/ paths used as fallback.
Phase 1: SCOPE — Derive Must-Haves from Goals
Work backward from the goal. Start with "what must be true when we are done?"
- State the goal. One sentence. What does the system do when this plan is complete?
1b. Load skill recommendations. After loading the spec, check for skill recommendations:
If docs/changes/<feature>/SKILLS.md exists alongside the spec: parse the Apply and Reference tiers. These inform task annotation in Phase 2.
If SKILLS.md is missing but a spec exists: run the advisor inline using advise_skills MCP tool to generate SKILLS.md.
If neither SKILLS.md nor a spec exists: emit a one-line note:
Note: No skill recommendations found. Run the advisor to discover
relevant design, framework, and knowledge skills:
harness advise-skills --spec-path <path>
Store the parsed skill list for use in Phase 2 task annotation.
Review prior decisions. Check decisions from the prior brainstorming session (loaded via sessions in gather_context). Do not re-decide what was already decided — build on those choices.
Derive observable truths. What can be observed (running a command, opening a browser, reading a file) that proves the goal is met? Be specific:
- BAD: "The API handles errors"
- GOOD: "GET /api/users/nonexistent returns 404 with
{ error: 'User not found' } body"
Derive required artifacts. For each truth, what files must exist? What functions? What tests pass? List exact file paths.
Identify key links. How do artifacts connect? What imports what? What calls what?
Apply YAGNI. For every artifact: "Is this required for an observable truth?" If not, cut it.
Surface uncertainties. Before proceeding to Phase 2, explicitly list what you do NOT know. For each uncertainty, classify it:
- Blocking: Cannot decompose tasks without resolving this. Escalate to user.
- Assumption: Can proceed with a stated assumption. Document it. If wrong, specific tasks will need revision.
- Deferrable: Does not affect task decomposition. Note for execution phase.
Format:
## Uncertainties
- [BLOCKING] How should the API handle partial failures? (Spec does not define.)
- [ASSUMPTION] Database supports transactions. (If not, Task 3 needs redesign.)
- [DEFERRABLE] Exact error message wording. (Can be finalized during implementation.)
Read-only constraint: Steps 1-6 above are research and analysis. Do not propose task structure, file organization, or implementation approaches during SCOPE. Record what must be true (observable truths) and what you do not know (uncertainties). Solutions belong in DECOMPOSE.
When scope is ambiguous, ask in plain text in your reply. Do NOT route this through emit_interaction, AskUserQuestion, or any tool. emit_interaction records the prompt but does not display it to the human — the client collapses the call to "Called harness" and the rendered text only returns to the model. AskUserQuestion is Claude-Code-only and caps headers at 12 chars / 4 options. Plain text in your own message is the only channel that reliably reaches the human across every tool (Claude Code, Cursor, Codex, Gemini CLI).
Present the choice as a markdown table so tradeoffs are scannable, state your recommendation, then STOP and wait for the human's reply:
### Decision needed: The spec mentions X but does not define behavior for Y. Should we:
| | A) Include Y in this plan | B) Defer Y to a follow-up plan | C) Update the spec first |
| ---------- | ------------------------------------------------------- | ---------------------------------------------------- | ----------------------------------------------------------------- |
| **Pros** | Complete feature in one pass; no follow-up coordination | Keeps current plan focused; ship sooner | Design is complete before planning; no surprises during execution |
| **Cons** | Increases scope and time; may delay delivery | Y remains unhandled; may need rework when Y is added | Blocks planning until spec is updated; extra round-trip |
| **Risk** | Medium | Low | Low |
| **Effort** | High | Low | Medium |
**Recommendation:** B) Defer Y to a follow-up plan (confidence: medium) — keeping the current plan focused reduces risk; Y can be addressed in a follow-up.
EARS Requirement Patterns
Use EARS (Easy Approach to Requirements Syntax) when writing observable truths. These patterns eliminate ambiguity via consistent grammatical structure.
| Pattern |
Template |
Use When |
| Ubiquitous |
The system shall [behavior]. |
Always applies, unconditionally |
| Event-driven |
When [trigger], the system shall [response]. |
Triggered by a specific event |
| State-driven |
While [state], the system shall [behavior]. |
Only during a certain state |
| Optional |
Where [feature is enabled], the system shall [behavior]. |
Gated by config or feature flag |
| Unwanted |
If [condition], then the system shall not [behavior]. |
Preventing undesirable behavior |
Worked Examples:
- Ubiquitous: "The system shall return JSON responses with
Content-Type: application/json header."
- Event-driven: "When a user submits an invalid form, the system shall display field-level error messages within 200ms."
- State-driven: "While the database connection is unavailable, the system shall serve cached responses and log reconnection attempts."
- Optional: "Where rate limiting is enabled, the system shall reject requests exceeding 100/minute per API key with HTTP 429."
- Unwanted: "If the request body exceeds 10MB, then the system shall not attempt to parse it — return HTTP 413 immediately."
Apply EARS for behavioral requirements, not structural checks (e.g., file existence does not need EARS framing).
Graph-Enhanced Context (when available)
When a knowledge graph exists at .harness/graph/, use graph queries for faster context:
query_graph — discover module dependencies for realistic task decomposition
get_impact — estimate which modules a feature touches
compute_blast_radius — simulate failure propagation from target files to understand scope
predict_failures — forecast which architectural constraints are at risk from planned changes, informing where extra test coverage or smaller tasks are needed
detect_anomalies — identify structural irregularities in the affected area before planning tasks around them
Fall back to file-based commands if no graph is available.
Intelligence Signals (when orchestrator is available)
If the orchestrator is running, request intelligence analysis via POST /api/analyze with the feature title/description before decomposing. The pipeline returns:
- SEL (Spec Enrichment) — affected systems and blast radius derived from the graph
- CML (Complexity Modeling) — structural, semantic, and historical complexity scores. Use
structuralComplexity > 0.7 to flag areas needing smaller, more cautious tasks.
- PESL (Pre-Execution Simulation) — simulated risk score. Use
riskScore > 0.6 to add extra checkpoints or split risky tasks further.
If no orchestrator, predict_failures and compute_blast_radius MCP tools provide equivalent directional signals.
Phase 1.5: KNOWLEDGE BASELINE — Materialize Domain Knowledge
Before decomposing into tasks, ensure domain knowledge from PRDs and specs is documented. Skip this phase when no PRDs, specs, or business domain documents exist in the project, or when rigor level is fast.
Run knowledge pipeline in detect mode. Execute harness knowledge-pipeline --domain <feature-domain> to produce a differential gap report comparing extracted business rules against documented knowledge in docs/knowledge/.
If gaps exist and --fix is appropriate, run harness knowledge-pipeline --fix --domain <feature-domain> to materialize docs/knowledge/{domain}/*.md files from extracted findings. This creates the knowledge baseline from PRDs before any tasks are written.
Cross-check uncertainties against materialized knowledge (from businessKnowledge loaded in gather_context and freshly materialized docs):
- Remove "assumptions" from the uncertainty list that are now documented facts in
docs/knowledge/
- Escalate if contradictions exist between PRDs and existing knowledge docs
- Use
business_fact nodes from the graph context to validate domain assumptions
Reference materialized knowledge in Phase 2 task decomposition. Tasks should reference specific knowledge docs they implement. Observable truths should map back to documented business rules. Use the businessKnowledge context (domains, tags, documented facts) loaded in Phase 1 to ground task instructions in verified domain knowledge rather than assumptions.
Phase 1.6: NFR ELICITATION — Turn Quality Targets into Verifiable Tasks
Non-functional requirements (NFRs) are quality targets — how fast, how safe, how far it scales, how it fails — that are cheapest to honor when treated as design inputs during planning, not as findings surfaced in review after the code is written. This phase elicits NFR targets across four dimensions (performance, security, scalability, resilience) and turns each stated target into a concrete plan task wired to machinery the harness already runs.
Additive by default. Elicitation is opt-in per dimension: every dimension offers a sensible default and an explicit skip. If the human skips every dimension, no NFR tasks are emitted and planning proceeds exactly as it would without this phase. Skip this phase entirely when rigor level is fast.
Elicitation protocol. Ask one dimension at a time, in plain text in your reply — do NOT route this through emit_interaction or AskUserQuestion (same channel constraint as the rest of this skill: neither reliably displays to the human, emit_interaction collapses to "Called harness" and AskUserQuestion is Claude-Code-only). For each dimension state the question, the default, and how to skip ("skip" / blank reply), then wait for the reply before asking the next. Keep answers plain-text and short.
| Dimension |
Prompt |
Default (if skipped) |
| Performance |
"Any latency/throughput target for a hot path? (e.g. p99 < 5ms for parseDocument)" |
No new benchmark; existing check-perf budgets stand. |
| Security |
"Any module handling untrusted input or secrets that must pass a clean scan?" |
check-security runs at its configured floor. |
| Scalability |
"Any target input size or concurrency the code must hold shape at? (e.g. 10k items)" |
No load-oriented benchmark; complexity budgets stand. |
| Resilience |
"Any failure mode that must degrade gracefully rather than crash? (e.g. DB down)" |
Failure paths covered by ordinary task tests only. |
Phrase each stated target as an EARS acceptance criterion (see the EARS table in Phase 1) and add it to the Observable Truths list so it traces to a task in Phase 2:
- Performance / Scalability → State-driven or Event-driven: "While processing
10k items, the system shall keep p99 latency under 200ms."
- Security → Unwanted: "If
check-security reports an error-severity finding in <module>, then the build shall not pass."
- Resilience → Unwanted / State-driven: "While the database is unreachable, the system shall serve cached data and shall not crash."
Wire each target to existing machinery. Do NOT invent new subsystems. Each dimension maps to a gate the harness already ships:
| Dimension |
Existing machinery |
Emitted plan task (template) |
Verifying command |
| Performance |
Perf baselines (.harness/perf/baselines.json) + harness perf bench |
"Add a *.bench.ts benchmark for <hot-path>; run harness perf bench; record the baseline with harness perf baselines update. Target: <p99 target>." |
harness perf bench + harness perf baselines show |
| Security |
Mechanical scan harness check-security (gate: FAIL on error-severity findings after the --severity filter) |
"harness check-security --severity error must report zero findings for the files under <module>. If the module warrants a stricter floor, set security.strict in harness.config.json." |
harness check-security --severity error |
| Scalability |
Perf benchmarks at target load + structural/coupling budgets (harness check-perf) |
"Add a *.bench.ts that exercises <component> at <target load>; assert it stays within budget. Keep complexity within budget via harness check-perf --structural --coupling." |
harness perf bench + harness check-perf --structural |
| Resilience |
Ordinary TDD task with a failure-path test + graph failure signals (predict_failures, compute_blast_radius) |
"Write a test that simulates <failure mode> and asserts graceful degradation (no crash, documented fallback). Use predict_failures to find where else this failure propagates." |
The failure-path test (npx vitest run <test-file>) |
Record the elicited targets in the constraints session section via manage_state — they are constraints the plan must honor — and carry them into Phase 2 as tasks tagged category: "nfr".
Honest scope. Performance, security, and scalability wire to mechanical gates (harness perf / harness check-perf / harness check-security) that pass or fail deterministically. Resilience has no dedicated scanner — it wires to the plan's own failure-path tests plus the graph failure-prediction signals this skill already uses. Do not claim a resilience "gate" the harness does not have; the verifiable artifact is the test.
Phase 2: DECOMPOSE — Map File Structure and Create Tasks
Report progress: **[Phase 2/4]** DECOMPOSE — mapping file structure and creating tasks
Map the file structure first. List every file to create or modify before writing tasks:
CREATE src/services/notification-service.ts
CREATE src/services/notification-service.test.ts
MODIFY src/services/index.ts (add export)
CREATE src/types/notification.ts
MODIFY src/api/routes/users.ts (add notification trigger)
Skeleton pass (rigor-gated). Lightweight skeleton (~200 tokens) validates direction before full expansion. Gating per Rigor Levels table.
Format: Numbered logical groups with task count and time. No file paths, code, or details.
1. Foundation types and interfaces (~3 tasks, ~10 min)
2. Core scoring module with TDD (~2 tasks, ~8 min)
3. CLI integration and flag parsing (~4 tasks, ~15 min)
**Estimated total:** 8 tasks, ~33 minutes
Approval gate: Ask in plain text in your reply — do NOT route this through emit_interaction or AskUserQuestion (they do not display to the human; emit_interaction collapses to "Called harness" and AskUserQuestion is Claude-Code-only). Ask directly: "Approve skeleton direction?" and wait. If approved, proceed to step 3. If rejected, revise and re-present.
Decompose into atomic tasks. Each task must:
- Be completable in 2-5 minutes, fit in a single context window
- Have a clear, testable outcome
- Follow TDD: write test, fail, implement, pass, commit
- Produce one atomic commit
Write complete instructions for each task. Not summaries — complete executable instructions:
- Exact file paths to create or modify
- Exact code to write (not "add validation logic" — write the actual code)
- Exact test commands (e.g.,
npx vitest run src/services/notification-service.test.ts)
- Exact commit message
harness validate as the final step
Skill annotations. If skill recommendations were loaded in Phase 1, annotate each task with relevant skills from the Apply and Reference tiers:
### Task 3: Implement dark mode toggle
**Skills:** `design-dark-mode` (apply), `a11y-color-contrast` (reference)
Match skills to tasks based on keyword and domain overlap between the task description and the skill's purpose/keywords. Only annotate when the match is relevant to the specific task.
Include checkpoints. Mark tasks requiring human input:
[checkpoint:human-verify] — Pause, show result, wait for confirmation
[checkpoint:decision] — Pause, present options, wait for choice
[checkpoint:human-action] — Pause, instruct human on required action
Derive integration tasks from the spec's Integration Points section. If the spec contains an Integration Points section, create tasks for each non-empty integration point. Skip subsections marked "None" — do not derive tasks from them. Integration tasks are normal plan tasks but tagged with category: "integration" in their description. They appear at the end of the task list, after all implementation tasks.
For each subsection of Integration Points, derive tasks:
| Integration Point |
Example Derived Task |
| Entry Points: "New CLI command" |
"Regenerate barrel exports. Verify new command appears in _registry.ts." |
| Registrations Required: "Skill at tier 2" |
"Add skill to tier list in AGENTS.md. Generate slash commands." |
| Documentation Updates: "AGENTS.md capabilities" |
"Update AGENTS.md to describe the feature." |
| Architectural Decisions: "ADR for approach X" |
"Write ADR docs/knowledge/decisions/NNNN-<slug>.md." |
| Knowledge Impact: "Domain concept Y" |
"Enrich knowledge graph with concept node." |
Integration tasks follow the same atomic task rules (2-5 minutes, exact file paths, exact code). Use the **Category:** integration tag in the task header, e.g.:
### Task N: Update AGENTS.md with new feature description
**Depends on:** Task N-1 | **Files:** `AGENTS.md` | **Category:** integration
If the spec has no Integration Points section, skip this step.
Emit NFR tasks from Phase 1.6. For each NFR target elicited in Phase 1.6 (skipped dimensions emit nothing), create one atomic task using the template from the NFR wiring table. Tag it **Category:** nfr in the task header. NFR tasks follow the same atomic rules (2-5 minutes, exact file paths, exact commands) and appear after implementation tasks, alongside integration tasks. Each NFR task's final step is its verifying command from the table, so the target is checkable at execution time rather than aspirational.
### Task N: Lock in p99 budget for parseDocument
**Depends on:** Task N-1 | **Files:** `src/parse.bench.ts` | **Category:** nfr
1. Create `src/parse.bench.ts` benchmarking `parseDocument`.
2. Run: `harness perf bench`
3. Record baseline: `harness perf baselines update`
4. Verify `p99Ms` in `.harness/perf/baselines.json` meets the target (`p99 < 5ms`).
5. Commit: `perf(parse): add parseDocument benchmark and baseline`
If Phase 1.6 was skipped, or every dimension was skipped, emit no NFR tasks — the plan is identical to one produced without NFR elicitation.
Phase 3: SEQUENCE — Order Tasks and Identify Dependencies
- Order by dependency. Types before implementations. Implementations before integrations. Integration tasks (tagged
category: "integration") after all implementation tasks. Tests alongside implementations (same task, TDD style).
- Identify parallel opportunities and record dependency edges. Tasks touching different subsystems with no shared state can run in parallel. Record each task's real dependencies as
dependsOn (the task IDs it must follow) in the task header **Depends on:** line, and keep the task's **Files:** list accurate — together these are exactly the edges plan_parallelization consumes (explicit dependsOn unioned with file-overlap edges) to build the wave DAG at execution time. A task with no dependencies records **Depends on:** none. Optionally, add an **Owns:** line listing glob paths a task claims ownership of (e.g. src/api/**); it feeds the cheap, deterministic, graph-free owns-overlap forecast in plan_parallelization — two tasks whose owned globs overlap are flagged in ownershipForecast and gain an implicit DAG edge, alongside the file-overlap edges. Owns is OPTIONAL: a task with no distinct ownership just omits it.
- Number tasks sequentially. Use
Task 1, Task 2, etc. Dependencies reference task numbers.
- Estimate total time. Sum 2-5 minutes per task. If total exceeds available time, identify a milestone boundary for pausing.
Phase 4: VALIDATE — Review and Finalize the Plan
Verify completeness. Every observable truth from Phase 1 must trace to specific task(s) that deliver it.
Verify task sizing. Could an agent complete each task in one context window without exploring or deciding? If not, split it.
Verify TDD compliance. Every code-producing task must include a test step. No "write tests later."
Run harness validate to verify project health before writing the plan.
Check failures log. Read .harness/failures.md. If planned approaches match known failures, flag them.
Run soundness review. Invoke harness-soundness-review --mode plan against the draft. Do not proceed until the review converges with no remaining issues.
Write the plan to docs/changes/<topic>/plans/. Naming: YYYY-MM-DD-<feature-name>-plan.md. Resolve <topic> from the spec path — if the spec lives at docs/changes/<topic>/proposal.md, the plan goes in the sibling plans/ directory. If the spec is not under docs/changes/, fall back to docs/plans/ and flag the spec location for human review. Create directories as needed.
Commit plan artifact. Immediately after writing the plan, commit it so the paper trail enters git history at planning time, not after execution:
git add docs/changes/<topic>/plans/YYYY-MM-DD-<feature-name>-plan.md
git commit -m "docs(<topic>): add plan"
Use docs/plans/ if the spec is not under docs/changes/<topic>/. Do not skip — harness-execution commits only implementation files, so a plan uncommitted here stays untracked until someone notices. If a pre-commit hook reformats the file, re-add and re-commit.
Write handoff. Write to the session-scoped path when session slug is known, otherwise fall back to global path:
- Session-scoped (preferred):
.harness/sessions/<session-slug>/handoff.json
- Global (fallback, deprecated):
.harness/handoff.json
[DEPRECATED] Writing to .harness/handoff.json is deprecated. In autopilot sessions, always use .harness/sessions/<slug>/handoff.json to prevent cross-session contamination.
Fields: fromSkill, phase, summary, completed, pending, concerns, decisions, contextKeywords.
Write session summary (if session is known). Call writeSessionSummary with skill, status, plan path, keyContext, nextStep. Skip if no session slug.
Request plan sign-off in plain text. Ask directly in your reply — do NOT route this through emit_interaction or AskUserQuestion (the human will not see it; emit_interaction collapses to "Called harness" and AskUserQuestion is Claude-Code-only). Present the plan path, task count, and time estimate, then ask: "Proceed? (yes/no)" and wait for the reply.
Suggest transition to execution. After approval, call emit_interaction with type: transition, completedPhase: "planning", suggestedNext: "execution", requiresConfirmation: true. Include qualityGate with checks: plan-written, harness-validate, observable-truths-traced, human-approved. If confirmed: invoke harness-execution. If declined: stop (handoff already written).
Plan Document Structure
# Plan: <Feature Name>
**Date:** YYYY-MM-DD | **Spec:** (if applicable) | **Tasks:** N | **Time:** N min | **Integration Tier:** small | medium | large
## Goal
One sentence.
## Observable Truths (Acceptance Criteria)
1. [observable truth]
## NFR Targets (if elicited)
- [Performance] EARS criterion → task N (verified by `harness perf bench`)
- [Security] EARS criterion → task N (verified by `harness check-security --severity error`)
_Omit this section entirely when no NFR dimension was elicited._
## File Map
- CREATE path/to/file.ts
- MODIFY path/to/other-file.ts
## Skeleton (if produced)
1. <group name> (~N tasks, ~N min)
_Skeleton approved: yes/no._
## Tasks
### Task 1: <descriptive name>
**Depends on:** none | **Files:** path/to/file.ts, path/to/file.test.ts | **Owns:** path/to/\*\*
<!-- `Depends on` lists the task IDs this task must follow (its `dependsOn` edges); `Files` is the independence-checking input; `Owns` (optional) lists glob paths this task claims, feeding the deterministic owns-overlap forecast. All three feed `plan_parallelization` during execution. -->
1. Create test file with exact test code
2. Run test — observe failure
3. Create implementation with exact code
4. Run test — observe pass
5. Run: `harness validate`
6. Commit: `feat(scope): descriptive message`
### Task 2: <descriptive name>
[checkpoint:human-verify] ...
Integration Tier Heuristics
When a spec contains an Integration Points section, set the plan's integrationTier field based on scope:
| Tier |
Signal |
Integration Requirements |
| small |
Bug fix, config change, < 3 files, no new exports |
Wiring checks only (defaults always run) |
| medium |
New feature within existing package, new exports, 3-15 files |
Wiring + project updates (roadmap, changelog, graph enrichment) |
| large |
New package, new skill, new public API surface, architectural change |
Wiring + project updates + knowledge materialization (ADRs, doc updates) |
If the spec has no Integration Points section, omit the integrationTier field from the plan header.
Session State
| Section |
Read |
Write |
Purpose |
| terminology |
yes |
no |
Consistent language in plan |
| decisions |
yes |
yes |
Brainstorming decisions; planning-phase decisions |
| constraints |
yes |
yes |
Existing constraints; constraints discovered during decomposition |
| risks |
yes |
yes |
Existing risks; implementation risks from task design |
| openQuestions |
yes |
yes |
Unresolved questions; new questions; resolve answered ones |
| evidence |
yes |
yes |
Prior evidence; file:line citations for task specs |
When to write: Phase 1 — constraints and risks. Phase 2 — decisions about task structure. Phase 4 — resolve questions.
When to read: Start of Phase 1 via gather_context with include: ["state", "learnings", "handoff", "graph", "businessKnowledge", "sessions", "validation"] to inherit brainstorming context and load documented business knowledge.
Evidence Requirements
When referencing existing code in task specs, cite evidence using file:line format, code pattern references, or test output. Write to evidence session section via manage_state.
When to cite: Phase 1 (existing files), Phase 2 (file paths and patterns), file map (existing files for modification).
Uncited claims: Prefix with [UNVERIFIED].
Harness Integration
harness validate — Run in Phase 4 (before writing plan) and included in every task.
harness check-deps — Referenced in tasks adding imports or creating modules.
harness perf bench / harness perf baselines update / harness check-perf — The verifying commands for performance and scalability NFR tasks emitted from Phase 1.6. Perf baselines live in .harness/perf/baselines.json; complexity/coupling/size budgets live in the performance block of harness.config.json.
harness check-security — The verifying command for security NFR tasks emitted from Phase 1.6. Its gate fails on error-severity findings after the --severity filter; security.strict in harness.config.json promotes warnings to errors for stricter modules.
- Plan commit — After writing the plan (Phase 4 Step 8), commit
docs/changes/<topic>/plans/YYYY-MM-DD-<feature-name>-plan.md so the paper trail enters git history at planning time. harness-execution does not backfill this commit.
- Plan location —
docs/changes/<topic>/plans/YYYY-MM-DD-<feature-name>-plan.md when the spec lives under docs/changes/<topic>/proposal.md; otherwise docs/plans/ as a fallback.
- Handoff — Once approved, invoke harness-execution for task-by-task implementation.
- Session directory — Session-scoped writes go to
.harness/sessions/<slug>/. Structure: handoff.json, state.json, artifacts.json (registry of spec/plan paths and produced file lists). Global .harness/handoff.json is deprecated for session-aware invocations.
emit_interaction — Call at end of Phase 4 to suggest transitioning to execution (confirmed transition).
- Rigor levels —
--fast/--thorough control skeleton pass. See Rigor Levels table.
- Two-pass planning — Skeleton (~200 tokens) before full expansion. Catches directional errors early.
Change Specifications
When planning changes to existing functionality (not greenfield), express requirements as deltas:
- [ADDED] — New behavior that does not exist today
- [MODIFIED] — Existing behavior that changes
- [REMOVED] — Existing behavior that goes away
Example:
## Changes to User Authentication
- [ADDED] OAuth2 refresh tokens with 7-day expiry
- [MODIFIED] Login endpoint returns `refreshToken` alongside `accessToken`
- [MODIFIED] Token validation accepts both JWT and OAuth2 tokens
- [REMOVED] Legacy API key authentication (deprecated in v2.1)
Only apply when modifying existing documented behavior. When docs/changes/ exists, produce docs/changes/<feature>/delta.md alongside the task plan.
Success Criteria
- Plan document exists at the resolved location (
docs/changes/<topic>/plans/ or docs/plans/ fallback) with all required sections
- Every task completable in 2-5 minutes (one context window)
- Every task includes exact file paths, exact code, and exact commands
- Every code-producing task follows TDD: test first, fail, implement, pass
- Observable truths trace to specific tasks
- File map lists every file to create or modify
- Checkpoints marked where human input is required
harness validate passes before plan is written and is in every task
- Human has reviewed and approved the plan
- Rigor level rules followed: fast skips skeleton; thorough always skeletons with approval; standard skeletons at >= 8 tasks
- NFR targets elicited in Phase 1.6 each trace to a task tagged
category: nfr whose final step is a verifying command, or the dimension was explicitly skipped (no NFR task emitted)
Red Flags
| Flag |
Corrective Action |
| "I know the implementation well enough to skip reading the spec" |
STOP. Phase 1 SCOPE starts by reading the spec. Assumptions about spec content lead to plans that implement the wrong thing. |
| "This task is self-explanatory, no need for exact file paths and commands" |
STOP. Iron Law: every task must contain exact file paths, exact commands, and complete code snippets. "Implement the service" is a wish, not a task. |
| "I'll plan the happy path now and add error handling tasks later" |
STOP. Error handling is not optional. The spec's success criteria include error scenarios. Plan them alongside the happy path. |
// detailed steps TBD or // expand during execution in task descriptions |
STOP. A task that defers detail to execution is a vague task. If you cannot write the exact steps now, you do not understand the task well enough to plan it. |
Rationalizations to Reject
| Rationalization |
Reality |
| "The task is conceptually clear so I do not need to include exact code in the plan" |
Every task must have exact file paths, exact code, and exact commands. If you cannot write the code in the plan, you do not understand the task well enough to plan it. |
| "This task touches 5 files but it is logically one unit of work, so splitting it would add overhead" |
Tasks touching more than 3 files must be split. The overhead of splitting is far less than the cost of a failed oversized task. |
…(truncated)
1---2name: harness-planning3description: Harness Planning4---5# Harness Planning67> Implementation planning with atomic tasks, goal-backward must-haves, and complete executable instructions. Every task fits in one context window.89## When to Use1011- After a design spec is approved (output of harness-brainstorming) and implementation needs planning12- When starting a new feature or project needing structured task decomposition13- When `on_new_feature` or `on_project_init` triggers fire and the work is non-trivial14- When resuming a stalled project that needs a fresh plan15- NOT for small tasks (under 15 minutes, single file — just do it)16- NOT for problem exploration (use harness-brainstorming)17- NOT when a plan exists and needs execution (use harness-execution)1819## Process2021### Iron Law2223**Every task in the plan must be completable in one context window (2-5 minutes). If a task is larger, split it.**2425A plan with vague tasks like "add validation" or "implement the service" is not a plan — it is a wish list. Every task must contain exact file paths, exact commands, and complete code snippets.2627---2829### Rigor Levels3031The `rigorLevel` is passed by autopilot (or set via `--fast`/`--thorough` flags). Default is `standard`.3233| Phase | `fast` | `standard` (default) | `thorough` |34| --------- | -------------------------------------------------- | ------------------------------------------ | --------------------------------------------------- |35| SCOPE | No change. | No change. | No change. |36| KNOWLEDGE | Skip entirely. | Run detect; fix if gaps found. | Run detect; fix if gaps found. |37| DECOMPOSE | Skip skeleton. Full tasks directly after file map. | Skeleton if tasks >= 8; full tasks if < 8. | Always skeleton. Require approval before expanding. |38| SEQUENCE | No change. | No change. | No change. |39| VALIDATE | No change. | No change. | No change. |4041The skeleton pass is the primary rigor lever. Fast mode goes straight to full detail. Thorough mode validates direction before investing tokens in expansion.4243---4445### Argument Resolution4647When invoked by autopilot (or with explicit arguments), resolve paths before starting:48491. **Session slug:** If `session-slug` argument provided, set `{sessionDir} = .harness/sessions/<session-slug>/`. Pass to `gather_context({ session: "<session-slug>", include: ["state", "learnings", "handoff", "graph", "businessKnowledge", "sessions", "validation"] })`. All handoff writes go to `{sessionDir}/handoff.json`.502. **Spec path:** If `spec-path` argument provided, read spec from that path. Otherwise, discover from `{sessionDir}/handoff.json` (read upstream brainstorming output) or prompt the user.513. **Rigor level:** If `fast`/`thorough` argument provided, use it. Otherwise default to `standard`.5253When no arguments are provided (standalone invocation), discover spec from context or prompt. Global `.harness/` paths used as fallback.5455---5657### Phase 1: SCOPE — Derive Must-Haves from Goals5859Work backward from the goal. Start with "what must be true when we are done?"60611. **State the goal.** One sentence. What does the system do when this plan is complete?62631b. **Load skill recommendations.** After loading the spec, check for skill recommendations:6465- If `docs/changes/<feature>/SKILLS.md` exists alongside the spec: parse the Apply and Reference tiers. These inform task annotation in Phase 2.66- If SKILLS.md is missing but a spec exists: run the advisor inline using `advise_skills` MCP tool to generate SKILLS.md.67- If neither SKILLS.md nor a spec exists: emit a one-line note:6869 ```70 Note: No skill recommendations found. Run the advisor to discover71 relevant design, framework, and knowledge skills:72 harness advise-skills --spec-path <path>73 ```7475Store the parsed skill list for use in Phase 2 task annotation.76772. **Review prior decisions.** Check `decisions` from the prior brainstorming session (loaded via `sessions` in gather_context). Do not re-decide what was already decided — build on those choices.78793. **Derive observable truths.** What can be observed (running a command, opening a browser, reading a file) that proves the goal is met? Be specific:80 - BAD: "The API handles errors"81 - GOOD: "GET /api/users/nonexistent returns 404 with `{ error: 'User not found' }` body"824. **Derive required artifacts.** For each truth, what files must exist? What functions? What tests pass? List exact file paths.835. **Identify key links.** How do artifacts connect? What imports what? What calls what?846. **Apply YAGNI.** For every artifact: "Is this required for an observable truth?" If not, cut it.85867. **Surface uncertainties.** Before proceeding to Phase 2, explicitly list what you do NOT know. For each uncertainty, classify it:87 - **Blocking:** Cannot decompose tasks without resolving this. Escalate to user.88 - **Assumption:** Can proceed with a stated assumption. Document it. If wrong, specific tasks will need revision.89 - **Deferrable:** Does not affect task decomposition. Note for execution phase.9091 Format:9293 ```94 ## Uncertainties95 - [BLOCKING] How should the API handle partial failures? (Spec does not define.)96 - [ASSUMPTION] Database supports transactions. (If not, Task 3 needs redesign.)97 - [DEFERRABLE] Exact error message wording. (Can be finalized during implementation.)98 ```99100 **Read-only constraint:** Steps 1-6 above are research and analysis. Do not propose task structure, file organization, or implementation approaches during SCOPE. Record what must be true (observable truths) and what you do not know (uncertainties). Solutions belong in DECOMPOSE.101102 When scope is ambiguous, ask in plain text in your reply. **Do NOT route this through `emit_interaction`, `AskUserQuestion`, or any tool.** `emit_interaction` records the prompt but does not display it to the human — the client collapses the call to "Called harness" and the rendered text only returns to the model. `AskUserQuestion` is Claude-Code-only and caps headers at 12 chars / 4 options. Plain text in your own message is the only channel that reliably reaches the human across every tool (Claude Code, Cursor, Codex, Gemini CLI).103104 Present the choice as a markdown table so tradeoffs are scannable, state your recommendation, then STOP and wait for the human's reply:105106 ```markdown107 ### Decision needed: The spec mentions X but does not define behavior for Y. Should we:108109 | | A) Include Y in this plan | B) Defer Y to a follow-up plan | C) Update the spec first |110 | ---------- | ------------------------------------------------------- | ---------------------------------------------------- | ----------------------------------------------------------------- |111 | **Pros** | Complete feature in one pass; no follow-up coordination | Keeps current plan focused; ship sooner | Design is complete before planning; no surprises during execution |112 | **Cons** | Increases scope and time; may delay delivery | Y remains unhandled; may need rework when Y is added | Blocks planning until spec is updated; extra round-trip |113 | **Risk** | Medium | Low | Low |114 | **Effort** | High | Low | Medium |115116 **Recommendation:** B) Defer Y to a follow-up plan (confidence: medium) — keeping the current plan focused reduces risk; Y can be addressed in a follow-up.117 ```118119#### EARS Requirement Patterns120121Use EARS (Easy Approach to Requirements Syntax) when writing observable truths. These patterns eliminate ambiguity via consistent grammatical structure.122123| Pattern | Template | Use When |124| ---------------- | -------------------------------------------------------- | ------------------------------- |125| **Ubiquitous** | The system shall [behavior]. | Always applies, unconditionally |126| **Event-driven** | When [trigger], the system shall [response]. | Triggered by a specific event |127| **State-driven** | While [state], the system shall [behavior]. | Only during a certain state |128| **Optional** | Where [feature is enabled], the system shall [behavior]. | Gated by config or feature flag |129| **Unwanted** | If [condition], then the system shall not [behavior]. | Preventing undesirable behavior |130131**Worked Examples:**1321331. **Ubiquitous:** "The system shall return JSON responses with `Content-Type: application/json` header."1342. **Event-driven:** "When a user submits an invalid form, the system shall display field-level error messages within 200ms."1353. **State-driven:** "While the database connection is unavailable, the system shall serve cached responses and log reconnection attempts."1364. **Optional:** "Where rate limiting is enabled, the system shall reject requests exceeding 100/minute per API key with HTTP 429."1375. **Unwanted:** "If the request body exceeds 10MB, then the system shall not attempt to parse it — return HTTP 413 immediately."138139Apply EARS for behavioral requirements, not structural checks (e.g., file existence does not need EARS framing).140141### Graph-Enhanced Context (when available)142143When a knowledge graph exists at `.harness/graph/`, use graph queries for faster context:144145- `query_graph` — discover module dependencies for realistic task decomposition146- `get_impact` — estimate which modules a feature touches147- `compute_blast_radius` — simulate failure propagation from target files to understand scope148- `predict_failures` — forecast which architectural constraints are at risk from planned changes, informing where extra test coverage or smaller tasks are needed149- `detect_anomalies` — identify structural irregularities in the affected area before planning tasks around them150151Fall back to file-based commands if no graph is available.152153### Intelligence Signals (when orchestrator is available)154155If the orchestrator is running, request intelligence analysis via `POST /api/analyze` with the feature title/description before decomposing. The pipeline returns:156157- **SEL** (Spec Enrichment) — affected systems and blast radius derived from the graph158- **CML** (Complexity Modeling) — structural, semantic, and historical complexity scores. Use `structuralComplexity > 0.7` to flag areas needing smaller, more cautious tasks.159- **PESL** (Pre-Execution Simulation) — simulated risk score. Use `riskScore > 0.6` to add extra checkpoints or split risky tasks further.160161If no orchestrator, `predict_failures` and `compute_blast_radius` MCP tools provide equivalent directional signals.162163---164165### Phase 1.5: KNOWLEDGE BASELINE — Materialize Domain Knowledge166167Before decomposing into tasks, ensure domain knowledge from PRDs and specs is documented. **Skip this phase** when no PRDs, specs, or business domain documents exist in the project, or when rigor level is `fast`.1681691. **Run knowledge pipeline in detect mode.** Execute `harness knowledge-pipeline --domain <feature-domain>` to produce a differential gap report comparing extracted business rules against documented knowledge in `docs/knowledge/`.1701712. **If gaps exist and `--fix` is appropriate,** run `harness knowledge-pipeline --fix --domain <feature-domain>` to materialize `docs/knowledge/{domain}/*.md` files from extracted findings. This creates the knowledge baseline from PRDs before any tasks are written.1721733. **Cross-check uncertainties against materialized knowledge** (from `businessKnowledge` loaded in gather_context and freshly materialized docs):174 - Remove "assumptions" from the uncertainty list that are now documented facts in `docs/knowledge/`175 - Escalate if contradictions exist between PRDs and existing knowledge docs176 - Use `business_fact` nodes from the graph context to validate domain assumptions1771784. **Reference materialized knowledge in Phase 2 task decomposition.** Tasks should reference specific knowledge docs they implement. Observable truths should map back to documented business rules. Use the `businessKnowledge` context (domains, tags, documented facts) loaded in Phase 1 to ground task instructions in verified domain knowledge rather than assumptions.179180---181182### Phase 1.6: NFR ELICITATION — Turn Quality Targets into Verifiable Tasks183184Non-functional requirements (NFRs) are quality targets — how fast, how safe, how far it scales, how it fails — that are cheapest to honor when treated as **design inputs during planning**, not as findings surfaced in review after the code is written. This phase elicits NFR targets across four dimensions (performance, security, scalability, resilience) and turns each stated target into a concrete plan task wired to machinery the harness already runs.185186**Additive by default.** Elicitation is opt-in per dimension: every dimension offers a sensible default and an explicit skip. If the human skips every dimension, no NFR tasks are emitted and planning proceeds exactly as it would without this phase. **Skip this phase entirely** when rigor level is `fast`.187188**Elicitation protocol.** Ask one dimension at a time, in plain text in your reply — do NOT route this through `emit_interaction` or `AskUserQuestion` (same channel constraint as the rest of this skill: neither reliably displays to the human, `emit_interaction` collapses to "Called harness" and `AskUserQuestion` is Claude-Code-only). For each dimension state the question, the default, and how to skip ("skip" / blank reply), then wait for the reply before asking the next. Keep answers plain-text and short.189190| Dimension | Prompt | Default (if skipped) |191| --------------- | -------------------------------------------------------------------------------------- | ------------------------------------------------------ |192| **Performance** | "Any latency/throughput target for a hot path? (e.g. `p99 < 5ms for parseDocument`)" | No new benchmark; existing `check-perf` budgets stand. |193| **Security** | "Any module handling untrusted input or secrets that must pass a clean scan?" | `check-security` runs at its configured floor. |194| **Scalability** | "Any target input size or concurrency the code must hold shape at? (e.g. `10k items`)" | No load-oriented benchmark; complexity budgets stand. |195| **Resilience** | "Any failure mode that must degrade gracefully rather than crash? (e.g. `DB down`)" | Failure paths covered by ordinary task tests only. |196197**Phrase each stated target as an EARS acceptance criterion** (see the EARS table in Phase 1) and add it to the Observable Truths list so it traces to a task in Phase 2:198199- **Performance / Scalability** → State-driven or Event-driven: "While processing `10k` items, the system shall keep `p99` latency under `200ms`."200- **Security** → Unwanted: "If `check-security` reports an error-severity finding in `<module>`, then the build shall not pass."201- **Resilience** → Unwanted / State-driven: "While the database is unreachable, the system shall serve cached data and shall not crash."202203**Wire each target to existing machinery.** Do NOT invent new subsystems. Each dimension maps to a gate the harness already ships:204205| Dimension | Existing machinery | Emitted plan task (template) | Verifying command |206| --------------- | --------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------- |207| **Performance** | Perf baselines (`.harness/perf/baselines.json`) + `harness perf bench` | "Add a `*.bench.ts` benchmark for `<hot-path>`; run `harness perf bench`; record the baseline with `harness perf baselines update`. Target: `<p99 target>`." | `harness perf bench` + `harness perf baselines show` |208| **Security** | Mechanical scan `harness check-security` (gate: FAIL on error-severity findings after the `--severity` filter) | "`harness check-security --severity error` must report zero findings for the files under `<module>`. If the module warrants a stricter floor, set `security.strict` in `harness.config.json`." | `harness check-security --severity error` |209| **Scalability** | Perf benchmarks at target load + structural/coupling budgets (`harness check-perf`) | "Add a `*.bench.ts` that exercises `<component>` at `<target load>`; assert it stays within budget. Keep complexity within budget via `harness check-perf --structural --coupling`." | `harness perf bench` + `harness check-perf --structural` |210| **Resilience** | Ordinary TDD task with a failure-path test + graph failure signals (`predict_failures`, `compute_blast_radius`) | "Write a test that simulates `<failure mode>` and asserts graceful degradation (no crash, documented fallback). Use `predict_failures` to find where else this failure propagates." | The failure-path test (`npx vitest run <test-file>`) |211212**Record the elicited targets** in the `constraints` session section via `manage_state` — they are constraints the plan must honor — and carry them into Phase 2 as tasks tagged `category: "nfr"`.213214> **Honest scope.** Performance, security, and scalability wire to mechanical gates (`harness perf` / `harness check-perf` / `harness check-security`) that pass or fail deterministically. Resilience has no dedicated scanner — it wires to the plan's own failure-path tests plus the graph failure-prediction signals this skill already uses. Do not claim a resilience "gate" the harness does not have; the verifiable artifact is the test.215216---217218### Phase 2: DECOMPOSE — Map File Structure and Create Tasks219220Report progress: `**[Phase 2/4]** DECOMPOSE — mapping file structure and creating tasks`2212221. **Map the file structure first.** List every file to create or modify before writing tasks:223224 ```225 CREATE src/services/notification-service.ts226 CREATE src/services/notification-service.test.ts227 MODIFY src/services/index.ts (add export)228 CREATE src/types/notification.ts229 MODIFY src/api/routes/users.ts (add notification trigger)230 ```2312322. **Skeleton pass (rigor-gated).** Lightweight skeleton (~200 tokens) validates direction before full expansion. Gating per Rigor Levels table.233234 **Format:** Numbered logical groups with task count and time. No file paths, code, or details.235236 ```237 1. Foundation types and interfaces (~3 tasks, ~10 min)238 2. Core scoring module with TDD (~2 tasks, ~8 min)239 3. CLI integration and flag parsing (~4 tasks, ~15 min)240 **Estimated total:** 8 tasks, ~33 minutes241 ```242243 **Approval gate:** Ask in plain text in your reply — do NOT route this through `emit_interaction` or `AskUserQuestion` (they do not display to the human; `emit_interaction` collapses to "Called harness" and `AskUserQuestion` is Claude-Code-only). Ask directly: "Approve skeleton direction?" and wait. If approved, proceed to step 3. If rejected, revise and re-present.2442453. **Decompose into atomic tasks.** Each task must:246 - Be completable in 2-5 minutes, fit in a single context window247 - Have a clear, testable outcome248 - Follow TDD: write test, fail, implement, pass, commit249 - Produce one atomic commit2502514. **Write complete instructions for each task.** Not summaries — complete executable instructions:252 - **Exact file paths** to create or modify253 - **Exact code** to write (not "add validation logic" — write the actual code)254 - **Exact test commands** (e.g., `npx vitest run src/services/notification-service.test.ts`)255 - **Exact commit message**256 - **`harness validate`** as the final step2572585. **Skill annotations.** If skill recommendations were loaded in Phase 1, annotate each task with relevant skills from the Apply and Reference tiers:259260 ```261 ### Task 3: Implement dark mode toggle262 **Skills:** `design-dark-mode` (apply), `a11y-color-contrast` (reference)263 ```264265 Match skills to tasks based on keyword and domain overlap between the task description and the skill's purpose/keywords. Only annotate when the match is relevant to the specific task.2662676. **Include checkpoints.** Mark tasks requiring human input:268 - `[checkpoint:human-verify]` — Pause, show result, wait for confirmation269 - `[checkpoint:decision]` — Pause, present options, wait for choice270 - `[checkpoint:human-action]` — Pause, instruct human on required action2712727. **Derive integration tasks from the spec's Integration Points section.** If the spec contains an Integration Points section, create tasks for each non-empty integration point. Skip subsections marked "None" — do not derive tasks from them. Integration tasks are normal plan tasks but tagged with `category: "integration"` in their description. They appear at the end of the task list, after all implementation tasks.273274 For each subsection of Integration Points, derive tasks:275276 | Integration Point | Example Derived Task |277 | ----------------------------------------------- | -------------------------------------------------------------------------- |278 | Entry Points: "New CLI command" | "Regenerate barrel exports. Verify new command appears in `_registry.ts`." |279 | Registrations Required: "Skill at tier 2" | "Add skill to tier list in `AGENTS.md`. Generate slash commands." |280 | Documentation Updates: "AGENTS.md capabilities" | "Update AGENTS.md to describe the feature." |281 | Architectural Decisions: "ADR for approach X" | "Write ADR `docs/knowledge/decisions/NNNN-<slug>.md`." |282 | Knowledge Impact: "Domain concept Y" | "Enrich knowledge graph with concept node." |283284 Integration tasks follow the same atomic task rules (2-5 minutes, exact file paths, exact code). Use the `**Category:** integration` tag in the task header, e.g.:285286 ```287 ### Task N: Update AGENTS.md with new feature description288289 **Depends on:** Task N-1 | **Files:** `AGENTS.md` | **Category:** integration290 ```291292 If the spec has no Integration Points section, skip this step.2932948. **Emit NFR tasks from Phase 1.6.** For each NFR target elicited in Phase 1.6 (skipped dimensions emit nothing), create one atomic task using the template from the NFR wiring table. Tag it `**Category:** nfr` in the task header. NFR tasks follow the same atomic rules (2-5 minutes, exact file paths, exact commands) and appear after implementation tasks, alongside integration tasks. Each NFR task's final step is its verifying command from the table, so the target is checkable at execution time rather than aspirational.295296 ```297 ### Task N: Lock in p99 budget for parseDocument298299 **Depends on:** Task N-1 | **Files:** `src/parse.bench.ts` | **Category:** nfr300301 1. Create `src/parse.bench.ts` benchmarking `parseDocument`.302 2. Run: `harness perf bench`303 3. Record baseline: `harness perf baselines update`304 4. Verify `p99Ms` in `.harness/perf/baselines.json` meets the target (`p99 < 5ms`).305 5. Commit: `perf(parse): add parseDocument benchmark and baseline`306 ```307308 If Phase 1.6 was skipped, or every dimension was skipped, emit no NFR tasks — the plan is identical to one produced without NFR elicitation.309310---311312### Phase 3: SEQUENCE — Order Tasks and Identify Dependencies3133141. **Order by dependency.** Types before implementations. Implementations before integrations. Integration tasks (tagged `category: "integration"`) after all implementation tasks. Tests alongside implementations (same task, TDD style).3152. **Identify parallel opportunities and record dependency edges.** Tasks touching different subsystems with no shared state can run in parallel. Record each task's real dependencies as `dependsOn` (the task IDs it must follow) in the task header `**Depends on:**` line, and keep the task's `**Files:**` list accurate — together these are exactly the edges `plan_parallelization` consumes (explicit `dependsOn` unioned with file-overlap edges) to build the wave DAG at execution time. A task with no dependencies records `**Depends on:** none`. Optionally, add an `**Owns:**` line listing glob paths a task claims ownership of (e.g. `src/api/**`); it feeds the cheap, deterministic, graph-free `owns`-overlap forecast in `plan_parallelization` — two tasks whose owned globs overlap are flagged in `ownershipForecast` and gain an implicit DAG edge, alongside the file-overlap edges. `Owns` is OPTIONAL: a task with no distinct ownership just omits it.3163. **Number tasks sequentially.** Use `Task 1`, `Task 2`, etc. Dependencies reference task numbers.3174. **Estimate total time.** Sum 2-5 minutes per task. If total exceeds available time, identify a milestone boundary for pausing.318319---320321### Phase 4: VALIDATE — Review and Finalize the Plan3223231. **Verify completeness.** Every observable truth from Phase 1 must trace to specific task(s) that deliver it.3242. **Verify task sizing.** Could an agent complete each task in one context window without exploring or deciding? If not, split it.3253. **Verify TDD compliance.** Every code-producing task must include a test step. No "write tests later."3264. **Run `harness validate`** to verify project health before writing the plan.3275. **Check failures log.** Read `.harness/failures.md`. If planned approaches match known failures, flag them.3286. **Run soundness review.** Invoke `harness-soundness-review --mode plan` against the draft. Do not proceed until the review converges with no remaining issues.3297. **Write the plan to `docs/changes/<topic>/plans/`.** Naming: `YYYY-MM-DD-<feature-name>-plan.md`. Resolve `<topic>` from the spec path — if the spec lives at `docs/changes/<topic>/proposal.md`, the plan goes in the sibling `plans/` directory. If the spec is not under `docs/changes/`, fall back to `docs/plans/` and flag the spec location for human review. Create directories as needed.3308. **Commit plan artifact.** Immediately after writing the plan, commit it so the paper trail enters git history at planning time, not after execution:331332 ```bash333 git add docs/changes/<topic>/plans/YYYY-MM-DD-<feature-name>-plan.md334 git commit -m "docs(<topic>): add plan"335 ```336337 Use `docs/plans/` if the spec is not under `docs/changes/<topic>/`. Do not skip — `harness-execution` commits only implementation files, so a plan uncommitted here stays untracked until someone notices. If a pre-commit hook reformats the file, re-add and re-commit.3383399. **Write handoff.** Write to the session-scoped path when session slug is known, otherwise fall back to global path:340 - Session-scoped (preferred): `.harness/sessions/<session-slug>/handoff.json`341 - Global (fallback, **deprecated**): `.harness/handoff.json`342343 > **[DEPRECATED]** Writing to `.harness/handoff.json` is deprecated. In autopilot sessions, always use `.harness/sessions/<slug>/handoff.json` to prevent cross-session contamination.344345 Fields: `fromSkill`, `phase`, `summary`, `completed`, `pending`, `concerns`, `decisions`, `contextKeywords`.34634710. **Write session summary (if session is known).** Call `writeSessionSummary` with skill, status, plan path, keyContext, nextStep. Skip if no session slug.34834911. **Request plan sign-off in plain text.** Ask directly in your reply — do NOT route this through `emit_interaction` or `AskUserQuestion` (the human will not see it; `emit_interaction` collapses to "Called harness" and `AskUserQuestion` is Claude-Code-only). Present the plan path, task count, and time estimate, then ask: "Proceed? (yes/no)" and wait for the reply.35035112. **Suggest transition to execution.** After approval, call `emit_interaction` with type: `transition`, `completedPhase: "planning"`, `suggestedNext: "execution"`, `requiresConfirmation: true`. Include `qualityGate` with checks: plan-written, harness-validate, observable-truths-traced, human-approved. If confirmed: invoke harness-execution. If declined: stop (handoff already written).352353---354355### Plan Document Structure356357```markdown358# Plan: <Feature Name>359360**Date:** YYYY-MM-DD | **Spec:** (if applicable) | **Tasks:** N | **Time:** N min | **Integration Tier:** small | medium | large361362## Goal363364One sentence.365366## Observable Truths (Acceptance Criteria)3673681. [observable truth]369370## NFR Targets (if elicited)371372- [Performance] EARS criterion → task N (verified by `harness perf bench`)373- [Security] EARS criterion → task N (verified by `harness check-security --severity error`)374 _Omit this section entirely when no NFR dimension was elicited._375376## File Map377378- CREATE path/to/file.ts379- MODIFY path/to/other-file.ts380381## Skeleton (if produced)3823831. <group name> (~N tasks, ~N min)384 _Skeleton approved: yes/no._385386## Tasks387388### Task 1: <descriptive name>389390**Depends on:** none | **Files:** path/to/file.ts, path/to/file.test.ts | **Owns:** path/to/\*\*391392<!-- `Depends on` lists the task IDs this task must follow (its `dependsOn` edges); `Files` is the independence-checking input; `Owns` (optional) lists glob paths this task claims, feeding the deterministic owns-overlap forecast. All three feed `plan_parallelization` during execution. -->3933941. Create test file with exact test code3952. Run test — observe failure3963. Create implementation with exact code3974. Run test — observe pass3985. Run: `harness validate`3996. Commit: `feat(scope): descriptive message`400401### Task 2: <descriptive name>402403[checkpoint:human-verify] ...404```405406### Integration Tier Heuristics407408When a spec contains an **Integration Points** section, set the plan's `integrationTier` field based on scope:409410| Tier | Signal | Integration Requirements |411| ---------- | -------------------------------------------------------------------- | ------------------------------------------------------------------------ |412| **small** | Bug fix, config change, < 3 files, no new exports | Wiring checks only (defaults always run) |413| **medium** | New feature within existing package, new exports, 3-15 files | Wiring + project updates (roadmap, changelog, graph enrichment) |414| **large** | New package, new skill, new public API surface, architectural change | Wiring + project updates + knowledge materialization (ADRs, doc updates) |415416If the spec has no Integration Points section, omit the `integrationTier` field from the plan header.417418## Session State419420| Section | Read | Write | Purpose |421| ------------- | ---- | ----- | ----------------------------------------------------------------- |422| terminology | yes | no | Consistent language in plan |423| decisions | yes | yes | Brainstorming decisions; planning-phase decisions |424| constraints | yes | yes | Existing constraints; constraints discovered during decomposition |425| risks | yes | yes | Existing risks; implementation risks from task design |426| openQuestions | yes | yes | Unresolved questions; new questions; resolve answered ones |427| evidence | yes | yes | Prior evidence; file:line citations for task specs |428429**When to write:** Phase 1 — constraints and risks. Phase 2 — decisions about task structure. Phase 4 — resolve questions.430431**When to read:** Start of Phase 1 via `gather_context` with `include: ["state", "learnings", "handoff", "graph", "businessKnowledge", "sessions", "validation"]` to inherit brainstorming context and load documented business knowledge.432433## Evidence Requirements434435When referencing existing code in task specs, cite evidence using `file:line` format, code pattern references, or test output. Write to `evidence` session section via `manage_state`.436437**When to cite:** Phase 1 (existing files), Phase 2 (file paths and patterns), file map (existing files for modification).438439**Uncited claims:** Prefix with `[UNVERIFIED]`.440441## Harness Integration442443- **`harness validate`** — Run in Phase 4 (before writing plan) and included in every task.444- **`harness check-deps`** — Referenced in tasks adding imports or creating modules.445- **`harness perf bench` / `harness perf baselines update` / `harness check-perf`** — The verifying commands for performance and scalability NFR tasks emitted from Phase 1.6. Perf baselines live in `.harness/perf/baselines.json`; complexity/coupling/size budgets live in the `performance` block of `harness.config.json`.446- **`harness check-security`** — The verifying command for security NFR tasks emitted from Phase 1.6. Its gate fails on error-severity findings after the `--severity` filter; `security.strict` in `harness.config.json` promotes warnings to errors for stricter modules.447- **Plan commit** — After writing the plan (Phase 4 Step 8), commit `docs/changes/<topic>/plans/YYYY-MM-DD-<feature-name>-plan.md` so the paper trail enters git history at planning time. `harness-execution` does not backfill this commit.448- **Plan location** — `docs/changes/<topic>/plans/YYYY-MM-DD-<feature-name>-plan.md` when the spec lives under `docs/changes/<topic>/proposal.md`; otherwise `docs/plans/` as a fallback.449- **Handoff** — Once approved, invoke harness-execution for task-by-task implementation.450- **Session directory** — Session-scoped writes go to `.harness/sessions/<slug>/`. Structure: `handoff.json`, `state.json`, `artifacts.json` (registry of spec/plan paths and produced file lists). Global `.harness/handoff.json` is deprecated for session-aware invocations.451- **`emit_interaction`** — Call at end of Phase 4 to suggest transitioning to execution (confirmed transition).452- **Rigor levels** — `--fast`/`--thorough` control skeleton pass. See Rigor Levels table.453- **Two-pass planning** — Skeleton (~200 tokens) before full expansion. Catches directional errors early.454455## Change Specifications456457When planning changes to existing functionality (not greenfield), express requirements as deltas:458459- **[ADDED]** — New behavior that does not exist today460- **[MODIFIED]** — Existing behavior that changes461- **[REMOVED]** — Existing behavior that goes away462463**Example:**464465```markdown466## Changes to User Authentication467468- [ADDED] OAuth2 refresh tokens with 7-day expiry469- [MODIFIED] Login endpoint returns `refreshToken` alongside `accessToken`470- [MODIFIED] Token validation accepts both JWT and OAuth2 tokens471- [REMOVED] Legacy API key authentication (deprecated in v2.1)472```473474Only apply when modifying existing documented behavior. When `docs/changes/` exists, produce `docs/changes/<feature>/delta.md` alongside the task plan.475476## Success Criteria477478- Plan document exists at the resolved location (`docs/changes/<topic>/plans/` or `docs/plans/` fallback) with all required sections479- Every task completable in 2-5 minutes (one context window)480- Every task includes exact file paths, exact code, and exact commands481- Every code-producing task follows TDD: test first, fail, implement, pass482- Observable truths trace to specific tasks483- File map lists every file to create or modify484- Checkpoints marked where human input is required485- `harness validate` passes before plan is written and is in every task486- Human has reviewed and approved the plan487- Rigor level rules followed: fast skips skeleton; thorough always skeletons with approval; standard skeletons at >= 8 tasks488- NFR targets elicited in Phase 1.6 each trace to a task tagged `category: nfr` whose final step is a verifying command, or the dimension was explicitly skipped (no NFR task emitted)489490## Red Flags491492| Flag | Corrective Action |493| ---------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------- |494| "I know the implementation well enough to skip reading the spec" | STOP. Phase 1 SCOPE starts by reading the spec. Assumptions about spec content lead to plans that implement the wrong thing. |495| "This task is self-explanatory, no need for exact file paths and commands" | STOP. Iron Law: every task must contain exact file paths, exact commands, and complete code snippets. "Implement the service" is a wish, not a task. |496| "I'll plan the happy path now and add error handling tasks later" | STOP. Error handling is not optional. The spec's success criteria include error scenarios. Plan them alongside the happy path. |497| `// detailed steps TBD` or `// expand during execution` in task descriptions | STOP. A task that defers detail to execution is a vague task. If you cannot write the exact steps now, you do not understand the task well enough to plan it. |498499## Rationalizations to Reject500501| Rationalization | Reality |502| ------------------------------------------------------------------------------------------------------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |503| "The task is conceptually clear so I do not need to include exact code in the plan" | Every task must have exact file paths, exact code, and exact commands. If you cannot write the code in the plan, you do not understand the task well enough to plan it. |504| "This task touches 5 files but it is logically one unit of work, so splitting it would add overhead" | Tasks touching more than 3 files must be split. The overhead of splitting is far less than the cost of a failed oversized task. 505506…(truncated)