# Epics

> Break a spec into epic documents with stories and tasks. Reads a specification (or a brief/description) and produces multiple epic docs — each with stories, sub-tasks, and a companion coverage matrix. Also offers a read-only dependency/readiness view, publishable as a shareable artifact, on request. Triggers on "/cpm:epics".

- Skill: `ninthspace/epics` (Agent Skill)
- Install (CLI): `npx skillmds@latest add ninthspace/epics`
- Raw SKILL.md: https://api.skillmd.com/api/skills/ninthspace/epics/raw
- Safety review: pending
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: ninthspace (https://skillmd.com/u/ninthspace)
- Updated: 2026-09-10
- Page: https://skillmd.com/skills/ninthspace/epics

---


# Work Breakdown into Epics

Turn a specification into a set of **epic documents** — each representing a major work area containing **stories** (meaningful deliverables with acceptance criteria) and **tasks** (implementation steps). Epic documents are the plan of record — task creation in Claude Code's native system happens later during execution via `cpm:do`.

## Input

Check for input in this order:

1. If `$ARGUMENTS` references a file path, read that file as the source.
2. If `$ARGUMENTS` contains a description, use that as the source.
3. Look for planning docs — check `docs/specifications/` first, then `docs/plans/`.
4. If nothing found, ask the user what work they want to break down.

**Dependency-view mode (on request).** If `$ARGUMENTS` asks for a **dependency / readiness / "what's ready to pick up" view** of the existing epics (e.g. contains `dependency view`, `dependencies`, `unblocked`, `ready to pick up`, or "what can I work on") rather than naming work to break down, do **not** run the production process below. Instead run the **Dependency View (on request)** section (after `## Output`) — a read-only projection over the epic docs that already exist. The two modes are mutually exclusive: breaking a spec into epics *writes* epic docs; the dependency view only *reads* them.

## Process

**State tracking**: Create the progress file before Step 1 and update it after each step completes. See State Management below for the format and rationale. Delete the file once the final epic docs have been saved.

### Termination

- **Success**: The user confirms the final task tree in Step 4 — save the epic documents and finish.
- **Blocker**: The user identifies a spec gap, missing requirement, or dependency that cannot be resolved in this session. Note the gap, save the epic documents with what's confirmed, and flag the gap for resolution via `cpm:pivot` or a spec update.
- **Ambiguity**: The user is uncertain about epic grouping, story scope, or task breakdown after one clarification round. Present a recommended structure with rationale. If the user still cannot decide, use the recommended default and note the decision as provisional — it can be revised before `cpm:do` begins execution.

**Facilitation depth**: Each presentation-and-refine gate (Step 2 epics, Step 3 stories, Step 3b tasks) converges in 1-2 rounds of AskUserQuestion. When the user approves, move on. Step 3d coverage matrix and Step 4 confirmation are single-pass gates — present once, refine once if needed, then proceed.

### Autonomous Mode

When the run is autonomous — this skill was invoked by a wrapper such as `cpm:ralph` and no human is present to answer — the gates listed below **do not block**. This section is the single source of that behaviour: `cpm:ralph`'s prompt references it rather than restating the dispositions, so a gate added or changed here needs no second edit there.

**Five of this skill's gates present a proposal the skill has just rendered.** For those, the autonomous disposition is **approve the rendered proposal and proceed** — the proposal is already the skill's own best judgement, and a second pass over it in the same run adds no information that the first pass did not have. Render the proposal exactly as the interactive path would, then continue without the gate:

| Gate | What it asks interactively | Autonomous disposition |
|---|---|---|
| Step 2 — Identify Epics | "Approve this grouping?" | Approve the rendered grouping and proceed to Step 3. No refine round |
| Step 3 — Break into Stories | "Approve these stories?" | Approve the rendered stories. Still render the tag distribution summary, and still flag a story with zero automated tags — that flag is a record, not a question |
| Step 3b — Identify Tasks within Stories | "Approve these tasks?" | Approve the rendered task list for each story |
| Step 3c — Integration Testing Story | Confirm the cross-story acceptance criteria | Accept the criteria as rendered. The *when warranted* test still decides whether the story exists at all; only the confirmation is skipped |
| Step 4 — Confirm | Final confirmation of the epic / story / task tree | Accept the tree and save. This is the **Termination — Success** condition, so an autonomous run reaches it by this disposition rather than by user confirmation |

**Rendering stays mandatory.** Each of the five still writes its proposal into the message body. Under an autonomous run nobody reads it at the moment it is produced — which is exactly why it has to be there to read afterwards.

**The sixth gate is absent from that table because it takes the opposite disposition.** Step 3's must-NOT clause proposal (*Must-NOT clause suggestion*) is the one gate where approving the skill's own proposal is a self-marking problem: the clause and its approval would come from the same pass, with nobody to check the judgement. Its autonomous disposition is **propagate, never invent**:

1. **Propagate every must-NOT line the source spec already carries.** *Must-NOT clause propagation* above is unchanged and needs no disposition of its own — copying a line the spec's own Section 6b probed for is transcription, and a reader can verify it afterwards against the spec.
2. **Attach nothing that cannot be quoted from the source spec.** A clause is citable when it appears as a `must NOT` line in the spec's Acceptance Criteria Coverage table, or in the requirement text that line is paired with. A clause whose subject the spec never raises is not citable, however reasonable it looks — that is judgement made in the moment, which is exactly what 43-02's citable-contradiction rule exists to refuse.
3. **Record what it would have proposed, rather than dropping it.** On the story the clause would have been attached to, write `**Must-NOT proposed (unreviewed)**: {clause} — {domain that triggered it} ({YYYY-MM-DD})`. Recorded is not attached: it sits on the story where a human reviewing the epic will see it, and it constrains no acceptance criterion until one accepts it. Nothing parses this field today — it is written for a reader, and saying so is the point, because a breadcrumb credited to a consumer that does not read it has now been shipped twice in this repo.

**Why not simply accept every proposal.** Auto-accepting reads as the maximally defensive choice and is not. A clause invented in the moment can be *unsatisfiable as written* — retro 21 recorded one that forbade a token its own explanatory prose had to use — and `cpm:do` would then be unable to close the story with nobody watching. A loop that cannot finish is a worse outcome than a missing boundary a human can still add on review.

**Write surface — three kinds of file, and no others.** An autonomous run writes epic documents under `docs/epics/`, their companion coverage matrices beside them, and its own progress file under `docs/plans/`. It writes nothing under `docs/specifications/`. The source document is the only artefact a human authored and the only fixed point the run is measured against, so a run able to edit it can move its own goalposts — and the coverage matrices it writes in the same pass would then agree with a target it had changed. A gap or contradiction found in the source mid-run is **recorded, not repaired**: state it in the epic's Notes and leave it for `/cpm:pivot`. That is **Termination — Blocker**'s record-and-flag half and only that half: the condition as written is triggered by a user who is not present here, and offers "or a spec update" as an alternative remedy — which is the one remedy an autonomous run may not take.

**Audit trail — every gate decision leaves a breadcrumb.** A run nobody watched is reviewable only from what it wrote down, so each disposition taken above records one line in the epic document's top-level metadata block, below `**Blocked by**`:

`**Autonomous gate**: {gate} · {what was chosen}`

for example `**Autonomous gate**: Step 2 — Identify Epics · approved the rendered grouping of 4 epics`. Gates that fire before any epic document exists — Step 2's grouping is the one that always does — are recorded on every epic that run produced, because the progress file is deleted when the run finishes and the epic documents are what outlive it. The `**Must-NOT proposed (unreviewed)**` lines from the sixth gate sit on their stories rather than here; both are breadcrumbs, but that one belongs beside the criterion it was almost attached to.

### Stale-Progress Check (Startup)

Follow the shared **Stale-Progress Check** procedure (from the CPM Shared Skill Conventions loaded at session start).

### Retro Check (Startup)

Follow the shared **Retro Awareness** procedure before beginning Step 1.

**Retro incorporation** (this skill):
- **Scope surprises**: Inform Step 3 (story sizing) — past stories that ran larger or smaller than expected suggest sizing rules to apply this round (e.g. "split stories that touch more than N components").
- **Patterns worth reusing**: Inform Step 3b (task lists) — surfaced abstractions and approaches become candidate tasks rather than re-discovered work.
- **Testing gaps**: Inform Step 3 acceptance criteria tagging — past untestable criteria become explicitly tagged this round.
- **Codebase discoveries**: Inform Step 3c (integration testing story) — surfaced integration points may need explicit cross-story coverage.

### Library Check (Startup)

Follow the shared **Library Check** procedure with scope keyword `epics`. Deep-read selectively during Step 2 epic grouping when architecture or coding-standards docs affect epic boundaries or dependency identification.

### Template Hint (Startup)

After startup checks and before Step 1, display:

> Output format is fixed (used by downstream skills). Run `/cpm:templates preview epics` to see the format.

### ADR Discovery (Startup)

After the Template Hint and before Step 1, discover existing Architecture Decision Records:

1. **Glob** `docs/architecture/[0-9]*-adr-*.md`. If no files found or directory doesn't exist, skip silently.
2. If ADRs exist, read each one and note the architectural decisions and their dependencies. Report to the user: "Found {N} existing ADRs: {titles}. I'll reference these when breaking down architectural work into epics and stories."
3. During Step 2 (Identify Epics) and Step 3 (Break into Stories), use ADR context to inform epic grouping — e.g. if an ADR identifies separate concerns or bounded contexts, these may map naturally to epics. Reference specific ADRs in story descriptions when they constrain or inform the implementation approach.

**Graceful degradation**: If ADRs are absent, epic breakdown works as before — deriving structure purely from the spec. The skill works with or without `cpm:architect` having been run.

### Step 1: Read Source

Read and understand the source document. Summarise the key work areas to the user.

### Step 2: Identify Epics

Analyse the source to identify major work areas. Each epic will become its own document at `docs/epics/{parent}-{seq}-epic-{slug}.md`. See the **Epic Filename Convention** subsection below for the parent-extraction and sub-number rules.

For each epic, determine:
- **Name**: A concise label for the work area (e.g. "Authentication System", "API Layer")
- **Filename prefix** (`{parent}-{seq}`): Assigned per the **Epic Filename Convention** subsection below — `{parent}` comes from the source spec number (or `00` for orphans), `{seq}` increments within that parent.
- **Slug** (`{slug}`): Kebab-case derived from the epic name (e.g. "authentication-system", "api-layer")
- **Summary**: One-sentence description of what this epic covers

Render the epic grouping (names, summaries, proposed filenames) in the message body so the user can see the full output plan. Then use AskUserQuestion as a short gate (e.g. "Approve this grouping?" with options `Approve` / `Request changes` / `Stop`) and refine. See the shared **Gate Presentation** convention.

Keep epics practical:
- 2-5 epics for a small feature
- 5-10 for a larger project
- Only create an epic when the work genuinely warrants one

### Epic Filename Convention

Epic documents use a **two-part numeric prefix**: `{parent}-{seq}-epic-{slug}.md`. This makes the parent spec discoverable at the filename level — a reader scanning `ls docs/epics/` can identify an epic's source without opening the file.

> **Transition note**: Epics created before the numbering update used flat `{nn}-epic-{slug}.md`. New epics use parent-scoped `{parent}-{seq}-epic-{slug}.md`. Both shapes coexist permanently — old epics are not migrated. Readers must handle both shapes; writers produce only the two-part shape.

**Parent extraction**:

- **Spec-linked epics** (the default — input is a spec): `{parent}` is the numeric prefix of the source spec's filename. For example, input `docs/specifications/28-spec-foo.md` yields `{parent} = 28`, producing `28-01-epic-foo.md`, `28-02-epic-bar.md`, and so on.
- **Orphan epics** (input is a brief, discussion, description, or bare prompt — no parent spec): `{parent}` is the literal string `00`. Orphan epics are written as `00-{seq}-epic-{slug}.md`. Every new epic produced by this skill uses the two-part shape.

**Sub-number assignment** (`{seq}`):

Assign `{seq}` per parent using the shared Numbering procedure — glob both `docs/epics/{parent}-[0-9]*-epic-*.md` and `docs/archive/epics/{parent}-[0-9]*-epic-*.md`, extract the second numeric field from each match, parse as integer (always integer comparison, never lexical), take `max + 1` across the union. If both sets are empty, start at `01`. Format using the shared Numbering width rule. Flat-shape files (e.g. `15-epic-consult-skill-core.md`) are naturally excluded from the scoped glob.

Sub-numbers are **identifiers**, not ordinals — gaps from deleted sub-numbers are preserved (the max-lookup skips them), keeping cross-references stable.

**General epic listings**: Operations that enumerate *all* epics (for dependency resolution, batch status checks, cross-epic references) must use the general glob `docs/epics/[0-9]*-epic-*.md`, which matches both flat and two-part shapes. Only the sub-number assignment above uses the scoped glob. Always use the general glob for listings — narrowing to the two-part shape would hide legacy epics from readers.

**Production loop**: After epics are confirmed, Steps 3, 3b, 3c, and 3d iterate per epic — break each epic into stories, then tasks, then assess integration testing, then verify requirement coverage. Save each epic document and its coverage matrix after they are finalised (before moving to the next epic). This means epic docs are written incrementally, not all at once at the end.

*Progress note: capture epic names, numbers, slugs, and output paths in the Step 2 summary.*

### Step 3: Break into Stories

For each epic, break into **stories** — meaningful deliverables that represent a coherent unit of value. A story answers "what are we delivering?" not "what file are we editing?"

Each story should have:
- A clear, actionable title (imperative form: "Set up compaction hook infrastructure")
- Acceptance criteria that describe the deliverable outcome
- An activeForm for progress display (present continuous: "Setting up compaction hook infrastructure")
- **Spec traceability** (when input is a spec): Which functional requirements from the spec this story satisfies. Use the requirement text or a short label. This enables verification that the spec is fully covered across all epics — every must-have requirement should appear in at least one story.

**Acceptance criteria fidelity** (when input is a spec): When deriving acceptance criteria from a spec, use the spec's language verbatim where it provides specific thresholds, values, or behaviours. Use the spec's language verbatim — preserve thresholds, values, and behaviours exactly. If the spec says "concurrent session limit of 3 per user", the story criterion says "concurrent session limit of 3 per user". The spec's specificity must survive intact into the stories.

**Testability standard**: Each acceptance criterion must be **testable as written** — it describes a specific, observable, verifiable outcome. Flag any criterion that relies on subjective judgement or cannot be verified through code, tests, or inspection. Criteria that fail this standard must be rewritten before the story is finalised.

Examples:
- **Fails**: "Users can log in" — no observable outcome defined
- **Passes**: "User with valid credentials receives a 200 response with a session token and is redirected to /dashboard"
- **Fails**: "Error handling works correctly" — subjective and unverifiable
- **Passes**: "Invalid OAuth token returns 401 with error body `{\"error\": \"invalid_token\"}` and no session is created"

**Reachability standard**: A criterion about what the system *returns* is a different claim from one about what a person can *do*, and a requirement covered only by the first is satisfiable by an endpoint with nothing wired to it. Where a requirement names an action a user takes — create, edit, delete, share, export, revoke — at least one criterion must name the affordance that reaches it, alongside any criterion about the response.

- **Incomplete**: "Submitting valid credentials returns 200 with a session token" — true of a route no page posts to
- **Complete**: that criterion, plus "The sign-in page presents an email and password form that posts to the session route and renders validation errors in place"

Both halves are needed and neither substitutes for the other: the response criterion does not say the capability is reachable, and the affordance criterion does not say the response is correct. The tests that follow inherit this — a suite whose every journey begins at an endpoint or a seeded session never exercises the path a person takes to get there, and reports full coverage of an application no one can enter.

**Test approach tag propagation** (when the input spec has a Testing Strategy section with tagged criteria):

1. Read the spec's Testing Strategy section — specifically the Acceptance Criteria Coverage table which maps requirements to criteria with `[tag]` annotations.
2. When writing story acceptance criteria, apply matching tags inline. For each acceptance criterion, append the appropriate tags from the spec's testing strategy: `[unit]`, `[integration]`, `[feature]`, `[manual]`, `[target]`, and `[tdd]`. Match by tracing the story's `**Satisfies**` field back to the spec requirement, then looking up that requirement's tag assignments. The `[tdd]` tag is a workflow mode tag (orthogonal to level tags) — propagate it alongside any level tag when present (e.g. `[tdd] [unit]`).
3. If a story's criteria go beyond a spec requirement's tagged criteria (e.g. the story introduces new criteria beyond the spec), default to an automated tag — `[unit]`, `[integration]`, or `[feature]`. Propose `[manual]` only when automation is genuinely infeasible, and when you do, include a one-line justification stating what blocks automation (e.g. "requires human visual judgement", "third-party OAuth flow we don't control"). See the **Default to automation** guideline below.
4. Tags appear at the end of the acceptance criteria line, e.g.: `- User can log in via OAuth [integration]` or `- Payment processor validates card [tdd] [integration]`

**Mockup-referencing criteria**: When a `spec` or `architect` artifact references an HTML **companion asset** (see the shared **HTML Output** convention), one rule holds for every asset and a second holds for only one kind of asset.

**Always: the asset is never a source of requirements.** The Markdown acceptance criteria are the machine-readable specification; the markup is not data and is never parsed for behaviour to implement. Never write an automated markup-parsing test against a companion asset.

**Then branch on what the asset is,** because the shared convention covers two kinds and only one of them is something you build:

- A **mockup** — the UI of the system being built, i.e. the asset represents deliverable functionality — is a **design target**. A criterion saying "build to match this asset" describes *visual conformance*, whose only oracle is human judgement, so tag it `[manual]` (or `[feature]` when the match is exercised through a user-facing workflow). This is the named case where `[manual]`/`[feature]` is correct *because* the oracle is visual — record the one-line justification the **Default to automation** guideline requires (e.g. "match the referenced mockup — visual conformance, no markup oracle").
- A **diagram** — a data-flow, sequence, or architecture illustration — explains the reasoning behind a decision. Nothing is built to look like it, so it gets **no criterion at all**: a visual-conformance criterion against a diagram has no deliverable to conform to and would be unverifiable in both directions. It is context a story may cite in its description; it is not a target. A diagram going unreferenced by every story is the correct outcome, not a coverage gap.

The test is whether the asset depicts something the work produces. If a reader could hold the built thing up beside it and judge the match, it is a mockup; if it depicts how parts relate, it is a diagram.

**Graceful degradation**: If the spec has no Testing Strategy section, no Acceptance Criteria Coverage table, or no tags, skip tag propagation entirely — write acceptance criteria without tags. The skill must work without `cpm:spec`'s enhanced Section 6 having been used.

**Must-NOT clause propagation** (when the input spec has `must NOT` lines from Section 6b):

1. Read the spec's Acceptance Criteria Coverage table for `must NOT` lines paired with positive criteria.
2. When writing story acceptance criteria, include the `must NOT` lines alongside their paired positive criteria. Preserve the spec's wording verbatim — these are defensive boundaries, not suggestions.
3. If a story's criteria go beyond the spec (new criteria without spec-originating must-NOTs), assess whether must-NOT clauses are warranted based on the story's domain (see below).

**Must-NOT clause suggestion** (when the spec has no must-NOT lines, or for criteria without them):

When a story touches any of these domains, propose `must NOT` clauses for the user to confirm:
- **Security**: authentication, authorization, session management, credential handling
- **Data integrity**: database writes, financial calculations, user data mutations
- **External systems**: API calls, webhook handling, third-party integrations

Propose 1-2 must-NOT clauses per relevant criterion. Present them via AskUserQuestion alongside the story's acceptance criteria for the user to accept, modify, or reject. If the user rejects all proposed must-NOTs, proceed without them — must-NOT clauses are advisory, not mandatory.

**Under an autonomous run this gate does not take the approve-your-own-proposal disposition** that the other five do. See **Autonomous Mode** above: propagate what the spec carries, attach nothing that cannot be quoted from it, and record the rest as `**Must-NOT proposed (unreviewed)**`.

**Graceful degradation**: If the spec has no must-NOT lines and the story does not touch security, data integrity, or external systems, skip must-NOT suggestion entirely.

**`[plan]` tag suggestion**: After defining a story's acceptance criteria, assess whether it warrants formal plan mode during execution. The `[plan]` tag forces an EnterPlanMode pause before implementation — it's a workflow lock, not a signal that the story is hard. Append `[plan]` to the story's `##` heading (e.g. `## Set up OAuth provider integration [plan]`) when the story involves:

- **Data model changes**: New or modified database schemas, entity relationships, or data structures that affect persistence
- **API contract changes**: New or modified public APIs, webhook schemas, or inter-service contracts where the design needs upfront agreement
- **Cross-system integration**: Coordination across multiple external systems, APIs, or services where the interaction design needs upfront thought

These are the default assignment categories. The user can also add `[plan]` manually to any story they want gated — these categories are defaults, not restrictions. Stories that follow existing patterns, are fully specified by their acceptance criteria, or are straightforward config/documentation changes use inline planning (the default in `cpm:do`).

**Stories vs tasks**: A story groups related implementation work under a single deliverable with shared acceptance criteria. If you find yourself writing a story title that describes a single file change or a single function — that's a task, not a story. Push it down to Step 3b.

Render the stories for each epic (titles, summaries, acceptance criteria) in the message body. After the stories, render a **tag distribution summary** showing per-story counts so any drift toward manual is visible at a glance. Format:

```
Tag distribution:
- Story 1 — {N} automated ([unit] x{a}, [integration] x{b}, [feature] x{c}), {M} manual
- Story 2 — {N} automated, {M} manual
- ...
```

If any story has zero automated tags, flag it explicitly under the table (e.g. "⚠ Story 2 is fully manual — confirm this is intentional"). Then use AskUserQuestion as a short gate (e.g. "Approve these stories?" with options `Approve` / `Request changes` / `Stop`). Refine before moving to the next epic. See the shared **Gate Presentation** convention.

*Progress note: record which tags were propagated to which stories in the Step 3 summary.*

### Step 3b: Identify Tasks within Stories

For each story, identify the **tasks** — concrete implementation steps needed to deliver the story. Each task should have:
- A clear, actionable title (imperative form: "Create hooks.json configuration")
- A dot-notation number linking it to its parent story (e.g. Task 1.1, 1.2, 1.3 for Story 1)
- A one-sentence **description** that scopes the task within its parent story — which acceptance criteria or concern this task addresses

**Task descriptions**: Write a description for every task in stories with multiple tasks. Descriptions eliminate the three-hop lookup (title → story criteria → spec) by anchoring each task to the acceptance criteria it addresses — e.g. "Covers the error handling criteria for roster loading" or "Produces the interface that Task 2.3 consumes." For single-task stories, omit the description when the title is self-evident — but when in doubt, write one.

Descriptions should state **scope boundaries**, not implementation steps. Good descriptions reference criteria, relationships, or constraints; bad descriptions prescribe how to build it.

- Good: "Add roster loading section — project override (`docs/agents/roster.yaml`) then plugin default (`../../agents/roster.yaml`), error if neither found. Present selected agent to the user."
- Good: "Addresses the error handling path, not the happy path — covers criteria 3 and 4."
- Bad: "Edit SKILL.md line 45 to add a YAML parsing block with error handling."
- Bad: "Create a function called loadRoster() that reads the file and returns an array."

Tasks are the actual work items. They should be specific enough that an implementer knows exactly what to do.

A single task per story is fine when the work is straightforward. Decompose only when it makes complex stories manageable.

**Auto-generated testing tasks**: After identifying implementation tasks for a story, check whether any of the story's acceptance criteria carry `[unit]`, `[integration]`, or `[feature]` tags. If at least one automated test tag is present:

1. Auto-generate a testing task titled "Write tests for {story title}".
2. **Placement depends on `[tdd]`**:
   - If the story's acceptance criteria include `[tdd]`, place the testing task **before** all implementation tasks (as the first task in the story). This enables the TDD red-green-refactor workflow — tests are written first, then implementation makes them pass. The testing task gets dot-notation number `{story}.1`, and implementation tasks follow sequentially.
   - If the story does **not** carry `[tdd]`, place the testing task **after** all implementation tasks (as the last task in the story, before the verification gate that `cpm:do` will create). The testing task's dot-notation number follows the last implementation task (e.g. if implementation tasks are 1.1, 1.2, 1.3, the testing task is 1.4).
3. Give it a description: "Write automated tests covering the story's acceptance criteria tagged `[unit]`, `[integration]`, or `[feature]`."

If **all** of a story's criteria are tagged `[manual]` (or have no tags), do **not** generate a testing task — there's nothing to automate.

**Graceful degradation**: If no tags were propagated during Step 3 (e.g. the spec had no testing strategy), skip testing task generation entirely.

Render the tasks for each story in the message body. Then use AskUserQuestion as a short gate (e.g. "Approve these tasks?" with options `Approve` / `Request changes` / `Stop`). Refine before moving to the next story. See the shared **Gate Presentation** convention.

### Step 3c: Integration Testing Story (when warranted)

After all implementation stories and their tasks are defined for an epic, assess whether the epic warrants a **dedicated integration testing story**. This is separate from per-story testing tasks (which test a single story's criteria) — an integration testing story verifies cross-story behaviour and integration points.

**When to generate**: Create an integration testing story when the epic has:
- Multiple stories with `[integration]` tagged criteria that interact with each other
- Cross-story data flows, API contracts, or event-driven interactions
- Stories that produce components which must work together as a system

**When to skip**: Skip if the epic has only 1-2 stories, no `[integration]` tags, or stories that are independent of each other.

**How to generate**:
1. Title: "Verify cross-story integration for {epic name}"
2. Story number: the next sequential number after the last implementation story
3. `**Blocked by**`: all implementation stories in the epic (comma-separated)
4. Acceptance criteria: specific cross-story integration points that need verification. These should describe observable behaviour that spans multiple stories — not just "everything works together." Confirm with the user via AskUserQuestion.
5. Tasks: typically a single task — "Write integration tests for {epic name}" — unless the integration points are complex enough to warrant separate tasks.

**Graceful degradation**: If no tags were propagated during Step 3, skip this step entirely.

*Progress note: capture whether an integration story was generated or skipped, with the rationale, in the Step 3c summary.*

### Step 3d: Requirement Coverage Matrix (when input is a spec)

After stories, tasks, and integration testing are defined for the current epic, verify that the spec requirements this epic claims to satisfy are properly covered. This runs per-epic as part of the production loop (after Step 3c, before saving the epic doc and moving to the next epic).

The coverage matrix is **procedural, not evaluative**. Your job is to extract and present verbatim text from both the spec and the stories so the user can judge fidelity. Present the side-by-side text and let the human judge fidelity — the assessment is theirs to make.

1. **Identify which spec requirements this epic covers.** Scan the epic's stories' `**Satisfies**` fields to find all referenced spec requirements.
2. **Read the source spec's requirement text.** For each referenced requirement, extract the number, label, and the specific text that defines thresholds, values, or behaviours. Quote from the **requirements section**, not the testing strategy table (these can differ — the requirements text is authoritative).
3. **Build a side-by-side coverage table.** For each referenced requirement, quote both the spec's requirement text and the matching story acceptance criterion text **verbatim** — preserve the exact wording on both sides.

If the spec's Testing Strategy section includes test approach tags, add a **Spec Test Approach** column showing the tag(s) from the spec's Acceptance Criteria Coverage table. This lets the user verify tag propagation in the same view.

Present the coverage matrix to the user:

```
| # | Spec Requirement | Spec Text (verbatim) | Story Criterion (verbatim) | Covered by | Spec Test Approach | Verified |
|---|------------------|----------------------|----------------------------|------------|--------------------|----------|
| 1 | {requirement label} | {exact text from spec requirements section} | {exact criterion text from story} | Story {N} | `[tag]` | |
```

The "Verified" column is empty for all rows at creation time — it is populated later by `/cpm:do` during verification gates.

**The Spec Requirement cell holds one bare label** — `FR7`, `ENV3`, `NFR10` — optionally followed by a parenthesised qualifier the spec itself uses, such as `FR1 (must NOT)`. Nothing else. A descriptive tail (`FR7 — Pence-exact remainder rule`) reads as more informative and is not: the roll-up joins this cell to the spec by exact match, so a decorated label traces to no requirement and the requirement reports as uncovered while a row on disk claims it.

**One requirement per row, never a range or a list.** `ENV1–ENV5` and `FR8, FR5` name several requirements against a single criterion and a single Verified cell, so one tick would mark all of them verified on one piece of evidence. Where one criterion genuinely covers several requirements, repeat the criterion on a row per requirement. The roll-up refuses a label it cannot resolve to exactly one requirement and reports the row as `UNRESOLVED` rather than guessing which was meant.

Where a single spec requirement maps to multiple story criteria, include one row per criterion so each mapping is independently visible.

If the user identifies a fidelity problem (story criterion is weaker than or contradicts spec text), update the affected story's acceptance criteria in the epic doc before proceeding.

**Persist the matrix**: After the user confirms, save the coverage matrix as `docs/epics/{parent}-{seq}-coverage-{slug}.md` — using the same two-part prefix and slug as the epic it covers (see Output section). Save the epic doc and coverage matrix together before moving to the next epic.

**Regeneration awareness**: Before saving, check whether a coverage matrix file already exists at the target path (e.g. from a previous `/cpm:epics` run). If an existing matrix is found and contains `✓` markers in the Verified column, the new matrix must clear verification for any rows whose "Story Criterion (verbatim)" text has changed — replace `✓` with empty for those rows. Rows whose criterion text is unchanged retain their `✓` status. If the existing matrix has no `✓` markers, or if no existing matrix is found, save the new matrix directly.

**Graceful degradation**: If the input is not a spec (e.g. a brief or description without structured requirements), skip this step — there's no requirement list to verify against.

*Progress note: capture the matrix presentation, any fidelity issues identified and corrected, and user confirmation in the Step 3d summary.*

### Step 4: Confirm

**Cross-epic gap check** (when input is a spec): Before presenting the task tree, read all per-epic coverage matrices produced in Step 3d and take the union of the requirements they cover. A requirement in either of these classes that appears in **no** coverage matrix is a **GAP** — flag it to the user:

- **Must Have** — a requirement under the spec's Must Have heading. The system fails without it.
- **Environmental** — a requirement labelled `ENVn` (something the target must provide) or `ENVXn` (something the work must not require), whatever MoSCoW heading it sits under.

A third class is a gap **despite** being covered, so finding it means reading the criteria rather than counting the rows:

- **Unreachable** — a Must Have requirement naming an action a user takes, whose covering rows are all system-response criteria with none describing the affordance that reaches it (the **Reachability standard** in Step 3). The requirement is ticked, the stories are honest, and no story owns the way in.

Gaps must be resolved (add to an existing epic, create a new story, or defer with justification) before proceeding. Should-have requirements not covered are warnings, not blockers.

An unreachable Must Have blocks on the same terms as an uncovered one, and for the same reason: the coverage matrix answers "is every requirement claimed by some story", which is a different question from "can the delivered system be used". Nothing downstream asks the second question — `cpm:do` verifies criteria as written, and the roll-up counts ticks — so this is the only gate that catches it.

An uncovered environmental constraint blocks for the same reason `coverage-rollup.sh` refuses to let a Scope deferral exclude one: it names something about the target host, and leaving it uncovered does not change the host — it only stops anyone saying so until the work is built and will not run. A design that cannot deploy is not a partial delivery.

Present the full task tree to the user showing:
- All epics with their stories and tasks (using story numbers and task dot-notation)
- Dependencies between epics (cross-epic) and between stories (cross-story)
- Suggested implementation order
- Cross-epic gap check result (all must-haves covered and reachable, or list unresolved gaps)

Use AskUserQuestion for final confirmation.

## Output

Follow the shared **Written Deliverable Length** convention — let the document's length match what the task needs, without padding, redundant summaries, or boilerplate sections.

Save each epic document to `docs/epics/{parent}-{seq}-epic-{slug}.md`. Create the `docs/epics/` directory if it doesn't exist.

- `{parent}-{seq}` is the two-part numeric prefix assigned during Step 2 — see the Epic Filename Convention subsection.
- `{slug}` is the kebab-case name derived from the epic name.

```markdown
# {Epic Name}

**Source spec**: {path to spec}
**Date**: {today's date}
**Status**: Pending
**Blocked by**: —

## {Story Title}
**Story**: {N}
**Status**: Pending
**Blocked by**: —
**Satisfies**: {spec requirement label(s) this story addresses — omit if no spec input}

**Acceptance Criteria**:
- {criterion}
- {criterion}

### {Task Title}
**Task**: {N.1}
**Description**: {Scope within parent story — which acceptance criteria this task covers, or what constraint/boundary it addresses}
**Status**: Pending

### {Task Title}
**Task**: {N.2}
**Description**: {Scope within parent story — which acceptance criteria this task covers, or what constraint/boundary it addresses}
**Status**: Pending

---
```

**Epic-level metadata**:
- `**Source spec**`: Back-reference to the specification that produced this epic. Enables traceability from implementation back to requirements.
- `**Status**`: Derived from stories — `Pending` when no stories are started, `In Progress` when any story is in progress, `Complete` when all stories are complete. (Readers accept `Done` as a synonym for `Complete`, but CPM always *writes* `Complete` — never emit `Done` into a new epic.) Two further values are **terminal and user-set, never auto-derived**: `Superseded` (the epic's work was replaced by other work) and `Withdrawn` (the work was dropped as no longer necessary). CPM never sets these on your behalf — you apply them by hand to retire an epic. A retired epic is excluded from actionable-work enumeration (state derivation, `/cpm:do`, `/cpm:ralph`, `/cpm:status`) and never nags a retro, but its stories still count as done toward progress; it is swept up by `/cpm:archive`. See the shared **CPM Status Model** (`Retired epics`) for the full contract.
- `**Blocked by**`: C

…(truncated)
