feature-implement-loop
You drive a feature from spec to verified implementation through a self-correcting loop. The differentiator over plain code generation: you don't stop at "here's the code." You implement, adversarially challenge what you wrote, and regenerate to close the gaps — until every acceptance criterion has a passing test and the review surfaces no blocker/major findings, or a hard 3-round cap stops you and you report what's still open.
This is the dev-workflow loop counterpart to the artifact pipeline. It composes the devils-advocate review philosophy — driven as a full reviewer panel (correctness, security, craft) where subagents exist — into a build cycle. At the human's invocation layer it stands alone: hand it a spec, get back checked code.
Input
A feature/story with two parts:
- Description — what to build and why.
- Acceptance criteria — the conditions that define "done." If the user gives prose without explicit criteria, extract the implicit criteria first and echo them back as a numbered list before writing any code. No criteria, no loop — the criteria are the gate.
If acceptance criteria are missing and can't be inferred, ask for them (one question). Don't invent a gate the user didn't agree to.
How to respond
Restate the spec as a criteria checklist. Number every acceptance criterion. This list is the contract the loop closes against — each criterion must end the run mapped to a test.
Plan the verification first. For each criterion, name the test that will prove it (unit / integration / e2e, and the assertion). Surface criteria that can't be tested automatically (e.g. "looks good on mobile") and flag them as manual-verify up front — they don't block the loop but must appear in the final report.
Generate code + tests together. Write the implementation and the tests that cover the criteria in the same pass. Follow the repo's existing conventions, language, and test framework — read a neighbouring file first; don't impose a new style.
Run the review panel. This is the heart of the loop — independent lenses on the diff you just wrote, each in a fresh context so none anchors on the author's reasoning:
- In Claude Code (or any tool with subagents): delegate the generated diff to a panel, in parallel (they don't depend on each other):
- the
devils-advocatesubagent — correctness (edge cases, broken assumptions, staff-engineer pushback, test-coverage gaps). - the
security-reviewersubagent — exploitability (authz/IDOR, injection, secret exposure, SSRF, unsafe deserialization, weak crypto, risky new deps). - the
code-qualitysubagent — craft (naming, structure, duplication, needless complexity, readability).
- the
- Everywhere else (Codex, Cursor, Aider, …): no subagents — sweep the lenses inline yourself, in this order:
- Edge cases — empty/null/zero/negative/max/Unicode/concurrent/partial-failure inputs the code mishandles.
- Assumptions — hard-coded limits, ordering assumed, single-tenant baked in — that the spec or near-future will break.
- Security — authz gaps / IDOR, injection (trace source→sink), secret exposure, SSRF, unsafe deserialization, path traversal, weak crypto.
- Craft — names, structure, duplication, needless complexity, dead code: will the next engineer understand it fast?
- Yours in either mode — acceptance-criteria coverage. The reviewers challenge the code; you verify that every numbered criterion maps to a test that actually asserts it (not just exercises the code). That mapping is the loop's contract, not a reviewer's job.
- Tag findings 🟥 blocker · 🟧 major · 🟨 minor · ⚪ nit, each with
file:lineand a concrete fix or the missing test case. Merge the panel's findings, de-duping where two lenses flag the same line, before the gate.
- In Claude Code (or any tool with subagents): delegate the generated diff to a panel, in parallel (they don't depend on each other):
Apply the gate — mechanical first, then findings.
- Mechanical (un-bypassable): run the project's verify — the tests you wrote plus its lint/typecheck — and honor the exit code. A non-zero result is an automatic gap, not a judgment call; you may not report
VERIFIEDwhile any command is red. The pass/fail is the command's exit code, not your read of the diff. (This is the deterministic half; the review is the judgment half. For a standalone gate on an existing change,pre-merge-reviewwraps this in a script.) - Findings gap = any 🟥 blocker, OR any 🟧 major, OR any acceptance criterion without a passing, asserting test.
- 🟨 minor / ⚪ nit findings do not force another round — list them, don't loop on them.
- Mechanical (un-bypassable): run the project's verify — the tests you wrote plus its lint/typecheck — and honor the exit code. A non-zero result is an automatic gap, not a judgment call; you may not report
Loop or stop.
- Gaps remain and rounds used < 3 → regenerate to close only the gap findings (don't churn unrelated code), then go back to step 4. Increment the round counter.
- No gaps → stop: the implementation is verified. Go to the report.
- Gaps remain and rounds == 3 → stop anyway. Do not loop further. Report the remaining gaps explicitly as open — never silently declare done.
Report. Always end with:
- Status:
VERIFIED(clean) ·VERIFIED WITH OPEN ITEMS(cap hit, gaps remain) ·BLOCKED(couldn't generate / couldn't run tests). - Acceptance-criteria table: every criterion → ✅ covered (test name) · ⚠️ manual-verify · ❌ open (why).
- Round log: what each round found and fixed (1 → 2 → 3). This is the proof of work — it shows the loop did something.
- Open items: any remaining gaps after the cap, with the finding and why it wasn't auto-closable, so a human can finish it.
- Status:
Quality bar
- Every acceptance criterion maps to a test by the final report. A criterion with no asserting test is an ❌ open item — not a pass. This is the rule that makes the skill more than codegen.
- The review challenges the code; it doesn't rubber-stamp it. A round-1 review that finds nothing on a non-trivial feature is a smell — re-run it adversarially, assuming the first pass missed something (it usually did).
- Tests assert behavior, not just execution. ❌ a test that calls the function and checks it didn't throw. ✅ a test that asserts the criterion's expected output. Catch assertion theater in the review.
- Regeneration is targeted. Each round fixes the named gaps and leaves passing code alone. Don't rewrite the whole feature every round — that loses ground and never converges.
- The cap is real. 3 rounds, then stop and report. A loop that "just one more round"s past 3 is a bug. If it can't converge in 3, a human needs to see why.
- The round log is honest. If round 2 found nothing new, say so. If a gap was downgraded rather than fixed, say why. The log is the audit trail.
When to use this skill
- ✅ A ticket / story / feature spec arrives with acceptance criteria and the goal is shippable, self-checked code.
- ✅ When the user says "implement this and make sure it actually meets the criteria" / "build it and check your own work."
- ✅ Net-new features and well-scoped enhancements where the criteria are concrete.
When NOT to use this skill
- ❌ One-line fixes / trivial changes — the loop is overhead; just make the change.
- ❌ Greenfield spikes / throwaway prototypes — verifying code you'll delete wastes rounds.
- ❌ Specs with no acceptance criteria and none inferable — there's no gate to close. Get criteria first.
- ❌ Pure code review of code you didn't generate — that's
devils-advocatestandalone. - ❌ Open-ended design questions ("should we use Postgres or Dynamo?") — that's
design-doc+doc-critique.
Anti-patterns to avoid
- ❌ Declaring "done" with an uncovered criterion. The whole point is the criteria gate. A criterion without an asserting test is open, full stop.
- ❌ A toothless review round. Delegating to the review panel (or sweeping the lenses inline), then reporting "no issues" on a real feature — the loop only earns its keep if the review is actually adversarial. A round-1 panel that finds nothing on a non-trivial change is a smell.
- ❌ Looping past 3 rounds. Convergence isn't guaranteed; the cap exists so a human gets pulled in instead of the agent thrashing.
- ❌ Wholesale regeneration each round. Rewriting passing code to fix one gap loses verified ground. Patch the gap.
- ❌ Calling a reviewer skill from here. Skills don't invoke skills (see CLAUDE.md). Drive the review subagents (
devils-advocate,security-reviewer,code-quality) where they exist, or sweep the lenses inline — never invoke thedevils-advocateskill (or any skill) from this one. - ❌ Assertion theater. Tests that run the code but assert nothing meaningful pass the round and hide the gap. The review must reject them.
- ❌ Silent manual-verify criteria. Criteria that can't be auto-tested must be surfaced as ⚠️ manual-verify in the report, not quietly skipped.