Set up the codebase harness
Harness engineering: the model is fixed — what you engineer is the scaffolding
around it (the environment, the docs, the feedback loops) so an agent can build and
verify software with minimal human attention. Humans steer; agents execute. Your
job is to make the repo legible, executable, and verifiable.
Work incrementally and depth-first: assess what exists, build the one missing
capability, use it to unlock the next. Don't boil the ocean — set up what the repo
actually needs. When the agent struggles, the fix is almost never "try harder" —
ask "what capability is missing, and how do I make it legible and enforceable?"
and add it.
This skill orchestrates the focused sub-skills: dev-local-setup,
e2e-setup, crabbox-setup (cloud/parallel), and pr.
0. Assess
Survey the repo: stack, package manager, services/ports, infra deps, existing
docs/tests/CI, and the implicit rules (buried in READMEs, PR comments, people's
heads). Note what's missing per pillar below.
1. Legible — the agent can reason about the repo
What the agent can't see doesn't exist. Knowledge in chat threads / heads is
invisible — push it into versioned, repo-local artifacts.
- a) Map, not manual. Shrink the root agent doc (
AGENTS.md / CLAUDE.md) to a
~100-line table of contents: one-line overview, project tree, golden rules
(the hard invariants), and a "where to look" table. Move the depth into a
structured docs/ system-of-record (architecture, frontend, testing, domain
topics) with a docs/index.md. A monolithic instruction file rots and crowds out
the task — keep the map small and stable, disclose detail progressively.
- b) Custom lints with remediation. Promote the prose golden rules into
mechanical checks — human taste captured once, enforced everywhere, every run.
One lint per invariant (layering / dependency direction, naming, no-
any,
forbidden imports, file-size, structured logging). Write the error message to
inject the fix ("X isn't allowed here — do Y") so the remediation lands in agent
context. Wire them into the repo's linter + CI.
- c) Queryable code graph. Index the repo with
codebase-memory-mcp
so the agent traces callers, data flow, and architecture from a knowledge graph
instead of blind grepping — faster, more precise navigation on large codebases.
- d) (later) Keep docs honest. A freshness / doc-gardening pass that flags docs
that no longer match the code and opens fix-up PRs.
2. Executable — the agent can run & drive the app
dev-local-setup → a one-command, reproducible local stack
(scripts/dev-local.sh up) running every service + infra.
- Make the app drivable: browser via the
playwright-cli skill; logs reachable.
crabbox-setup → an isolated cloud box per agent — the parallel-safe
counterpart to dev-local. Reach for it when loops run concurrently: one laptop
can't host N full stacks (fixed ports, one Docker daemon, one DB), and per-worktree
local doesn't fix it — the worktrees still share the host. crabbox gives each agent
its own stack + an in-box browser, so parallel verification never collides.
- Advanced: a local, ephemeral observability stack (queryable logs/metrics) for
perf/reliability prompts.
3. Verifiable — the agent can prove it works
e2e-setup → a trustworthy e2e gate: real flows (not bypass), a reusable
auth/session helper, layered client → server → product assertions, video/trace
evidence, sandbox-only external services.
pr → the verify-before-ship loop: a fresh verifier sub-agent drives the
real app to confirm the just-built feature works; the main agent fixes until
green, runs the codified regression sweep, and opens a PR with a reviewable proof
link. Add the session helper so the verifier can reach login-gated features.
4. Others — keep it coherent over time
- Commit hygiene: conventional commits + format/lint on commit (e.g. husky
lint-staged + commitlint). Keep merge gates light — at high agent throughput,
corrections are cheap and waiting is expensive.
- Garbage collection: encode "golden principles", then run periodic cleanup
passes that open small refactor PRs — pay tech debt down continuously, not in
painful bursts. Human taste captured once, enforced on every line.
- Agent-to-agent review for correctness-critical changes (independent reviewers,
not self-review).
Order & what you leave behind
1a (map) → 2 (dev-local) → 3 (e2e + /pr), then 1b (lints) and 4 as the
repo matures. The artifacts — slim map + docs/, scripts/dev-local.sh, an e2e/
suite, the /pr skill, and custom lints — are each a reusable, legible capability
that compounds. Prefer "boring", composable, stable tech the agent can fully model.
1---2name: setup-codebase-harness3description: Sets up a codebase for reliable agent-driven development by making it legible (structured docs, custom lints, code graph), executable (one-command dev stack, cloud sandbox), and verifiable (e2e gate, verify-before-ship loop).4---56# Set up the codebase harness78**Harness engineering:** the model is fixed — what you engineer is the *scaffolding*9around it (the environment, the docs, the feedback loops) so an agent can build and10verify software with minimal human attention. Humans steer; agents execute. Your11job is to make the repo **legible, executable, and verifiable.**1213Work **incrementally and depth-first**: assess what exists, build the one missing14capability, use it to unlock the next. Don't boil the ocean — set up what the repo15actually needs. When the agent struggles, the fix is almost never "try harder" —16ask *"what capability is missing, and how do I make it legible and enforceable?"*17and add it.1819This skill orchestrates the focused sub-skills: **`dev-local-setup`**,20**`e2e-setup`**, **`crabbox-setup`** (cloud/parallel), and **`pr`**.2122## 0. Assess2324Survey the repo: stack, package manager, services/ports, infra deps, existing25docs/tests/CI, and the *implicit* rules (buried in READMEs, PR comments, people's26heads). Note what's missing per pillar below.2728## 1. Legible — the agent can reason about the repo2930> What the agent can't see doesn't exist. Knowledge in chat threads / heads is31> invisible — push it into versioned, repo-local artifacts.3233- **a) Map, not manual.** Shrink the root agent doc (`AGENTS.md` / `CLAUDE.md`) to a34 ~100-line **table of contents**: one-line overview, project tree, **golden rules**35 (the hard invariants), and a "where to look" table. Move the depth into a36 structured **`docs/` system-of-record** (architecture, frontend, testing, domain37 topics) with a `docs/index.md`. A monolithic instruction file rots and crowds out38 the task — keep the map small and stable, disclose detail progressively.39- **b) Custom lints with remediation.** Promote the prose golden rules into40 **mechanical checks** — human taste captured once, enforced everywhere, every run.41 One lint per invariant (layering / dependency direction, naming, no-`any`,42 forbidden imports, file-size, structured logging). **Write the error message to43 inject the fix** ("X isn't allowed here — do Y") so the remediation lands in agent44 context. Wire them into the repo's linter + CI.45- **c) Queryable code graph.** Index the repo with [`codebase-memory-mcp`](https://github.com/DeusData/codebase-memory-mcp)46 so the agent traces callers, data flow, and architecture from a knowledge graph47 instead of blind grepping — faster, more precise navigation on large codebases.48- **d) (later) Keep docs honest.** A freshness / doc-gardening pass that flags docs49 that no longer match the code and opens fix-up PRs.5051## 2. Executable — the agent can run & drive the app5253- **`dev-local-setup`** → a one-command, reproducible local stack54 (`scripts/dev-local.sh up`) running every service + infra.55- Make the app **drivable**: browser via the `playwright-cli` skill; logs reachable.56- **`crabbox-setup`** → an **isolated cloud box per agent** — the parallel-safe57 counterpart to dev-local. Reach for it when loops run **concurrently**: one laptop58 can't host N full stacks (fixed ports, one Docker daemon, one DB), and per-worktree59 local doesn't fix it — the worktrees still share the host. crabbox gives each agent60 its own stack + an in-box browser, so parallel verification never collides.61- *Advanced:* a local, ephemeral observability stack (queryable logs/metrics) for62 perf/reliability prompts.6364## 3. Verifiable — the agent can prove it works6566- **`e2e-setup`** → a trustworthy e2e gate: real flows (not bypass), a reusable67 auth/session helper, layered client → server → product assertions, video/trace68 evidence, sandbox-only external services.69- **`pr`** → the verify-before-ship loop: a fresh **verifier sub-agent drives the70 real app** to confirm the just-built feature works; the main agent fixes until71 green, runs the codified regression sweep, and opens a PR with a reviewable proof72 link. Add the session helper so the verifier can reach login-gated features.7374## 4. Others — keep it coherent over time7576- **Commit hygiene**: conventional commits + format/lint on commit (e.g. husky77 lint-staged + commitlint). Keep merge gates **light** — at high agent throughput,78 corrections are cheap and waiting is expensive.79- **Garbage collection**: encode "golden principles", then run periodic cleanup80 passes that open small refactor PRs — pay tech debt down continuously, not in81 painful bursts. Human taste captured once, enforced on every line.82- **Agent-to-agent review** for correctness-critical changes (independent reviewers,83 not self-review).8485## Order & what you leave behind8687**1a (map) → 2 (dev-local) → 3 (e2e + /pr)**, then **1b (lints)** and **4** as the88repo matures. The artifacts — slim map + `docs/`, `scripts/dev-local.sh`, an `e2e/`89suite, the `/pr` skill, and custom lints — are each a reusable, legible capability90that compounds. Prefer "boring", composable, stable tech the agent can fully model.