# Deepen

> Use when growing an already-playable Godot game along a depth axis (systemic / content / run-meta) without regressing proven behavior. The first ITERATION skill in the GameForge loop — it EXTENDS the rules engine (the deliberate inverse of the asset re-skin's frozen-logic rule), TDD-ing each new sub-system against selftest.gd as a regression guard, records manifest.depth_pass, loops back through validator, and does NOT advance status.

- Skill: `qmertesdorf/deepen` (Agent Skill)
- Install (CLI): `npx skillmds@latest add qmertesdorf/deepen`
- Raw SKILL.md: https://api.skillmd.com/api/skills/qmertesdorf/deepen/raw
- Safety review: pending
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: qmertesdorf (https://skillmd.com/u/qmertesdorf)
- Updated: 2026-09-17
- Page: https://skillmd.com/skills/qmertesdorf/deepen

---


# deepen

Take an already-playable game and grow it along ONE depth axis — systemic (new
interacting mechanics), content (more of the same), or run-meta (map / events /
economy / progression) — without regressing what already works, and **prove** the new
depth landed. Every other GameForge skill is build-once; `deepen` is the loop's first
in-place ITERATION skill.

## Loop position

```
prompt → concept → builder → validator → playtest → ( deepen → validator → playtest )*  → asset → visual-audit → audio → packager
```

`deepen` operates in place on a `validated`/`playable` game and loops back through the
validator. It does **not** advance status — a deeper game is the same status, just
bigger.

## When `deepen` is REQUIRED (the `*` is not zero)

The loop writes `( deepen → validator → playtest )*`, and the POC read the `*` as
"optional" — so it ran **once across eleven titles** (only `deckbuilder-0001` has a
`depth_pass`). That is the single biggest reason the catalogue plays "weak": every other
game shipped at *first-playable depth* — a thin prototype with a solved loop and a content
ceiling a minute in. **A `playable` game is not a finished game; it is a prototype that
proved its loop runs.** At least **one** depth pass is REQUIRED before a title is a
candidate for `asset`/polish — the `asset` skill now gates on the presence of a
`depth_pass` and bounces a never-deepened game back here. Spend iterations *deepening one
good loop*, not restarting near-duplicates (the match3-survival 0001→0002→0003 churn re-
attempted a flawed concept three times instead of deepening one). Run `deepen` until the
**content-ceiling** and **dominant-strategy** tests below both pass at least once.

## Inputs
- A game at status ≥ `validated` with a working `games/<id>/selftest.gd`.
- A chosen depth axis + scope (spec-given, or assessed per the method below).

## Outputs
- Extended game code; a grown `selftest.gd`; a `manifest.depth_pass` record.
- Durable lessons folded back into this skill.

## The method

1. **Assess depth — diagnose *where* it's shallow before choosing what to add.** Run
   these two tests on the current game; they pinpoint the leak nearly every weak GameForge
   title shares:
   - **Content-ceiling test:** play (or read the loop) and ask *"on what beat does the game
     stop introducing anything new?"* — the last new mechanic, enemy, recipe, or rule the
     player meets. Both shipped POC games hit their ceiling early (shopkeep introduces nothing
     after day 4; match3-survival's only escalation is a shrinking timer on one static threat).
     Everything past the ceiling is the same system replayed — that's the "weak" the player
     feels. Your job is to push the ceiling out: stage in newness over the session.
   - **Progression-payoff test (reject DEPTH-AS-MULTIPLIER).** "Newness" must be new
     *gameplay*, not a bigger number on the same decision. For every progression vector the
     game offers — going deeper, levelling, scoring, a higher tier/day/wave — name **what the
     player DOES at tier N that they could not do at tier 1.** If the only answer is "the same
     action, for more points/value," the ceiling has **not** moved — that progression is a
     score multiplier wearing a costume, and it is the single most common reason a game with
     "progression" still feels shallow (diver-0001 shipped exactly this: depth only scaled
     treasure *value*, so there was no gameplay reason to descend — owner-rejected on
     playtest). A progression vector earns its keep only when reaching it **unlocks a new
     decision, a new interaction, or gated content that plays differently** — a destination,
     not a dial.
   - **Dominant-strategy test:** name the single move a skilled player repeats every beat. If
     it's strictly best (most reward *and* safest), the loop is **solved** and no amount of
     content fixes it — you must add a **cost / tradeoff** that makes the dominant move
     situational (this is the systemic axis, and the inverse of `concept`'s tradeoff gate:
     when `concept` let a solved loop through, `deepen` is where it gets repaired). match3
     -survival shipped solved (purge = both the safe move and the high-score move); fixing that
     is higher-leverage than any new content on top.
   - **Core-before-meta gate (don't bolt a wrapper onto a shallow core).** A run-meta or
     content layer *multiplies whatever the core loop already is*. If the core loop is the thin
     part — fails the dominant-strategy test, or offers one real decision — adding upgrades / a
     map / an economy on top just yields a **longer shallow game**, and the meta layer reads as
     busywork because the thing it wraps isn't worth repeating. Before choosing `run-meta` or
     `content`, confirm the **core loop itself** clears the dominant-strategy + progression-payoff
     bars. If it doesn't, the axis is **systemic** (fix the loop first) no matter how tempting the
     shiny meta layer is. (diver-0001's first deepen got this wrong: it added a run-meta upgrade
     economy over a one-decision core, so the upgrades had nothing meaningful to deepen.)
   - **Upgrade / reward-coherence test (for every unlock, upgrade, or reward you add).** Two
     questions per item: (1) *"what does the player DO differently after getting this?"* — if the
     answer is only "the same thing, slightly better," it's a dead stat, not a decision; prefer
     upgrades that **open new play or change a choice** (gate access to new content, enable a new
     tactic, flip a risk calculus) over flat ±X nudges. (2) *"can the player FEEL it on the very
     next run/dive?"* — a buff that needs 3–4 stacked levels before it's perceptible is invisible
     and reads as pointless (diver-0001: +16 air/level barely changed reachable depth, so the shop
     felt meaningless). Tune so one purchase visibly changes what the player can do.
   Then pick **ONE axis** — the one that addresses the leak the tests found — and the single
   highest-leverage expansion on it. Don't widen three axes at once. Per-axis playbook:
   - **systemic** (fixes a *solved/shallow* loop): add an *interacting* mechanic that creates a
     new decision — a cost on the dominant move, a second resource that contends with the first,
     a threat/opportunity that rewards a *different* response than the default. The test: after
     it lands, the dominant-strategy test must no longer have a single answer.
   - **content** (fixes a *low* ceiling on a loop that's *already* a real decision): more of the
     same *kind* — new recipes/enemies/cards/tiles — **staged in over the session**, not all at
     start, so the ceiling moves. Only reach for this once the dominant-strategy test passes;
     piling content onto a solved loop just makes a longer solved loop.
   - **run-meta** (fixes *"no reason to start run 2"*): a wrapper that makes sessions differ and
     accrue — a map/event/economy/unlock track, escalating modifiers, a persisted best or
     prestige. Gives the loop somewhere to *go* across plays.
2. **Decompose into sub-systems.** Each one purpose, with a clean interface
   (data layer + logic + screen). Find the **extension seams**: where the existing code
   already supports growth (e.g. a static-func data table) vs. where you must
   **refactor to create a seam first** (e.g. a hardcoded linear sequence → a
   data-driven state machine). **Refactor-for-seam before adding content.**
3. **TDD each new system on the self-test — the validation spine:**
   - **`deepen` EXTENDS the logic; it does NOT freeze it.** This is the deliberate
     inverse of the `asset`/re-skin "logic FROZEN" rule. Confusing the two is the
     classic mistake — re-skinning must not touch rules; deepening is *all about*
     touching them, safely.
   - Existing assertions are the **regression guard**. `SELFTEST OK` must hold after
     every change — and `UITEST OK` too if a `uitest.gd` exists: deepening adds
     screens and controls, and a new view that renders fine can still swallow taps
     or skip its rebuild event (invisible to selftest, which bypasses the view).
     New tappable screens get new `uitest.gd` checks, same RED→GREEN discipline.
   - **`PLAYTEST OK` too if a `games/<id>/playtest.gd` exists — deepening changes TUNING,
     which is exactly what breaks winnability.** A depth pass that retunes costs, gate
     depths, spawn geometry, or the ramp can make the game unwinnable while every logic
     assertion stays green (the `playtest-audit` skill exists because of exactly this).
     Re-run the balance bot after the pass; a `PLAYTEST FAIL` on re-validation is attributed
     to **this `deepen` pass**, and the fix is tuning, never weakening `selftest`.
   - **Keep the generate-and-verify gate green if the game has one (`make_verified` +
     `Solver`, per `builder`).** Adding content on the **content axis** is the classic way to
     silently introduce unsolvable instances: every new tile / recipe / wave type / map piece
     *widens the instance space the generator can deal*, and the solver guarantee only held
     over the *old* space. Re-run the generate-and-verify selftest assertions (every
     `make_verified` solvable over K seeds; fallback rate still rare; fallback still solvable)
     after the pass — a regression here is attributed to **this `deepen` pass**, and the fix is
     the generator / new content, **never weakening `Solver.is_solvable`**. If the depth pass
     adds discrete generated content to a game that *didn't* have the gate (e.g. the content
     axis turns a fixed layout into a procedural one), that is exactly when to **introduce**
     `make_verified` — treat it as a new sub-system with its own RED→GREEN assertions.
   - For each new system, **write its assertion first (RED) → implement → GREEN.**
     Prove new mechanics the same deterministic, headless way the original logic was.
   - **Never weaken or delete an existing assertion to make room.** If a new system
     genuinely changes old behavior, *surface it* — call it out and confirm it's
     intended — never silently overwrite the guard.
   - **A pure refactor adds no new behavior.** If the behavior it restructures isn't
     already covered, pin it with a **characterization assertion first** (one that
     passes both before and after the refactor), then refactor. Pure refactors add no
     *new-behavior* assertions.

### Balance tuning (parameter search) — propose a config, don't hand-guess

A pass that changes **tuning** (systemic/run-meta retunes costs, gate depths, drain,
spawn geometry, the ramp) is exactly what makes a game unwinnable or unfair while every
logic assertion stays green. Instead of hand-guessing constants and re-running the bot,
**search** the tuning space against `playtest-audit`'s metrics:

1. **Add the `GF_TUNE`/`GF_SEED` seam** (a `preload`-able `Tune` static, per
   `games/diver-0001/Tune.gd`): the data layer reads each tunable from `Tune.num(...)`
   defaulting to its `const`. UNSET env → identical production behavior. Only seam the
   constants you intend to search.
2. **Emit the metrics contract:** `playtest.gd` must print one `PLAYTEST METRICS {json}`
   line (see `playtest-audit`) carrying the numbers it already computes + the invariant
   booleans.
3. **Declare a `balance.spec.json`** (search space + objective). The objective HARD-REJECTS
   any config failing the playtest invariants, then scores survivors by **distance OUTSIDE
   target BANDS** — never by maximization (maximizing earnings/clear-rate yields a trivially
   easy game). Use **single-player** metrics + a **retention/engagement proxy** (low
   time-to-first-goal, accruing-but-not-instant economy, a smooth/tight difficulty curve via
   the air/HP margin, did pushing pay off). **Pick floor vs. two-sided band per metric's
   meaning** (see Lesson 1): a *cautious-bot solvency* rate (does a careful player reliably
   succeed?) is a **hard floor in `require`**, NOT a two-sided band — banding it penalizes the
   very robustness you want. Two-sided bands fit metrics whose value reflects *difficulty/pacing*
   (time-to-first-goal, air margin, commissions filled), or a win-rate that genuinely reflects
   challenge (a roguelike clear-rate). NOT win-rate disparity. The realistic retention bar is top-quartile
   ~7-8% D7 (GameAnalytics def) — do NOT anchor on the old unverified "20%"; and we do not
   literally measure D7, so the proxy is a heuristic.
4. **Run** `node tools/balance.mjs <game-dir> <spec.json>` (each candidate is run across K
   seeds so "clear-rate" is meaningful and the config isn't seed-lucky).
5. **READ the per-focus-point breakdown + the non-dominated shortlist and CHOOSE** — weigh
   the tradeoffs yourself (great pacing vs. borderline economy); do not blindly take the
   lowest composite. The composite is a heuristic sort key, not a verdict.
6. **Apply** the chosen config to the defaults, then re-run the **full** gate set
   (`SELFTEST` / `UITEST` / `PLAYTEST`) with env unset.

**Honesty rule (load-bearing):** the tool **proposes**; the **human playtest decides fun**.
No validated automated fun proxy exists — the search guarantees winnable/fair/well-paced,
never *fun*. An owner "this isn't fun" verdict overrides any proxy win. Record the chosen
config + why in `depth_pass.notes`.

**Lessons from first use (diver-0001 dogfood):**
- **Lesson 1 — solvency is a FLOOR, not a band.** The first objective two-sided-banded
  `clear_rate` at `[0.6, 0.9]` and the search penalised the diver for the cautious bot
  *always* banking (100%). But for push-your-luck (and most solo games) a careful player
  reliably succeeding is exactly the property you want — the *risk* lives in the player
  *choosing* to push deep, captured by the commission / margin metrics. Model cautious-bot
  solvency as `require: { clear_rate: ≥X }` and leave it out of `bands`.
- **Lesson 2 — a result that's FLAT across the whole space is a finding, not a failure.**
  When every config ties on a residual penalty (diver: `commissions_filled = 1` for all 36
  configs), the gap is **not tuning-fixable** — it is bot-skill- or structure-limited (here
  the competent bot can't grab sparse deep qualifying treasures, exactly the limitation
  `playtest-audit` says to *report, not gate*). **Do NOT change constants to chase a
  bot-unreachable metric** — that is the maximize-the-proxy mistake. Report it as a
  human-playtest item and move on; "no change warranted" is a legitimate, honest outcome of a
  balance pass.

4. **One sub-system at a time**, each independently self-tested and committed. Don't
   batch five then debug the soup. Keep a playable game at every step.
5. **Grow the UI per system**, reusing established chrome. Hand composited-screen
   judgment to `visual-audit` and correctness to `validator`. `deepen` owns *systems &
   content* — not pixels, not the gate mechanics. **New screens aren't self-test-gated**
   (the headless self-test never instantiates the scene tree), so gate them two other
   ways: a **headless boot check** (`godot --headless --path … --quit-after N`) that
   proves the router/view code parses and runs with no `SCRIPT ERROR`, plus a
   **throwaway real-renderer harness** (a `SceneTree` script that builds the relevant
   state, instantiates the view, waits ~200 frames, saves a PNG) for a visual sanity
   glance. Delete the throwaway; keep the PNG as a probe-data artifact.
6. **Verify the depth landed — INDEPENDENT design-depth audit (REQUIRED).** The agent that
   did the deepening cannot grade its own depth: it knows what it *intended* to add and reads
   the diff as proof, so it ships "bigger" believing it shipped "deeper" (diver-0001's first
   deepen passed its own assessment and was still owner-rejected as shallow). Fix it the way
   `visual-audit` fixes the screen — with **fresh, adversarial eyes**, but pointed at the
   *systems* instead of the pixels. **Dispatch a fresh subagent** (no knowledge of what you
   set out to add — give it only the running game + the concept) to play/read it and answer,
   bluntly:
   - **Is each progression vector a destination or a dial?** For going deeper / levelling /
     scoring: *what does the player DO at the top that they couldn't at the bottom?* "Same
     action, more points" = FAIL (depth-as-multiplier).
   - **Do the new systems change decisions?** Name a concrete moment the new mechanic/upgrade
     made the player choose differently. If none, it's inert.
   - **Is it more fun, or just more?** One sentence: did this make the game deeper, or longer?
   Treat a "just bigger / just longer" verdict as a **failed pass** — iterate (often the real
   fix is a different axis: the auditor saying "the upgrades are meaningless because the core
   loop is one decision" means you picked run-meta when the answer was systemic). Record the
   auditor's verdict in `depth_pass.notes`. Scale the audit to the change: one skeptic for a
   small content add, a fuller play-and-critique for a systemic/run-meta pass.
7. **Record + codify.** Write `manifest.depth_pass` (axis, systems added, new-assertion
   count, **and the independent audit's verdict**). Fold durable lessons back into this skill.

## manifest.depth_pass

```json
"depth_pass": {
  "axis": "run-meta | systemic | content",
  "systems_added": ["..."],
  "selftest_assertions_added": 0,
  "notes": "what changed, and any surfaced behavior-changes to previously-frozen logic"
}
```

## Boundaries / non-goals
- Not a re-skin (`asset` + `visual-audit`) and not the audio pass.
- Does not invent a new status or touch the packaging gate.
- Does not redesign from scratch — it grows what exists along one axis.

## Project gotchas (carry these)
- Headless `godot --script` does NOT instantiate autoloads → data layers via
  `preload` + `static func`.
- Seed every RNG; Fisher–Yates, never `Array.shuffle()`.
- Reset `user://save.json` before asserting on meta writes (stale-file false positives).
- A growing `selftest.gd` runs all stages in **one function scope** → give each stage's
  locals **unique names** (e.g. suffix with the stage number) or you get redeclaration
  parse errors as you append.
- Avoid GDScript method names that collide with `Object` built-ins (`connect`, `draw`,
  `set`, …) on your data/model classes — they parse-error or shadow silently.
- A new `manifest.depth_pass` field is **not free**: the manifest schema is
  `additionalProperties: false`, so add the field to `schema/manifest.schema.json` and
  re-run `node tools/manifest.mjs validate <id>` + the vitest suite before committing.

## Lessons from first use (run-layer dogfood)
- **The "surface, don't swallow" rule earns its keep.** Two changes touched
  previously-frozen combat logic — threading run-persistent HP through `setup()`, and
  giving a relic that had silently been a no-op a real effect. Both were named in
  `depth_pass.notes` rather than slipped in. When deepening forces a change to old
  behavior, that is normal — make it loud.
- **Prove the system headless first, wire the screen second.** Every sub-system landed
  its self-test assertion *before* any view existed, so the regression gate never
  depended on rendering. This ordering is what let view work stay a separate, lower-risk
  concern handed to `visual-audit`.
- **Check the acquisition path, not just the hook.** A hook that only fires at one
  moment (e.g. run-start) is dormant for anything acquired *after* that moment. When you
  add hook points, confirm the real in-game path that grants the thing actually triggers
  the hook — or record the limitation explicitly instead of shipping a dead feature.

## Lessons from second use (shopkeep-0001 systemic dogfood — a "triage" loop)

The stated hook (a cashier triage: "who do I serve next?") was INERT, and it took THREE
iterations under the independent audit to fix. Each round failed for a *different* reason
the implementer couldn't see — the audit is what caught them. The durable lessons:

- **A "choice" loop needs CONTENTION — a scarce resource the action itself consumes —
  before any decision exists.** `serve()` was an instant, free, exact-match action on
  independent shelves, so "who next?" had a trivially optimal answer (serve whoever's about
  to leave) and was a *reflex*. **Adding more options or values does NOT help while you can
  satisfy everyone.** Decoupling value from urgency (tourist = cheap+impatient vs. regular =
  rich+patient) was completely inert until a **serve-time cost** (a register cooldown =
  throughput) made serving A literally spend a resource B needed. Diagnose missing contention
  FIRST when a "decision" loop feels flat; a serve/triage/allocation loop with no action
  budget, cooldown, or contested stock is a reflex no matter how many patron types you add.
- **A "destination" must add a VERB or gated content — a price multiplier is a dial in a
  costume,** even pre-announced or flavored. Reputation "unlocking Regulars" who wanted the
  same demand items you'd already craft, paying 2×, was just a multiplier. It only became a
  real destination when Regulars placed **pre-announced standing orders for top-tier goods
  you must deliberately pre-stock** — i.e. it changed the *CRAFT* decision, a new verb.
- **A reward/commitment with no penalty for ignoring it is optional flavor, not a decision.**
  Standing orders changed nothing until an **unfilled order cost reputation** — only then did
  spending scarce materials to pre-stock them become a genuine bet.
- **Structure can LAND while the verdict stays "conditional on tuning" — and that is the
  STOP signal, not a cue to keep twiddling constants.** After three iterations the audit went
  from "inert" to "genuinely well-constructed decision … JUST BIGGER/LONGER, conditional on
  tuning": the *mechanism* was right, but whether the contention BITES often enough is a
  tuning question (serve time vs. patience fuses vs. queue size vs. spawn rate). That belongs
  to a **`playtest.gd` bot + `tools/balance.mjs` search + the human fun check**, NOT to blind
  hand-tuning toward the proxy auditor (that is the maximize-the-proxy mistake). Recognize the
  boundary: once the structure is sound, stop iterating it and hand tuning to the balance pass.
- **The independent audit pays for itself every single round.** Three rounds, three "just
  bigger" verdicts, three *different* real flaws pinpointed (dial-reputation → no-contention →
  toothless-orders). Self-assessment would have shipped after round 1 believing it was deep.
  Re-dispatch a FRESH subagent each iteration — a re-used one anchors on its prior read.
- **A packaged game with no `depth_pass` and no `playtest.gd` is the loud symptom of the
  original POC gap** (polish shipped on an unverified-deep, unverified-winnable core). When you
  re-open one, expect the depth pass to *also* surface the missing winnability bot — record it
  as a required follow-up even if you don't build it in the same pass.

## Lessons from the shopkeep-0001 BALANCE pass (building the playtest bot + running the search)

The required-next from the systemic pass — build `playtest.gd`, add the `GF_TUNE` seam, run
`tools/balance.mjs` — and what the bot actually *found* turned a "conditional on tuning"
verdict into a live decision. The durable lessons:

- **A new contention mechanic can be silently MASKED by a co-located scarcity — measure WHICH
  constraint binds before assuming your mechanic drives the decision.** The serve-time cooldown
  was completely inert at the shipped tuning, not because the cooldown was too short, but
  because a *different* scarce resource bound first: the shop is shelf-stock-limited and the
  "empty shelves end the day" rule is a soft landing that sells out before the register ever
  saturates (`blocked_ticks = 0`, `soldout_days = every day`). The fix needed dense arrivals
  (spawn_interval) to make the register the bottleneck. **Instrument the candidate bottlenecks**
  (a "register-busy-while-a-servable-patron-waits" counter, a "sold-out" flag) and confirm your
  new mechanic is the one that binds — a green winnability gate hides which constraint is live.
- **`forced_walkouts` was the WRONG oracle; prove "it's a real decision" by POLICY DIVERGENCE.**
  The intuitive metric (a patron times out with their item still on a shelf) ~never fired,
  because the stockout ended the day first. The honest test: run TWO reasonable policies on the
  **same seed** (here value-first vs urgent-first cashier) and compare outcomes. Convergence
  (Δ≈0) = the choice is inert; **bidirectional** divergence (policy A wins some seeds, B wins
  others — *no dominant policy*) is the signature of a genuine tradeoff. This is stronger and
  more honest than any single bot's score, and it directly answers the audit's "do the new
  systems change decisions?" — empirically, by playing, not on paper.
- **Decision-pressure is often SEED-DEPENDENT — judge it in AGGREGATE (a banded mean), never as
  a per-seed hard gate.** Whether a given run pits value against urgency depends on the random
  patron mix; ~half the seeds were inert even on a good config. Gating `no_trivial_dominant` on
  per-seed divergence flaked the whole search to "0 configs survive." Fix: **hard-floor the
  ROBUST winnability invariants per-seed** (solvent / first-goal / no-death-spiral / excess-
  demand-exists), and put the **intermittent decision-quality metric in a banded mean** the
  search optimises across seeds. (Pairs with Lesson 1 from the diver dogfood: floors vs bands.)
- **A balance search can park a real decision on a fragile tuning KNIFE-EDGE — record that as a
  limitation, not a win, and do NOT chase it with more tuning.** The independent audit confirmed
  the triage is real but only in `serve_time ≈ 2.5–3.5` (it vanishes at 2.0), and that the
  sharpest tension turned out to be a *strategic* gold-now-vs-reputation-later axis, not the
  split-second patience triage the concept advertised. That fragility is honest signal for the
  HUMAN playtest and a possible *future structural* iteration (e.g. remove the sellout soft-
  landing so stock-scarcity stops masking the register) — not a cue to keep twiddling constants
  toward the proxy. Once the structure is proven real, the remaining "is it FUN / does the
  knife-edge feel good" is the human's call, full stop.

