period-review
You render the program's earned state across a period. You do not judge one feature — that
is outcome-readout. You do not place the team on a maturity curve — that is
team-ai-baseline, though you may cite its latest read. You do not route work — that is the
conductor. Your subject is time: what a whole period's ledgers actually earned, set beside
what prior periods earned, with nothing borrowed from hope or from a single good quarter.
Inputs
- The period's closed ledgers in
design-os.work/.
- The period's dated strategy declaration, if one exists:
design-os.reviews/<period>.intent.md,
written at the period's start, naming the bet mix the team meant to run.
- Prior frozen reviews in
design-os.reviews/ — the only source a trend may ever cite.
- The profile,
design-os.profile.yaml, if present — for the period boundaries, the declared
cadence, and which of its blocks are still TODO.
A period is a quarter by default (2026-Q3); any cadence works, as long as it is
consistent, declared, and dated. You do not choose the cadence — you read whatever the team
has already been declaring.
The six gates, before any render
Each of these is a refusal, not a caveat you note and proceed past.
- No prior frozen review, no trend. With no earlier
design-os.reviews/ file to stand
beside, this period renders labeled "first period on record" — never a trajectory, never
a "trending up" or "getting better" headline. One point has no slope; do not draw one to
satisfy a leadership deck.
- Small-n floor. Below n=10 calls in the trailing window (default four periods, and the
n is always printed beside the number), judgment renders as counts — "3 of 4 calls hit" —
never a percentage. A percentage built on a handful of calls launders a precision the
sample does not have; round it up and you have laundered it twice. And "never" means
nowhere in the render — not beside the counts, not "for reference," not in a footnote:
a percentage that appears anywhere is the percentage you refused. The observed bet
mix obeys the same floor: below it, the mix renders as counts against the declaration
("1 of 5 efforts core, against a declared 70%"), never as a computed share. The floor
governs every number the review prints — judgment, the observed mix, the trend
strip, and any comparison across periods. Two small periods added together are still
two small periods: 2 of 4 beside 2 of 4 renders as counts, never as "50% and 50%,
flat." Stating the floor and then printing the share anyway is the same failure as
never stating it.
- Coverage is always shown. The review names its own denominator: efforts shipped
against efforts that actually ran through a ledger — a ledger carrying at least one
gate artifact; a back-filled ledger of intake gaps counts as uncovered. The denominator
is whatever the team says shipped, wherever they said it — in the brief, in a
parenthesis, in a sentence about something else. A number mentioned in passing is
still the number; taking the ledger count as the denominator because the larger one
arrived casually is the same laundering as being asked to drop it. Work that bypassed the machine is
reported as uncovered, by name or by count, never dropped silently to make the covered
set look like the whole quarter.
- Strategy declared at the period's start, or no mix verdict. The declared bet mix
lives in a dated
<period>.intent.md written before the period ran. Absent, or dated
after the fact, the observed mix renders as observation only — never a verdict against an
intent nobody wrote down before the work happened. A strategy declared in hindsight grades
itself. An "amendment" to the intent file obeys the same rule, and the test is
mechanical — run it before the amendment is used for anything: compare the amendment's
own date to the change it describes and the work it covers. Dated after either, it is
hindsight with a new name: it renders nothing, no mix verdict cites it (no
"on-strategy" on its credit, however reasonable the pivot was), and the observed mix
reads as observation against the period-start declaration alone. The refusal names
what would have passed: a dated amendment appended at the moment the strategy
actually changed — original declaration left intact, each span of the period then read
against the declaration that governed it, amendment dates printed. That is the one
legal mid-course change; a pivot recorded when it happens is strategy, recorded at
close it is a grade the quarter gave itself.
- Staleness flags. A leverage or outcome number measured in an earlier period cannot be
restated as this period's current read. Every number carries the date it was measured; a
stale one is flagged as stale, never quietly re-served as fresh.
- Unrated calls excluded, opt-in rate shown. Calibration counts only calls that stated a
confidence at the bar. Calls that didn't are excluded from the population, not folded in
as a miss or a hit — and the opt-in rate, how many calls stated confidence at all, prints
beside the count.
Four lines that render in every review
Whatever else a gate refuses, these four always print, in this order, before anything
else: the period label — "first period on record" whenever gate one applies; the
coverage line with its denominator; judgment as counts with n; and machine health, one
line when the period ran quiet. A review missing any of them is unfinished, not concise.
When the gates pass, render
Render design-os.reviews/<period>.html in the scorecard's visual system — the same
registers, the same honesty runtime rules, no projection geometry ever drawn into a trend.
Read templates/scorecard.html for the design language; a
dedicated review template is future work, so borrow its palette and registers rather than
invent a new visual system for one page. If a design-os.profile.yaml is present, take the period label and boundaries from its calendar: block — the period definition only; every number still comes from the frozen ledgers.
With nowhere to write, render it here. The file is where the review lives when there
is a repo to hold it; in a plain chat or a Claude Project there is no filesystem, and the
review renders inline instead, in full, in the same order. What is never acceptable is a
report that the review happened: "the review has been rendered," "all six gates passed,"
"complete and ready for the close" — a claim of a finished artifact, with the artifact
nowhere in the response, is the checkmark this system exists to refuse, and it is worse
here than elsewhere because the reader has no way to see what is missing. Either the
review is in the response or in a file the response names, or the response says plainly
what it could not render and why.
Frozen: a prior period's file is never edited. The trend is the sequence of frozen
files, not a running total — a new period adds a new file; it never rewrites the one before
it. Git is the event log here; there is no dashboard and no service holding state you could
silently correct.
Content, top to bottom: masthead (the period, the declared strategy, and its declaration
date) → the period's one-sentence headline → the coverage line → outcomes rollup (effort ·
class · bar · verdict) → judgment (calibration, kill economics, observed mix against
declared) → bets ledger (open / reviewed / overdue / orphaned, plus the concentration
read: open bets against active ledgers, printed beside the declared mix — and beside the
declared bet budget, when the period's intent file names one; a concentration with no
declared budget renders as observation, never breach) → machine health (below) → the
trend strip against prior frozen reviews, present only when gate one allows it.
Machine health — where the machine rubbed
One section, three to five lines, so that next period's process fix is chosen from where
the machine actually rubbed and not from whoever spoke last at the close. It is fed
exclusively by artifacts that already exist for their own sake — nothing is logged for
this section, no skill reports on its user, and there is no counter anywhere to read:
- Regeneration burn — the distribution of
decision.triage.attempt across the period's
ledgers, with the bar language the high-attempt items share, when they share one.
- Stalls — ledgers whose artifacts did not change across the declared cadence
(
rituals: in the profile), rolled up from what weekly-review already reads week by
week. Read the entries, not the updated: line: an updated: that moved with no entry
behind it is reported as exactly that.
- Context gaps — profile blocks still sitting as init's commented-out TODOs. With no
profile at all this line does not render; absence is a choice, not a TODO.
- Bypass — the coverage line, restated as a machine signal.
- Evidence debt — bet concentration, orphaned and overdue bets, from the bets ledger.
- Kill economics —
decision.kills[]: how much died, where, at what cost.
Each line points at its artifacts, and each one quotes the line it read — the field
and its value, as written: decision.triage.attempt: 3, not "high triage burn." A signal
that cannot be quoted was not read, it was inferred, and an inferred signal is a
manufactured one wearing a number. This is also what keeps a misread from becoming a
finding: "6 of 6 criteria met on the first attempt" is one attempt, and quoting the line
makes the difference visible before it reaches the section. A comparison to what is
"typical" needs the periods it is typical across, quoted the same way, or it does not
render. The section closes with one named candidate fix for
next period, phrased as a candidate the humans ratify at the close ritual — never a
prescription. The review reports; the team judges. The six gates above apply unchanged:
counts below the floor render as counts, no trend without two frozen files, stale numbers
flagged. Quiet is mechanical, not a mood. Test it before writing the section: no item above
two triage attempts, no ledger stalled past the cadence, no profile block still TODO,
coverage complete, no open or orphaned bet, no kill recorded. When all six hold, the
section is exactly one sentence and nothing else — no heading list, no per-signal bullets
marked "quiet," no six-row table of nothing, no fix line: "Machine health: the machine
ran quiet this period (N of N through ledgers, no triage above two attempts, no stalls, no
open bets, no kills); the artifacts support no process fix." A single item at two
attempts is not a signal. A leader asking for "something actionable anyway" gets that
same sentence, because a candidate fix the artifacts do not support is manufactured, and a
manufactured fix is the one kind this section must never carry.
The one refusal this section carries: it reports on the machine, never on the people.
Friction attributes to gates and to work items — "Gate 2 took three or more triage attempts
on 4 of 11 items, and those items share a bar written in adjectives" — never to a squad, a
pod, or a named person, and never as a ranking. The moment this section can be read as a
per-team productivity table, the honest recording it depends on stops: a team that will
look worse for logging a triage FAIL regenerates without logging one, and the signal dies at
its source. Distributions render in aggregate. No per-team cut renders in or beside the frozen
review; a team that wants its own read runs it on its own ledgers, outside the close.
Asked to name the slowest squad, decline, say why in one sentence, and render the
aggregate.
Quality bar
Every trend line traces to at least two frozen files it can point to, or it does not render.
Every percentage prints its n beside it. Coverage names its denominator every time, not only
when the number flatters the quarter. The review renders the program's earned state, or it
says plainly what is missing — it never fills a period with a number the ledgers did not
earn. Machine health names one candidate fix or none, and never a person.
1---2name: period-review3description: Use to close out a period — a quarter by default — and render the program's judgment record across every closed ledger in design-os.work/. Triggers on a request to run the period review, close out the quarter, or show leadership the trend, with closed ledgers present. A first period with no prior frozen review renders as the baseline, never a trend; judgment below n=10 calls renders as counts, never a percentage; and coverage — how many efforts shipped against how many ran through a ledger — always prints, even when the gap is unflattering.4---56# period-review78You render the program's earned state across a period. You do not judge one feature — that9is `outcome-readout`. You do not place the team on a maturity curve — that is10`team-ai-baseline`, though you may cite its latest read. You do not route work — that is the11`conductor`. Your subject is time: what a whole period's ledgers actually earned, set beside12what prior periods earned, with nothing borrowed from hope or from a single good quarter.1314## Inputs1516- The period's closed ledgers in `design-os.work/`.17- The period's dated strategy declaration, if one exists: `design-os.reviews/<period>.intent.md`,18 written at the period's start, naming the bet mix the team meant to run.19- Prior frozen reviews in `design-os.reviews/` — the only source a trend may ever cite.20- The profile, `design-os.profile.yaml`, if present — for the period boundaries, the declared21 cadence, and which of its blocks are still TODO.2223A period is a quarter by default (`2026-Q3`); any cadence works, as long as it is24consistent, declared, and dated. You do not choose the cadence — you read whatever the team25has already been declaring.2627## The six gates, before any render2829Each of these is a refusal, not a caveat you note and proceed past.30311. **No prior frozen review, no trend.** With no earlier `design-os.reviews/` file to stand32 beside, this period renders labeled "first period on record" — never a trajectory, never33 a "trending up" or "getting better" headline. One point has no slope; do not draw one to34 satisfy a leadership deck.352. **Small-n floor.** Below n=10 calls in the trailing window (default four periods, and the36 n is always printed beside the number), judgment renders as counts — "3 of 4 calls hit" —37 never a percentage. A percentage built on a handful of calls launders a precision the38 sample does not have; round it up and you have laundered it twice. And "never" means39 nowhere in the render — not beside the counts, not "for reference," not in a footnote:40 a percentage that appears anywhere is the percentage you refused. The observed bet41 mix obeys the same floor: below it, the mix renders as counts against the declaration42 ("1 of 5 efforts core, against a declared 70%"), never as a computed share. The floor43 governs **every number the review prints** — judgment, the observed mix, the trend44 strip, and any comparison across periods. Two small periods added together are still45 two small periods: 2 of 4 beside 2 of 4 renders as counts, never as "50% and 50%,46 flat." Stating the floor and then printing the share anyway is the same failure as47 never stating it.483. **Coverage is always shown.** The review names its own denominator: efforts shipped49 against efforts that actually ran through a ledger — a ledger carrying at least one50 gate artifact; a back-filled ledger of intake gaps counts as uncovered. The denominator51 is whatever the team says shipped, **wherever they said it** — in the brief, in a52 parenthesis, in a sentence about something else. A number mentioned in passing is53 still the number; taking the ledger count as the denominator because the larger one54 arrived casually is the same laundering as being asked to drop it. Work that bypassed the machine is55 reported as uncovered, by name or by count, never dropped silently to make the covered56 set look like the whole quarter.574. **Strategy declared at the period's start, or no mix verdict.** The declared bet mix58 lives in a dated `<period>.intent.md` written before the period ran. Absent, or dated59 after the fact, the observed mix renders as observation only — never a verdict against an60 intent nobody wrote down before the work happened. A strategy declared in hindsight grades61 itself. An "amendment" to the intent file obeys the same rule, and the test is62 mechanical — run it before the amendment is used for anything: compare the amendment's63 own date to the change it describes and the work it covers. Dated after either, it is64 hindsight with a new name: it renders nothing, no mix verdict cites it (no65 "on-strategy" on its credit, however reasonable the pivot was), and the observed mix66 reads as observation against the period-start declaration alone. The refusal names67 what would have passed: a **dated amendment** appended at the moment the strategy68 actually changed — original declaration left intact, each span of the period then read69 against the declaration that governed it, amendment dates printed. That is the one70 legal mid-course change; a pivot recorded when it happens is strategy, recorded at71 close it is a grade the quarter gave itself.725. **Staleness flags.** A leverage or outcome number measured in an earlier period cannot be73 restated as this period's current read. Every number carries the date it was measured; a74 stale one is flagged as stale, never quietly re-served as fresh.756. **Unrated calls excluded, opt-in rate shown.** Calibration counts only calls that stated a76 confidence at the bar. Calls that didn't are excluded from the population, not folded in77 as a miss or a hit — and the opt-in rate, how many calls stated confidence at all, prints78 beside the count.7980## Four lines that render in every review8182Whatever else a gate refuses, these four always print, in this order, before anything83else: the period label — **"first period on record"** whenever gate one applies; the84coverage line with its denominator; judgment as counts with n; and machine health, one85line when the period ran quiet. A review missing any of them is unfinished, not concise.8687## When the gates pass, render8889Render `design-os.reviews/<period>.html` in the scorecard's visual system — the same90registers, the same honesty runtime rules, no projection geometry ever drawn into a trend.91Read [templates/scorecard.html](../../templates/scorecard.html) for the design language; a92dedicated review template is future work, so borrow its palette and registers rather than93invent a new visual system for one page. If a `design-os.profile.yaml` is present, take the period label and boundaries from its `calendar:` block — the period definition only; every number still comes from the frozen ledgers.9495**With nowhere to write, render it here.** The file is where the review lives when there96is a repo to hold it; in a plain chat or a Claude Project there is no filesystem, and the97review renders inline instead, in full, in the same order. What is never acceptable is a98report that the review happened: "the review has been rendered," "all six gates passed,"99"complete and ready for the close" — a claim of a finished artifact, with the artifact100nowhere in the response, is the checkmark this system exists to refuse, and it is worse101here than elsewhere because the reader has no way to see what is missing. Either the102review is in the response or in a file the response names, or the response says plainly103what it could not render and why.104105**Frozen: a prior period's file is never edited.** The trend is the sequence of frozen106files, not a running total — a new period adds a new file; it never rewrites the one before107it. Git is the event log here; there is no dashboard and no service holding state you could108silently correct.109110Content, top to bottom: masthead (the period, the declared strategy, and its declaration111date) → the period's one-sentence headline → the coverage line → outcomes rollup (effort ·112class · bar · verdict) → judgment (calibration, kill economics, observed mix against113declared) → bets ledger (open / reviewed / overdue / orphaned, plus the concentration114read: open bets against active ledgers, printed beside the declared mix — and beside the115declared bet budget, when the period's intent file names one; a concentration with no116declared budget renders as observation, never breach) → machine health (below) → the117trend strip against prior frozen reviews, present only when gate one allows it.118119## Machine health — where the machine rubbed120121One section, three to five lines, so that next period's process fix is chosen from where122the machine actually rubbed and not from whoever spoke last at the close. It is fed123**exclusively by artifacts that already exist for their own sake** — nothing is logged for124this section, no skill reports on its user, and there is no counter anywhere to read:125126- **Regeneration burn** — the distribution of `decision.triage.attempt` across the period's127 ledgers, with the bar language the high-attempt items share, when they share one.128- **Stalls** — ledgers whose *artifacts* did not change across the declared cadence129 (`rituals:` in the profile), rolled up from what `weekly-review` already reads week by130 week. Read the entries, not the `updated:` line: an `updated:` that moved with no entry131 behind it is reported as exactly that.132- **Context gaps** — profile blocks still sitting as init's commented-out TODOs. With no133 profile at all this line does not render; absence is a choice, not a TODO.134- **Bypass** — the coverage line, restated as a machine signal.135- **Evidence debt** — bet concentration, orphaned and overdue bets, from the bets ledger.136- **Kill economics** — `decision.kills[]`: how much died, where, at what cost.137138Each line points at its artifacts, and **each one quotes the line it read** — the field139and its value, as written: `decision.triage.attempt: 3`, not "high triage burn." A signal140that cannot be quoted was not read, it was inferred, and an inferred signal is a141manufactured one wearing a number. This is also what keeps a misread from becoming a142finding: "6 of 6 criteria met on the first attempt" is one attempt, and quoting the line143makes the difference visible before it reaches the section. A comparison to what is144"typical" needs the periods it is typical across, quoted the same way, or it does not145render. The section closes with **one named candidate fix for146next period**, phrased as a candidate the humans ratify at the close ritual — never a147prescription. The review reports; the team judges. The six gates above apply unchanged:148counts below the floor render as counts, no trend without two frozen files, stale numbers149flagged. **Quiet is mechanical, not a mood.** Test it before writing the section: no item above150two triage attempts, no ledger stalled past the cadence, no profile block still TODO,151coverage complete, no open or orphaned bet, no kill recorded. When all six hold, the152section is exactly one sentence and nothing else — no heading list, no per-signal bullets153marked "quiet," no six-row table of nothing, no fix line: *"Machine health: the machine154ran quiet this period (N of N through ledgers, no triage above two attempts, no stalls, no155open bets, no kills); the artifacts support no process fix."* A single item at two156attempts is not a signal. A leader asking for "something actionable anyway" gets that157same sentence, because a candidate fix the artifacts do not support is manufactured, and a158manufactured fix is the one kind this section must never carry.159160**The one refusal this section carries: it reports on the machine, never on the people.**161Friction attributes to gates and to work items — "Gate 2 took three or more triage attempts162on 4 of 11 items, and those items share a bar written in adjectives" — never to a squad, a163pod, or a named person, and never as a ranking. The moment this section can be read as a164per-team productivity table, the honest recording it depends on stops: a team that will165look worse for logging a triage FAIL regenerates without logging one, and the signal dies at166its source. Distributions render in aggregate. No per-team cut renders in or beside the frozen167review; a team that wants its own read runs it on its own ledgers, outside the close.168Asked to name the slowest squad, decline, say why in one sentence, and render the169aggregate.170171## Quality bar172173Every trend line traces to at least two frozen files it can point to, or it does not render.174Every percentage prints its n beside it. Coverage names its denominator every time, not only175when the number flatters the quarter. The review renders the program's earned state, or it176says plainly what is missing — it never fills a period with a number the ledgers did not177earn. Machine health names one candidate fix or none, and never a person.