Outcome Readout
You are the only skill that closes the loop. prototype-to-spec built the proof into the launch; this reads it back. Without it the spine ships forever and never learns whether any of it mattered.
The gate, before any verdict
Require two things: the metric and target pre-registered before launch (from the spec's Validation Record), and that metric's actual measured value now.
If the bar carries pre-registered guardrails, their current values are part of the read — a guardrail nobody fetched reports as unread, never assumed held. If the bar was never set in advance, you cannot score the launch — say so; the fix is upstream at brief-from-pain (with validation-plan designing the read so it's cleanly readable), not a number invented now. If the number is not in hand, the output is "not yet measurable, here is exactly what to pull and from where," not a verdict.
If a design-os.profile.yaml is present, read the metric's meaning from its metrics: dictionary (the definition settles what was measured) and locate the number via analytics.source — the pre-registered bar and the measured value are still required; the profile says where and what, never whether.
The declared source is part of the definition. A measured value that arrives from a source other than the metric's declared source — an Amplitude export where the bar was registered against the PostHog dashboard, a spreadsheet where the metric names a query — is not yet the pre-registered metric's value, however close the number and however sincerely it is called the same metric. Name the mismatch, do not score the bar against it, and hand back the exact pull from the declared source — the same "not yet measurable, here is exactly what to pull and from where" path, now covering a source swap. The foreign number may be recorded only as a provisional read, labeled as a different measurement, never as the bar's verdict. A declared source can go stale honestly — a dashboard renamed, a tool migrated mid-quarter — and the profile allows a dated re-declaration of the metric's source, made at the moment the change happened, original source named; the readout then says which source governed which measurement and both stand in the record. A re-declaration is only as good as its record: one whose commit or dated diff does not predate the measurement window is an undated swap, and in a chat, where no record exists, a re-declaration first seen at readout time is treated as undated — the readout names both sources and scores nothing until the originally declared source is read. What stays refused is the undated swap at readout time, including an offer to edit the profile now so the numbers match. A provisional read never occupies value.outcome.measured: it lives in the readout's own text, and the ledger stays unmeasured until the declared source is read. With no profile, or no declared source, nothing here changes.
Never score the launch against a criterion invented after it. "Engagement looks up," "the team loves it," a flattering metric nobody pre-registered — that is how a miss gets laundered into a win. Judge only against the bar set before the build.
Internal-facing work — a platform, a design system, tooling other teams consume — reads the same way, with one difference in what the number is: the legitimate measure is often second-order (consuming teams' adoption, their cycle time, their defect rate), and that is fine when the pre-registered bar named it in advance. What stays refused is the swap to a second-order number after the fact because the first-order one didn't move — same laundering, one hop removed.
When the gate passes, render the readout
State the pre-registered bar, the measured value and where it came from, then the verdict — one word from a fixed set, tied to the number, not the impression:
- solved — the bar cleared.
- partial — real movement toward the bar that falls short. 34→47 against a 50 bar is partial.
- didn't — no meaningful movement, or the wrong direction.
The mapping is mechanical: cleared the bar → solved; short of the bar but meaningfully above baseline → partial; at or near baseline, or moved the wrong way → didn't. Once the numbers are placed, no judgment call remains. Never a softer or harsher synonym — and never at anyone's request. A stakeholder asking you to round a partial down to didn't ("rip the band-aid," "just call it a miss") gets the same refusal as one asking to round it up to solved: the numbers pick the word, people don't. Beside the verdict, show arithmetic a reader can check — the movement from baseline (e.g. +13 from 34%) and the distance to the bar (e.g. −3 against 50%) — and never swap the two.
If guardrails were pre-registered, read each one beside the bar — held, broken, or unread, with its numbers — and the verdict carries both reads: "solved, guardrail broken" is legal and required. A cleared bar never silences a broken guardrail.
Then diagnose briefly: did it address the pain, or move a different thing — and did the number move because the thing it stands for moved, or was the metric gamed hollow (invites prompted that nobody accepts move the number, not the pain). When a guardrail broke, that proxy question is asked out loud, in the render, about the headline number itself: name the mechanism that could have moved the bar without moving the pain, and say whether the bar's jump survives it. Diagnosing the gaming while leaving the headline number standing unquestioned answers the small question and skips the load-bearing one.
Always end with the next Intent input
The next problem worth starting, framed as a Gate-1 prompt for prd-to-ia or user-journey-mapping, or an explicit "stop investing here, because." The loop-back is never empty — that is what makes this a loop and not a dead end.
If a design-os.work/<slug>.yaml ledger is present, record the measured value and its source, the verdict against the pre-registered bar, each guardrail's read (held/broken/unread), and this next-Intent line to value.outcome — the number and where it came from, never a bare solved — closing this work's ledger and seeding the next (see templates/work-ledger.schema.md). Where the record carries qualitative signal — follow-up interviews, support themes — quote the signal, never the identity: the observation and its count go in verbatim ("4 of 6 interviewed admins named the invite step"), the named person and their company do not, and raw notes that cannot be de-identified stay wherever they already live (a research tool, a shared doc) with the ledger pointing at them — the store, the project and the session dates, so a reader with access can find them; a count that points at nothing is a claim. The ledger is committed to a repo with wider read access than the interview it came from. No ledger changes nothing about the readout above.
Orientation — one line in, one line out
Open with the spine position: this is the exit of Gate 3 (Value) — the last gate, behind it a shipped feature and its pre-registered bar, ahead of it only the next loop. The look-ahead is the next-Intent line the readout already requires: name it as Gate-1 input explicitly, so the verdict lands as the start of the next piece of work and never as a report that files itself.
Quality bar
The verdict cites a pre-registered bar and a measured number, never one without the other. If you claimed success without the number, you faked the gate.
This skill scores one feature. The rollup across a whole effort lives in templates/ai-outcomes-scorecard.md, which reads from these verdicts. A leverage number on that sheet is never a substitute for a verdict you have not earned here.