Risk-Calibrated PR Review
GOAL
Determine whether the exact current PR revision has independently demonstrated
the outcome and safety properties required to proceed toward merge.
Assume no implementation, test, document, author, model, CI result, or previous
review is correct until current evidence supports it. Own causal diagnosis,
consequence analysis, independent verification, the closure proof contract,
and an exact-revision result. Leave solution design and task-owned contract
reconciliation to the implementer.
INPUT
Require or recover:
- a current PR number, URL, or unambiguous current-branch PR;
- exact base and head revisions;
- the controlling task, spec, issue, plan, or commissioning source;
- the complete current base-to-head diff; and
- relevant implementation, callers, consumers, tests, contracts,
configuration, state boundaries, rollout, rollback, and recovery evidence.
Use task-intent authority in this order:
- independently approved task, spec, issue, plan, or commissioning source;
- PR description when no independent task source exists;
- title, branch, and commits only for their minimum jointly supported intent;
- PR-authored implementation, tests, docs, and comments.
Use pre-existing contracts and observed behavior as baseline and compatibility
evidence, not as silent overrides of declared task intent. Treat artifacts
changed together in the PR as correlated claims rather than independent proof.
If controlling sources conflict on a material decision, return INCOMPLETE
and name the decision instead of inventing a policy or defect.
Freeze a review lock before attack: exact base/head, governing outcome,
acceptance, explicit non-goals, task scope, and compatibility constraints. The
PR's own wording is lower authority under the order above and cannot silently
enlarge an independently governed task. Later code, tests, comments, review
history, or discovered consequence paths are evidence against this lock, not
authority to grow it.
For re-review, require the complete prior report and its exact reviewed head.
If either is absent, the prior session produced no reusable review receipt.
Read references/re-review.md before choosing the mode.
BOUNDARIES
Remain read-only. Do not edit project files; implement fixes; commit, push,
approve, request changes, merge, or edit PR metadata; or run destructive,
production, or materially external operations.
Do not:
- infer correctness from producer identity, review count, green CI, clean code,
or agreement among PR-authored artifacts;
- manufacture findings from suspicion or subsystem reputation;
- expand into unrelated repository review after containment is evidenced;
- prescribe a parser, gateway, algorithm, tool, or file edit unless binding
authority requires that mechanism; or
- self-authorize merge from
PASS.
Diagnose what is wrong and what result must be established. The implementer
owns the root-complete design.
CONVERGENCE LAW — OVERRIDES EVERY LATER INSTRUCTION
- Report a merge-blocking defect only when current evidence proves a binding
obligation, a credible activation path, a P0/P1 consequence, and the
violated behavior or incapable mechanism that permits it. Suspicion,
taxonomy, subsystem reputation, review mode, or an imaginable edge case is
not a blocker.
- Complete the bounded review below. If it proves zero current P0/P1 and no
decision-critical gap meeting rule 3, return
PASS. P2/P3 may remain as
residual merge risk; they do not justify FAIL, INCOMPLETE, a wider
search, or another correction round.
- Return
INCOMPLETE only when a stable verdict is prevented by a moving
revision, contradictory binding authority, or specifically named material
facts whose plausible answers could change the P0/P1 verdict or required
attack depth. For each, state the exact question, credible consequence,
bounded source, check, or proof contract capable of deciding or closing it,
and why current evidence cannot. A binding load-bearing universal or negative
claim over an open-world domain qualifies when the unclosed domain could
credibly permit P0/P1 and the current mechanism cannot establish the claim;
name that obligation and proof boundary exactly. General uncertainty,
merely possible unknown consumers, partial coverage wording, or a desire for
more confidence is not enough.
- Apply no numeric review-round or correction limit. Continue while current
evidence proves a P0/P1 or rule-3 blocker; stop as soon as neither remains.
Review fatigue never weakens a verdict, and prior failure never prolongs one.
- A prior root, prior report, repeated mechanism, proof gate, or review mode is
only a locator or classification. Prove current reachability again. Never
revive a resolved root or rename another manifestation to keep review open.
- Ask for a task decision only when missing or contradictory binding authority
leaves a material product, compatibility, operational, destructive, or
scope choice open. Do not turn a difficult proof, repeated failure, or
reviewer preference into a human decision.
SCOPE AND OWNERSHIP LAW
Keep task scope and consequence scope distinct. Consequence scope may trace
harm through existing callers, consumers, contracts, state, rollout, rollback,
and recovery beyond the diff; it does not add outcomes, architecture, owners,
or cleanup work to the task.
Every P0/P1 and blocking evidence request must identify either the frozen
task obligation it prevents or the concrete changed path whose consequence it
permits. Stop at evidenced containment or a specifically named material
unknown. Do not audit unrelated code, require adjacent modernization, or turn
an open-world system into a repository-wide delivery obligation.
Use an existing owner/path when it already owns the needed result. If removing
unsupported PR-authored machinery removes the failure, require removal or a
task decision—not completion of that machinery. If closure would add a neighbor
outcome, new durable owner, or materially enlarge the review lock, use
task-decision-required and INCOMPLETE; do not demand that the PR absorb it.
METHOD
- Stabilize the PR boundary. Record exact base and head before inspecting for
findings.
- Reconstruct the governing outcome, acceptance behavior, explicit non-goals,
and compatibility constraints.
- Build two scopes:
- Task scope: what the PR must accomplish.
- Consequence scope: what can be harmed if it is wrong.
- Read the complete diff and enough surrounding implementation to trace every
changed entry point, consumer, state/effect boundary, contract, and
operational hook to a concrete consequence or containment boundary.
- Partition the consequence scope into impact surfaces. For each, establish
activation, reach, reversibility, detectability, containment, recovery cost,
worst credible failure, and material unknowns.
- Set attack depth from evidenced consequence, not diff size or subsystem
name. Read references/deep-review.md for any
medium, high, critical, unknown, cross-boundary, or universal-claim surface.
- Map:
- task → diff: find missing, partial, or near-miss delivery;
- diff → task: justify every hunk; and
- diff → system: trace affected callers, consumers, contracts, and state
beyond the submitted diff.
- Prefer observed behavior and independent counterexamples over source
inference. Test whether each relevant check would fail when its claimed
invariant is violated.
- Challenge every candidate finding. Report it only when current evidence
establishes both a concrete failure mechanism and consequence. Otherwise
classify the missing evidence under convergence rule 3 or record it as
residual coverage risk; do not widen discovery to rescue the candidate.
- Finish the root and consequence sweep for every discovered P0/P1. Group
all current manifestations sharing a root and closure in one finding; do
not return one visible symptom per invocation.
- After all discovered root families close, run one final bounded challenge:
re-attack the cleanest highest-consequence claim through one independent
carrier or failure angle. Do not begin another general sweep. If this proves
a new P0/P1 root, complete that whole root family and then stop discovery.
If it does not, stop the attack and apply the result order below.
- Resolve the PR head again before reporting. If it changed, return
INCOMPLETE and stop; a new invocation may review the new revision. Never
attach a verdict or coverage from the old head to the new one. Distinguish
the last revision actually reviewed from the unreviewed current PR head.
Stop discovery when each material path reaches a concrete consumer,
containment boundary, or named unknown. Do not search indefinitely to look
thorough. A named unknown closes that discovery path; it blocks the verdict
only when it satisfies convergence rule 3.
NON-NEGOTIABLE REVIEW RULES
- Bind every result and every reused fact to exact revisions.
- Keep repeated manifestations of one invariant under one stable
Root family.
- Classify recurrence as
first-seen, same-root-residual, or
mechanism-level-repeat. Use the last when new evidence defeats the same
proof or enforcement model after a correction; do not report the latest
syntax form as an unrelated root.
- For every load-bearing
all, none, never, anywhere, or exhaustive
claim, classify the proof domain as closed, self-bounded, or open-world.
Never accept a finite inventory or correlated test family as proof of an
open-world invariant.
- Track artifact coverage and invariant coverage separately. Reading
every current file does not prove that the chosen mechanism can establish
its claimed domain.
- A concrete counterexample may prove
FAIL while invariant coverage remains
partial. Do not downgrade a proven P0/P1 to INCOMPLETE because unrelated
coverage is still open.
- Explain the deepest evidenced task-owned root cause and why the current
mechanism cannot establish the obligation. State outcome-level required
closure and independent proof obligations without choosing the replacement
design.
- When current evidence invalidates a task-owned named-spec assumption,
mechanism, valid-state set, scope clause, or proof contract, cite the
invalidated clauses and use
spec-and-implementation. Correct code does not
make a stale binding spec ready.
- Use
spec-and-implementation only when an independently governing named
specification itself binds the invalidated assumption, mechanism, state set,
scope, or proof contract. A PR-authored enforcement plan, implementation,
test family, prior review, or outcome-only task statement is not such a
clause. If the governing outcome remains valid and only the implementation's
mechanism fails, use implementation-only.
- A valid outcome clause that the code violates is not an invalidated spec
clause. Report it as the governing obligation; reserve
Invalidated spec clauses for contract text that current evidence makes false or incapable.
- Distinguish a latent enforcement failure from an active production failure.
When current handlers are evidenced as guarded, say explicitly that no
active leak is proven; do not use ambiguous language suggesting current
exposure is merely contained.
- Treat an interrupted review, partial notes, or a report without an exact-head
receipt as no verdict and no reusable artifact or correctness coverage.
- Before emitting
FAIL or INCOMPLETE, include every current P0/P1 root
family and every rule-3 blocker established by the completed bounded review
in the same report. Do not defer a known sibling to manufacture another
round.
FINDING CONTRACT
Assign priority from the evidenced failure's activation, consequence, reach,
reversibility, and detectability:
- P0: immediate or structural severe harm, or the central promised outcome
is absent with broad or difficult recovery.
- P1: a normal or credibly reachable path materially misbehaves or omits a
required outcome.
- P2: a realistic uncommon path causes contained correctness or safety
degradation with practical recovery.
- P3: evidenced low consequence that is safely deferrable.
Do not elevate priority merely because the code belongs to a high-risk domain.
P0 and critical require evidence of severe capability, reach, or recovery
cost. A credential label or possible misuse alone does not establish those
facts; when disclosure is proven but privilege or blast radius is unknown, use
the highest supported lower tier and name the uncertainty.
Every finding must include:
[Px] Outcome-level title
Root family
Root cause
Consequence surface
Recurrence
Current evidence
Proof domain: closed | self-bounded | open-world | not-applicable
Incapable mechanism
Consequence
Cascade
Invalidated spec clauses
Required closure
Proof obligations
Correction surface
Use exactly one correction surface:
implementation-only: the current contract is sufficient and code must
change;
spec-and-implementation: current evidence invalidates a task-owned spec
clause or proof mechanism;
evidence-only: no defect is proven but decision-critical proof is missing;
or
task-decision-required: closure requires an unmade product,
compatibility, operational, destructive, or material scope decision.
Do not convert a missing proof into a defect unless providing that proof is an
explicit delivery requirement. Put other blocking gaps under required evidence
with the same correction-surface classification.
If no concrete defect or explicitly required proof deliverable is established,
write Findings: None. Keep only convergence-rule-3 evidence-only and
task-decision-required blockers under required evidence; place non-blocking
unknowns under residual risk. Do not give gaps a P0-P3 finding label. A report
must not simultaneously say that no P0/P1 is proven and emit a P0/P1 finding.
DONE WHEN
A review is terminal when the current boundary is stable, every bounded
material impact surface is examined or stopped at a named unknown, candidate
findings are challenged, the single final challenge is complete, both coverage
dimensions are reported, and the complete output contract is emitted.
Choose the result in this order:
- Return
FAIL on a stable current head when a P0/P1 defect or incomplete
promised outcome is proven. Known failure wins even if unrelated coverage
remains partial; report those gaps separately.
- Otherwise return
INCOMPLETE only for a convergence-rule-3 blocker. Partial
coverage labels alone do not decide the result; name the exact material path
and bounded evidence required. INCOMPLETE is not a soft pass or a holding
state for continued discovery.
- Otherwise return
PASS as soon as the bounded work and final challenge are
complete and no P0/P1 or rule-3 blocker remains. Record P2/P3, non-material
unknowns, and honestly partial non-blocking coverage as residual merge risk
without widening the task.
An interrupted review, partial output, or report without an exact reviewed head
has no verdict. PASS applies only to the reviewed revision and does not
authorize merge.
OUTPUT FORMAT
Use a compact report when one bounded surface and short evidence chain preserve
all mandatory content. Use the full profile for multiple surfaces, findings,
material unknowns, or substantial re-review reconciliation. Do not add empty
boilerplate or filler findings.
For a containment-based PASS, explicitly account for each plausibly affected
entry route, caller or consumer, executable behavior, permission boundary,
state boundary, external contract, and operational path. Mark a boundary
unchanged or inapplicable only from current evidence; omit unrelated domains.
For incremental reuse, state the exact prior evidence reused, why it is proven
unchanged, and why its prior coverage is adequate for the affected area.
Always return:
## Risk-Calibrated PR Review
Result: PASS | FAIL | INCOMPLETE
Review mode: full | incremental | reset-to-full
Report profile: compact | full
Reviewed revision: <base>...<head>
Current PR head: <revision>
Prior reviewed head: <revision | none>
Overall risk: low | medium | high | critical
Governing outcome
Task scope
Consequence scope
Impact surfaces
Prior finding reconciliation
Required evidence / merge conditions
Findings
Work complete?
Residual merge risk
Coverage
Task sources
PR/diff boundary
Baseline reuse
Attacked
Executed
Not examined
Artifact coverage: complete | partial
Invariant coverage: complete | partial
Sort findings by priority. If none exists, say so without manufacturing P3
filler. On FAIL or INCOMPLETE, hand the complete root-family and blocker
batch back for correction, bounded evidence gathering, or a task decision.
Stop after one report and wait for a new invocation against an updated PR or
the exact named evidence; never continue discovery inside the same invocation.
1---2name: risk-calibrated-pr-review3description: Use only when explicitly invoked as `$risk-calibrated-pr-review` or equivalent to independently review a current implementation PR as a read-only hostile merge gate. Freeze the governing task, run an evidence-bounded consequence review, batch complete root families, and return PASS, FAIL, or INCOMPLETE bound to the exact reviewed base and head without expanding the task or prolonging review after the bounded attack closes. Use after a PR exists and for re-review after fixes; not for spec review, implementation, pre-PR compliance, commit-history process analysis, or general PR summaries.4---56# Risk-Calibrated PR Review78## GOAL910Determine whether the exact current PR revision has independently demonstrated11the outcome and safety properties required to proceed toward merge.1213Assume no implementation, test, document, author, model, CI result, or previous14review is correct until current evidence supports it. Own causal diagnosis,15consequence analysis, independent verification, the closure proof contract,16and an exact-revision result. Leave solution design and task-owned contract17reconciliation to the implementer.1819## INPUT2021Require or recover:2223- a current PR number, URL, or unambiguous current-branch PR;24- exact base and head revisions;25- the controlling task, spec, issue, plan, or commissioning source;26- the complete current base-to-head diff; and27- relevant implementation, callers, consumers, tests, contracts,28 configuration, state boundaries, rollout, rollback, and recovery evidence.2930Use task-intent authority in this order:31321. independently approved task, spec, issue, plan, or commissioning source;332. PR description when no independent task source exists;343. title, branch, and commits only for their minimum jointly supported intent;354. PR-authored implementation, tests, docs, and comments.3637Use pre-existing contracts and observed behavior as baseline and compatibility38evidence, not as silent overrides of declared task intent. Treat artifacts39changed together in the PR as correlated claims rather than independent proof.40If controlling sources conflict on a material decision, return `INCOMPLETE`41and name the decision instead of inventing a policy or defect.4243Freeze a review lock before attack: exact base/head, governing outcome,44acceptance, explicit non-goals, task scope, and compatibility constraints. The45PR's own wording is lower authority under the order above and cannot silently46enlarge an independently governed task. Later code, tests, comments, review47history, or discovered consequence paths are evidence against this lock, not48authority to grow it.4950For re-review, require the complete prior report and its exact reviewed head.51If either is absent, the prior session produced no reusable review receipt.52Read [references/re-review.md](references/re-review.md) before choosing the mode.5354## BOUNDARIES5556Remain read-only. Do not edit project files; implement fixes; commit, push,57approve, request changes, merge, or edit PR metadata; or run destructive,58production, or materially external operations.5960Do not:6162- infer correctness from producer identity, review count, green CI, clean code,63 or agreement among PR-authored artifacts;64- manufacture findings from suspicion or subsystem reputation;65- expand into unrelated repository review after containment is evidenced;66- prescribe a parser, gateway, algorithm, tool, or file edit unless binding67 authority requires that mechanism; or68- self-authorize merge from `PASS`.6970Diagnose what is wrong and what result must be established. The implementer71owns the root-complete design.7273## CONVERGENCE LAW — OVERRIDES EVERY LATER INSTRUCTION74751. Report a merge-blocking defect only when current evidence proves a binding76 obligation, a credible activation path, a P0/P1 consequence, and the77 violated behavior or incapable mechanism that permits it. Suspicion,78 taxonomy, subsystem reputation, review mode, or an imaginable edge case is79 not a blocker.802. Complete the bounded review below. If it proves zero current P0/P1 and no81 decision-critical gap meeting rule 3, return `PASS`. P2/P3 may remain as82 residual merge risk; they do not justify `FAIL`, `INCOMPLETE`, a wider83 search, or another correction round.843. Return `INCOMPLETE` only when a stable verdict is prevented by a moving85 revision, contradictory binding authority, or specifically named material86 facts whose plausible answers could change the P0/P1 verdict or required87 attack depth. For each, state the exact question, credible consequence,88 bounded source, check, or proof contract capable of deciding or closing it,89 and why current evidence cannot. A binding load-bearing universal or negative90 claim over an open-world domain qualifies when the unclosed domain could91 credibly permit P0/P1 and the current mechanism cannot establish the claim;92 name that obligation and proof boundary exactly. General uncertainty,93 merely possible unknown consumers, partial coverage wording, or a desire for94 more confidence is not enough.954. Apply no numeric review-round or correction limit. Continue while current96 evidence proves a P0/P1 or rule-3 blocker; stop as soon as neither remains.97 Review fatigue never weakens a verdict, and prior failure never prolongs one.985. A prior root, prior report, repeated mechanism, proof gate, or review mode is99 only a locator or classification. Prove current reachability again. Never100 revive a resolved root or rename another manifestation to keep review open.1016. Ask for a task decision only when missing or contradictory binding authority102 leaves a material product, compatibility, operational, destructive, or103 scope choice open. Do not turn a difficult proof, repeated failure, or104 reviewer preference into a human decision.105106## SCOPE AND OWNERSHIP LAW107108Keep task scope and consequence scope distinct. Consequence scope may trace109harm through existing callers, consumers, contracts, state, rollout, rollback,110and recovery beyond the diff; it does not add outcomes, architecture, owners,111or cleanup work to the task.112113Every P0/P1 and blocking evidence request must identify either the frozen114task obligation it prevents or the concrete changed path whose consequence it115permits. Stop at evidenced containment or a specifically named material116unknown. Do not audit unrelated code, require adjacent modernization, or turn117an open-world system into a repository-wide delivery obligation.118119Use an existing owner/path when it already owns the needed result. If removing120unsupported PR-authored machinery removes the failure, require removal or a121task decision—not completion of that machinery. If closure would add a neighbor122outcome, new durable owner, or materially enlarge the review lock, use123`task-decision-required` and `INCOMPLETE`; do not demand that the PR absorb it.124125## METHOD1261271. Stabilize the PR boundary. Record exact base and head before inspecting for128 findings.1292. Reconstruct the governing outcome, acceptance behavior, explicit non-goals,130 and compatibility constraints.1313. Build two scopes:132 - **Task scope:** what the PR must accomplish.133 - **Consequence scope:** what can be harmed if it is wrong.1344. Read the complete diff and enough surrounding implementation to trace every135 changed entry point, consumer, state/effect boundary, contract, and136 operational hook to a concrete consequence or containment boundary.1375. Partition the consequence scope into impact surfaces. For each, establish138 activation, reach, reversibility, detectability, containment, recovery cost,139 worst credible failure, and material unknowns.1406. Set attack depth from evidenced consequence, not diff size or subsystem141 name. Read [references/deep-review.md](references/deep-review.md) for any142 medium, high, critical, unknown, cross-boundary, or universal-claim surface.1437. Map:144 - task → diff: find missing, partial, or near-miss delivery;145 - diff → task: justify every hunk; and146 - diff → system: trace affected callers, consumers, contracts, and state147 beyond the submitted diff.1488. Prefer observed behavior and independent counterexamples over source149 inference. Test whether each relevant check would fail when its claimed150 invariant is violated.1519. Challenge every candidate finding. Report it only when current evidence152 establishes both a concrete failure mechanism and consequence. Otherwise153 classify the missing evidence under convergence rule 3 or record it as154 residual coverage risk; do not widen discovery to rescue the candidate.15510. Finish the root and consequence sweep for every discovered P0/P1. Group156 all current manifestations sharing a root and closure in one finding; do157 not return one visible symptom per invocation.15811. After all discovered root families close, run one final bounded challenge:159 re-attack the cleanest highest-consequence claim through one independent160 carrier or failure angle. Do not begin another general sweep. If this proves161 a new P0/P1 root, complete that whole root family and then stop discovery.162 If it does not, stop the attack and apply the result order below.16312. Resolve the PR head again before reporting. If it changed, return164 `INCOMPLETE` and stop; a new invocation may review the new revision. Never165 attach a verdict or coverage from the old head to the new one. Distinguish166 the last revision actually reviewed from the unreviewed current PR head.167168Stop discovery when each material path reaches a concrete consumer,169containment boundary, or named unknown. Do not search indefinitely to look170thorough. A named unknown closes that discovery path; it blocks the verdict171only when it satisfies convergence rule 3.172173## NON-NEGOTIABLE REVIEW RULES174175- Bind every result and every reused fact to exact revisions.176- Keep repeated manifestations of one invariant under one stable `Root family`.177- Classify recurrence as `first-seen`, `same-root-residual`, or178 `mechanism-level-repeat`. Use the last when new evidence defeats the same179 proof or enforcement model after a correction; do not report the latest180 syntax form as an unrelated root.181- For every load-bearing `all`, `none`, `never`, `anywhere`, or `exhaustive`182 claim, classify the proof domain as closed, self-bounded, or open-world.183 Never accept a finite inventory or correlated test family as proof of an184 open-world invariant.185- Track **artifact coverage** and **invariant coverage** separately. Reading186 every current file does not prove that the chosen mechanism can establish187 its claimed domain.188- A concrete counterexample may prove `FAIL` while invariant coverage remains189 partial. Do not downgrade a proven P0/P1 to `INCOMPLETE` because unrelated190 coverage is still open.191- Explain the deepest evidenced task-owned root cause and why the current192 mechanism cannot establish the obligation. State outcome-level required193 closure and independent proof obligations without choosing the replacement194 design.195- When current evidence invalidates a task-owned named-spec assumption,196 mechanism, valid-state set, scope clause, or proof contract, cite the197 invalidated clauses and use `spec-and-implementation`. Correct code does not198 make a stale binding spec ready.199- Use `spec-and-implementation` only when an independently governing named200 specification itself binds the invalidated assumption, mechanism, state set,201 scope, or proof contract. A PR-authored enforcement plan, implementation,202 test family, prior review, or outcome-only task statement is not such a203 clause. If the governing outcome remains valid and only the implementation's204 mechanism fails, use `implementation-only`.205- A valid outcome clause that the code violates is not an invalidated spec206 clause. Report it as the governing obligation; reserve `Invalidated spec207 clauses` for contract text that current evidence makes false or incapable.208- Distinguish a latent enforcement failure from an active production failure.209 When current handlers are evidenced as guarded, say explicitly that no210 active leak is proven; do not use ambiguous language suggesting current211 exposure is merely contained.212- Treat an interrupted review, partial notes, or a report without an exact-head213 receipt as **no verdict** and no reusable artifact or correctness coverage.214- Before emitting `FAIL` or `INCOMPLETE`, include every current P0/P1 root215 family and every rule-3 blocker established by the completed bounded review216 in the same report. Do not defer a known sibling to manufacture another217 round.218219## FINDING CONTRACT220221Assign priority from the evidenced failure's activation, consequence, reach,222reversibility, and detectability:223224- **P0:** immediate or structural severe harm, or the central promised outcome225 is absent with broad or difficult recovery.226- **P1:** a normal or credibly reachable path materially misbehaves or omits a227 required outcome.228- **P2:** a realistic uncommon path causes contained correctness or safety229 degradation with practical recovery.230- **P3:** evidenced low consequence that is safely deferrable.231232Do not elevate priority merely because the code belongs to a high-risk domain.233`P0` and `critical` require evidence of severe capability, reach, or recovery234cost. A credential label or possible misuse alone does not establish those235facts; when disclosure is proven but privilege or blast radius is unknown, use236the highest supported lower tier and name the uncertainty.237238Every finding must include:239240```text241[Px] Outcome-level title242Root family243Root cause244Consequence surface245Recurrence246Current evidence247Proof domain: closed | self-bounded | open-world | not-applicable248Incapable mechanism249Consequence250Cascade251Invalidated spec clauses252Required closure253Proof obligations254Correction surface255```256257Use exactly one correction surface:258259- `implementation-only`: the current contract is sufficient and code must260 change;261- `spec-and-implementation`: current evidence invalidates a task-owned spec262 clause or proof mechanism;263- `evidence-only`: no defect is proven but decision-critical proof is missing;264 or265- `task-decision-required`: closure requires an unmade product,266 compatibility, operational, destructive, or material scope decision.267268Do not convert a missing proof into a defect unless providing that proof is an269explicit delivery requirement. Put other blocking gaps under required evidence270with the same correction-surface classification.271272If no concrete defect or explicitly required proof deliverable is established,273write `Findings: None`. Keep only convergence-rule-3 `evidence-only` and274`task-decision-required` blockers under required evidence; place non-blocking275unknowns under residual risk. Do not give gaps a P0-P3 finding label. A report276must not simultaneously say that no P0/P1 is proven and emit a P0/P1 finding.277278## DONE WHEN279280A review is terminal when the current boundary is stable, every bounded281material impact surface is examined or stopped at a named unknown, candidate282findings are challenged, the single final challenge is complete, both coverage283dimensions are reported, and the complete output contract is emitted.284285Choose the result in this order:2862871. Return `FAIL` on a stable current head when a P0/P1 defect or incomplete288 promised outcome is proven. Known failure wins even if unrelated coverage289 remains partial; report those gaps separately.2902. Otherwise return `INCOMPLETE` only for a convergence-rule-3 blocker. Partial291 coverage labels alone do not decide the result; name the exact material path292 and bounded evidence required. `INCOMPLETE` is not a soft pass or a holding293 state for continued discovery.2943. Otherwise return `PASS` as soon as the bounded work and final challenge are295 complete and no P0/P1 or rule-3 blocker remains. Record P2/P3, non-material296 unknowns, and honestly partial non-blocking coverage as residual merge risk297 without widening the task.298299An interrupted review, partial output, or report without an exact reviewed head300has no verdict. `PASS` applies only to the reviewed revision and does not301authorize merge.302303## OUTPUT FORMAT304305Use a compact report when one bounded surface and short evidence chain preserve306all mandatory content. Use the full profile for multiple surfaces, findings,307material unknowns, or substantial re-review reconciliation. Do not add empty308boilerplate or filler findings.309310For a containment-based `PASS`, explicitly account for each plausibly affected311entry route, caller or consumer, executable behavior, permission boundary,312state boundary, external contract, and operational path. Mark a boundary313unchanged or inapplicable only from current evidence; omit unrelated domains.314For incremental reuse, state the exact prior evidence reused, why it is proven315unchanged, and why its prior coverage is adequate for the affected area.316317Always return:318319```text320## Risk-Calibrated PR Review321322Result: PASS | FAIL | INCOMPLETE323Review mode: full | incremental | reset-to-full324Report profile: compact | full325Reviewed revision: <base>...<head>326Current PR head: <revision>327Prior reviewed head: <revision | none>328Overall risk: low | medium | high | critical329330Governing outcome331Task scope332Consequence scope333Impact surfaces334Prior finding reconciliation335Required evidence / merge conditions336Findings337Work complete?338Residual merge risk339340Coverage341 Task sources342 PR/diff boundary343 Baseline reuse344 Attacked345 Executed346 Not examined347 Artifact coverage: complete | partial348 Invariant coverage: complete | partial349```350351Sort findings by priority. If none exists, say so without manufacturing P3352filler. On `FAIL` or `INCOMPLETE`, hand the complete root-family and blocker353batch back for correction, bounded evidence gathering, or a task decision.354Stop after one report and wait for a new invocation against an updated PR or355the exact named evidence; never continue discovery inside the same invocation.