PR review orchestrator
Three reviews of the same diff merged into one report: correctness, type safety, comments. Correctness runs as two independent passes, the harness's code-review skill invoked by the orchestrator and a stage subagent reviewing by hand. The orchestrator resolves the refs, builds a structural brief, dispatches everything, and consolidates what comes back.
The orchestrator keeps its context lean. It reads briefs and subagent reports and nothing else. It never opens a diff, a changed file, or a git show, however small the PR looks. Everything that needs the diff happens inside a subagent.
Input
/review:pr-review <pr> [<pr> ...]
One PR number, or several forming a stack in merge order. Everything else is resolved from gh.
Step 1: resolve the stack
For each PR number:
gh pr view <n> --json number,title,baseRefName,headRefName,url
Each PR is reviewed against its own baseRefName, never against the trunk. In a stack, PR two's base is PR one's head branch, so reviewing it against main would re-review PR one's diff.
Fetch both refs so the diff resolves in this clone:
git fetch origin <baseRefName> <headRefName>
The base and head passed to every stage are then origin/<baseRefName> and origin/<headRefName>.
State the resolved pairs before starting, one line per PR: number, title, base, head, url. If gh is unavailable, or a number does not resolve to a PR, stop and say which number failed. Do not guess a base ref.
Step 2: structural brief
One brief per PR, plain text, pasted whole into every subagent prompt for that PR.
If a skill named investigate is available, invoke it for that PR's changed paths and use its brief.
If it is not, build a lighter one from these commands:
git diff --stat <BASE>...<HEAD>
git diff -U0 <BASE>...<HEAD> | grep -E '^@@'
git diff <BASE>...<HEAD> | grep -E '^\+(export )?(async )?(function|class|const|interface|type) |^\+[[:space:]]*(public|protected|private) function'
git grep -nw <symbol> <HEAD>
<BASE> and <HEAD> are the two refs resolved in step 1, the same slots step 3 fills.
The first gives the shape of the change. The next two give the added or changed functions, classes, and methods, from the hunk headers and from the added lines. The last runs once per changed symbol and gives its callers on the head ref, which is what tells a stage whether a signature change has call sites the PR missed.
Both anchors in the third command carry a +, so it matches added lines and not the context lines around them. Dropping the + inverts the result: every unchanged declaration in the hunk matches and every added one does not, and the per-symbol git grep then has nothing to run on.
Three dots, so only what the branch adds is in scope.
The commands run here produce the brief. Their output goes into the brief, not into a reading pass by the orchestrator. Keep each brief to roughly a page: the stat block, the symbol list, and the caller list.
Step 3: run the passes
Per PR: one code-review invocation and three stage subagents, all dispatched in the same turn so they run in parallel where the harness allows. A three PR stack is three invocations and nine subagents.
Correctness, first pass: code-review
If a skill named code-review is listed, the orchestrator invokes it once per PR, from its own context, with the PR number as the target:
code-review <number>
The skill runs its review in a background subagent and returns only that subagent's name. Its report arrives later as a task completion notification addressed to the session that invoked it. Invoked from the orchestrator, that is the orchestrator. Invoked from inside a stage subagent, the harness parents the background subagent to the top-level session anyway, so the notification bypasses the stage, and the stage waits for a result that never reaches it. That is why the stage template below reviews by hand and does not mention code-review.
The notification carries the report as text. Hold it for step 4. Consolidation waits until it has arrived; the correctness section is never written from what the skill was expected to find. If the skill is not listed, or its subagent stops without a report, the hand review is the only correctness pass, and ## Not available in this run says so.
The same parenting applies one level down. At effort levels where code-review spawns its own finder subagents, those finders are registered under the orchestrator's session as well, so their results arrive at the orchestrator while the fork idles with nothing left to wait on. The sign is a completion notification from the fork saying its finders are still running, followed by idle notices from agents named finder-*. Collect every finder result that arrives, then send them verbatim in one message to the fork by its name, code-review. It resumes, runs its verification pass, and returns the consolidated report as a second completion notification. Finder results are candidates, not findings; only the fork's consolidated report enters step 4.
Stage subagents
Each prompt is one of the templates below with these slots filled:
| Slot | Value |
|---|---|
<BRIEF> |
the whole brief from step 2, for this PR |
<BASE> |
origin/<baseRefName> from step 1 |
<HEAD> |
origin/<headRefName> from step 1 |
<PR_TITLE> |
the PR title |
<PR_URL> |
the PR url |
Correctness stage, the second pass:
Review PR "<PR_TITLE>" (<PR_URL>) for correctness. Base ref <BASE>, head ref <HEAD>.
Only what the branch adds is in scope: git diff <BASE>...<HEAD>, three dots.
Structural brief:
<BRIEF>
Review by hand: read the diff and the changed files, and look for behaviour changes,
error handling, boundary conditions, and missing tests. Validate any claim you can by
running the project's tests or linters; record what you ran.
Skills that run in a background subagent, `code-review` among them, deliver their
result to the session that spawned you, not to you. The orchestrator runs
`code-review` itself as a separate pass, so your findings are your own reading of
the diff.
Return exactly these three sections and nothing else:
## Findings
One entry per finding, each starting with `path:line` on the HEAD side, then the
claim in one or two sentences. No finding without a `path:line`.
## Validated
Every check you ran, with the command and its result.
## Skipped
Every skill, tool, or check you could not use, with the reason.
Type safety stage:
Review PR "<PR_TITLE>" (<PR_URL>) for type safety. Base ref <BASE>, head ref <HEAD>.
Structural brief:
<BRIEF>
Invoke `review:type-safety-review base=<BASE> head=<HEAD>`. That skill returns its
report as its response and writes no file. Relay its findings in the shape below,
keeping each finding's rule id (`PHP-1` to `PHP-5`, `TS-1` to `TS-4`), its verbatim
quote, and its proposed shape.
That skill inherits comment-audit's `base=`/`head=` detection and its ask. Both are
already given above, so if it asks for anything else, you have no one to ask: record
the miss under `## Skipped` and carry on with the rest of the review rather than
stopping.
Return exactly these three sections and nothing else:
## Findings
One entry per finding, each starting with `path:line` on the HEAD side, then the
rule id, the quote, the reason, and the proposed shape.
## Validated
Every check you ran, with the command and its result.
## Skipped
Every skill, tool, or check you could not use, with the reason. Include the paths
the skill reported as skipped for being generated or vendored.
Comments stage:
Review PR "<PR_TITLE>" (<PR_URL>) for comments and documentation. Base ref <BASE>,
head ref <HEAD>.
Structural brief:
<BRIEF>
Invoke `review:comment-audit base=<BASE> head=<HEAD>` in report mode. Never pass
`--apply`: this run reports and changes nothing. That skill writes its report to a
file. Read that file and return its summary and findings in the shape below.
That skill's step 0 asks for `readme=` when a MOVE verdict needs one, and for a
`base=` it cannot resolve. You have no one to ask. Both are already given above
except `readme=`, so if it asks for that, record the miss under `## Skipped`, take
the proposed section it writes without a path, and carry on with the rest of the
audit rather than stopping.
Return exactly these three sections and nothing else:
## Findings
The audit report lists findings as `**L123, VERDICT**` under a `### path` heading;
prefix each one you relay with the path from its file heading, so it reads
`path:line`. One entry per finding, each starting with `path:line` on the HEAD side,
then the verdict (DELETE, TRIM, MOVE, KEEP, UNSURE, WRONG), the verbatim comment,
and the rewrite or pointer where the verdict has one. KEEP verdicts may be one line
each.
## Validated
Every check you ran, with the command and its result. Include the report file path
and the comment line count the audit opened with.
## Skipped
Every skill, tool, or check you could not use, with the reason.
Subagents may use IDE MCP tools, a local database, or a REPL when those exist in the setup. They record what they used under ## Validated and what they could not under ## Skipped. A stage that returns no ## Skipped entries had everything it needed.
A subagent that cannot invoke its skill falls back to reviewing by hand against the same criteria and says so under ## Skipped. It does not return an empty report.
Step 4: consolidate
One report, written by the orchestrator from the subagent responses. It is the whole output of this skill:
# Review: <PR list>
## <number> <title>
<url>
### Correctness
### Type safety
### Comments
## Not available in this run
Rules for the body:
- One section per PR, in the merge order given as input, holding the three stage sections.
- Every finding keeps its
path:linefrom the HEAD side, so it can be pasted as a PR review comment. - A finding two stages both reported appears once, under the stage that ruled on it most precisely, labelled with both stage names.
### Correctnessmerges the two passes. A finding both passes reported appears once, labelledboth passes. A finding one pass reported keeps its label,code-review onlyorhand review only, so the reader knows how much weight it carries. Where the passes disagree, both claims are listed under the finding; the orchestrator does not pick a side, since it has not read the diff.- A stage with no findings gets its heading and one line saying so. An empty heading reads as a lost subagent.
## Not available in this runnames every skill or tool any subagent listed under## Skipped, once each, with the stages that wanted it. Ifcode-reviewwas not listed or returned no report, it is named here with the reason, and the correctness findings carry thehand review onlylabel. When every stage had everything, the section says so in one line rather than being dropped.
Style
The consolidated report follows ../writing-comments/references/style.md.