Review PR
Review the current PR's exact HEAD, fix the findings that belong to it, and return a machine-readable verdict. Two rules run through the whole loop: keep polling deterministic — never spend model turns waiting for external state — and keep the diff inside the scope the PR description already claims.
This skill takes no flags. Its only input is the PR on the current branch (or a PR number the caller names).
Review gate
CodeRabbit (coderabbitai[bot]) is the only gating reviewer. It runs on
push; the fallback trigger is a @coderabbitai review comment. The verdict waits
on CodeRabbit and on the ruleset-required status checks — nothing else.
reviewDecision carries CodeRabbit's standing answer across cycles wherever the
repository enables reviews.request_changes_workflow. It holds
CHANGES_REQUESTED while any of its comments is open, and turns APPROVED once
they are resolved and the pre-merge checks pass. That approval is a real GitHub
review object, submitted for a clean review with no unresolved CodeRabbit threads
and no blocking pre-merge checks.
CodeRabbit approves a pull request once. The approval stands over every later
commit, so reviewDecision keeps reading APPROVED for commits the reviewer has
yet to see. It tracks the commit under review where branch protection dismisses
stale approvals on push. Everywhere else, read it as a fact about the pull
request, and take the exact-HEAD answer from Section 6.
Read which case applies from the snapshot Section 3 already fetches, and see Repository expectations below for what each costs the run.
Never post @coderabbitai approve. It submits an approving review on demand,
with an empty body, whatever the state of the code — so it forges the exact
signal this skill waits on. @coderabbitai review and @coderabbitai full review are the only commands this skill posts.
Unresolved human review threads block pass: triage every unresolved thread in
Section 4 regardless of who opened it. The verdict itself waits on CodeRabbit.
Repository expectations
This skill runs against whatever configuration a repository already has and changes none of it. Three conditions decide how well the loop runs, and all three are read from evidence the run already produces. Name every one that is missing in the verdict, so the person reading it can decide whether to set it.
| Condition | What it buys the loop | How this run detects it |
|---|---|---|
reviews.request_changes_workflow: true, with stale approvals dismissed on push |
reviewDecision becomes CodeRabbit's standing answer, so convergence is read from the decision |
reviewDecision is APPROVED or CHANGES_REQUESTED. null with CodeRabbit reviews present means it is off. An APPROVED review whose commit_id names an older commit means the dismissal half is missing |
reviews.auto_review.auto_pause_after_reviewed_commits: 0 |
Automatic review survives a loop longer than five commits | The default is 5. A reviewer that answered early cycles and goes quiet on a later one has likely reached it |
| A non-bot identity for the loop's own comments | CodeRabbit reads the fix rationale, answers a declined finding, and resolves the thread itself | The run knows who its token authenticates as, and a Skipped: comment is from another GitHub bot reply is the positive proof |
None is required. With the first off, Section 6's acceptance rule carries the
verdict alone and CodeRabbit's reviews land as COMMENTED. With the second at its
default, the tag path in Section 2 recovers the paused reviewer. With the third
missing, every review thread is a monologue and some cycles end needs-changes on
a commit nothing reviewed. Each costs the run time or certainty, which is what
the verdict line reports.
Set the first as a pair. The workflow enabled while stale approvals survive a
push leaves a standing APPROVED over every later commit, which reads as
convergence the loop has not earned. Report that shape by name.
Check two things before enabling the first. Where GitHub's Code Review Limits
are on, only accounts with explicitly granted read access or higher may submit a
review that approves or requests changes; ordinary comments still land. Under
that restriction, request_changes_workflow stops CodeRabbit posting review
comments at all, with Failed to post review comments. Look at Settings →
Access → Moderation options → Code review limits first, and either turn off
"Limit to users explicitly granted read or higher access" or grant CodeRabbit
read access.
Then turn on dismiss stale pull request approvals when new commits are
pushed, in the branch rule protecting the base branch. It is what makes the
approval arrive once per commit, and so what makes reviewDecision track the
commit under review.
The third is a property of the token this skill runs with, and CodeRabbit exposes
no setting for it — see
reviewer-edge-cases.md. The remedy is a
machine-user PAT for the commenting calls. An actions/create-github-app-token
installation token authenticates as a Bot, which the skip applies to.
0. Preflight: pick one transport, or stop
This skill needs bash, jq, git, and exactly one working GitHub API
transport. Probe in this order, before anything else:
for c in bash jq git; do command -v "$c" >/dev/null || echo "MISSING: $c"; done
# Transport A: gh — prove it END TO END with a real API call. In a macOS command
# sandbox, `gh auth status` can pass while every API call dies on TLS
# (`x509: OSStatus …`), because Go's platform verifier needs the keychain.
gh api rate_limit >/dev/null 2>&1 && echo "TRANSPORT: gh"
# Transport B: curl + token — a complete equivalent.
[ -n "${GH_TOKEN:-${GITHUB_TOKEN:-}}" ] && \
curl -fsS -H "Authorization: Bearer ${GH_TOKEN:-$GITHUB_TOKEN}" \
https://api.github.com/rate_limit >/dev/null 2>&1 && echo "TRANSPORT: curl"
Pick the first transport that works and use it for every call in this run —
the same REST/GraphQL endpoints either way (the gh api commands below name the
canonical paths; in curl mode hit the same paths with the token header). Never
mix transports call-by-call within a cycle, and never downgrade to partial
substitutes.
If bash/jq/git is missing, or neither transport works, stop the skill
immediately. Report exactly what is unavailable and end the turn. A partial
review produces a confident verdict from data it could not read. Do not work
around it. Specifically, never:
- install, download, or build
gh,jq, or any other dependency; - parse command output with
sed/awk/grepbecausejqis missing; - substitute local
git log, the PR page's HTML, or your own reading of the diff for the review surfaces in Section 3; - emit a verdict — including
needs-changes— from an incomplete snapshot.
The stop message is the deliverable in that case. Say which tool is missing and
what the operator needs to do (install it, run gh auth login, export
GH_TOKEN, or run the skill outside the sandbox).
1. Anchor one cycle
Resolve the current branch PR with gh pr view --json number,headRefOid. With no
PR on the branch, stop and tell the user to open one; this skill never creates a
PR.
Refuse a dirty tree before reviewing. Push, then require PR headRefOid to equal
local HEAD. Record:
HEAD_SHAand its short form.REVIEW_STARTinZUTC.COMMIT_DATEfromgh api repos/$OWNER/$REPO/commits/$HEAD_SHA --jq .commit.committer.date.- The
git status --porcelainbaseline for later fix ownership. SCOPE: the PR title and body fromgh pr view --json title,body, plus the changed-file list fromgh pr diff --name-only. This is the scope contract for Section 4.
Use the API committer date; local git log can preserve a non-UTC zone. A
finding belongs to this cycle when its commit_id is HEAD_SHA, its timestamp
is at or after either anchor, or its body cites the full/short SHA.
If an API call fails here — auth expired, repository not visible, network blocked — re-run the Section 0 probe once and switch to the other transport for the whole cycle if it passes; if neither transport works, stop as in Section 0.
2. Wait without model turns
Resolve scripts/wait-for-reviews.sh relative to this skill. Skill directories
are often symlinks into a path the command sandbox cannot read (e.g.
~/.claude/skills/x -> ~/.agents/skills/x); if invoking the script fails with
a permission error, copy it byte-for-byte to $TMPDIR with the harness file
tools (Read → Write, which are not command-sandboxed) and run the copy — same
arguments, same contract. Run it with the PR,
HEAD, anchors, and base branch. It silently checks CodeRabbit's artifacts and the
ruleset-required status checks, preserves the last good snapshot, and emits one
terminal JSON object.
Use its --once mode for a read-only diagnostic probe or validation. The normal
bounded wait runs without it.
On Claude Code, invoke the helper through one main-session Monitor call
with timeout_ms slightly above the helper's 900-second budget. Do not delegate
to an Agent subagent. After Monitor started, end the assistant turn; the task
notification resumes the main session.
On another harness, run the helper in one long-lived execution and use that harness's process-wait primitive. Never implement an LLM polling loop.
The helper selects its own transport the same way Section 0 does: gh when a
real API call works, otherwise curl with GH_TOKEN/GITHUB_TOKEN. In a
sandbox where only direct gh invocations are exempted, the helper's nested
gh calls are still confined — export the token so its curl path engages:
WATCHER=skills/review-pr/scripts/wait-for-reviews.sh
GH_TOKEN="$GH_TOKEN" "$WATCHER" \
--owner "$OWNER" --repo "$REPO" --pr "$PR" --base "$BASE" \
--head "$HEAD_SHA" --review-start "$REVIEW_START" \
--commit-date "$COMMIT_DATE" \
--interval 50 --timeout 900
On every cycle after the first — the wait that follows pushing fixes for
findings from an exact-HEAD review at the previous anchor — add
--re-review-cycle. Automatic incremental review picks up every push, so on a
re-review cycle the wait is the mechanism: post nothing, and let it run. The
flag changes only what happens when the wait runs dry, because a clean
incremental review can emit no artifact at all — see needs_full_review below.
A review takes minutes. Measure the repository's own push-to-review latency
before shortening any delay: --tag-after and --full-review-after default to
values that sit above a several-minute review, and a command posted inside that
window interrupts the automatic review it was meant to provoke.
At this decision point, enforce all of these:
- Do not call
sleepdirectly; sleeping is inside the helper/Monitor. - Do not call
true,date,echo waiting, tail a watcher log, manually poll, start a second monitor, or narrate heartbeats while the helper runs. - Keep the monitor silent. Only its terminal JSON should wake the model.
- If state is
needs_tag— cycle 1 only — post one@coderabbitai reviewcomment, then restart the helper once with--tagged-coderabbit. Never retag. If you filter comments by time afterwards, derivesince=from the posted tag comment'screated_atin the API response. CodeRabbit can reply within seconds, and a window derived from your own clock misses it. - If state is
needs_full_review, post one@coderabbitai full reviewcomment, then restart the helper once with--forced-full-review. Never escalate twice. Usefull reviewhere;reviewis itself incremental and returns an acknowledgement while automatic review is un-paused. This state also arrives on a first cycle whose rate-limit window has run out: the refused request was dropped, so asking again ends the wait. - If state is
failed,timeout, orsnapshot_fetch_failing, read reviewer-edge-cases.md before judging it. - If the helper exits
2with no JSON, a dependency is missing or both transports failed. Stop as in Section 0 — do not retry it and do not review by hand.
ready means waiting is complete. The final snapshot and triage decide whether
the PR passes, and the acceptance rule in Section 6 says what counts as
CodeRabbit's exact-HEAD response.
unresolved_coderabbit_threads in the helper's JSON is reported and never gated
on. Section 4 resolves those threads itself, so a zero count describes this run's
own actions.
3. Fetch one final snapshot
After the waiter stops, fetch the full bodies once from all three REST surfaces:
{ gh api --paginate repos/$OWNER/$REPO/issues/$PR/comments --jq '.[] | {surface:"issue",login:.user.login,id,ts:.created_at,commit:"",path:null,line:null,url:.html_url,body}'
gh api --paginate repos/$OWNER/$REPO/pulls/$PR/reviews --jq '.[] | {surface:"review",login:.user.login,id,ts:.submitted_at,commit:.commit_id,path:null,line:null,url:.html_url,body}'
gh api --paginate repos/$OWNER/$REPO/pulls/$PR/comments --jq '.[] | {surface:"inline",login:.user.login,id,ts:.created_at,commit:(.original_commit_id // .commit_id),path,line,url:.html_url,body}'
} | jq -s 'sort_by(.ts // "")'
If the issue-comments REST call fails, fetch the same full top-level comment
bodies through the paginated GraphQL pull-request comments connection. A
successful equivalent fallback is a valid surface. An empty response caused by a
failed command fails closed.
REST review payloads have no severity field. Read CodeRabbit's severity label
from its body. Only inline findings carry path:line. Scope inline findings with
original_commit_id; GitHub may rewrite commit_id as the diff moves, which can
make stale feedback look current.
Fetch every unresolved thread, whoever opened it. Keep the thread node ID for
resolution and fullDatabaseId for replies:
gh api graphql --paginate -f query='
query($owner:String!,$repo:String!,$pr:Int!,$endCursor:String){
repository(owner:$owner,name:$repo){ pullRequest(number:$pr){
reviewThreads(first:100,after:$endCursor){
nodes{id isResolved isOutdated path line
comments(first:50){nodes{fullDatabaseId body createdAt url author{login}}}}
pageInfo{hasNextPage endCursor}}}}}
' -f owner=$OWNER -f repo=$REPO -F pr=$PR \
--jq '.data.repository.pullRequest.reviewThreads.nodes[] | select(.isResolved==false)'
Also fetch the base ruleset and
gh pr view --json isDraft,mergeStateStatus,mergeable,reviewDecision,statusCheckRollup.
A failed snapshot or ruleset lookup fails closed; never interpret an empty
payload as clean.
Take the verdict from comment and review bodies. CodeRabbit's own review-progress
check is bound to the head commit and reports that the run finished;
reviews.fail_commit_status defaults to false, so it goes green with findings
open, and no acceptance rule below rests on it. Where the repository enables it,
read it once for completion: a run that finished and said nothing is the silence
needs_full_review answers.
A review in progress note counts as engagement. So does a green CodeRabbit
check with no review body and no inline finding. A review whose state is
APPROVED or CHANGES_REQUESTED at HEAD_SHA is a verdict at any body length,
and so is a walkthrough comment reporting no actionable comments over a range
ending at HEAD_SHA, which on a clean pull request is routinely the only
artifact present. A walkthrough carrying neither is engagement. Section 6's
acceptance rule holds on every cycle.
4. Triage against the scope contract, then fix
Dedupe by thread/URL. If an inline finding is an unresolved thread, handle it once as the thread.
Classify every finding against SCOPE before deciding how to answer it. A
finding is in scope only when it is one of:
- a defect in the diff — the code this PR added or changed is wrong, unsafe, or breaks a caller;
- a gap the PR description itself promises to close;
- a required-check failure this PR causes.
Everything else is out of scope, however reasonable it sounds: adjacent pre-existing bugs, refactors of code the PR merely touched, new abstractions, extra features, wider test coverage than the change needs, renames, and style rewrites beyond the changed lines. Reviewers and bots suggest these constantly. Do not implement them in this PR. Reply with one sentence naming the reason (out of scope for this PR), then resolve. Carry every declined finding into the Section 6 verdict so the user sees what was turned down and can decide whether it deserves its own PR.
When a finding's scope is genuinely unclear, or a reviewer argues the PR's
approach is architecturally wrong, read docs/architecture/INTENT.md and
docs/architecture/adr/* before answering. If a recorded decision already
settles it, reply citing that document by path and decline the change. If those
files do not settle it, ask the user — never widen the diff to end an argument.
Then, for the findings that survive triage:
- Fix in-scope bugs, correctness/security issues, and CodeRabbit
⚠️/🔴/🟠at the root. Still triage low-severity findings even when the overall verdict says the patch is correct. - Apply cheap, correct CodeRabbit
🧹/🔵suggestions that land inside the changed lines; otherwise explain the skip. Ignore🤖 Prompt for AI Agentsand🧩 Analysisscaffolding. - Ask the user on a design judgment call. Reply to a false positive with evidence.
- For consistency findings, trace the real resolver and make both paths read the same runtime source.
Never edit a file outside the PR's changed-file list to satisfy a finding. If a correct fix genuinely requires touching a new file, say so and get the user's agreement first. That is a scope change.
Always reply before resolving. Resolve only a finding you fixed or explicitly
declined as out of scope; anything still open or awaiting the user stays
unresolved. Reply inline through
pulls/$PR/comments/{id}/replies, top-level through issues/$PR/comments, and
resolve with:
gh api graphql -f query='mutation($t:ID!){resolveReviewThread(input:{threadId:$t}){thread{isResolved}}}' -f t=$THREAD_ID
For an outdated thread, reply and resolve only when the current diff already covers it; otherwise leave it unresolved and tell the user.
5. Verify, push, and repeat
Every ruleset-required check must be SUCCESS/SKIPPED; a missing required
context is pending. A ruleset lookup failure is needs-changes — with one
carve-out: HTTP 403 with "Upgrade to GitHub Pro or make this repository public"
means the rules feature is unavailable on this plan, so no rules can exist. Read
that as an empty ruleset with no required checks. A required pull_request
approval gate only evaluates once the PR is ready for review; on a draft PR,
report it as deferred.
BLOCKED/DIRTY/UNKNOWN blocks unless only that ready-only gate remains with
zero pending checks and unresolved threads.
Commit only files introduced relative to the cycle baseline. Stop if unrelated
edits appear. Before pushing, diff the new changed-file list against the SCOPE
list from Section 1: a file that review fixes added to the PR is a scope change
the user has not seen — stop and confirm it. Run the repo's own
lint/format/typecheck entry point on the changed files, reading
CONTRIBUTING.md, AGENTS.md, CLAUDE.md, or the package/Makefile scripts to
find the command.
If the accepted fixes made the PR description inaccurate, update the description to match what the PR now does. Never let the diff drift ahead of its stated scope.
Push fixes, re-anchor the new HEAD, and run another deterministic wait with
--re-review-cycle set (Section 2). Cap at five cycles; each cycle must make a
concrete fix. Stop on repeated or non-actionable churn.
Every cycle spends a review. Automatic incremental review of a pushed fix
counts against the hourly allowance exactly as a first review does, and so does
each @coderabbitai review or full review this skill posts. Five cycles are
therefore five reviews before either escalation, against ten an hour on Pro+ and
five on Pro — the accounting behind the rate limits in Section 5a. A cycle that
pushes a one-character fix costs the same as one that pushes a rewrite, so batch
a cycle's fixes into a single push.
The approval trails the resolution that triggers it, so a wait can return while
reviewDecision still reads CHANGES_REQUESTED from the review you just
answered. With every finding fixed and the decision still standing, let it settle
through the same turn-free mechanism as Section 2, and judge afterwards. Never
hand-roll a while loop around gh api, including when the helper returns
something you doubt. Report a waiter that stops too early as a bug.
5a. Re-arm on only-waiting
Before returning needs-changes, check whether the only blocker is external
progress — CodeRabbit still in-progress, a rate limit with time left to run, or a
required check still pending-not-failed — with zero actionable findings/threads,
zero failed checks, and zero pending local fixes. Any actionable finding or
failed check returns needs-changes immediately; never mask a real problem
behind "still waiting".
A rate limit belongs here because it expires on its own. The quota is
per-developer and charged to whichever identity pushed, so a loop running as a
bot spends a different allowance than the repository owner. When the helper
reports coderabbit: unavailable, coderabbit_retry_after carries the seconds
still left on the window CodeRabbit published — re-arm on that number.
Published windows run from seconds to tens of minutes, so any fixed delay misses.
failed never re-arms: CodeRabbit has stopped, and the cause it prints (most
often a pull request closed under the review) persists through any wait.
unavailable means a live notice: CodeRabbit's own refusal, with time still
on its window and nothing said since. Once that window runs out the helper stops
reporting it, and the next wait escalates through needs_full_review, because a
refused request is dropped. Read a refusal only from CodeRabbit's own notice
wording. It summarises every diff in its own words, so a pull request about rate
limits earns a walkthrough containing "rate limit" while the review proceeds
normally.
On that only-waiting state, re-arm, so the operator never re-runs the helper by hand:
- Claude Code / dynamic-loop harness: schedule a self-paced wake-up
(
ScheduleWakeup, dynamic-mode loop) with a long fallback interval (~20–30 min) re-arming this skill for the same PR, then end the turn. Idle must cost nothing — no assistant turns between wakes, never a short busy poll. This re-arm is outer only: never wrap the Section 2 turn-free watcher wait in a recurring loop. - Each wake: re-resolve the PR HEAD and re-run Section 1 exact-HEAD anchoring before judging — a human may have pushed between wakes — then re-enter from Section 2.
- Always terminate. Stop on a real verdict (
passor realneeds-changes), on operator interrupt, or after a few consecutive wakes with no reviewer/check movement — then return the lastneeds-changes(still waiting). Never re-arm forever. - Other harness with no self-paced-loop primitive: emit
needs-changes(still waiting) and tell the operator to re-arm with/loop /review-pr <PR#>(dynamic mode, long fallback).
6. Return the verdict
GITHUB_REVIEW_RESULT:
- PR: <url or number>
- CodeRabbit: <responded/in-progress/unavailable/failed; reason. On unavailable,
name the account whose quota ran out and the published retry delay>
- Review cycles: <count>
- Issues found / fixed in scope / declined out of scope: <n> / <n> / <n>
- Declined as out of scope: <one line each: finding, thread URL, reason; or none>
- Files added to the PR by review fixes: <list or none>
- Unresolved actionable threads: <count>
- Review decision: <APPROVED/CHANGES_REQUESTED/null; null where the
request-changes workflow is off. On APPROVED, name the commit the approving
review carries>
- Repository configuration to set: <each setting from Repository expectations
this run found missing, with what it cost; or none>
- Pending required checks: <list or none; include mergeStateStatus>
- Ready-only deferred gates: <list or none>
- Verdict: <pass/needs-changes>
- Summary: <1-2 sentences>
Return pass only when CodeRabbit delivered an exact-HEAD response,
every in-scope finding is fixed, every out-of-scope one is answered and listed in
the verdict, all threads are resolved, reviewDecision reads APPROVED wherever
the request-changes workflow is enabled, and every required check is
SUCCESS/SKIPPED.
reviewDecision is necessary and never sufficient. Auto-approval fires once
CodeRabbit's comments are resolved, and Section 4 puts resolution in this skill's
own hands, so the flag confirms that CodeRabbit agrees with what the skill
already did — which is why Section 4 resolves only what it fixed or declined. The
field also carries no commit: CodeRabbit approves a pull request once, so an
APPROVED read on cycle two can be the approval earned on cycle one.
Take the exact-HEAD response from the acceptance rule below. Check the
approving review's own commit_id against HEAD_SHA before reading the decision
as this cycle's answer, and where it names an older commit, report the missing
dismissal under Repository configuration. The substantive conditions above are
what make the resolution honest. Read the decision as the last check on a case
built elsewhere. A PR that ships exactly what its description promised is the
goal, and declining scope creep is a pass.
CodeRabbit's exact-HEAD response is a substantive response at HEAD_SHA, which
is any one of a nonzero-body review, inline findings, a review carrying the
state APPROVED or CHANGES_REQUESTED on that exact commit, or the walkthrough
comment carrying No actionable comments were generated in the recent review
over a reviewed range ending at HEAD_SHA. The state counts on its own:
CodeRabbit approves with an empty body when it has nothing to say, which is the
shape of every clean review.
The walkthrough counts because a clean review can produce no review object and no
inline comment at all — that one sentence, edited into the walkthrough comment,
is the entire verdict. Require both halves on the same comment. The sentence
alone names no commit, and CodeRabbit rewrites the comment in place on each
review, so the anchor is the end of the range it reports. A range that
starts at HEAD_SHA describes a review of the commit after it.
The rule is the same on every cycle. Automatic incremental review covers each new push, so a fix commit earns the same artifact a first commit does. These count as engagement on any cycle:
- an empty-body
COMMENTEDreview — CodeRabbit wraps its thread replies in one, so it marks a conversation; Review finishedorAction performed— the acknowledgement of a command, emitted within seconds of the command itself;No new commits to review— a statement about CodeRabbit's own bookkeeping;- zero unresolved CodeRabbit threads — Section 4 resolves those threads itself.
Silence is the hard case. Read the walkthrough comment before concluding a review
is silent; a clean review usually lives there. Where even that is absent,
needs_full_review exists: one @coderabbitai full review, and if that still
yields no artifact bound to HEAD_SHA, the verdict is needs-changes naming the
commit as unreviewed. Asked to redo a review already reported in the walkthrough,
that escalation answers Action performed / Full review finished within seconds
and produces nothing.
pass means reviewed and ready to merge. Its scope ends there: building,
deploying, and working live all sit outside it. Never merge from this skill; hand
back to the operator.
Return needs-changes for an unresolved finding/thread, a failed or pending
required check, an engaged-but-incomplete CodeRabbit review, CodeRabbit being
unavailable (quota/rate limit), or snapshot_fetch_failing. When the only
remaining blocker is a still-in-progress CodeRabbit review or a
pending-not-failed required check with nothing actionable left, route through 5a
and re-arm.