Codex Review
Orchestrate codex exec review to review branch changes and iterate on fixes.
The root
The caller supplies the root — the ref the change under review is measured from — on diff-root's consumer contract. Apply that contract here, halt included. Every --base command below resolves the merge base in its own block, as that skill's per-command conversion requires.
Core commands
Review branch diff against base
mb=$(git merge-base <root-rev> HEAD)
test -n "$mb" || { echo "no merge base for <root>"; exit 1; }
codex exec review --base "$mb" -o <output-file> </dev/null
This runs git diff <base-SHA> internally and reviews the entire diff. The review covers all committed changes on the current branch relative to the base — both the original work and any subsequent fix commits.
Other review modes
| Command | Scope |
|---|---|
codex exec review --uncommitted </dev/null |
Staged + unstaged + untracked changes |
codex exec review --commit <SHA> </dev/null |
A single commit's diff |
codex exec "<inline prompt>" </dev/null |
Free-form prompt with inline context (see "Switching to exec with inline context" below) |
Output options
| Flag | Effect |
|---|---|
-o <file> |
Write final review message to file |
--json |
Emit JSONL event stream to stdout |
"custom prompt" |
Positional arg — additional review instructions |
Switching to exec with inline context
codex exec review --base ... operates on git diff alone — it cannot read GitHub Issues, ADRs, or any design intent encoded outside the diff. For self-contained changes this is fine; for phased rollouts where the judgement criteria live outside the diff, review mode systematically misjudges intentional design decisions as regressions. Switch to free-form codex exec "<inline prompt>" to attach the context.
Trigger — before every codex review, check:
- Is this commit part of a numbered phase in a tracked Issue?
- Does the relevant ADR contain phrases like "unmeasured → sentinel X", "tracked separately in a follow-up", "intentional default until ", "out of this phase's scope"?
- Are there thresholds, defaults, or scoped-off paths that look like bugs from the diff but are ADR / Issue-intentional?
If any of the above is yes, skip codex exec review and use codex exec "<inline prompt>" instead.
Why the workaround is needed: --base and the positional [PROMPT] are mutually exclusive on codex exec review, so phase context cannot be attached to a review invocation. The free-form codex exec "..." accepts an arbitrary prompt that can direct codex to read the surrounding context first.
Inline prompt template — the prompt should tell codex to:
- Run
git log <root-rev>..HEAD --onelineandgit show <sha>to read the commits. - Read the tracking Issue (
gh issue view <n>) and the relevant ADRs (pass absolute paths so codex doesn't have to search). - Enumerate explicitly which states are intentional (sentinel values, scoped-off paths, deferred behaviors) and must NOT be flagged as regressions.
- List what the review should focus on (forwarding correctness, default / expert contract consistency, docstring drift, missed call sites, etc. — project-specific).
- List what to ignore — the enumerated intentional states from step 3.
For non-phased self-contained changes, keep using vanilla codex exec review --base ... in the form Core commands gives. The inline-context workaround is only needed when judgement criteria live outside the diff.
Triaging review output
codex review operates on git diff output alone — it has no access to the broader project context, test results, runtime behavior, or design rationale. This means a significant fraction of its findings will be false positives: technically plausible concerns that don't apply given information the reviewer can't see.
Typical false positive patterns:
- Assumed standard behavior: "this regex won't match standard X format" when the actual data uses a project-specific format (verified by tests)
- Missing context on intentional decisions: flagging a design choice as a bug when it was deliberate and tested
- Hypothetical edge cases: warning about inputs that can't occur given the system's constraints
When presenting review output, triage each finding:
- Read the review output and identify each distinct finding (usually formatted as
[P1/P2] summary — file:line) - Cross-check against project context you already have — test results, prior conversation, code you've read. You have far more context than the reviewer did.
- Classify each finding under the
finding-triageSSOT dispositions, applying each per its definition there. The common codex-review cases areactionable,false-positive, anduncertain-validity. - Present the triage to the user, not the raw output. Lead with actionable items, note dismissed items with reasoning.
Review-fix loop
Each iteration runs a full, unbiased review of the entire diff against base. Do NOT inject previous review comments into the prompt — this narrows the reviewer's focus and risks missing new regressions introduced by the fix. The reviewer should always see the code with fresh eyes.
Procedure
Run review
mb=$(git merge-base <root-rev> HEAD) test -n "$mb" || { echo "no merge base for <root>"; exit 1; } codex exec review --base "$mb" -o /tmp/codex-review.md </dev/nullTriage the output using the process above. Present classified findings to the user.
If actionable issues are found, the user (or Claude) fixes them and commits.
Re-run the review — same command, same flags. The new diff includes the fix commits, so the reviewer sees the full picture: original changes plus fixes.
Re-triage — a finding dismissed as false positive in iteration N may become relevant in iteration N+1 if the fix changed the surrounding code. Don't carry over dismissals blindly.
Repeat until no actionable findings remain, or the user explicitly waives the remainder with reasoning — a waived finding no longer stands as an open actionable finding, in this loop and in any loop wrapping it.
What "clean" means
No open actionable findings after triage. A review with only false positives counts as clean; a valid but minor finding keeps whichever SSOT disposition it earns — minor-ness alone never makes it clean.
Important constraints
- Stdin redirect: every
codex execinvocation needs</dev/null; hang signal is absence of theOpenAI Codex v...banner afterReading additional input from stdin.... - Run fresh: run each review with fresh context — fresh reviews are the correct approach for iteration.
- Non-interactive only: Always use
codex exec review, notcodex review, when running from scripts or automation. Theexecvariant runs non-interactively and exits when done. - Timeout: Set timeout to 600000ms (10 minutes) when calling from Bash. Reviews of large diffs can take several minutes.