Panel review
A workflow skill. The user has just changed code (or explicitly asked for a review), and the change needs to be reviewed across four orthogonal axes — spec adherence, project standards, independent test adequacy, and specialist domain expertise — with each axis able to disagree with the others without being silenced.
Proactive trigger discipline. This skill is meant to auto-fire
after implementation, not on every chat message. Fire when there's a
plausible code-change checkpoint — a finished feature, a fixed bug, a
landed refactor, an "I think that's done" — but not on conversation,
questions, planning, or tiny one-line edits the user clearly didn't
want reviewed. When in doubt, run a quick git diff --stat and ask:
"Want me to run a panel review on this?" rather than firing
unilaterally on an ambiguous message.
Each axis catches a different failure mode:
- Spec catches "implements the wrong thing well". Standards and Domain both pass; the diff just doesn't do what the issue asked.
- Standards catches "implements the right thing the wrong way for this repo". Spec passes; the conventions are broken.
- Test adequacy catches "the implementation and tests agree on a weaker contract". It checks observable AC coverage, negative-test mutation value, boundary cases, and whether the tests exercise the real seam.
- Domain catches "implements the right thing with a hidden landmine". Spec and Standards pass; a security, K8s, library, or duplication issue lurks.
Reporting them separately is deliberate. Synthesised together, one
axis masks another. Inspired by Matt Pocock's two-axis /review
skill (Standards + Spec), generalised with a Domain axis powered by
the project's installed specialist reviewer agents.
Operating modes
Interactive mode (default): keep the confirmation behavior below: ask when no fixed point/spec is available, the diff is large, or the Domain panel has 5+ agents.
Orchestrator mode: /plan-orchestrate supplies the fixed point and spec. Do
not pause for confirmation because the diff or panel is large; announce the
scope and proceed. Keep the same bounded briefs and selected panel, record any
missing source or reviewer failure, and always finish with the stable PANEL:
line. All other interactive behavior is unchanged outside this mode.
Prerequisites
- Working directory is a git repository (or files were explicitly named).
- At least one of: a spec source (issue / PRD), project standards
docs (CLAUDE.md, etc.), or one or more reviewer agents installed
at
.claude/agents/or~/.claude/agents/. One axis is enough to run. ghCLI installed if reviewing a PR by number or fetching issue bodies.
Step 1 — Pin the fixed point
The fixed point is what we're diffing against. Whatever the user
said is the fixed point — a commit SHA, branch name, tag,
origin/main, HEAD~5, a PR number, an explicit file list. Don't
be opinionated; pass it through.
If the user didn't specify, ask before proceeding:
Review against what — a branch (
main?), a commit, "since I started this branch", or a PR number?
Do not silently auto-detect. The whole review hangs on this and a wrong guess wastes the panel's effort.
Once pinned, capture once and reuse:
- Diff command:
git diff <fp>...HEAD(three-dot — compares against the merge-base, which is what "what did this branch change?" means). - Commit list:
git log <fp>..HEAD --oneline(used to find issue/spec references). - File list:
git diff <fp>...HEAD --name-only.
Print to the user:
Reviewing against : N files, M insertions, L deletions, K commits. <file list, capped at 20 with "+N more">
If the diff is empty, ask for explicit file paths or a PR number.
If the diff is > 50 files or > 3000 lines, interactive mode asks whether to split into smaller reviews or proceed. Orchestrator mode proceeds without confirmation and keeps every reviewer brief bounded.
Step 2 — Read project context
In one parallel batch, read the repository's instruction and standards sources:
CLAUDE.md/AGENTS.md (including parents), .claude/rules/*.md, relevant
architecture/design/ADR indexes, SECURITY.md, CONTRIBUTING.md, and any
CONTEXT.md/CONTEXT-MAP.md files that exist.
Step 3 — Detect the spec source
The Spec axis needs to know what was asked for. Look in this order:
Issue references in commit messages. Parse
git log <fp>..HEAD --format=%Bfor#123,Closes #45,Fixes GH-67,gitlab !89,JIRA-100,LIN-1234. Fetch the issue body viagh issue view <N>(or the project'sdocs/agents/issue-tracker.mdworkflow if present).A path the user passed as an argument — "review against spec at docs/specs/foo.md".
PRD / spec files under conventional locations matching the branch name or feature:
docs/specs/<name>.md,docs/prd/<name>.md,specs/<name>.md,.scratch/<name>.md,docs/acceptance/<name>.md.If nothing is found, ask the user:
I don't see a spec or issue reference for this branch. Path to a spec, or skip the Spec axis?
If the Spec axis is skipped, note it in the final output as "Spec axis: no source available — skipped" with the reason. Do not fabricate a spec.
Step 4 — Inventory standards sources
The Standards axis needs to know what conventions this repo documents. Collect filenames (the Standards subagent will read them):
CLAUDE.md,AGENTS.md,CONTRIBUTING.mdCONTEXT.md,CONTEXT-MAP.md, per-directoryCONTEXT.mdfilesdocs/adr/*.md(architectural decisions ARE standards)STYLE.md,STANDARDS.md,STYLEGUIDE.mdat repo root or underdocs/.claude/rules/*.mddocs/design/principles.mdif present
Explicit skip rule (inherited from Matt's design): tell the Standards
subagent not to re-check anything enforced by detected formatter, linter,
type-checker, compiler, or project tooling configs (for example
.editorconfig, ESLint/Biome/Prettier/TypeScript, .golangci.yml, pyproject
Ruff/Black/mypy, rustfmt/clippy). Re-flagging tool-enforced rules wastes tokens.
If no standards docs are found, note "Standards axis: no project standards docs found — axis returned an empty report" in output; don't skip the axis silently.
Step 5 — Inventory the available domain panel
List .claude/agents/*.md and ~/.claude/agents/*.md. Parse each
file's description field (frontmatter).
Build a table:
| Agent | Scope (first line of description) | Defer-to |
|---|
This is the domain panel of available specialist reviewers. Don't fabricate agents that aren't installed — if a dimension has no matching agent, note the gap rather than skip the dimension silently.
Step 6 — Classify the diff (Domain axis)
For each file in the diff, identify the domain dimensions that apply. The mapping below is a default; project-specific agents take precedence over generic ones when their scope matches.
| Dimension | File / content signals | Default agent |
|---|---|---|
| Security (cross-cutting) | HTTP handlers, route registration, Authorization, crypto., tls, fetch( / http.Get with user input (SSRF), deserialisation, SQL/template assembly, JWT, OAuth, password handling |
secure-code-reviewer |
| Architecture (cross-language) | New module / package; new public API; cross-module imports; renamed exported type; new interface; new top-level directory; >3 files touched across distinct modules; new layering boundary; or any non-trivial *.go / *.ts / *.tsx / *.py / *.rs / *.java change with structural shape |
software-architect |
| K8s in-cluster | *.yaml / *.yml with kind:, Chart.yaml, values.yaml, kustomization.yaml, templates/*.yaml |
kubernetes-deployment-expert |
| K8s controller / CRD | api/v*alpha* / api/v*beta* / api/v1, controllers/, internal/controller/, imports of sigs.k8s.io/controller-runtime, controller-gen markers |
kubernetes-operator-expert |
| DevOps / CI / IaC | .github/workflows/*, .gitlab-ci.yml, *.tf, *.tfvars, Dockerfile, *.dockerfile, cloudbuild.yaml, Jenkinsfile, Pulumi.yaml, cdk.json, Taskfile.yml, Makefile |
devops-expert |
| Duplication | Three+ touched files with similar shape; new files that look like copies of existing | code-duplication-reviewer |
| Library reuse / dep audit | go.mod, package.json, requirements.txt, pyproject.toml, new top-level dependency, hand-rolled utility shapes |
library-reuse-reviewer |
| Reinvention / over-build | Any non-trivial code diff (default-on, NOT signal-gated — see below) | library-reuse-reviewer + code-duplication-reviewer |
| Project-specific surfaces | (varies — read each available agent's description frontmatter to learn its scope) |
Any project-level agent in .claude/agents/ that names a domain not covered above — typically architects for a specific framework, protocol, API surface, or UI workspace |
If no Domain agent is available (including a pure-docs diff), mark
Domain axis: unavailable — no applicable installed specialist and still fan
out Spec, Standards, and Test adequacy. Domain unavailability is not a reviewer
failure.
Classification rules:
- Always include
secure-code-reviewerfor diffs touching application code with external trust boundaries. software-architectis the default architecture reviewer for any non-trivial code diff — cross-language design plus the language-architecture role when no language-specific architect is installed.- Project-specific architects in the repo's
.claude/agents/compose withsoftware-architectrather than replacing it. Both can run on the same diff: the project-specific one carries domain-loaded invariants and ADR knowledge,software-architectcarries the cross-cutting design lens. code-duplication-reviewerandlibrary-reuse-reviewerare DEFAULT-ON for any non-trivial code diff (any diff touching application logic — not pure-docs / pure-config). Do NOT gate them on "a new dependency was added" or "actual duplication shape is already visible". Their entire job is to FIND the duplication and the reinvented stdlib that ISN'T obvious from the diff surface; gating them on the signal already being visible skips them on exactly the diffs where they add the most value. When in doubt, include both — they default to silence when they find nothing. The reuse pair is cheap insurance against the most common over-build failure mode an agent introduces (reinvented stdlib, speculative abstraction, unrequested layer). See the reuse-ladder brief in Step 8.- Pure-docs diffs skip the Domain axis (Spec axis may still run).
- Pure-test diffs: a single Domain reviewer (project's test-review agent if present, else language architect).
Step 7 — Announce the four-axis plan
Reviewing against <fp>: N files, M insertions, L deletions, K commits.
Spec axis: checking against #123 ("Add /preview endpoint")
Standards axis: reading CLAUDE.md, .claude/rules/, docs/adr/
skipping tooling: golangci-lint, biome, prettier
Test adequacy: independently tracing requirements to assertions and seams
Domain axis (running in parallel):
- secure-code-reviewer — auth/HTTP/SSRF surface in api/handlers/
- kubernetes-deployment-expert — deploy/staging/ manifests
- devops-expert — .github/workflows/release.yml changes
- library-reuse-reviewer — default-on: reinvented stdlib / over-build
- code-duplication-reviewer — default-on: duplication across the diff
Skipping (no diff in scope):
- kubernetes-operator-expert
Gaps (dimension detected, no matching agent installed):
- (none)
Wait for user pushback only in interactive mode when the panel is large (5+ agents across the Domain axis) or they asked for a dry-run. Orchestrator mode proceeds immediately.
Step 8 — Fan out (PARALLEL)
In a single assistant turn, issue all of:
- 1 ×
Agentcall for the Spec axis (general-purpose subagent with brief). - 1 ×
Agentcall for the Standards axis (general-purpose subagent with brief). - 1 ×
Agentcall for the Test adequacy axis (general-purpose subagent with brief). - N ×
Agentcalls for the Domain panel (named specialist agents).
They run concurrently, separate contexts, no order dependencies.
Spec subagent
subagent_type:general-purposedescription: "Spec adherence check for "prompt:Diff to review:
git diff <fp>...HEADCommit list: Spec source: <path or inline contents of the issue/PRD>Read the spec carefully, then read the diff. Report:
- Missing — requirements the spec asked for that the diff doesn't implement, or only partially implements.
- Scope creep — behaviour added by the diff that the spec didn't ask for.
- Wrong — requirements that look implemented but where the implementation appears incorrect against the spec.
Quote the specific spec line / requirement for each finding. Classify every finding explicitly as
blocker,important, oradvisory: blocker = must fix before merge; important = concrete non-blocking fix; advisory = judgement/polish. Under 400 words. Default to silence when uncertain — only flag concrete mismatches.Root-cause discipline: when a finding names a symptom, note whether the diff fixes the root cause or only the path the ticket names — a sibling caller may still be broken.
If no spec source was found in Step 3, skip this call and note in final output.
Standards subagent
subagent_type:general-purposedescription: "Standards conformance check for "prompt:Diff to review:
git diff <fp>...HEADStandards source files (read these): <list from Step 4> Tooling-enforced configs to SKIP (tooling already runs on every commit; do not re-flag what they cover):Read the standards docs, then the diff. Report — per file / hunk where relevant — every place the diff violates a documented standard. Cite the standard (file + the rule). Classify every finding explicitly as
blocker,important, oradvisory: blocker = must fix before merge; important = concrete non-blocking fix; advisory = judgement/polish. Under 400 words. Default to silence when uncertain.
Test-adequacy subagent
subagent_type:general-purposedescription: "Independent test-adequacy check for "prompt:Diff to review:
git diff <fp>...HEADSpec source: <path or inline contents; state unavailable if skipped> Acceptance plan / verify contract:Review tests independently of the implementation. Trace each requirement or AC to an assertion at the lowest adequate layer. Flag missing or weaker coverage, tests that cannot fail on regression, negative assertions without a planted violation, fake-only tests that bypass the real seam, and missing boundary/failure cases. Do not re-report style or implementation findings. Classify findings as blocker, important, or advisory and cite test paths and requirement/AC ids. Under 400 words; default to silence when adequate.
For a pure-docs or pure-config diff with no executable behavior contract, mark this axis not applicable; that is a deliberate skip, not a reviewer failure.
Domain agents
For each specialist in the panel:
subagent_type: the agent'snamefield.description: " review of ".prompt:- The exact diff scope (
git diff <fp>...HEAD -- <paths>). - A one-line statement of what the agent should focus on, derived from its own description.
- The project-context items the agent's body needs.
- Reminder: "Respect your calibration discipline — flag only findings that affect correctness, security, or stated requirements. Classify every finding as blocker, important, or advisory. Map your native severity deterministically: Critical/High = blocker, Medium = important, Low/Info = advisory. Default action when uncertain is silence."
- The exact diff scope (
Reuse pair — extra brief (library-reuse-reviewer +
code-duplication-reviewer). When fanning out the default-on reuse
pair, append the reuse-ladder brief from
references/reuse-ladder.md to each of
their prompts so they review systematically rather than ad-hoc. The
brief gives them the 7-rung reuse ladder (Does this need to exist?
→ Already in codebase? → Stdlib? → Native platform? → Installed
dep? → One line? → Minimum), the delete:/stdlib:/native:/
yagni:/shrink: tag vocabulary, the net: -<N> lines possible.
score, and root-cause discipline (grep every caller, fix the shared
function once). Read the reference file and paste the brief block
into each reuse agent's prompt.
Set run_in_background: false (default) so synthesis blocks on
completion.
If the environment or the user has expressed a model preference, honour it on each Agent call.
Step 9 — Four-axis output + machine result
Do NOT merge findings across axes. They are deliberately orthogonal — one passing while another fails is exactly the information you want to preserve. Cross-axis merging is the failure mode Matt's design exists to prevent.
For the Spec, Standards, and Test adequacy axes, present each subagent's explicitly classified findings verbatim or lightly cleaned. Don't rerank or infer a classification from prose. An unclassified finding is a reviewer failure to retry, not a countable finding.
For the Domain axis only, synthesise across the panel:
Synthesis (Domain axis only)
- Combined table — every finding from every domain agent, preserving its
source severity and its required panel classification. Domain mapping is
exhaustive: Critical/High →
blocker, Medium →important, Low/Info →advisory; missing or unknown severity/classification is a reviewer failure. - Dedup — two findings are duplicates when they share location
AND underlying issue. Merge with
Sources: [A, B]and tag as cross-confirmed (highest-confidence — multiple specialists independently reached the same conclusion). - Prioritise — Critical → High → Medium → Low → Info; within severity, cross-confirmed before single-source; within that, by file path.
- Tag for action class without changing the explicit panel class:
- Ship-blockers — Domain
blockerfindings; must address before merge. - Mechanical fixes — clear path, no judgement.
- Judgement calls — architectural trade-offs.
- Polish — optional.
- Ship-blockers — Domain
Final output structure
Render the report following the template in references/output-template.md. The prose rules above (don't merge across axes; synthesis within Domain only) govern; the template is the shape.
End every report with exactly one machine-readable line, with no prose after it:
PANEL: ship_blockers=<n> important=<n> advisory=<n> reviewer_failures=<n>
Count only explicit classifications. Each Spec, Standards, and Test-adequacy
finding counts once in its declared class. Domain duplicates are deduplicated
as above and each resulting Domain finding counts once by the exhaustive native
severity mapping; cross-axis findings remain distinct because the axes are
orthogonal. Never infer severity from category names or prose.
reviewer_failures counts agents that errored, timed out, or returned findings
without the required classification; a documented unavailable/not-applicable/
skipped axis is not a failure. Values are non-negative base-10 integers. This
line is the stable automation contract; prose is not.
Step 10 — Offer follow-up
Offer one short follow-up immediately before emitting the final PANEL: line
(the machine line remains the last line):
- "Apply mechanical fixes now?" → suggest /simplify
- "Drill into a specific finding?" → user names one
- "Post these as inline PR comments?" → suggest /pr-review-post if installed
- "Address the Spec misalignment first?" → user redirects work
- "Ship as-is" → user accepts the report
Don't loop on synthesis. The panel ran once; the report stands.
What this skill does NOT do
- Doesn't modify code. Synthesis is a report.
- Doesn't merge findings across axes — Spec, Standards, Test adequacy, and Domain are orthogonal by design; one masking another is the failure mode this skill exists to prevent.
- Doesn't fan out to every available Domain agent regardless of diff — irrelevant reviewers waste tokens and add noise. (The reuse pair is the exception: default-on for non-trivial code diffs, because over-engineering is the most common agent failure mode and the signal is rarely visible on the diff surface.)
- Doesn't duplicate the built-in
/code-reviewskill's single-pass logic. - Doesn't override the user's choice of agents.
- Doesn't auto-detect the fixed point — always pin it explicitly (asking if needed).
- Doesn't fabricate a spec if none exists. Skips the Spec axis with a noted reason instead.
- Doesn't re-check rules that project tooling (linters, formatters, type-checkers) already enforces.
- Doesn't fire on every chat message just because trigger keywords appear. Plausible code-change checkpoint required: a finished feature, a fixed bug, a landed refactor. Pure-question messages, planning conversations, and trivial one-line edits don't trigger. When ambiguous, ask the user before fanning out the panel.
Error handling
| Symptom | Action |
|---|---|
| No fixed point provided | Ask before proceeding; don't auto-detect. |
| No spec source found | Skip Spec axis; note in output. |
| No standards docs found | Standards subagent returns empty; note in output. |
| No executable behavior or tests in scope | Mark Test adequacy not applicable; do not count a reviewer failure. |
| No domain agents installed | Run Spec + Standards + Test adequacy; mark Domain unavailable, not failed. |
git diff returns nothing |
Ask for explicit scope (files / PR / branch). |
gh pr diff N / gh issue view N fails |
Tell the user gh isn't authenticated or the resource doesn't exist. |
| An agent times out / errors | Note the gap in synthesis; don't fail the whole panel. |
| Two domain agents conflict | Surface both views as a judgement-call entry. |
| Spec and Domain disagree | Present both; neither is wrong — they're asking different questions. |
| Diff exceeds practical size (>50 files / >3000 lines) | Ask the user whether to split or proceed. |
See also
- Reuse-ladder brief
- Output template
- Matt Pocock's
/review - Claude Code best practices
- Sub-agents reference
- Bundled
/code-review(single pass) and/code-review ultra(cloud panel)