Perform Adversarial Review
Try to falsify a proposed optimization experiment before it consumes GPU time. Review the proposal; do not generate a
second one, edit its draft, deploy it, or run AIPerf.
Inputs
Require:
EXP_ROOT and the current source iteration;
- exact
EXP_ROOT/user_workload.yaml path and SHA256 supplied by the parent;
- exact active benchmark-plan path, SHA256, and series ID returned by
perf-analyzer;
- current
DEPLOY_ROOT/deployment_ledger.json;
- current
DEPLOY_ROOT/applied_manifests/deploy.yaml;
- current
DEPLOY_ROOT/benchmark/benchmark_audit.json;
- current
DEPLOY_ROOT/benchmark/benchmark_summary.json;
- current
DEPLOY_ROOT/benchmark/performance_analysis.json;
- current
DEPLOY_ROOT/next-candidate/knowledge-consult.md;
- current
DEPLOY_ROOT/next-candidate/deploy-draft.yaml (proposal reviews only; a stop-request carries no draft
and its absence is not an objection);
- prior deployment and benchmark artifacts;
EXP_ROOT/analysis/search-calibration.md (and, for a stop-request, the submitted ledger SHA256 cited in
knowledge-consult.md); and
EXP_ROOT/analysis/hypothesis-backlog.jsonl and EXP_ROOT/analysis/challenger-reviews.jsonl when present; and
EXP_ROOT/manifest.yaml (session start time, for stop-request budget arithmetic); and
- for a stop-request: the three Finalize paths under
EXP_ROOT/final/ (recommended_config.md,
reproduced_commands.sh, known_limitations.md), derived from the submitted EXP_ROOT, whose on-disk
existence the validation gate verifies.
Review only a consultation whose decision is proposed and whose draft materialization completed successfully. For
no-proposal or blocked that carries no stop-request, return without writing a candidate verdict. For a
stop-request (a no-proposal consultation whose delta cites the search-calibration ledger path and a
submitted SHA256 — a blocked consultation never carries one), run the Stop-Request Validation below instead of
returning.
Stop-Request Validation
- Verify the SHA256 cited in
knowledge-consult.md matches the on-disk
EXP_ROOT/analysis/search-calibration.md. On mismatch, reject: the ledger moved after submission.
- Validate completeness and evidence class against the ledger, not the consult file (which carries only the
delta): every lever family carries a terminal disposition (
tested, ruled-out, not-applicable, or deferred — an answered ask resolves its family into one of these; untested-promising and reopened-by-new-evidence are non-terminal); every ruled-out row cites a measurement, a sourced hard constraint, a confirmed incompatibility,
or an explicit operator decision; every deferred row is terminal on a recorded ground - upside below the primary series' measured minimum
detectable effect (read from series_noise_floor and minimum_detectable_effect in the current
performance_analysis.json, the authoritative source per run-artifacts.md), or a cited cost-estimate-vs-remaining-budget arithmetic (reject when an above-MDE deferred
family lacks that arithmetic, or when its estimate fits the remaining budget; return that family as the
required follow-up); for a throughput-class objective the recommendation
carries saturation evidence or a recorded budget/operator reason in known_limitations.md ("still rising at
the top of the measured grid" is not terminal); and all three Finalize files EXIST ON DISK at EXP_ROOT/final/ — recommended_config.md (carrying its
required Correctness status: line), reproduced_commands.sh, and known_limitations.md — verified by
path, not by the submitter's claim (require those paths as inputs for stop-request validation; a
recommendation that exists only in conversation is a blocking objection).
- Verify the stop-request delta cites derived budget consumption (wall clock from
manifest.yaml session start;
failed-deploy count from the deployment ledgers; GPU-hours from GPU allocation time, per deployment ledger
allocated_at-to-torn_down_at span — or to now, for a live deployment — times its gpus_requested,
summed across deployments) and that the cited sources
support the arithmetic, whenever any granted budget is non-null.
- Append the verdict to
EXP_ROOT/analysis/challenger-reviews.jsonl as for any review, binding it to the
submitted ledger SHA256, and state in it that this is procedural validation, not independent adversarial
assurance. A validated stop-request returns to the PARENT with state STOP_REQUESTED for operator grant; it is
never a deployment handoff, and return_to is the parent, not recipe-deployer.
Read The Applicable Rules
Always read:
agent-docs/rules/benchmarking/evidence-eligibility.md;
agent-docs/rules/benchmarking/comparison-uncertainty.md;
agent-docs/rules/benchmarking/series-boundaries.md;
agent-docs/rules/optimization/evidence-before-spend.md;
agent-docs/rules/optimization/one-variable.md;
agent-docs/rules/verification/config-engagement.md;
agent-docs/rules/verification/implausible-speedup.md;
agent-docs/rules/verification/overlap.md;
agent-docs/rules/verification/stack-verdict.md;
agent-docs/rules/execution/user-workload.md (the resources.pinned and resources.gpu_ceiling semantics);
agent-docs/guides/knob-tuning/tuning-hierarchy.md; and
- the sources and repository guides cited for the proposal's central mechanism.
Read the Dynamo catalog for a Dynamo-owned knob, only the active engine guide for an engine-owned knob, the model-sizing
guides for a topology or memory-fit proposal, and the rate-matching guide for a disaggregated allocation proposal. Do
not invoke consult-perf-knowledge again or reconstruct a new shortlist. When relevant to the attached evidence, also
read the proxy-workload, concurrency-grid, or benchmark-isolation rule.
Establish Review Integrity
Before judging the idea:
- Require
benchmark_audit.json to report valid or valid_with_recovery.
- Confirm the plan, audit, summary, and analysis identify the same active series and every direct comparison is
same-series. Record material resource or placement differences as limitations.
- Recompute the source manifest, consultation, and draft SHA256 hashes.
- Require the source and draft hashes to match
Materialization Handoff.
- Parse both manifests and independently compute their semantic diff.
- Require the consultation, materialized diff, user workload, and performance analysis to identify the same model,
engine, deployment, objective, and source operating region.
- Require a concrete performance question, target operating region, and expected measurable effect for the candidate.
Treat a missing, stale, contradictory, or non-comparable input as a blocking objection. Do not repair it inside the
review.
Challenge The Proposal
Attack the proposal from these directions:
- Evidence: Does it contain at least three distinct qualifying categories, including AIPerf profiler data? Does
each source support the stated mechanism, or has contextual evidence been promoted beyond its limits?
- Uncertainty: Are changes at or below the measured noise floor of the active benchmark series classified as
noise (per
comparison-uncertainty.md)? For larger changes, does the conclusion match the
available single-run or multi-run evidence? Were confidence intervals used only when deliberate repetitions made
them useful, and is every requested repeat necessary enough to justify its GPU cost? Are degraded single-run
statistics or surprising gains treated cautiously?
- Redundancy: Has the same semantic configuration already been tested, rejected, or left inconclusive under the
same workload? If so, is there specific new evidence that makes this attempt different?
- Frontier: When the proposal selects or rejects a topology/config family, is the judgment made at each
family's own SLO frontier per
tuning-hierarchy.md? Reject a family selection argued from a single shared
operating point when the families' frontiers differ or the winner's frontier is unmeasured. A candidate whose
stated purpose is to MEASURE a family's not-yet-measured frontier is exploratory and admissible; what this
gate rejects is an adoption or rejection CLAIM resting on a frontier that has not been measured.
- Priority: Compared with the existing backlog, does this candidate have competitive information value, likely
impact on the primary objective, reversibility, GPU cost, and risk at the target operating region?
- Attribution: Does the complete diff express one independently testable knob? For a coupled bundle, is every field
required for one mechanism or supported by prior interaction evidence, with an ablation where needed?
- Provenance: When
deployment.origin is recipe-confirmed or agent-authored, reject any framing of the
baseline as a production reference; iteration 0 characterizes an unvalidated starting point, and topology
families inherited from it are open questions, not settled decisions.
- Mechanism: Does the proposed lever address a plausible reducible gap at the target operating region, or merely
move work that evidence suggests is already bounded? Are internal causes still labeled as hypotheses?
- Evaluation: Is the expected effect tied to the primary objective or failed SLO? Does the proposal state what the
candidate should teach us without prescribing favorable benchmark settings or assuming the current series must be
reused? Could another declared metric or target operating point regress enough to defeat the experiment's value?
- Feasibility: Is the knob valid for the active Dynamo and engine versions? Reject outright any candidate that
changes a knob listed in the contract's
resources.pinned or whose deployment would exceed
resources.gpu_ceiling; these are blocking objections regardless of evidence quality. Check GPU and replica
arithmetic, memory headroom, startup and OOM risk, topology consistency, and whether engagement can be proven
after deployment.
- Correctness: Does the draft preserve target-fixed model, framework, precision, hardware, and workload
constraints? If the change can alter output behavior, is there a concrete correctness check and rollback criterion?
- Spend: Is the expected information value worth the deployment and benchmark cost? Could a cheaper evidence
check resolve the uncertainty before GPU time is spent?
Use the attached consultation as the proposal's evidence boundary. Verify its claims, but do not reject it merely for
using concise prose or a flexible section layout.
Choose A Verdict
Return exactly one:
approve: no blocking objection remains; the existing draft is ready to enter deployment unchanged.
revise: the same experiment is worth testing after a small, explicit correction.
reject: the experiment is redundant, invalid, out of scope, unsafe, weakly supported, or unlikely to answer the
target performance question.
An approval cannot contain a blocking objection. For revise, give the smallest useful revision. Return every
revise or reject verdict to hypothesis-generator; the generator decides which of its skills to rerun.
Never edit knowledge-consult.md or deploy-draft.yaml during review.
Write The Review
Append one compact JSON object to:
<EXP_ROOT>/analysis/challenger-reviews.jsonl
Use this contract:
{
"review_id": "deploy-iter-<NNN>-<draft-sha256-prefix>",
"source_iteration": 0,
"candidate_iteration": 1,
"consult_path": "artifacts/deploy-iter-<NNN>/next-candidate/knowledge-consult.md",
"consult_sha256": "",
"candidate_path": "artifacts/deploy-iter-<NNN>/next-candidate/deploy-draft.yaml",
"candidate_sha256": "",
"verdict": "approve",
"return_to": "recipe-deployer",
"summary": "",
"objections": [
{
"severity": "blocking",
"check": "",
"finding": "",
"evidence": [],
"required_resolution": ""
}
],
"revised_experiment_plan": null,
"supersedes_review_id": null,
"reviewed_at": ""
}
Order objections by severity and impact. Cite exact iteration IDs, paths, hashes, metrics, or source files. For
approve, set objections to only non-blocking cautions or an empty list and identify the exact approved candidate
path and hash. Set return_to to recipe-deployer only for an approved PROPOSAL; for a validated stop-request set it to the
parent (operator grant); otherwise set it to hypothesis-generator.
For revise, make revised_experiment_plan concise. Do not turn a rejection into an unrelated replacement hypothesis.
A stop-request validation returns stop-validated or stop-rejected instead of the proposal verdicts, and its
review ID binds to the submitted ledger SHA256 prefix (deploy-iter-<NNN>-stop-<ledger-sha256-prefix>).
Use a stable review ID bound to the draft hash. If that exact draft already has a review, return the existing record
instead of appending a duplicate. A revised draft receives a new review ID and names the prior record in
supersedes_review_id.
Return
For approve (proposals only), return the review ID, exact candidate path and SHA256, performance question, and target operating region
to the parent and recipe-deployer. For stop-validated or stop-rejected, return the verdict and review ID to
the parent only (STOP_REQUESTED awaits operator grant; a rejection names the required follow-up families). The parent must carry the question and operating region into the candidate's
perf-analyzer assignment. For revise or reject, return the verdict, strongest objections, and any minimal revision
or required follow-up to hypothesis-generator. Do not create the next DEPLOY_ROOT for candidate iteration
<NNN + 1>.
1---2name: perform-adversarial-review3description: Adversarially reviews an evidence-backed Dynamo optimization proposal and DGD draft for comparability, duplication, attribution, correctness, feasibility, and worthwhile GPU spend. Use after hypothesis-generator writes knowledge-consult.md and deploy-draft.yaml and before recipe-deployer creates the next deployment iteration.4license: Apache-2.05---67# Perform Adversarial Review89<!--10SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.11SPDX-License-Identifier: Apache-2.012-->1314Try to falsify a proposed optimization experiment before it consumes GPU time. Review the proposal; do not generate a15second one, edit its draft, deploy it, or run AIPerf.1617## Inputs1819Require:2021- `EXP_ROOT` and the current source iteration;22- exact `EXP_ROOT/user_workload.yaml` path and SHA256 supplied by the parent;23- exact active benchmark-plan path, SHA256, and series ID returned by `perf-analyzer`;24- current `DEPLOY_ROOT/deployment_ledger.json`;25- current `DEPLOY_ROOT/applied_manifests/deploy.yaml`;26- current `DEPLOY_ROOT/benchmark/benchmark_audit.json`;27- current `DEPLOY_ROOT/benchmark/benchmark_summary.json`;28- current `DEPLOY_ROOT/benchmark/performance_analysis.json`;29- current `DEPLOY_ROOT/next-candidate/knowledge-consult.md`;30- current `DEPLOY_ROOT/next-candidate/deploy-draft.yaml` (proposal reviews only; a stop-request carries no draft31 and its absence is not an objection);32- prior deployment and benchmark artifacts;33- `EXP_ROOT/analysis/search-calibration.md` (and, for a stop-request, the submitted ledger SHA256 cited in34 `knowledge-consult.md`); and35- `EXP_ROOT/analysis/hypothesis-backlog.jsonl` and `EXP_ROOT/analysis/challenger-reviews.jsonl` when present; and36- `EXP_ROOT/manifest.yaml` (session start time, for stop-request budget arithmetic); and37- for a stop-request: the three Finalize paths under `EXP_ROOT/final/` (`recommended_config.md`,38 `reproduced_commands.sh`, `known_limitations.md`), derived from the submitted `EXP_ROOT`, whose on-disk39 existence the validation gate verifies.4041Review only a consultation whose decision is `proposed` and whose draft materialization completed successfully. For42`no-proposal` or `blocked` that carries no stop-request, return without writing a candidate verdict. For a43stop-request (a `no-proposal` consultation whose delta cites the search-calibration ledger path and a44submitted SHA256 — a `blocked` consultation never carries one), run the Stop-Request Validation below instead of45returning.4647## Stop-Request Validation48491. Verify the SHA256 cited in `knowledge-consult.md` matches the on-disk50 `EXP_ROOT/analysis/search-calibration.md`. On mismatch, reject: the ledger moved after submission.512. Validate completeness and evidence class against the ledger, not the consult file (which carries only the52 delta): every lever family carries a terminal disposition (`tested`, `ruled-out`, `not-applicable`, or `deferred` — an answered ask resolves its family into one of these; `untested-promising` and `reopened-by-new-evidence` are non-terminal); every `ruled-out` row cites a measurement, a sourced hard constraint, a confirmed incompatibility,53 or an explicit operator decision; every `deferred` row is terminal on a recorded ground - upside below the primary series' measured minimum54 detectable effect (read from `series_noise_floor` and `minimum_detectable_effect` in the current55 `performance_analysis.json`, the authoritative source per `run-artifacts.md`), or a cited cost-estimate-vs-remaining-budget arithmetic (reject when an above-MDE deferred56 family lacks that arithmetic, or when its estimate fits the remaining budget; return that family as the57 required follow-up); for a throughput-class objective the recommendation58 carries saturation evidence or a recorded budget/operator reason in `known_limitations.md` ("still rising at59 the top of the measured grid" is not terminal); and all three Finalize files EXIST ON DISK at `EXP_ROOT/final/` — `recommended_config.md` (carrying its60 required `Correctness status:` line), `reproduced_commands.sh`, and `known_limitations.md` — verified by61 path, not by the submitter's claim (require those paths as inputs for stop-request validation; a62 recommendation that exists only in conversation is a blocking objection).633. Verify the stop-request delta cites derived budget consumption (wall clock from `manifest.yaml` session start;64 failed-deploy count from the deployment ledgers; GPU-hours from GPU allocation time, per deployment ledger65 `allocated_at`-to-`torn_down_at` span — or to now, for a live deployment — times its `gpus_requested`,66 summed across deployments) and that the cited sources67 support the arithmetic, whenever any granted budget is non-null.684. Append the verdict to `EXP_ROOT/analysis/challenger-reviews.jsonl` as for any review, binding it to the69 submitted ledger SHA256, and state in it that this is procedural validation, not independent adversarial70 assurance. A validated stop-request returns to the PARENT with state `STOP_REQUESTED` for operator grant; it is71 never a deployment handoff, and `return_to` is the parent, not `recipe-deployer`.7273## Read The Applicable Rules7475Always read:7677- `agent-docs/rules/benchmarking/evidence-eligibility.md`;78- `agent-docs/rules/benchmarking/comparison-uncertainty.md`;79- `agent-docs/rules/benchmarking/series-boundaries.md`;80- `agent-docs/rules/optimization/evidence-before-spend.md`;81- `agent-docs/rules/optimization/one-variable.md`;82- `agent-docs/rules/verification/config-engagement.md`;83- `agent-docs/rules/verification/implausible-speedup.md`;84- `agent-docs/rules/verification/overlap.md`;85- `agent-docs/rules/verification/stack-verdict.md`;86- `agent-docs/rules/execution/user-workload.md` (the `resources.pinned` and `resources.gpu_ceiling` semantics);87- `agent-docs/guides/knob-tuning/tuning-hierarchy.md`; and88- the sources and repository guides cited for the proposal's central mechanism.8990Read the Dynamo catalog for a Dynamo-owned knob, only the active engine guide for an engine-owned knob, the model-sizing91guides for a topology or memory-fit proposal, and the rate-matching guide for a disaggregated allocation proposal. Do92not invoke `consult-perf-knowledge` again or reconstruct a new shortlist. When relevant to the attached evidence, also93read the proxy-workload, concurrency-grid, or benchmark-isolation rule.9495## Establish Review Integrity9697Before judging the idea:98991. Require `benchmark_audit.json` to report `valid` or `valid_with_recovery`.1002. Confirm the plan, audit, summary, and analysis identify the same active series and every direct comparison is101 same-series. Record material resource or placement differences as limitations.1023. Recompute the source manifest, consultation, and draft SHA256 hashes.1034. Require the source and draft hashes to match `Materialization Handoff`.1045. Parse both manifests and independently compute their semantic diff.1056. Require the consultation, materialized diff, user workload, and performance analysis to identify the same model,106 engine, deployment, objective, and source operating region.1077. Require a concrete performance question, target operating region, and expected measurable effect for the candidate.108109Treat a missing, stale, contradictory, or non-comparable input as a blocking objection. Do not repair it inside the110review.111112## Challenge The Proposal113114Attack the proposal from these directions:115116- **Evidence**: Does it contain at least three distinct qualifying categories, including AIPerf profiler data? Does117 each source support the stated mechanism, or has contextual evidence been promoted beyond its limits?118- **Uncertainty**: Are changes at or below the measured noise floor of the active benchmark series classified as119 noise (per `comparison-uncertainty.md`)? For larger changes, does the conclusion match the120 available single-run or multi-run evidence? Were confidence intervals used only when deliberate repetitions made121 them useful, and is every requested repeat necessary enough to justify its GPU cost? Are degraded single-run122 statistics or surprising gains treated cautiously?123- **Redundancy**: Has the same semantic configuration already been tested, rejected, or left inconclusive under the124 same workload? If so, is there specific new evidence that makes this attempt different?125- **Frontier**: When the proposal selects or rejects a topology/config family, is the judgment made at each126 family's own SLO frontier per `tuning-hierarchy.md`? Reject a family selection argued from a single shared127 operating point when the families' frontiers differ or the winner's frontier is unmeasured. A candidate whose128 stated purpose is to MEASURE a family's not-yet-measured frontier is exploratory and admissible; what this129 gate rejects is an adoption or rejection CLAIM resting on a frontier that has not been measured.130- **Priority**: Compared with the existing backlog, does this candidate have competitive information value, likely131 impact on the primary objective, reversibility, GPU cost, and risk at the target operating region?132- **Attribution**: Does the complete diff express one independently testable knob? For a coupled bundle, is every field133 required for one mechanism or supported by prior interaction evidence, with an ablation where needed?134- **Provenance**: When `deployment.origin` is `recipe-confirmed` or `agent-authored`, reject any framing of the135 baseline as a production reference; iteration 0 characterizes an unvalidated starting point, and topology136 families inherited from it are open questions, not settled decisions.137- **Mechanism**: Does the proposed lever address a plausible reducible gap at the target operating region, or merely138 move work that evidence suggests is already bounded? Are internal causes still labeled as hypotheses?139- **Evaluation**: Is the expected effect tied to the primary objective or failed SLO? Does the proposal state what the140 candidate should teach us without prescribing favorable benchmark settings or assuming the current series must be141 reused? Could another declared metric or target operating point regress enough to defeat the experiment's value?142- **Feasibility**: Is the knob valid for the active Dynamo and engine versions? Reject outright any candidate that143 changes a knob listed in the contract's `resources.pinned` or whose deployment would exceed144 `resources.gpu_ceiling`; these are blocking objections regardless of evidence quality. Check GPU and replica145 arithmetic, memory headroom, startup and OOM risk, topology consistency, and whether engagement can be proven146 after deployment.147- **Correctness**: Does the draft preserve target-fixed model, framework, precision, hardware, and workload148 constraints? If the change can alter output behavior, is there a concrete correctness check and rollback criterion?149- **Spend**: Is the expected information value worth the deployment and benchmark cost? Could a cheaper evidence150 check resolve the uncertainty before GPU time is spent?151152Use the attached consultation as the proposal's evidence boundary. Verify its claims, but do not reject it merely for153using concise prose or a flexible section layout.154155## Choose A Verdict156157Return exactly one:158159- `approve`: no blocking objection remains; the existing draft is ready to enter deployment unchanged.160- `revise`: the same experiment is worth testing after a small, explicit correction.161- `reject`: the experiment is redundant, invalid, out of scope, unsafe, weakly supported, or unlikely to answer the162 target performance question.163164An approval cannot contain a blocking objection. For `revise`, give the smallest useful revision. Return every165`revise` or `reject` verdict to `hypothesis-generator`; the generator decides which of its skills to rerun.166Never edit `knowledge-consult.md` or `deploy-draft.yaml` during review.167168## Write The Review169170Append one compact JSON object to:171172```text173<EXP_ROOT>/analysis/challenger-reviews.jsonl174```175176Use this contract:177178```json179{180 "review_id": "deploy-iter-<NNN>-<draft-sha256-prefix>",181 "source_iteration": 0,182 "candidate_iteration": 1,183 "consult_path": "artifacts/deploy-iter-<NNN>/next-candidate/knowledge-consult.md",184 "consult_sha256": "",185 "candidate_path": "artifacts/deploy-iter-<NNN>/next-candidate/deploy-draft.yaml",186 "candidate_sha256": "",187 "verdict": "approve",188 "return_to": "recipe-deployer",189 "summary": "",190 "objections": [191 {192 "severity": "blocking",193 "check": "",194 "finding": "",195 "evidence": [],196 "required_resolution": ""197 }198 ],199 "revised_experiment_plan": null,200 "supersedes_review_id": null,201 "reviewed_at": ""202}203```204205Order objections by severity and impact. Cite exact iteration IDs, paths, hashes, metrics, or source files. For206`approve`, set `objections` to only non-blocking cautions or an empty list and identify the exact approved candidate207path and hash. Set `return_to` to `recipe-deployer` only for an approved PROPOSAL; for a validated stop-request set it to the208parent (operator grant); otherwise set it to `hypothesis-generator`.209For `revise`, make `revised_experiment_plan` concise. Do not turn a rejection into an unrelated replacement hypothesis.210211A stop-request validation returns `stop-validated` or `stop-rejected` instead of the proposal verdicts, and its212review ID binds to the submitted ledger SHA256 prefix (`deploy-iter-<NNN>-stop-<ledger-sha256-prefix>`).213214Use a stable review ID bound to the draft hash. If that exact draft already has a review, return the existing record215instead of appending a duplicate. A revised draft receives a new review ID and names the prior record in216`supersedes_review_id`.217218## Return219220For `approve` (proposals only), return the review ID, exact candidate path and SHA256, performance question, and target operating region221to the parent and `recipe-deployer`. For `stop-validated` or `stop-rejected`, return the verdict and review ID to222the parent only (`STOP_REQUESTED` awaits operator grant; a rejection names the required follow-up families). The parent must carry the question and operating region into the candidate's223`perf-analyzer` assignment. For `revise` or `reject`, return the verdict, strongest objections, and any minimal revision224or required follow-up to `hypothesis-generator`. Do not create the next `DEPLOY_ROOT` for candidate iteration225`<NNN + 1>`.