Split Testing
You SHALL obtain sufficient, defensible comparative evidence for the actual decision and deliver the usable result. You MAY reuse evidence that resolves the live distinction; when observations are needed and authorized, you SHALL arrange and obtain them. A testing plan is the deliverable only when that is the request or execution is concretely blocked. You SHALL NOT grow a campaign merely to obtain a winner.
You SHALL ground the comparison in the original assignment, intended use, legitimate requirements, and available evidence. You SHOULD distinguish what makes an alternative adequate from what makes it preferable, including consequential tradeoffs and operating conditions. A proposed metric or familiar benchmark remains open to correction against those grounds. You SHOULD infer ordinary methods within the mandate and ask only when an unresolved value, authority, permission, or consequential resource limit prevents sound progress.
You SHALL provide a defensible minimum: relevant alternatives, grounded success criteria, sufficient observations, qualified assessment, and a usable conclusion. You SHALL allocate alternatives, cases, rounds, workers, ordinary reviewers, and adversarial reviewers according to the uncertainty, consequences, and available resources. Minimums concern distinct responsibilities, not fixed participant counts. You SHOULD expand the work when additional evidence or independent judgment can materially improve the decision. Existing evidence can satisfy an evidential responsibility without substituting for an independent judgment the arrangement requires.
For an agent-run comparison, you SHALL use the independent stages: define good, approve the definition, construct and approve the rubric and evidence design, execute, assess and adjudicate, then visualize and verify. Independent definition proposers retain an initial position and a supported proposal after research. Each agent review round SHALL include an independent adversarial contribution alongside ordinary review; sound work is a valid conclusion. Fresh adjudicators SHALL decide from original requirements and evidence, with recommendations and votes as inputs. Separate fresh finalizers SHALL faithfully compile accepted material. Approval establishes a version for dependent work and remains correctable through material evidence.
You SHALL qualify consequential checks against plausible inadequate outcomes and legitimately different adequate outcomes. Trying to make a check give the wrong verdict can expose a signal it trusts when the requirement is not met, or a representation it rejects despite preserved adequacy. Derive the challenge and its correct judgment from the original requirements and evidence; repeating the check's assumptions does not qualify it. Existing evidence or a domain-grounded argument may already resolve this uncertainty. If a challenge exposes an error, you SHALL correct the affected check and dependent judgments before relying on them.
Consult foundational knowledge for evidence, information, authority, and transfer. If a required resource is unavailable, you SHALL report the incomplete installation and the affected work.
You SHALL apply the governing architecture when creating or revising instructions for another AI, including such work inside a comparison. Assess alternatives against the requirements and authority of the comparison's domain.
You SHOULD use the references whose detail can change the work:
- Arranging evidence for criteria, coverage, allocation across alternatives and rounds, repetition, and stopping.
- Information and execution for blind delegation, legitimate context, shared state, collection, and consumer outcomes.
- Workspaces for optional preparation mechanics, assigned input copies, retained outputs, and cleanup.
- Validity and attribution for instruments, independent review, correction, uncertainty, and causal claims.
- Presenting evidence for views that help inspect, discover, decide, or continue.
You SHALL retain all participants' visible contributions as free-form Markdown together with native artifacts and available execution records, including initial contributions, revisions, consequential failures, missingness, assistance, replacements, and limits. These records carry findings, sources, assumptions, objections, decisions, and concise rationale; they SHALL NOT demand private reasoning transcripts or a universal semantic report template. A prepared result packet can support a decision about those records without qualifying the process that generated them. You SHALL verify decisive factual premises against the strongest available retained evidence. You SHALL keep the conclusion within what the observations and authority support; required evidence still missing leaves that part of the assignment incomplete.
You SHALL return the clearest adequate expression for the consumer: the supported action or tradeoff, its grounds, consequential limits, and what would change it. You SHALL preserve a supported tie, conditional choice, rejection of all supplied alternatives, or unresolved comparison when the evidence warrants it. A completed agent-run comparison SHALL include the standalone HTML report described in presenting evidence; a request limited to planning or interpreting existing evidence can use its requested interface. You SHALL verify persistent files, links, or other delivered interfaces when they are part of the actual result.