Write-up SOP
This is an agent-gated publication workflow, not a filesystem security boundary. Hooks block only a narrower set of obvious unsafe writes.
Inputs and final deliverables
Require the target audience/venue, output path and format, intended claims, linked hypothesis/proposition ids, metric pins, and whether each result is confirmatory or exploratory. Deliver the manuscript/report, a claim manifest, reviewer JSON, and a short list of unresolved limitations.
1. Build the claim manifest before drafting
List every meaningful claim with:
- exact intended wording;
- kind:
result_metric,statistical_claim,context, ortheorem; - role: central, supporting, or background;
- mode: confirmatory, exploratory, or not applicable;
- linked hypothesis/proposition id;
- required evidence and current status.
Dates, versions, seed counts, baseline counts, model sizes, and timeouts are usually context. They must be accurate but are not automatic provenance gates.
2. Close the empirical evidence chain
For each publication-critical numeric or statistical claim:
- Call
mcp__verify__check_provenance. Missing provenance means rerun, remove the claim, or clearly downgrade it to exploratory. - Require a real
pin_idfor central metrics. - Call
mcp__verify__refresh_claim. Any stale code, data, config, Git state, dependency lock, runtime, or tracked environment blocks the central claim. Legacyuncheckedevidence must be disclosed and is insufficient as the only support for a headline result. - Require a current stable seed verdict for central experimental metrics, or narrow the wording to an unstable/exploratory observation.
- For method-versus-baseline claims, require a fair
baseline_fairnessverdict or disclose the resource mismatch next to the comparison. - For confirmatory claims, require the matching preregistration to be
met. Check its fixedfamily_idandfamily_size; an open or missed row blocks confirmatory wording.
Never convert an observed exploratory run into a confirmatory claim after the fact.
3. Describe BT rankings honestly
When rankings matter, report strength, comparison count, and uncertainty only
as a joint batch MAP Bradley-Terry result. The compatibility fields lcb and
ucb are uncalibrated approximate posterior intervals. Do not call them strict
95% confidence intervals, and do not use overlap/non-overlap as proof that one
hypothesis is truly superior.
4. Close theorem and proof claims
For every theorem, lemma, proposition, corollary, or "we prove" statement:
- link a real proposition node;
- find the latest proof draft and diagnostic manifest;
- require the final manifest to be
empty; anopenmanifest blocks writing; - cite a verified Lean attempt when available;
- when Lean verification is absent, add an explicit
unverifiedannotation and state what evidence the natural-language proof has received; - refresh any provenance-backed references used in the proof.
Use $prove-sop to repair missing proof evidence before continuing.
5. Draft with evidence-local wording
Write each central claim close to its scope, metric definition, uncertainty, dataset, baseline conditions, and limitation. Keep exploratory language visibly distinct from confirmatory language. Do not bury failed preregistration, unstable seeds, stale evidence, or resource imbalance in an appendix.
A useful report order is:
- question and contribution;
- related evidence and gap;
- method and preregistered/exploratory status;
- experiment and run-manifest details;
- results as findings, not table narration;
- verification, failures, and robustness;
- proof evidence when applicable;
- limitations and conclusion.
6. Run the adversarial reviewer
Use the reviewer role when available; otherwise apply its checklist inline. The required JSON keys are:
verdict:accept,revise, orreject;numeric_claims;theorem_claims;provenance_trace;blockers;notes.
accept is forbidden when blockers is non-empty, a central numeric claim
lacks a pin, central evidence is stale, a confirmatory preregistration is not
met, a theorem manifest is open, or an unformalized theorem lacks an explicit
unverified flag.
7. Respond to the verdict
accept: write/finalize the requested artifact and optionally export a closure report.revise: address every blocker, refresh the claim manifest, and rerun the reviewer. Do not silently weaken checks.reject: do not publish through this workflow. Return the blocking evidence and the minimum rerun/removal needed for reconsideration.
Completion criteria
The write-up is complete only when every central claim maps to fresh evidence, exploratory and confirmatory wording is accurate, statistical uncertainty is described with its real calibration status, theorem claims pass the proof branch or are explicitly unverified, reviewer JSON is complete, and the final artifact contains no unresolved blocker disguised as prose.