EXECUTE NOW
Target: $ARGUMENTS
Parse immediately:
- If target is a run_id: synthesize that specific run
- If target is empty: list recent runs without experiment notes and ask which to synthesize
- If target is
--all: synthesize all runs that lack experiment notes - If target is
--batch N: synthesize the N most recent unsynthesized runs
Execute steps 0 through 6 in order. Do not ask for confirmation between steps.
Step 0: Validate Target
- Check that
runs/{run_id}/exists - Check that
runs/{run_id}/run.yamlexists — if not, runpython scripts/backfill-run-yaml.py --run-id {run_id}to generate it - Check if
vault/experiments/{run_id}.mdalready exists — if so, report "Already synthesized" and stop (unless--force)
Step 1: Read Run Metadata
Read runs/{run_id}/run.yaml and extract:
slug,created_utcexperiment.type,experiment.hypothesis,experiment.swept_parameters,experiment.seeds,experiment.total_runsresults.status,results.primary_metric,results.primary_result,results.significant_findingsartifacts.*— what files are availabletagslinks.claims— any pre-linked claims
If experiment.type is sweep or study, also read the summary JSON file referenced in artifacts.summary.
Step 2: Read Summary Data
Based on experiment.type:
For sweep: Read summary.json. Extract:
total_runs,total_hypothesesn_bonferroni_significantorbonferroni_survivorsswept_parameters- Top 3 most significant results (by effect size)
p_hacking_audit
For redteam: Read report.json. Extract:
robustness_score,gradeattacks_tested,attacks_successfulvulnerabilities(critical and high severity)
For study: Read analysis/summary.json or summary.json. Extract:
descriptivestatistics per conditionpairwise_tests— significant comparisonsbonferroni_survivors
For single: Read history.json. Extract:
- Final epoch metrics: welfare, toxicity, acceptance rate
- Agent-type breakdown if available
For calibration: Read summary.json and recommendation.json. Extract:
- Parameter values tested
- Recommended configuration
Step 3: Generate Experiment Note
Create vault/experiments/{run_id}.md using the template at vault/templates/experiment-note.md.
Critical rules:
- description must be ≤200 chars, must add info beyond the title, no trailing period
- Title (H1) must be a prose proposition describing the experiment's purpose and outcome
- Claims affected section: scan
vault/claims/for claims whose topics overlap with this run's tags (≥2 tag overlap). List them with relationship context. - Key results section: include effect sizes, p-values, and correction methods. Use connective words.
- Reproduction section: construct the CLI command from run.yaml metadata
Quality gate: Can you complete "This experiment showed that [title]"?
Step 4: Check Claims for New Evidence
For each claim file in vault/claims/:
- Read the claim's frontmatter (
domain,evidence, topics footer) - Check tag overlap between the claim's topics and the run's tags
- If overlap ≥ 2 tags:
a. Read the run's summary data
b. Determine if the run's results support or weaken the claim
c. If the run provides new evidence not already listed in
evidence.supportingorevidence.weakening:- Report the finding but do not auto-edit the claim
- Instead, output a structured recommendation:
### Claim update recommendation: {claim_id}
**Action**: Add supporting evidence / Add weakening evidence / Add boundary condition
**Run**: {run_id}
**Detail**: {effect size, p-value, correction}
**Suggested edit**: Add to evidence.supporting: ...
Critical: Never automatically modify claim files. Claims require human judgment. Only recommend updates.
Step 5: Rebuild Index
- Run
python scripts/index-runs.pyto updaterun-index.yaml - Regenerate
vault/_index.md:- Re-scan all vault subdirectories for notes
- Update counts and listings
- Preserve any manually-added content in the index
Step 6: Report
Print a structured summary:
## Synthesis Complete: {run_id}
### Experiment note
✓ vault/experiments/{run_id}.md
Type: {experiment_type}
Key finding: {one-line summary}
### Claims checked
- {N} claims scanned for tag overlap
- {M} claims have overlapping evidence:
{list of claim_id + recommended action}
### Index
✓ run-index.yaml rebuilt ({total} runs)
✓ vault/_index.md updated
### Next steps
- Review claim update recommendations above
- Edit claim cards manually if evidence warrants it
- Run /verify to check vault health
Batch Mode (--all or --batch N)
When processing multiple runs:
- Glob
runs/*/run.yamlto find all runs - Glob
vault/experiments/*.mdto find already-synthesized runs - Compute the unsynthesized set
- Sort by date (newest first)
- If
--batch N, take only the first N - For each run, execute Steps 1–4 (skip Step 5 until all runs are done)
- Run Step 5 once at the end
- Print a batch summary:
## Batch Synthesis Complete
- {N} runs synthesized
- {M} claim update recommendations
- {K} runs skipped (already synthesized)
Critical Constraints
Never auto-edit claims. Claims require human judgment. Only recommend updates.
description field is mandatory on every generated note. It must add info beyond the title. Max 200 chars. No trailing period.
Topics footer is mandatory on every generated note. Link to
[[_index]]at minimum.Prose-as-title convention. The H1 heading is a complete proposition, not a label.
Effect sizes and correction methods must be included in Key Results. Do not report raw p-values without specifying the correction method (Bonferroni, Holm, BH).
Wiki-links inline as prose. Example: "This replicates the finding from [[baseline governance v2 sweep]] with finer granularity."
Idempotent. Running /synthesize twice on the same run should produce no changes (unless --force).