Reproducibility packaging
You take a completed analysis and produce the artifact a journal, a data repository, and a skeptical replicator all need: a self-contained archive that regenerates every table and figure from raw data, in one command, on a clean machine — plus the documentation that makes it usable.
This is the final packaging step. The per-stage handoff audits inside
r-analyst / stata-analyst / text-analyst (their handoff_audit technique)
are what make this step cheap: if each stage was already reproducible, you're
assembling and verifying, not rescuing. Do not duplicate those audits here — build
on them.
Project Integration
Reads project.yaml paths; writes the archive under a replication/ (or
submissions/) directory. Updates progress.yaml:
status:
replication_check: done
checks:
results_reproducible: true
artifacts:
replication_archive: replication/
Core Principles
- From raw, not from clean. The archive must rebuild starting from the raw
inputs (or a documented data-access step), not from a mid-pipeline
.rds/.dtayou happen to have. If a step can't run from raw, it isn't reproducible yet. - One command. A single master script runs every stage in order. If a human has to "also run this by hand," document it or script it — don't leave it implicit.
- Respect data licenses. Never bundle restricted microdata (GSS/ANES weights
aside, ICPSR-restricted, DUA data) into a shareable archive. Ship a data-access
statement instead — the DOI/version and how to obtain it — drawn from the
PROVENANCE.mdthedata-acquisitionskill wrote. This is both a legal and an ethical line. - Every exhibit is traceable. Each table and figure in the paper maps to the
script (and ideally line) that produces it. A reviewer should never wonder where
Table 3 came from. Every number in the report must likewise trace to the
results ledger — run the
stat-checkreconciliation and confirm it's clean (no orphans) before archiving the write-up. - Capture the environment. Record software versions and packages so the code still runs in two years. Reproducibility that depends on "whatever was installed that day" isn't reproducibility.
Workflow
Phase 1 — Clean-room rebuild (prove it works)
Run the whole pipeline from raw in a fresh session and confirm every output regenerates. Use the analysis skill's final reproducibility check as the engine.
- R:
Rscript replication/code/00_run_all.Rfrom a clean session; ideally underrenv::restore()so package versions are locked. - Stata:
stata -b do replication/code/00_run_all.do(Mac/Linux) orStataMP-64.exe /e do ...(Windows), starting each stage withclear all. - Python:
uv run replication/code/00_run_all.py(uv resolves the pinned deps).
Then confirm the regenerated tables/figures match the ones in the paper. Compare values, not bytes (PDF/PNG timestamps differ) — check the numbers and the figure content. Any mismatch is a blocker: fix it (or the paper) before packaging. If a stage is slow (bootstraps, big text models), note expected runtime; don't silently skip it.
If this fails, stop and route back to the analysis skill's handoff audit — the package can't be built on an analysis that doesn't reproduce.
Phase 2 — Assemble the archive
Lay out a clean, self-contained tree (see templates/README_archive.md for the
README to drop in):
replication/
README.md # how to run, expected runtime, software versions, exhibit map
data/
raw/ # raw data IF licensing permits; else a DATA_ACCESS.md
DATA_ACCESS.md # DOI/version + how to obtain any data not bundled
code/
00_run_all.(R|do|py) # master: runs every stage from raw, in order
01_... # stages
output/
tables/ figures/ # regenerated exhibits
codebook/ # variable definitions (from the source + any constructed vars)
environment/ # version + package snapshot (see below)
Assemble each piece:
- Master runner (
00_run_all.*) that sources/does/imports every stage in order from raw to final output. No manual steps between stages. - README from the template: one-command run instructions, per-stage description, expected total runtime, software + version, and the output→exhibit map.
- Codebook for both source variables and every constructed variable (the recodes live in the cleaning scripts — summarize them here).
- DATA_ACCESS.md for anything not bundled: pull the source, DOI, version, and
retrieval instructions straight from
data/raw/<dataset>/PROVENANCE.md. State clearly which files a replicator must obtain themselves and where they go. - Environment snapshot:
- R:
renv::snapshot()(commitrenv.lock), or at minimum savesessionInfo()/sessioninfo::session_info()toenvironment/R_session.txt. - Stata: record the Stata version (
about) and installed user commands with their versions (which reghdfe,which esttab, …) toenvironment/stata_env.txt; note theversion XXline each do-file uses. - Python: ship the
uvPEP-723 headers / lockfile; also writeuv pip freeze(orpip freeze) andpython --versiontoenvironment/py_env.txt.
- R:
Phase 3 — Verify the archive is self-contained
The archive must reproduce on its own, not because the rest of your project is lying around.
- Copy
replication/to a fresh location (outside the project) and run the master script there. It must find every input via relative paths and regenerate outputs. - Grep the code for absolute paths (
/Users/,C:\\,~/), leftover interactive state, and reads from files outside the archive — fix any hit. - Confirm no restricted data slipped in (check against the licenses noted in
PROVENANCE.md); confirm large/restricted files are excluded and documented in DATA_ACCESS.md instead. - Build a
MANIFESTwith a checksum per file so the archive's integrity is verifiable.
Output — handoff note
Report to the user:
- Reproduced from raw in a clean session: yes / the blocker.
- Every paper exhibit maps to a script (list any that don't).
- What's bundled vs. what a replicator must obtain (the data-access story).
- Environment captured (which lockfile/version file).
- Archive verified self-contained from a fresh location.
- Where the archive is and what to deposit (Dataverse/OSF/journal system).
Depositing on Dataverse? The archive you built here is exactly what goes up; cite the resulting DOI in the paper.