AISTATS Reproducibility
Use this before submission and again before camera-ready. Reopen the current CFP and
OpenReview forms to confirm whether a reproducibility checklist is required.
Evidence map
- Map each theorem, algorithmic claim, simulation claim, and empirical claim to a verifiable
location in the paper, appendix, supplement, or artifact package.
- For theory, state assumptions, proof dependencies, convergence conditions, constants, and
failure modes clearly enough for statistical readers.
- For experiments, report datasets, splits, preprocessing, evaluation metrics, baselines,
hyperparameter ranges, final selected settings, seeds, repeated runs, compute, and runtime.
- For small performance differences, add uncertainty estimates: standard errors, confidence
intervals, paired tests, bootstrap intervals, or repeated trials as appropriate.
- Explain missing code/data honestly and describe how a reader could reproduce the analysis
in principle.
- Keep the checklist consistent with the manuscript; contradictions between checklist and
paper are review-risk multipliers.
Checklist-to-claim audit table
| Checklist item |
Pure-theory answer |
Theory-plus-experiments answer |
| Code availability |
NA only if there is literally no computation |
Anonymous archive, or an honest stated reason |
| Assumptions stated |
Every theorem lists its conditions inline |
Plus a note on which experiments satisfy them |
| Error bars |
NA for deterministic results |
Required for every stochastic figure and table |
| Compute resources |
NA |
Hardware, runtime, and total number of runs |
Marking NA on an item the paper actually triggers is a recognizable AISTATS red flag,
because reviewers cross-check checklist answers against the PDF and read contradictions as
carelessness about the rest of the paper.
Vignette: a rates-plus-simulation paper
Consider a submission proving posterior contraction rates for a Bayesian nonparametric model,
validated by MCMC simulation. Its reproducibility spine: prior hyperparameters and their
selection rule, chain length, burn-in, convergence diagnostics, replication seeds, and a
statement of which contraction-theorem conditions the simulated model satisfies — plus one
honest sentence about the condition it does not.
Degrees of reproducibility
- Turnkey: one command regenerates each figure from logged seeds.
- Scripted: scripts exist but require documented manual steps or external data access.
- Descriptive: prose detailed enough that a competent reader could rebuild the pipeline.
For AISTATS, simulations should be turnkey because statistician reviewers actually rerun
them; large real-data pipelines may stay scripted with deviations documented. Stating the
achieved level honestly beats overpromising turnkey behavior that fails on a clean machine.
Output format
[Claim inventory] <claim -> evidence location>
[Checklist status] complete / inconsistent / missing
[Statistical reproducibility gaps] <assumptions/seeds/uncertainty/hyperparameters/compute>
[Paper fixes] <must appear in main PDF>
[Supplement fixes] <appendix or artifact additions>
1---2name: aistats-reproducibility3description: Use when strengthening AISTATS reproducibility evidence, including the official reproducibility checklist, statistical assumptions, proofs, datasets, hyperparameters, random seeds, compute, uncertainty estimates, baselines, code/data release statements, and checklist-to-claim consistency audits.4---56# AISTATS Reproducibility78Use this before submission and again before camera-ready. Reopen the current CFP and9OpenReview forms to confirm whether a reproducibility checklist is required.1011## Evidence map1213- Map each theorem, algorithmic claim, simulation claim, and empirical claim to a verifiable14 location in the paper, appendix, supplement, or artifact package.15- For theory, state assumptions, proof dependencies, convergence conditions, constants, and16 failure modes clearly enough for statistical readers.17- For experiments, report datasets, splits, preprocessing, evaluation metrics, baselines,18 hyperparameter ranges, final selected settings, seeds, repeated runs, compute, and runtime.19- For small performance differences, add uncertainty estimates: standard errors, confidence20 intervals, paired tests, bootstrap intervals, or repeated trials as appropriate.21- Explain missing code/data honestly and describe how a reader could reproduce the analysis22 in principle.23- Keep the checklist consistent with the manuscript; contradictions between checklist and24 paper are review-risk multipliers.2526## Checklist-to-claim audit table2728| Checklist item | Pure-theory answer | Theory-plus-experiments answer |29|---|---|---|30| Code availability | NA only if there is literally no computation | Anonymous archive, or an honest stated reason |31| Assumptions stated | Every theorem lists its conditions inline | Plus a note on which experiments satisfy them |32| Error bars | NA for deterministic results | Required for every stochastic figure and table |33| Compute resources | NA | Hardware, runtime, and total number of runs |3435Marking NA on an item the paper actually triggers is a recognizable AISTATS red flag,36because reviewers cross-check checklist answers against the PDF and read contradictions as37carelessness about the rest of the paper.3839## Vignette: a rates-plus-simulation paper4041Consider a submission proving posterior contraction rates for a Bayesian nonparametric model,42validated by MCMC simulation. Its reproducibility spine: prior hyperparameters and their43selection rule, chain length, burn-in, convergence diagnostics, replication seeds, and a44statement of which contraction-theorem conditions the simulated model satisfies — plus one45honest sentence about the condition it does not.4647## Degrees of reproducibility4849- Turnkey: one command regenerates each figure from logged seeds.50- Scripted: scripts exist but require documented manual steps or external data access.51- Descriptive: prose detailed enough that a competent reader could rebuild the pipeline.5253For AISTATS, simulations should be turnkey because statistician reviewers actually rerun54them; large real-data pipelines may stay scripted with deviations documented. Stating the55achieved level honestly beats overpromising turnkey behavior that fails on a clean machine.5657## Output format5859```text60[Claim inventory] <claim -> evidence location>61[Checklist status] complete / inconsistent / missing62[Statistical reproducibility gaps] <assumptions/seeds/uncertainty/hyperparameters/compute>63[Paper fixes] <must appear in main PDF>64[Supplement fixes] <appendix or artifact additions>65```