NeurIPS Artifact Evaluation
NeurIPS main track does not reduce artifact quality to a generic badge workflow. It expects code,
data, and execution details when they are needed to support the scientific claim, and its checklist
and code/data guidance make artifact quality visible to reviewers.
Artifact decision
- If the contribution is a method, include training and evaluation code or justify why it cannot be
shared.
- If the contribution is a dataset or benchmark, provide metadata, license, preservation plan,
representative-use discussion, and access restrictions.
- If the contribution depends on a model, include weights, prompts, decoding settings, compute
resources, or a precise explanation of unavailable components.
- If the contribution is theoretical, artifact focus may shift to proof checks, symbolic scripts,
experiment notebooks, or counterexample generation.
Anonymous review package
- Keep the ZIP within the current official size limit and anonymize filenames, repository URLs,
usernames, commit history, model cards, dataset cards, comments, notebooks, and logs.
- Include a short
README with exact commands, environment, expected runtime, hardware assumptions,
and which experiments are reproducible from the package.
- Do not require reviewers to run unsafe code outside a secure environment.
- Avoid external links unless the current policy allows them and anonymous browsing is guaranteed.
Public release package
- De-anonymize accepted artifacts.
- Add licenses for code, data, model weights, and generated outputs.
- Archive code in a durable service when appropriate; NeurIPS MLRC guidance recommends Software
Heritage for reproducibility papers.
- Keep a mapping from paper claims to commands or notebooks so users can reproduce headline results.
Output format
[Artifact role] method / dataset / benchmark / model / demo / proof / none
[Review package] sufficient / insufficient
[Anonymity risks] <paths, metadata, URLs, usernames>
[Reproducibility gaps] <commands, environment, data, hardware, licenses>
[Public-release plan] <archive, DOI, license, docs>
Source: brycewang-stanford/Awesome-Journal-Skills → NeurIPS-Skills/skills/neurips-artifact-evaluation/SKILL.md
1---2name: neurips-artifact-evaluation3description: Use when packaging NeurIPS code, data, models, demos, benchmarks, or other research artifacts for anonymous review, reproducibility, public release, or MLRC-style artifact scrutiny.4---567# NeurIPS Artifact Evaluation89NeurIPS main track does not reduce artifact quality to a generic badge workflow. It expects code,10data, and execution details when they are needed to support the scientific claim, and its checklist11and code/data guidance make artifact quality visible to reviewers.1213## Artifact decision1415- If the contribution is a method, include training and evaluation code or justify why it cannot be16 shared.17- If the contribution is a dataset or benchmark, provide metadata, license, preservation plan,18 representative-use discussion, and access restrictions.19- If the contribution depends on a model, include weights, prompts, decoding settings, compute20 resources, or a precise explanation of unavailable components.21- If the contribution is theoretical, artifact focus may shift to proof checks, symbolic scripts,22 experiment notebooks, or counterexample generation.2324## Anonymous review package2526- Keep the ZIP within the current official size limit and anonymize filenames, repository URLs,27 usernames, commit history, model cards, dataset cards, comments, notebooks, and logs.28- Include a short `README` with exact commands, environment, expected runtime, hardware assumptions,29 and which experiments are reproducible from the package.30- Do not require reviewers to run unsafe code outside a secure environment.31- Avoid external links unless the current policy allows them and anonymous browsing is guaranteed.3233## Public release package3435- De-anonymize accepted artifacts.36- Add licenses for code, data, model weights, and generated outputs.37- Archive code in a durable service when appropriate; NeurIPS MLRC guidance recommends Software38 Heritage for reproducibility papers.39- Keep a mapping from paper claims to commands or notebooks so users can reproduce headline results.4041## Output format4243```text44[Artifact role] method / dataset / benchmark / model / demo / proof / none45[Review package] sufficient / insufficient46[Anonymity risks] <paths, metadata, URLs, usernames>47[Reproducibility gaps] <commands, environment, data, hardware, licenses>48[Public-release plan] <archive, DOI, license, docs>49```5051---5253**Source:** [`brycewang-stanford/Awesome-Journal-Skills`](https://github.com/brycewang-stanford/Awesome-Journal-Skills) → `NeurIPS-Skills/skills/neurips-artifact-evaluation/SKILL.md`