Replication & Data Policy (jpube-replication-and-data-policy)
When to trigger
- You are preparing a data-availability statement or research-data declaration
- Your results rest on restricted tax/health/register data and you must document access
- You want a reproducible package that pre-empts referee replication requests
- You need to choose between repository deposit and a restricted-data access statement
JPubE / Elsevier requirement
JPubE's Guide for Authors applies Elsevier Option C research-data instructions: deposit research
data in a relevant repository, cite and link the dataset in the article, or provide a statement
explaining why research data cannot be shared. For public-economics work using restricted tax, health,
or register microdata, the practical route is usually a data statement plus a code package and exact
access path rather than public release of protected rows.
Build the package around Option C
Public-economics referees frequently ask to see the elasticity/bunching/RD pipeline, so build a clean
package around the data route:
- Data-availability statement matching reality: open data -> link the repository; restricted
tax/health/register data -> state the access path, the agency, and why microdata cannot be shared.
- Dataset citation/linking. Cite open datasets in the reference list with repository, version, year,
persistent identifier, and the
[dataset] marker where applicable.
- Programs even when microdata are proprietary. Supply all cleaning and estimation code so the
workflow is auditable even if IRS/SSA/CMS or register microdata cannot leave the enclave.
- Disclosure compliance. Document cell-size suppression and output-clearance for restricted data;
never embed suppressed cells in shared outputs.
- One master script (
run_all) regenerating every table and figure from inputs; pin
software/package versions (renv.lock, requirements.txt, recorded ssc versions); set and report
seeds for bootstrap / randomization inference.
- README mapping each exhibit to the script that produces it.
Checklist
Anti-patterns
- Treating Option C as optional boilerplate rather than a repository link or a concrete reason data
cannot be shared
- A data-availability statement that does not match what was actually used
- Sharing restricted-data outputs without documented disclosure clearance
- A package with no master script, unpinned versions, or unreported seeds
Data-availability routing by source
Public-finance papers lean on restricted microdata more than most fields, so the availability statement
is rarely "open repository." Route by what you actually used.
| Data source |
Availability statement says |
What you still ship |
| Public tax/SOI tabulations, survey extracts |
Link repository (Mendeley/openICPSR/Zenodo) and cite dataset |
Data + all code |
| IRS/SSA/CMS enclave microdata |
Access path + agency + why microdata cannot leave |
All cleaning + estimation code |
| European whole-population registers |
Application route, custodian, approval ID |
Code + non-disclosive aggregates |
| Mixed (public + restricted) |
Split the statement by component |
Repository for the open part, access note otherwise |
Worked vignette: a register-DID package referees can trust
A social-insurance reform evaluated on a national register cannot share person-level rows. The package
still makes the DID pipeline auditable: run_all regenerates every exhibit from cleared aggregates; the
README maps Table 3 (the moral-hazard wedge) and Figure 2 (the event study) to their scripts;
renv.lock pins versions; the bootstrap seed is fixed so the SEs on the MVPF = 1.4 statistic
(illustrative) replicate. The availability statement names the custodian, the approval ID, and the
cell-suppression rule (min count 10), so a referee sees the workflow without touching protected
microdata.
Calibration anchors
- The reproducibility bar a JPubE referee imagines: could a second analyst, given the same authorized
access, rebuild every elasticity/MVPF/bunching number? Code completeness and exact access
documentation are what you control.
- Option C is about deposit/citation/linking or a sharing explanation; it is not the same as promising a
named AEA-style data-editor code run.
Output format
【Data type】open / restricted-administrative / register / mixed
【Availability statement】drafted + consistent? [Y/N]
【Restricted access】path + agency + limits documented? [Y/N]
【Programs supplied】all cleaning + estimation code? [Y/N]
【Reproducibility】run_all + pinned versions + seeds? [Y/N]
【Policy check】Option C route and Editorial Manager data fields checked? [Y/N]
【Next step】jpube-review-process
1---2name: jpube-replication-and-data-policy3description: Use when assembling data and code for a Journal of Public Economics (JPubE) manuscript under Elsevier's Option C research-data framework — data-availability statements, repository deposit/citation/linking, restricted administrative data, and a reproducible package. It does not run the analysis.4---56# Replication & Data Policy (jpube-replication-and-data-policy)78## When to trigger910- You are preparing a data-availability statement or research-data declaration11- Your results rest on restricted tax/health/register data and you must document access12- You want a reproducible package that pre-empts referee replication requests13- You need to choose between repository deposit and a restricted-data access statement1415## JPubE / Elsevier requirement1617JPubE's Guide for Authors applies Elsevier **Option C** research-data instructions: deposit research18data in a relevant repository, cite and link the dataset in the article, or provide a statement19explaining why research data cannot be shared. For public-economics work using restricted tax, health,20or register microdata, the practical route is usually a data statement plus a code package and exact21access path rather than public release of protected rows.2223## Build the package around Option C2425Public-economics referees frequently ask to see the elasticity/bunching/RD pipeline, so build a clean26package around the data route:2728- **Data-availability statement** matching reality: open data -> link the repository; restricted29 tax/health/register data -> state the access path, the agency, and why microdata cannot be shared.30- **Dataset citation/linking.** Cite open datasets in the reference list with repository, version, year,31 persistent identifier, and the `[dataset]` marker where applicable.32- **Programs even when microdata are proprietary.** Supply all cleaning and estimation code so the33 workflow is auditable even if IRS/SSA/CMS or register microdata cannot leave the enclave.34- **Disclosure compliance.** Document cell-size suppression and output-clearance for restricted data;35 never embed suppressed cells in shared outputs.36- **One master script** (`run_all`) regenerating every table and figure from inputs; pin37 software/package versions (`renv.lock`, `requirements.txt`, recorded `ssc` versions); set and report38 seeds for bootstrap / randomization inference.39- **README** mapping each exhibit to the script that produces it.4041## Checklist4243- [ ] Data-availability statement drafted and consistent with the data used44- [ ] Open data deposited, cited, and linked; restricted data explained with access path45- [ ] Restricted-data access path, agency, and sharing limits documented46- [ ] All cleaning + estimation programs supplied (even if microdata are restricted)47- [ ] Disclosure / cell-suppression compliance documented for shared outputs48- [ ] `run_all` master script regenerates all exhibits; versions and seeds pinned49- [ ] README maps exhibits -> scripts50- [ ] Current JPubE/Elsevier data fields checked in Editorial Manager before upload5152## Anti-patterns5354- Treating Option C as optional boilerplate rather than a repository link or a concrete reason data55 cannot be shared56- A data-availability statement that does not match what was actually used57- Sharing restricted-data outputs without documented disclosure clearance58- A package with no master script, unpinned versions, or unreported seeds5960## Data-availability routing by source6162Public-finance papers lean on restricted microdata more than most fields, so the availability statement63is rarely "open repository." Route by what you actually used.6465| Data source | Availability statement says | What you still ship |66|-------------|------------------------------|----------------------|67| Public tax/SOI tabulations, survey extracts | Link repository (Mendeley/openICPSR/Zenodo) and cite dataset | Data + all code |68| IRS/SSA/CMS enclave microdata | Access path + agency + why microdata cannot leave | All cleaning + estimation code |69| European whole-population registers | Application route, custodian, approval ID | Code + non-disclosive aggregates |70| Mixed (public + restricted) | Split the statement by component | Repository for the open part, access note otherwise |7172## Worked vignette: a register-DID package referees can trust7374A social-insurance reform evaluated on a national register cannot share person-level rows. The package75still makes the DID pipeline auditable: `run_all` regenerates every exhibit from cleared aggregates; the76README maps Table 3 (the moral-hazard wedge) and Figure 2 (the event study) to their scripts;77`renv.lock` pins versions; the bootstrap seed is fixed so the SEs on the **MVPF = 1.4** statistic78(illustrative) replicate. The availability statement names the custodian, the approval ID, and the79cell-suppression rule (min count 10), so a referee sees the workflow without touching protected80microdata.8182## Calibration anchors8384- The reproducibility bar a JPubE referee imagines: could a second analyst, given the same authorized85 access, rebuild every elasticity/MVPF/bunching number? Code completeness and exact access86 documentation are what you control.87- Option C is about deposit/citation/linking or a sharing explanation; it is not the same as promising a88 named AEA-style data-editor code run.8990## Output format9192```93【Data type】open / restricted-administrative / register / mixed94【Availability statement】drafted + consistent? [Y/N]95【Restricted access】path + agency + limits documented? [Y/N]96【Programs supplied】all cleaning + estimation code? [Y/N]97【Reproducibility】run_all + pinned versions + seeds? [Y/N]98【Policy check】Option C route and Editorial Manager data fields checked? [Y/N]99【Next step】jpube-review-process100```