Research Design (eursr-research-design)
ESR is a quantitative journal exacting about whether the comparative or longitudinal design actually
identifies the mechanism from eursr-theory-building and rules out the leading confound. The design
must connect the cross-level hypothesis to evidence that a single cross-section could not provide.
When to trigger
- Specifying the comparative frame, the panel structure, sampling, or the identification strategy
- A reviewer questioned causal claims, generalization, selection, measurement comparability, or a confound
- Justifying why your design adjudicates the rival account from
eursr-literature-positioning
Comparative / cross-national
- Justify the country set by design logic (institutional contrast, regime types, most/least-similar),
not by data availability alone; say what variation each context contributes.
- Measurement equivalence is the first reviewer demand: establish that constructs mean the same
across countries (configural/metric/scalar invariance for latent scales; harmonized coding for
education via ISCED/CASMIN, occupation via ISCO/ISEI/EGP).
- Macro N is small. With ~20-30 countries, country-level effects rest on few degrees of freedom —
design the macro hypothesis so it does not over-claim from a handful of clusters (see
eursr-data-analysis).
Panel / longitudinal / event-history
- State what the panel buys. Within-person change (fixed effects), duration/timing (event history),
or growth (latent growth) — match the estimator to the theoretical quantity.
- Attrition and selection into and out of the panel must be addressed (weights, IPW, sensitivity).
- For staggered policy exposure, use heterogeneity-robust DiD (Callaway-Sant'Anna, Sun-Abraham,
Borusyak et al.), not naive TWFE.
Causal inference where feasible
- Much of ESR is observational; distinguish description, association, and causation honestly. If
causal, state the assumptions (ignorability, parallel trends, exclusion) and defend them; report a
sensitivity bound (how strong an unobserved confounder would have to be).
Multilevel / SEM
- Specify the level structure (individuals in countries/regions/cohorts), the random effects, and why a
multilevel model is warranted; for measurement, build the latent model before the structural one.
The adjudication test (ESR-specific)
For the single strongest rival explanation: "If the rival were true rather than my argument, the
cross-national (or over-time) pattern would look like ___; instead it looks like ___." If you cannot
write it, the comparative/panel design does not yet identify the contribution.
What ESR referees demand of each design
| Design |
Referee's first demand |
Satisfying move |
| Comparative cross-national |
"Are the measures equivalent?" |
invariance / harmonized coding; justified country set |
| Panel / fixed-effects |
"What does within-person change identify?" |
match estimator to the quantity; handle attrition |
| Event history |
"Right risk set and time scale?" |
defined onset, censoring, time-varying covariates |
| Causal (DiD/IV/RDD) |
"Assumption defended?" |
state + test the assumption; sensitivity bound |
| Multilevel / SEM |
"Enough clusters; measurement first?" |
macro df honesty; fit the latent model before structure |
Worked micro-example (illustrative)
A comparative study argues that vocational specificity smooths the school-to-work transition.
Country set: most-different welfare/training regimes (e.g., dual-system vs. general-education systems),
chosen for institutional contrast, not convenience
Measurement: education harmonized via ISCED; vocational specificity coded from program-level data
Design: cross-national + cohort variation; cross-level interaction (specificity × individual track)
Disconfirming pattern sought: if signaling (not skills) drove it, the advantage would vanish once firms
learn quality → instead it persists across the early career, as the specificity argument predicts
Macro-N caution: ~24 countries → country-level claim kept modest; SEs / df handled in data-analysis
The country set is design-driven, the measures are comparable, and the design specifies what pattern would falsify the argument.
Referee pushback → ESR-specific fix
- "Measures aren't comparable across countries." → Test invariance; report partial invariance and
what it permits; use harmonized coding schemes.
- "You infer too much from ~20 countries." → Re-state the macro claim modestly; use df-appropriate
inference (see
eursr-data-analysis).
- "Association dressed as causation." → Restate what the design identifies; add a sensitivity bound or
placebo; drop causal verbs you cannot defend.
Calibration anchors
- Measurement equivalence is the comparative gate. A cross-national claim built on non-equivalent
scales is the most common fatal design flaw at ESR.
- The adjudication sentence is the test. If you can't write "if the rival were true the pattern
would look like ___," the comparison/panel does not yet earn the contribution.
- Identification honesty travels. Stating plainly what observational European data can and cannot
establish reads as strength to a quantitative panel.
Execution bridge (StatsPAI / Stata MCP)
Estimate and audit the design, don't only describe it. Full map:
execution-with-mcp. ESR is comparative quantitative sociology; cross-country panels with confounded institutions — foreground fixed effects and clustering.
detect_design → recommend → fit with as_handle=true → audit_result.
- Observational causal claims: staggered DiD (
callaway_santanna / sun_abraham +
bacon_decomposition + honest_did_from_result); IV (effective_f_test +
anderson_rubin_ci); RDD (rdrobust + mccrary_test).
- Experiments: randomization-based inference,
romano_wolf for many-outcome
family-wise control, and mediate for mediation (not naive controlling-away).
- Sensitivity:
oster_delta / sensemakr for observational claims.
Report the effect size in interpretable units; route the full battery to the
appendix/supplement. A run end-to-end (synthetic data, real returns) is in the
JF execution walkthrough.
Anti-patterns
- A country set chosen by data availability and dressed up as theory-driven
- Cross-national latent comparisons with no measurement-invariance check
- Over-claiming country-level effects from a handful of clusters
- Naive TWFE on staggered policy timing; ignoring panel attrition
- A design that cannot distinguish your mechanism from the leading alternative
Output format
【Design】comparative / panel / event-history / causal / multilevel-SEM
【What it identifies】description / association / causation
【Comparability / assumption】invariance or key assumption + how defended
【Rival ruled out】the adjudication sentence
【Macro-N / attrition / sensitivity】planned
【Next】eursr-data-analysis
Supplementary resources
Source: brycewang-stanford/Awesome-Journal-Skills → European-Sociological-Review-Skills/skills/eursr-research-design/SKILL.md
1---2name: eursr-research-design3description: Use when defending the research design of a European Sociological Review (ESR) manuscript — comparative cross-national designs, panel/longitudinal and event-history designs, multilevel structures, and causal inference where feasible on harmonized survey or register data. ESR judges whether the design lets the comparison or panel identify the mechanism. Strengthens the design; it does not write code.4---567# Research Design (eursr-research-design)89ESR is a quantitative journal exacting about whether the **comparative or longitudinal design actually10identifies the mechanism** from `eursr-theory-building` and rules out the leading confound. The design11must connect the cross-level hypothesis to evidence that a single cross-section could not provide.1213## When to trigger1415- Specifying the comparative frame, the panel structure, sampling, or the identification strategy16- A reviewer questioned causal claims, generalization, selection, measurement comparability, or a confound17- Justifying why your design adjudicates the rival account from `eursr-literature-positioning`1819## Comparative / cross-national20- **Justify the country set** by design logic (institutional contrast, regime types, most/least-similar),21 not by data availability alone; say what variation each context contributes.22- **Measurement equivalence** is the first reviewer demand: establish that constructs mean the same23 across countries (configural/metric/scalar invariance for latent scales; harmonized coding for24 education via ISCED/CASMIN, occupation via ISCO/ISEI/EGP).25- **Macro N is small.** With ~20-30 countries, country-level effects rest on few degrees of freedom —26 design the macro hypothesis so it does not over-claim from a handful of clusters (see27 `eursr-data-analysis`).2829## Panel / longitudinal / event-history30- **State what the panel buys.** Within-person change (fixed effects), duration/timing (event history),31 or growth (latent growth) — match the estimator to the theoretical quantity.32- **Attrition and selection** into and out of the panel must be addressed (weights, IPW, sensitivity).33- For staggered policy exposure, use **heterogeneity-robust DiD** (Callaway-Sant'Anna, Sun-Abraham,34 Borusyak et al.), not naive TWFE.3536## Causal inference where feasible37- Much of ESR is observational; **distinguish description, association, and causation honestly**. If38 causal, state the assumptions (ignorability, parallel trends, exclusion) and defend them; report a39 **sensitivity bound** (how strong an unobserved confounder would have to be).4041## Multilevel / SEM42- Specify the level structure (individuals in countries/regions/cohorts), the random effects, and why a43 multilevel model is warranted; for measurement, build the latent model before the structural one.4445## The adjudication test (ESR-specific)4647For the **single strongest rival explanation**: *"If the rival were true rather than my argument, the48cross-national (or over-time) pattern would look like ___; instead it looks like ___."* If you cannot49write it, the comparative/panel design does not yet identify the contribution.5051## What ESR referees demand of each design5253| Design | Referee's first demand | Satisfying move |54|--------|------------------------|------------------|55| Comparative cross-national | "Are the measures equivalent?" | invariance / harmonized coding; justified country set |56| Panel / fixed-effects | "What does within-person change identify?" | match estimator to the quantity; handle attrition |57| Event history | "Right risk set and time scale?" | defined onset, censoring, time-varying covariates |58| Causal (DiD/IV/RDD) | "Assumption defended?" | state + test the assumption; sensitivity bound |59| Multilevel / SEM | "Enough clusters; measurement first?" | macro df honesty; fit the latent model before structure |6061## Worked micro-example (illustrative)6263A comparative study argues that vocational specificity smooths the school-to-work transition.6465```66Country set: most-different welfare/training regimes (e.g., dual-system vs. general-education systems),67 chosen for institutional contrast, not convenience68Measurement: education harmonized via ISCED; vocational specificity coded from program-level data69Design: cross-national + cohort variation; cross-level interaction (specificity × individual track)70Disconfirming pattern sought: if signaling (not skills) drove it, the advantage would vanish once firms71 learn quality → instead it persists across the early career, as the specificity argument predicts72Macro-N caution: ~24 countries → country-level claim kept modest; SEs / df handled in data-analysis73```7475The country set is design-driven, the measures are comparable, and the design specifies what pattern would falsify the argument.7677## Referee pushback → ESR-specific fix7879- *"Measures aren't comparable across countries."* → Test invariance; report partial invariance and80 what it permits; use harmonized coding schemes.81- *"You infer too much from ~20 countries."* → Re-state the macro claim modestly; use df-appropriate82 inference (see `eursr-data-analysis`).83- *"Association dressed as causation."* → Restate what the design identifies; add a sensitivity bound or84 placebo; drop causal verbs you cannot defend.8586## Calibration anchors8788- **Measurement equivalence is the comparative gate.** A cross-national claim built on non-equivalent89 scales is the most common fatal design flaw at ESR.90- **The adjudication sentence is the test.** If you can't write "if the rival were true the pattern91 would look like ___," the comparison/panel does not yet earn the contribution.92- **Identification honesty travels.** Stating plainly what observational European data can and cannot93 establish reads as strength to a quantitative panel.9495## Execution bridge (StatsPAI / Stata MCP)9697Estimate and audit the design, don't only describe it. Full map:98[`execution-with-mcp`](../../../shared-resources/empirical-methods/execution-with-mcp.md). ESR is comparative quantitative sociology; cross-country panels with confounded institutions — foreground fixed effects and clustering.99100- `detect_design` → `recommend` → fit with `as_handle=true` → `audit_result`.101- **Observational causal claims:** staggered DiD (`callaway_santanna` / `sun_abraham` +102 `bacon_decomposition` + `honest_did_from_result`); IV (`effective_f_test` +103 `anderson_rubin_ci`); RDD (`rdrobust` + `mccrary_test`).104- **Experiments:** randomization-based inference, `romano_wolf` for many-outcome105 family-wise control, and `mediate` for mediation (not naive controlling-away).106- **Sensitivity:** `oster_delta` / `sensemakr` for observational claims.107108Report the effect size in interpretable units; route the full battery to the109appendix/supplement. A run end-to-end (synthetic data, real returns) is in the110[JF execution walkthrough](../../../Journal-of-Finance-Skills/resources/worked-examples/02-execution-walkthrough.md).111## Anti-patterns112113- A country set chosen by data availability and dressed up as theory-driven114- Cross-national latent comparisons with no measurement-invariance check115- Over-claiming country-level effects from a handful of clusters116- Naive TWFE on staggered policy timing; ignoring panel attrition117- A design that cannot distinguish your mechanism from the leading alternative118119## Output format120121```122【Design】comparative / panel / event-history / causal / multilevel-SEM123【What it identifies】description / association / causation124【Comparability / assumption】invariance or key assumption + how defended125【Rival ruled out】the adjudication sentence126【Macro-N / attrition / sensitivity】planned127【Next】eursr-data-analysis128```129130## Supplementary resources131132- [`../../resources/external_tools.md`](../../resources/external_tools.md) — multilevel / SEM / event-history / DiD tooling133- [`../../resources/code/`](../../resources/code/) — reproducible Stata + Python causal-inference skeleton (DiD/IV/RDD/DML)134- [`../../resources/official-source-map.md`](../../resources/official-source-map.md) — ESR methodological expectations135136---137138**Source:** [`brycewang-stanford/Awesome-Journal-Skills`](https://github.com/brycewang-stanford/Awesome-Journal-Skills) → `European-Sociological-Review-Skills/skills/eursr-research-design/SKILL.md`