Data Analysis & Inference for JAE (jae-data-analysis)
When to trigger
- The sample is built and it is time to estimate and report
- You are unsure how to specify fixed effects or cluster standard errors
- Reviewers will probe endogeneity, correlated omitted variables, or sample selection
- A reviewer says "the standard errors are understated" or "this is not identified"
Build and document the archival sample first
JAE reviewers expect a transparent sample-construction waterfall: starting population (e.g., Compustat firm-years), each merge (CRSP, I/B/E/S, Execucomp, DealScan, Audit Analytics via WRDS), each exclusion (financials/utilities, missing data, penny stocks), and the final N at every step. Report descriptive statistics and a correlation table. Winsorize continuous variables (commonly at 1%/99%) and say so.
Specify the estimator to match the panel and the design
| Data structure / claim |
Estimator / specification |
| Firm panel with unobserved heterogeneity |
Firm and year fixed effects (e.g., reghdfe) |
| Inference with within-firm correlation |
Standard errors clustered by firm; often two-way (firm & year) |
| Regulatory shock / treatment |
Difference-in-differences; report pre-trends |
| Endogenous regressor |
2SLS/IV with first-stage diagnostics (F-stat, exclusion) |
| Self-selection |
Heckman (report inverse Mills) or PSM (report balance) |
| Information event |
Short-window CARs; cross-sectional regression of returns |
| Binary/limited outcome |
Logit/probit/Tobit as the outcome dictates |
Match the clustering to where correlation lives in the data; a single firm-clustered SE may understate inference when shocks are common across firms in a year — two-way clustering is the JAE norm for many panels.
Execute the identification, not just the regression
- DiD: plot/test parallel pre-trends; report the dynamic (event-time) coefficients, not only the average treatment effect.
- IV: report the first stage, the instrument's strength, and defend the exclusion restriction in words.
- Matching/Heckman: report covariate balance or the selection equation; show results are not an artifact of the procedure.
- Cross-sectional partitions: the theory's mechanism test — show the effect concentrates where the friction (information asymmetry, weak governance, tight covenants) is severe.
Robustness (expected, not optional)
- Alternative proxies for the key construct (e.g., different discretionary-accruals or conservatism measures).
- Alternative specifications (controls in/out, alternative fixed effects, subsamples).
- Placebo/falsification tests and, for DiD, a non-event window.
- Sensitivity to correlated omitted variables (e.g., bounding / coefficient-stability arguments).
- Address economically plausible alternative explanations empirically.
Execution bridge (StatsPAI / Stata MCP)
Run the battery, don't just enumerate it. Full map:
execution-with-mcp. JAE is empirical accounting with an economics lens; treat identification and weak-IV-robust inference as the binding constraints.
- Many outcomes / specifications:
romano_wolf (step-down FWER) or
benjamini_hochberg — report the adjusted threshold.
- OVB sensitivity:
oster_delta / sensemakr.
- Inference:
wild_cluster_bootstrap (few clusters), twoway_cluster / conley;
multilevel data → cluster at the right level.
- Re-fit off one handle:
audit_result(result_id) lists the missing checks and the
exact suggest_function for each.
- Exhibits:
etable / did_summary_to_latex from the handle — no retyped numbers.
Keep the decisive checks in the body and the exhaustive battery in the appendix. See the
executed chain in the JF execution walkthrough.
Checklist
Anti-patterns
- Pooled OLS with no fixed effects or clustering on a firm panel.
- One-way clustering when shocks are common across firms within a year.
- Reporting an IV with no first stage or no exclusion-restriction defense.
- DiD with no pre-trend evidence.
- Significance with no economic magnitude ("statistically significant" but trivially small).
- Selective controls that make the result appear.
Output format
【Sample】population → merges → exclusions → final N; winsorized at ...
【Specification】FE (firm/year); SE clustering (firm / two-way)
【Identification executed】pre-trends / first-stage F / balance ...
【Main result】coefficient, t-stat, economic magnitude
【Mechanism (cross-section)】effect concentrated where friction severe
【Robustness】alt proxies / specs / placebo / sensitivity
【Open issues for reviewers】...
【Next step】jae-contribution-framing
Source: brycewang-stanford/Awesome-Journal-Skills → Journal-of-Accounting-and-Economics-Skills/skills/jae-data-analysis/SKILL.md
1---2name: jae-data-analysis3description: Use when running and reporting the empirical analysis for a Journal of Accounting and Economics (JAE) manuscript — building the archival sample, choosing fixed effects and clustered standard errors, executing the identification design, and demonstrating robustness for large-sample capital-markets/contracting/disclosure data. Executes and reports the analysis; it does not design the study (jae-methods) or frame the contribution (jae-contribution-framing).4---5
6
7# Data Analysis & Inference for JAE (jae-data-analysis)
8
9## When to trigger
10
11- The sample is built and it is time to estimate and report
12- You are unsure how to specify fixed effects or cluster standard errors
13- Reviewers will probe endogeneity, correlated omitted variables, or sample selection
14- A reviewer says "the standard errors are understated" or "this is not identified"
15
16## Build and document the archival sample first
17
18JAE reviewers expect a transparent **sample-construction waterfall**: starting population (e.g., Compustat firm-years), each merge (CRSP, I/B/E/S, Execucomp, DealScan, Audit Analytics via WRDS), each exclusion (financials/utilities, missing data, penny stocks), and the final N at every step. Report descriptive statistics and a correlation table. **Winsorize** continuous variables (commonly at 1%/99%) and say so.
19
20## Specify the estimator to match the panel and the design
21
22| Data structure / claim | Estimator / specification |
23|-----------------------------------------------|-------------------------------------------------------------|
24| Firm panel with unobserved heterogeneity | Firm and year fixed effects (e.g., `reghdfe`) |
25| Inference with within-firm correlation | Standard errors clustered by firm; often **two-way** (firm & year) |
26| Regulatory shock / treatment | Difference-in-differences; report pre-trends |
27| Endogenous regressor | 2SLS/IV with first-stage diagnostics (F-stat, exclusion) |
28| Self-selection | Heckman (report inverse Mills) or PSM (report balance) |
29| Information event | Short-window CARs; cross-sectional regression of returns |
30| Binary/limited outcome | Logit/probit/Tobit as the outcome dictates |
31
32Match the **clustering** to where correlation lives in the data; a single firm-clustered SE may understate inference when shocks are common across firms in a year — two-way clustering is the JAE norm for many panels.
33
34## Execute the identification, not just the regression
35
36- **DiD**: plot/test parallel pre-trends; report the dynamic (event-time) coefficients, not only the average treatment effect.
37- **IV**: report the first stage, the instrument's strength, and defend the exclusion restriction in words.
38- **Matching/Heckman**: report covariate balance or the selection equation; show results are not an artifact of the procedure.
39- **Cross-sectional partitions**: the theory's mechanism test — show the effect concentrates where the friction (information asymmetry, weak governance, tight covenants) is severe.
40
41## Robustness (expected, not optional)
42
43- Alternative proxies for the key construct (e.g., different discretionary-accruals or conservatism measures).
44- Alternative specifications (controls in/out, alternative fixed effects, subsamples).
45- Placebo/falsification tests and, for DiD, a non-event window.
46- Sensitivity to correlated omitted variables (e.g., bounding / coefficient-stability arguments).
47- Address economically plausible alternative explanations empirically.
48
49## Execution bridge (StatsPAI / Stata MCP)
50
51Run the battery, don't just enumerate it. Full map:
52[`execution-with-mcp`](../../../shared-resources/empirical-methods/execution-with-mcp.md). JAE is empirical accounting with an economics lens; treat identification and weak-IV-robust inference as the binding constraints.
53
54- **Many outcomes / specifications:** `romano_wolf` (step-down FWER) or
55 `benjamini_hochberg` — report the adjusted threshold.
56- **OVB sensitivity:** `oster_delta` / `sensemakr`.
57- **Inference:** `wild_cluster_bootstrap` (few clusters), `twoway_cluster` / `conley`;
58 multilevel data → cluster at the right level.
59- **Re-fit off one handle:** `audit_result(result_id)` lists the missing checks and the
60 exact `suggest_function` for each.
61- **Exhibits:** `etable` / `did_summary_to_latex` from the handle — no retyped numbers.
62
63Keep the decisive checks in the body and the exhaustive battery in the appendix. See the
64executed chain in the [JF execution walkthrough](../../../Journal-of-Finance-Skills/resources/worked-examples/02-execution-walkthrough.md).
65## Checklist
66
67- [ ] Sample waterfall with N at each step; winsorization stated
68- [ ] Descriptives and correlation table reported
69- [ ] Fixed effects and **clustered (often two-way)** SEs match the design
70- [ ] Identification executed (pre-trends / first stage / balance), not assumed
71- [ ] Cross-sectional partition supports the economic mechanism
72- [ ] Robustness: alternative proxies, specifications, placebos, sensitivity
73- [ ] Economic magnitude (not only significance) reported
74
75## Anti-patterns
76
77- **Pooled OLS with no fixed effects or clustering** on a firm panel.
78- **One-way clustering** when shocks are common across firms within a year.
79- **Reporting an IV with no first stage** or no exclusion-restriction defense.
80- **DiD with no pre-trend evidence.**
81- **Significance with no economic magnitude** ("statistically significant" but trivially small).
82- **Selective controls** that make the result appear.
83
84## Output format
85
86```
87【Sample】population → merges → exclusions → final N; winsorized at ...
88【Specification】FE (firm/year); SE clustering (firm / two-way)
89【Identification executed】pre-trends / first-stage F / balance ...
90【Main result】coefficient, t-stat, economic magnitude
91【Mechanism (cross-section)】effect concentrated where friction severe
92【Robustness】alt proxies / specs / placebo / sensitivity
93【Open issues for reviewers】...
94【Next step】jae-contribution-framing
95```
96
97---
98
99**Source:** [`brycewang-stanford/Awesome-Journal-Skills`](https://github.com/brycewang-stanford/Awesome-Journal-Skills) → `Journal-of-Accounting-and-Economics-Skills/skills/jae-data-analysis/SKILL.md`