Data Analysis Workflow
Run an end-to-end data analysis in R: load, explore, analyze, and produce publication-ready output.
Input: $ARGUMENTS — a dataset path (e.g., data/county_panel.csv) or a description of the analysis goal (e.g., "regress wages on education with state fixed effects using CPS data").
Constraints
- Follow R code conventions in
.claude/rules/r-code-conventions.md
- Save all scripts to
scripts/R/ with descriptive names
- Save all outputs (figures, tables, RDS) to
output/
- Use
saveRDS() for every computed object — Quarto slides may need them
- Use project theme for all figures (check for custom theme in
.claude/rules/)
- Run r-reviewer on the generated script before presenting results
Workflow Phases
Phase 1: Setup and Data Loading
- Read
.claude/rules/r-code-conventions.md for project standards
- Create R script with proper header (title, author, purpose, inputs, outputs)
- Load required packages at top (
library(), never require())
- Set seed once at top:
set.seed(42)
- Load and inspect the dataset
Phase 2: Exploratory Data Analysis
Generate diagnostic outputs:
- Summary statistics:
summary(), missingness rates, variable types
- Distributions: Histograms for key continuous variables
- Relationships: Scatter plots, correlation matrices
- Time patterns: If panel data, plot trends over time
- Group comparisons: If treatment/control, compare pre-treatment means
Save all diagnostic figures to output/diagnostics/.
Phase 3: Main Analysis
Based on the research question:
- Regression analysis: Use
fixest for panel data, lm/glm for cross-section
- Standard errors: Cluster at the appropriate level (document why)
- Multiple specifications: Start simple, progressively add controls
- Effect sizes: Report standardized effects alongside raw coefficients
Phase 4: Publication-Ready Output
Tables:
- Use
modelsummary for regression tables (preferred) or stargazer
- Include all standard elements: coefficients, SEs, significance stars, N, R-squared
- Export as
.tex for LaTeX inclusion and .html for quick viewing
Figures:
- Use
ggplot2 with project theme
- Set
bg = "transparent" for Beamer compatibility
- Include proper axis labels (sentence case, units)
- Export with explicit dimensions:
ggsave(width = X, height = Y)
- Save as both
.pdf and .png
Phase 5: Save and Review
saveRDS() for all key objects (regression results, summary tables, processed data)
- Create
output/ subdirectories as needed with dir.create(..., recursive = TRUE)
- Run the r-reviewer agent on the generated script:
Delegate to the r-reviewer agent:
"Review the script at scripts/R/[script_name].R"
- Address any Critical or High issues from the review.
Script Structure
Follow this template:
# ============================================================
# [Descriptive Title]
# Author: [from project context]
# Purpose: [What this script does]
# Inputs: [Data files]
# Outputs: [Figures, tables, RDS files]
# ============================================================
# 0. Setup ----
library(tidyverse)
library(fixest)
library(modelsummary)
set.seed(42)
dir.create("output/analysis", recursive = TRUE, showWarnings = FALSE)
# 1. Data Loading ----
# [Load and clean data]
# 2. Exploratory Analysis ----
# [Summary stats, diagnostic plots]
# 3. Main Analysis ----
# [Regressions, estimation]
# 4. Tables and Figures ----
# [Publication-ready output]
# 5. Export ----
# [saveRDS for all objects, ggsave for all figures]
Important
- Reproduce, don't guess. If the user specifies a regression, run exactly that.
- Show your work. Print summary statistics before jumping to regression.
- Check for issues. Look for multicollinearity, outliers, perfect prediction.
- Use relative paths. All paths relative to repository root.
- No hardcoded values. Use variables for sample restrictions, date ranges, etc.
1---2name: data-analysis-33description: End-to-end R data analysis workflow from exploration through regression to publication-ready tables and figures4---5
6# Data Analysis Workflow
7
8Run an end-to-end data analysis in R: load, explore, analyze, and produce publication-ready output.
9
10**Input:** `$ARGUMENTS` — a dataset path (e.g., `data/county_panel.csv`) or a description of the analysis goal (e.g., "regress wages on education with state fixed effects using CPS data").
11
12---
13
14## Constraints
15
16- **Follow R code conventions** in `.claude/rules/r-code-conventions.md`
17- **Save all scripts** to `scripts/R/` with descriptive names
18- **Save all outputs** (figures, tables, RDS) to `output/`
19- **Use `saveRDS()`** for every computed object — Quarto slides may need them
20- **Use project theme** for all figures (check for custom theme in `.claude/rules/`)
21- **Run r-reviewer** on the generated script before presenting results
22
23---
24
25## Workflow Phases
26
27### Phase 1: Setup and Data Loading
28
291. Read `.claude/rules/r-code-conventions.md` for project standards
302. Create R script with proper header (title, author, purpose, inputs, outputs)
313. Load required packages at top (`library()`, never `require()`)
324. Set seed once at top: `set.seed(42)`
335. Load and inspect the dataset
34
35### Phase 2: Exploratory Data Analysis
36
37Generate diagnostic outputs:
38- **Summary statistics:** `summary()`, missingness rates, variable types
39- **Distributions:** Histograms for key continuous variables
40- **Relationships:** Scatter plots, correlation matrices
41- **Time patterns:** If panel data, plot trends over time
42- **Group comparisons:** If treatment/control, compare pre-treatment means
43
44Save all diagnostic figures to `output/diagnostics/`.
45
46### Phase 3: Main Analysis
47
48Based on the research question:
49- **Regression analysis:** Use `fixest` for panel data, `lm`/`glm` for cross-section
50- **Standard errors:** Cluster at the appropriate level (document why)
51- **Multiple specifications:** Start simple, progressively add controls
52- **Effect sizes:** Report standardized effects alongside raw coefficients
53
54### Phase 4: Publication-Ready Output
55
56**Tables:**
57- Use `modelsummary` for regression tables (preferred) or `stargazer`
58- Include all standard elements: coefficients, SEs, significance stars, N, R-squared
59- Export as `.tex` for LaTeX inclusion and `.html` for quick viewing
60
61**Figures:**
62- Use `ggplot2` with project theme
63- Set `bg = "transparent"` for Beamer compatibility
64- Include proper axis labels (sentence case, units)
65- Export with explicit dimensions: `ggsave(width = X, height = Y)`
66- Save as both `.pdf` and `.png`
67
68### Phase 5: Save and Review
69
701. `saveRDS()` for all key objects (regression results, summary tables, processed data)
712. Create `output/` subdirectories as needed with `dir.create(..., recursive = TRUE)`
723. Run the r-reviewer agent on the generated script:
73
74```
75Delegate to the r-reviewer agent:
76"Review the script at scripts/R/[script_name].R"
77```
78
794. Address any Critical or High issues from the review.
80
81---
82
83## Script Structure
84
85Follow this template:
86
87```r
88# ============================================================
89# [Descriptive Title]
90# Author: [from project context]
91# Purpose: [What this script does]
92# Inputs: [Data files]
93# Outputs: [Figures, tables, RDS files]
94# ============================================================
95
96# 0. Setup ----
97library(tidyverse)
98library(fixest)
99library(modelsummary)
100
101set.seed(42)
102
103dir.create("output/analysis", recursive = TRUE, showWarnings = FALSE)
104
105# 1. Data Loading ----
106# [Load and clean data]
107
108# 2. Exploratory Analysis ----
109# [Summary stats, diagnostic plots]
110
111# 3. Main Analysis ----
112# [Regressions, estimation]
113
114# 4. Tables and Figures ----
115# [Publication-ready output]
116
117# 5. Export ----
118# [saveRDS for all objects, ggsave for all figures]
119```
120
121---
122
123## Important
124
125- **Reproduce, don't guess.** If the user specifies a regression, run exactly that.
126- **Show your work.** Print summary statistics before jumping to regression.
127- **Check for issues.** Look for multicollinearity, outliers, perfect prediction.
128- **Use relative paths.** All paths relative to repository root.
129- **No hardcoded values.** Use variables for sample restrictions, date ranges, etc.