Data Analysis
Generate rigorous statistical analysis code with multi-round review.
Input
$0 — Data source (CSV, JSON, pickle, or experiment logs)
$1 — Research goal or hypothesis to test
References
- 4-round code review prompts:
~/.claude/skills/data-analysis/references/review-prompts.md
Scripts
Statistical summary and comparison
python ~/.claude/skills/data-analysis/scripts/stat_summary.py --input results.csv --compare method --metric accuracy --output summary.json
python ~/.claude/skills/data-analysis/scripts/stat_summary.py --input results.csv --describe
Detects data types, recommends tests, runs comparisons, outputs effect sizes and significance stars. Requires numpy, scipy.
Format p-values
python ~/.claude/skills/data-analysis/scripts/format_pvalue.py --values "0.001 0.05 0.23" --format stars
python ~/.claude/skills/data-analysis/scripts/format_pvalue.py --csv results.csv --column pvalue --format latex
Formats p-values with stars, LaTeX notation, or plain text. Stdlib-only.
Workflow
Step 1: Generate Analysis Code
Structure the code with these sections:
# IMPORT — pandas, numpy, scipy, statsmodels, sklearn
# LOAD DATA — Load from original data files
# DATASET PREPARATIONS — Missing values, units, exclusion criteria
# DESCRIPTIVE STATISTICS — Summary tables if needed
# PREPROCESSING — Dummy variables, normalization
# ANALYSIS — Statistical tests per hypothesis
# SAVE ADDITIONAL RESULTS — Extra results to pickle
Step 2: 4-Round Code Review
- Round 1 — Code Flaws: Mathematical/statistical errors, wrong calculations, trivial tests
- Round 2 — Data Handling: Missing values, units, preprocessing, test choice
- Round 3 — Per-Table: Sensible values, measures of uncertainty, missing data
- Round 4 — Cross-Table: Completeness, consistency, missing variables
Step 3: Produce Results
- Every nominal value must have uncertainty (CI, STD, or p-value)
- Statistical tests must be appropriate for the data type
- Results must match actual data — never hallucinate
Allowed Packages
pandas, numpy, scipy, statsmodels, sklearn, pickle
Statistical Test Selection
| Data Type |
Test |
| Two groups, normal |
Independent t-test |
| Two groups, non-normal |
Mann-Whitney U |
| Paired samples |
Paired t-test / Wilcoxon |
| Multiple groups |
ANOVA / Kruskal-Wallis |
| Categorical |
Chi-square / Fisher's exact |
| Correlation |
Pearson / Spearman |
| Regression |
OLS / Logistic / Mixed effects |
Rules
- Always report p-values for statistical tests
- Account for relevant confounding variables
- Use inherent package functionality (e.g.,
formula = "y ~ a * b" for interactions)
- Do not manually implement available statistical functions
- Access dataframes using string-based column names, not integer indices
Related Skills
1---2name: data-analysis-23description: Generate statistical analysis code with 4-round review. Select appropriate statistical tests, interpret results, and produce analysis reports with p-values, effect sizes, and confidence intervals. Use when analyzing experimental data for a paper.4---5
6# Data Analysis
7
8Generate rigorous statistical analysis code with multi-round review.
9
10## Input
11
12- `$0` — Data source (CSV, JSON, pickle, or experiment logs)
13- `$1` — Research goal or hypothesis to test
14
15## References
16
17- 4-round code review prompts: `~/.claude/skills/data-analysis/references/review-prompts.md`
18
19## Scripts
20
21### Statistical summary and comparison
22```bash
23python ~/.claude/skills/data-analysis/scripts/stat_summary.py --input results.csv --compare method --metric accuracy --output summary.json
24python ~/.claude/skills/data-analysis/scripts/stat_summary.py --input results.csv --describe
25```
26
27Detects data types, recommends tests, runs comparisons, outputs effect sizes and significance stars. Requires numpy, scipy.
28
29### Format p-values
30```bash
31python ~/.claude/skills/data-analysis/scripts/format_pvalue.py --values "0.001 0.05 0.23" --format stars
32python ~/.claude/skills/data-analysis/scripts/format_pvalue.py --csv results.csv --column pvalue --format latex
33```
34
35Formats p-values with stars, LaTeX notation, or plain text. Stdlib-only.
36
37## Workflow
38
39### Step 1: Generate Analysis Code
40Structure the code with these sections:
411. `# IMPORT` — pandas, numpy, scipy, statsmodels, sklearn
422. `# LOAD DATA` — Load from original data files
433. `# DATASET PREPARATIONS` — Missing values, units, exclusion criteria
444. `# DESCRIPTIVE STATISTICS` — Summary tables if needed
455. `# PREPROCESSING` — Dummy variables, normalization
466. `# ANALYSIS` — Statistical tests per hypothesis
477. `# SAVE ADDITIONAL RESULTS` — Extra results to pickle
48
49### Step 2: 4-Round Code Review
501. **Round 1 — Code Flaws**: Mathematical/statistical errors, wrong calculations, trivial tests
512. **Round 2 — Data Handling**: Missing values, units, preprocessing, test choice
523. **Round 3 — Per-Table**: Sensible values, measures of uncertainty, missing data
534. **Round 4 — Cross-Table**: Completeness, consistency, missing variables
54
55### Step 3: Produce Results
56- Every nominal value must have uncertainty (CI, STD, or p-value)
57- Statistical tests must be appropriate for the data type
58- Results must match actual data — never hallucinate
59
60## Allowed Packages
61
62`pandas`, `numpy`, `scipy`, `statsmodels`, `sklearn`, `pickle`
63
64## Statistical Test Selection
65
66| Data Type | Test |
67|-----------|------|
68| Two groups, normal | Independent t-test |
69| Two groups, non-normal | Mann-Whitney U |
70| Paired samples | Paired t-test / Wilcoxon |
71| Multiple groups | ANOVA / Kruskal-Wallis |
72| Categorical | Chi-square / Fisher's exact |
73| Correlation | Pearson / Spearman |
74| Regression | OLS / Logistic / Mixed effects |
75
76## Rules
77
78- Always report p-values for statistical tests
79- Account for relevant confounding variables
80- Use inherent package functionality (e.g., `formula = "y ~ a * b"` for interactions)
81- Do not manually implement available statistical functions
82- Access dataframes using string-based column names, not integer indices
83
84## Related Skills
85- Upstream: [experiment-code](../experiment-code/), [experiment-design](../experiment-design/)
86- Downstream: [table-generation](../table-generation/), [figure-generation](../figure-generation/), [backward-traceability](../backward-traceability/)
87- See also: [math-reasoning](../math-reasoning/)