Survival Analysis (Kaplan-Meier)
Kaplan-Meier survival analysis tool for clinical and biological research. Generates publication-ready survival curves with statistical tests.
Features
- Kaplan-Meier Curve Generation: Publication-quality survival plots with confidence intervals
- Statistical Tests: Log-rank test, Wilcoxon test, Peto-Peto test
- Hazard Ratios: Cox proportional hazards regression with 95% CI
- Summary Statistics: Median survival time, restricted mean survival time (RMST)
- Multi-group Analysis: Supports 2+ comparison groups
- Risk Tables: Optional at-risk table below curves
Usage
Python Script
python scripts/main.py --input data.csv --time time_col --event event_col --group group_col --output results/
Arguments
| Argument |
Description |
Required |
--input |
Input CSV file path |
Yes |
--time |
Column name for survival time |
Yes |
--event |
Column name for event indicator (1=event, 0=censored) |
Yes |
--group |
Column name for grouping variable |
Optional |
--output |
Output directory for results |
Yes |
--conf-level |
Confidence level (default: 0.95) |
Optional |
--risk-table |
Include risk table in plot |
Optional |
Input Format
CSV with columns:
- Time column: Numeric, time to event or censoring
- Event column: Binary (1 = event occurred, 0 = censored/right-censored)
- Group column: Categorical variable for stratification
Example:
patient_id,time_months,death,treatment_group
P001,24.5,1,Drug_A
P002,36.2,0,Drug_A
P003,18.7,1,Placebo
Output Files
km_curve.png: Kaplan-Meier survival curve
km_curve.pdf: Vector version for publications
survival_stats.csv: Statistical summary (median survival, confidence intervals)
hazard_ratios.csv: Cox regression results with HR and 95% CI
- `logrank_test.csv**: Pairwise comparison p-values
- `report.txt**: Human-readable summary report
Technical Details
Statistical Methods
Kaplan-Meier Estimator: Non-parametric maximum likelihood estimate of survival function
- Product-limit estimator: Ŝ(t) = Π(tᵢ≤t) (1 - dᵢ/nᵢ)
- Greenwood's formula for variance estimation
Log-Rank Test: Most widely used test for comparing survival curves
- Null hypothesis: No difference between groups
- Weighted by number at risk at each event time
Cox Proportional Hazards: Semi-parametric regression model
- h(t|X) = h₀(t) × exp(β₁X₁ + β₂X₂ + ...)
- Proportional hazards assumption checked via Schoenfeld residuals
Dependencies
lifelines: Core survival analysis library
matplotlib, seaborn: Visualization
pandas, numpy: Data handling
scipy: Statistical tests
Technical Difficulty: High ⚠️
This skill involves advanced statistical modeling. Results should be reviewed by a biostatistician, especially for:
- Proportional hazards assumption violations
- Small sample sizes (< 30 per group)
- Heavy censoring (> 50%)
- Time-varying covariates
References
See references/ folder for:
- Kaplan EL, Meier P (1958) original paper
- Cox DR (1972) regression models paper
- Sample datasets for testing
- Clinical reporting guidelines (ATN, CONSORT)
Parameters
| Parameter |
Type |
Default |
Description |
--input |
str |
Required |
Input CSV file path |
--time |
str |
Required |
Column name for survival time |
--event |
str |
Required |
|
--group |
str |
Required |
|
--output |
str |
Required |
Output directory for results |
--conf-level |
float |
0.95 |
|
--risk-table |
str |
Required |
Include risk table in plot |
--figsize |
str |
'10 |
|
--dpi |
int |
300 |
|
Example
# Basic survival curve
python scripts/main.py \
--input clinical_data.csv \
--time overall_survival_months \
--event death \
--group treatment_arm \
--output ./results/ \
--risk-table
Output includes:
- Survival curves with 95% confidence bands
- Median survival: Drug A = 28.4 months (95% CI: 24.1-32.7), Placebo = 18.2 months (95% CI: 15.3-21.1)
- Log-rank test p-value: 0.0023
- Hazard ratio: 0.62 (95% CI: 0.45-0.85), p = 0.003
Risk Assessment
| Risk Indicator |
Assessment |
Level |
| Code Execution |
Python/R scripts executed locally |
Medium |
| Network Access |
No external API calls |
Low |
| File System Access |
Read input files, write output files |
Medium |
| Instruction Tampering |
Standard prompt guidelines |
Low |
| Data Exposure |
Output files saved to workspace |
Low |
Security Checklist
Prerequisites
# Python dependencies
pip install -r requirements.txt
Evaluation Criteria
Success Metrics
Test Cases
- Basic Functionality: Standard input → Expected output
- Edge Case: Invalid input → Graceful error handling
- Performance: Large dataset → Acceptable processing time
Lifecycle Status
- Current Stage: Draft
- Next Review Date: 2026-03-06
- Known Issues: None
- Planned Improvements:
- Performance optimization
- Additional feature support
1---2name: survival-analysis-km3description: Generates Kaplan-Meier survival curves, calculates survival statistics (log-rank test, median survival time), and estimates hazard ratios for clinical and biological survival data analysis. Triggered when user requests survival analysis, Kaplan-Meier plots, time-to-event analysis, or asks about survival statistics in biomedical contexts.4license: MIT5---6
7# Survival Analysis (Kaplan-Meier)
8
9Kaplan-Meier survival analysis tool for clinical and biological research. Generates publication-ready survival curves with statistical tests.
10
11## Features
12
13- **Kaplan-Meier Curve Generation**: Publication-quality survival plots with confidence intervals
14- **Statistical Tests**: Log-rank test, Wilcoxon test, Peto-Peto test
15- **Hazard Ratios**: Cox proportional hazards regression with 95% CI
16- **Summary Statistics**: Median survival time, restricted mean survival time (RMST)
17- **Multi-group Analysis**: Supports 2+ comparison groups
18- **Risk Tables**: Optional at-risk table below curves
19
20## Usage
21
22### Python Script
23
24```bash
25python scripts/main.py --input data.csv --time time_col --event event_col --group group_col --output results/
26```
27
28### Arguments
29
30| Argument | Description | Required |
31|----------|-------------|----------|
32| `--input` | Input CSV file path | Yes |
33| `--time` | Column name for survival time | Yes |
34| `--event` | Column name for event indicator (1=event, 0=censored) | Yes |
35| `--group` | Column name for grouping variable | Optional |
36| `--output` | Output directory for results | Yes |
37| `--conf-level` | Confidence level (default: 0.95) | Optional |
38| `--risk-table` | Include risk table in plot | Optional |
39
40### Input Format
41
42CSV with columns:
43- **Time column**: Numeric, time to event or censoring
44- **Event column**: Binary (1 = event occurred, 0 = censored/right-censored)
45- **Group column**: Categorical variable for stratification
46
47Example:
48```csv
49patient_id,time_months,death,treatment_group
50P001,24.5,1,Drug_A
51P002,36.2,0,Drug_A
52P003,18.7,1,Placebo
53```
54
55### Output Files
56
57- `km_curve.png`: Kaplan-Meier survival curve
58- `km_curve.pdf`: Vector version for publications
59- `survival_stats.csv`: Statistical summary (median survival, confidence intervals)
60- `hazard_ratios.csv`: Cox regression results with HR and 95% CI
61- `logrank_test.csv**: Pairwise comparison p-values
62- `report.txt**: Human-readable summary report
63
64## Technical Details
65
66### Statistical Methods
67
681. **Kaplan-Meier Estimator**: Non-parametric maximum likelihood estimate of survival function
69 - Product-limit estimator: Ŝ(t) = Π(tᵢ≤t) (1 - dᵢ/nᵢ)
70 - Greenwood's formula for variance estimation
71
722. **Log-Rank Test**: Most widely used test for comparing survival curves
73 - Null hypothesis: No difference between groups
74 - Weighted by number at risk at each event time
75
763. **Cox Proportional Hazards**: Semi-parametric regression model
77 - h(t|X) = h₀(t) × exp(β₁X₁ + β₂X₂ + ...)
78 - Proportional hazards assumption checked via Schoenfeld residuals
79
80### Dependencies
81
82- `lifelines`: Core survival analysis library
83- `matplotlib`, `seaborn`: Visualization
84- `pandas`, `numpy`: Data handling
85- `scipy`: Statistical tests
86
87### Technical Difficulty: High ⚠️
88
89This skill involves advanced statistical modeling. Results should be reviewed by a biostatistician, especially for:
90- Proportional hazards assumption violations
91- Small sample sizes (< 30 per group)
92- Heavy censoring (> 50%)
93- Time-varying covariates
94
95## References
96
97See `references/` folder for:
98- Kaplan EL, Meier P (1958) original paper
99- Cox DR (1972) regression models paper
100- Sample datasets for testing
101- Clinical reporting guidelines (ATN, CONSORT)
102
103
104## Parameters
105
106| Parameter | Type | Default | Description |
107|-----------|------|---------|-------------|
108| `--input` | str | Required | Input CSV file path |
109| `--time` | str | Required | Column name for survival time |
110| `--event` | str | Required | |
111| `--group` | str | Required | |
112| `--output` | str | Required | Output directory for results |
113| `--conf-level` | float | 0.95 | |
114| `--risk-table` | str | Required | Include risk table in plot |
115| `--figsize` | str | '10 | |
116| `--dpi` | int | 300 | |
117
118## Example
119
120```bash
121# Basic survival curve
122python scripts/main.py \
123 --input clinical_data.csv \
124 --time overall_survival_months \
125 --event death \
126 --group treatment_arm \
127 --output ./results/ \
128 --risk-table
129```
130
131Output includes:
132- Survival curves with 95% confidence bands
133- Median survival: Drug A = 28.4 months (95% CI: 24.1-32.7), Placebo = 18.2 months (95% CI: 15.3-21.1)
134- Log-rank test p-value: 0.0023
135- Hazard ratio: 0.62 (95% CI: 0.45-0.85), p = 0.003
136
137## Risk Assessment
138
139| Risk Indicator | Assessment | Level |
140|----------------|------------|-------|
141| Code Execution | Python/R scripts executed locally | Medium |
142| Network Access | No external API calls | Low |
143| File System Access | Read input files, write output files | Medium |
144| Instruction Tampering | Standard prompt guidelines | Low |
145| Data Exposure | Output files saved to workspace | Low |
146
147## Security Checklist
148
149- [ ] No hardcoded credentials or API keys
150- [ ] No unauthorized file system access (../)
151- [ ] Output does not expose sensitive information
152- [ ] Prompt injection protections in place
153- [ ] Input file paths validated (no ../ traversal)
154- [ ] Output directory restricted to workspace
155- [ ] Script execution in sandboxed environment
156- [ ] Error messages sanitized (no stack traces exposed)
157- [ ] Dependencies audited
158## Prerequisites
159
160```bash
161# Python dependencies
162pip install -r requirements.txt
163```
164
165## Evaluation Criteria
166
167### Success Metrics
168- [ ] Successfully executes main functionality
169- [ ] Output meets quality standards
170- [ ] Handles edge cases gracefully
171- [ ] Performance is acceptable
172
173### Test Cases
1741. **Basic Functionality**: Standard input → Expected output
1752. **Edge Case**: Invalid input → Graceful error handling
1763. **Performance**: Large dataset → Acceptable processing time
177
178## Lifecycle Status
179
180- **Current Stage**: Draft
181- **Next Review Date**: 2026-03-06
182- **Known Issues**: None
183- **Planned Improvements**:
184 - Performance optimization
185 - Additional feature support