Source: https://github.com/aipoch/medical-research-skills
Clinical Data Cleaner
Clean, validate, and standardize clinical trial data to meet CDISC SDTM standards for regulatory submissions to FDA or EMA.
When to Use
- Use this skill when the task needs Use when cleaning clinical trial data, preparing data for FDA/EMA submission, standardizing SDTM datasets, handling missing values in clinical studies, detecting outliers in lab results, or converting raw CRF data to CDISC format. Cleans and standardizes clinical trial data for regulatory compliance with audit trails.
- Use this skill for data analysis tasks that require explicit assumptions, bounded scope, and a reproducible output format.
- Use this skill when you need a documented fallback path for missing inputs, execution errors, or partial evidence.
Key Features
- Scope-focused workflow aligned to: Use when cleaning clinical trial data, preparing data for FDA/EMA submission, standardizing SDTM datasets, handling missing values in clinical studies, detecting outliers in lab results, or converting raw CRF data to CDISC format. Cleans and standardizes clinical trial data for regulatory compliance with audit trails.
- Packaged executable path(s):
scripts/main.py.
- Reference material available in
references/ for task-specific guidance.
- Structured execution path designed to keep outputs consistent and reviewable.
Dependencies
Python: 3.10+. Repository baseline for current packaged skills.
numpy: unspecified. Declared in requirements.txt.
pandas: unspecified. Declared in requirements.txt.
scipy: unspecified. Declared in requirements.txt.
Example Usage
cd "20260318/scientific-skills/Data Analytics/clinical-data-cleaner"
python -m py_compile scripts/main.py
python scripts/main.py --help
Example run plan:
- Confirm the user input, output path, and any required config values.
- Edit the in-file
CONFIG block or documented parameters if the script uses fixed settings.
- Run
python scripts/main.py with the validated inputs.
- Review the generated output and return the final artifact with any assumptions called out.
Implementation Details
See ## Workflow above for related details.
- Execution model: validate the request, choose the packaged workflow, and produce a bounded deliverable.
- Input controls: confirm the source files, scope limits, output format, and acceptance criteria before running any script.
- Primary implementation surface:
scripts/main.py.
- Reference guidance:
references/ contains supporting rules, prompts, or checklists.
- Parameters to clarify first: input path, output path, scope filters, thresholds, and any domain-specific constraints.
- Output discipline: keep results reproducible, identify assumptions explicitly, and avoid undocumented side effects.
Quick Check
Use this command to verify that the packaged script entry point can be parsed before deeper execution.
python -m py_compile scripts/main.py
Audit-Ready Commands
Use these concrete commands for validation. They are intentionally self-contained and avoid placeholder paths.
python -m py_compile scripts/main.py
python scripts/main.py --help
python scripts/main.py --input "Audit validation sample with explicit symptoms, history, assessment, and next-step plan."
Workflow
- Confirm the user objective, required inputs, and non-negotiable constraints before doing detailed work.
- Validate that the request matches the documented scope and stop early if the task would require unsupported assumptions.
- Use the packaged script path or the documented reasoning path with only the inputs that are actually available.
- Return a structured result that separates assumptions, deliverables, risks, and unresolved items.
- If execution fails or inputs are incomplete, switch to the fallback path and state exactly what blocked full completion.
Quick Start
from scripts.main import ClinicalDataCleaner
# Initialize for Demographics domain
cleaner = ClinicalDataCleaner(domain='DM')
# Clean data with default settings
cleaned = cleaner.clean(raw_data)
# Save with audit trail
cleaner.save_report('output.csv')
Core Capabilities
1. SDTM Domain Validation
cleaner = ClinicalDataCleaner(domain='DM') # or 'LB', 'VS'
is_valid, missing = cleaner.validate_domain(data)
Required Fields:
- DM: STUDYID, USUBJID, SUBJID, RFSTDTC, RFENDTC, SITEID, AGE, SEX, RACE
- LB: STUDYID, USUBJID, LBTESTCD, LBCAT, LBORRES, LBORRESU, LBSTRESC, LBDTC
- VS: STUDYID, USUBJID, VSTESTCD, VSORRES, VSORRESU, VSSTRESC, VSDTC
2. Missing Value Handling
cleaner = ClinicalDataCleaner(
domain='DM',
missing_strategy='median' # mean, median, mode, forward, drop
)
cleaned = cleaner.handle_missing_values(data)
3. Outlier Detection
cleaner = ClinicalDataCleaner(
domain='LB',
outlier_method='domain', # iqr, zscore, domain
outlier_action='flag' # flag, remove, cap
)
flagged = cleaner.detect_outliers(data)
Clinical Thresholds:
| Parameter |
Range |
Unit |
| Glucose |
50-500 |
mg/dL |
| Hemoglobin |
5-20 |
g/dL |
| Systolic BP |
70-220 |
mmHg |
4. Date Standardization
standardized = cleaner.standardize_dates(data)
# Converts to ISO 8601: 2023-01-15T09:30:00
5. Complete Pipeline
cleaner = ClinicalDataCleaner(
domain='DM',
missing_strategy='median',
outlier_method='iqr',
outlier_action='flag'
)
cleaned_data = cleaner.clean(data)
cleaner.save_report('output.csv')
Output Files:
output.csv - Cleaned SDTM data
output.report.json - Audit trail for regulatory submission
CLI Usage
# Clean demographics
python scripts/main.py \
--input dm_raw.csv \
--domain DM \
--output dm_clean.csv \
--missing-strategy median \
--outlier-method iqr \
--outlier-action flag
# Clean lab data with clinical thresholds
python scripts/main.py \
--input lb_raw.csv \
--domain LB \
--output lb_clean.csv \
--outlier-method domain
Common Patterns
See references/common-patterns.md for detailed examples:
- Regulatory Submission Preparation
- Interim Analysis Data Preparation
- Database Migration Cleanup
- External Lab Data Integration
Troubleshooting
See references/troubleshooting.md for solutions to:
- Validation failures
- Date parsing errors
- Memory errors with large datasets
- Outlier detection issues
Quality Checklist
Pre-Cleaning:
Post-Cleaning:
References
references/sdtm_ig_guide.md - CDISC SDTM Implementation Guide
references/domain_specs.json - Domain-specific field requirements
references/outlier_thresholds.json - Clinical outlier thresholds
references/common-patterns.md - Detailed usage patterns
references/troubleshooting.md - Problem-solving guide
Skill ID: 189 | Version: 2.0 | License: MIT
Output Requirements
Every final response should make these items explicit when they are relevant:
- Objective or requested deliverable
- Inputs used and assumptions introduced
- Workflow or decision path
- Core result, recommendation, or artifact
- Constraints, risks, caveats, or validation needs
- Unresolved items and next-step checks
Error Handling
- If required inputs are missing, state exactly which fields are missing and request only the minimum additional information.
- If the task goes outside the documented scope, stop instead of guessing or silently widening the assignment.
- If
scripts/main.py fails, report the failure point, summarize what still can be completed safely, and provide a manual fallback.
- Do not fabricate files, citations, data, search results, or execution outcomes.
Input Validation
This skill accepts requests that match the documented purpose of clinical-data-cleaner and include enough context to complete the workflow safely.
Do not continue the workflow when the request is out of scope, missing a critical input, or would require unsupported assumptions. Instead respond:
clinical-data-cleaner only handles its documented workflow. Please provide the missing required inputs or switch to a more suitable skill.
Response Template
Use the following fixed structure for non-trivial requests:
- Objective
- Inputs Received
- Assumptions
- Workflow
- Deliverable
- Risks and Limits
- Next Checks
If the request is simple, you may compress the structure, but still keep assumptions and limits explicit when they affect correctness.
1---2name: clinical-data-cleaner3description: Use when cleaning clinical trial data, preparing data for FDA/EMA submission, standardizing SDTM datasets, handling missing values in clinical studies, detecting outliers in lab results, or converting raw CRF data to CDISC format. Cleans and standardizes clinical trial data for regulatory compliance with audit trails.4license: MIT5---6> **Source**: [https://github.com/aipoch/medical-research-skills](https://github.com/aipoch/medical-research-skills)
7# Clinical Data Cleaner
8
9Clean, validate, and standardize clinical trial data to meet CDISC SDTM standards for regulatory submissions to FDA or EMA.
10
11## When to Use
12
13- Use this skill when the task needs Use when cleaning clinical trial data, preparing data for FDA/EMA submission, standardizing SDTM datasets, handling missing values in clinical studies, detecting outliers in lab results, or converting raw CRF data to CDISC format. Cleans and standardizes clinical trial data for regulatory compliance with audit trails.
14- Use this skill for data analysis tasks that require explicit assumptions, bounded scope, and a reproducible output format.
15- Use this skill when you need a documented fallback path for missing inputs, execution errors, or partial evidence.
16
17## Key Features
18
19- Scope-focused workflow aligned to: Use when cleaning clinical trial data, preparing data for FDA/EMA submission, standardizing SDTM datasets, handling missing values in clinical studies, detecting outliers in lab results, or converting raw CRF data to CDISC format. Cleans and standardizes clinical trial data for regulatory compliance with audit trails.
20- Packaged executable path(s): `scripts/main.py`.
21- Reference material available in `references/` for task-specific guidance.
22- Structured execution path designed to keep outputs consistent and reviewable.
23
24## Dependencies
25
26- `Python`: `3.10+`. Repository baseline for current packaged skills.
27- `numpy`: `unspecified`. Declared in `requirements.txt`.
28- `pandas`: `unspecified`. Declared in `requirements.txt`.
29- `scipy`: `unspecified`. Declared in `requirements.txt`.
30
31## Example Usage
32
33```bash
34cd "20260318/scientific-skills/Data Analytics/clinical-data-cleaner"
35python -m py_compile scripts/main.py
36python scripts/main.py --help
37```
38
39Example run plan:
401. Confirm the user input, output path, and any required config values.
412. Edit the in-file `CONFIG` block or documented parameters if the script uses fixed settings.
423. Run `python scripts/main.py` with the validated inputs.
434. Review the generated output and return the final artifact with any assumptions called out.
44
45## Implementation Details
46
47See `## Workflow` above for related details.
48
49- Execution model: validate the request, choose the packaged workflow, and produce a bounded deliverable.
50- Input controls: confirm the source files, scope limits, output format, and acceptance criteria before running any script.
51- Primary implementation surface: `scripts/main.py`.
52- Reference guidance: `references/` contains supporting rules, prompts, or checklists.
53- Parameters to clarify first: input path, output path, scope filters, thresholds, and any domain-specific constraints.
54- Output discipline: keep results reproducible, identify assumptions explicitly, and avoid undocumented side effects.
55
56## Quick Check
57
58Use this command to verify that the packaged script entry point can be parsed before deeper execution.
59
60```bash
61python -m py_compile scripts/main.py
62```
63
64## Audit-Ready Commands
65
66Use these concrete commands for validation. They are intentionally self-contained and avoid placeholder paths.
67
68```bash
69python -m py_compile scripts/main.py
70python scripts/main.py --help
71python scripts/main.py --input "Audit validation sample with explicit symptoms, history, assessment, and next-step plan."
72```
73
74## Workflow
75
761. Confirm the user objective, required inputs, and non-negotiable constraints before doing detailed work.
772. Validate that the request matches the documented scope and stop early if the task would require unsupported assumptions.
783. Use the packaged script path or the documented reasoning path with only the inputs that are actually available.
794. Return a structured result that separates assumptions, deliverables, risks, and unresolved items.
805. If execution fails or inputs are incomplete, switch to the fallback path and state exactly what blocked full completion.
81
82## Quick Start
83
84```python
85from scripts.main import ClinicalDataCleaner
86
87# Initialize for Demographics domain
88cleaner = ClinicalDataCleaner(domain='DM')
89
90# Clean data with default settings
91cleaned = cleaner.clean(raw_data)
92
93# Save with audit trail
94cleaner.save_report('output.csv')
95```
96
97## Core Capabilities
98
99### 1. SDTM Domain Validation
100
101```python
102cleaner = ClinicalDataCleaner(domain='DM') # or 'LB', 'VS'
103is_valid, missing = cleaner.validate_domain(data)
104```
105
106**Required Fields:**
107- **DM**: STUDYID, USUBJID, SUBJID, RFSTDTC, RFENDTC, SITEID, AGE, SEX, RACE
108- **LB**: STUDYID, USUBJID, LBTESTCD, LBCAT, LBORRES, LBORRESU, LBSTRESC, LBDTC
109- **VS**: STUDYID, USUBJID, VSTESTCD, VSORRES, VSORRESU, VSSTRESC, VSDTC
110
111### 2. Missing Value Handling
112
113```python
114cleaner = ClinicalDataCleaner(
115 domain='DM',
116 missing_strategy='median' # mean, median, mode, forward, drop
117)
118cleaned = cleaner.handle_missing_values(data)
119```
120
121### 3. Outlier Detection
122
123```python
124cleaner = ClinicalDataCleaner(
125 domain='LB',
126 outlier_method='domain', # iqr, zscore, domain
127 outlier_action='flag' # flag, remove, cap
128)
129flagged = cleaner.detect_outliers(data)
130```
131
132**Clinical Thresholds:**
133| Parameter | Range | Unit |
134|-----------|-------|------|
135| Glucose | 50-500 | mg/dL |
136| Hemoglobin | 5-20 | g/dL |
137| Systolic BP | 70-220 | mmHg |
138
139### 4. Date Standardization
140
141```python
142standardized = cleaner.standardize_dates(data)
143
144# Converts to ISO 8601: 2023-01-15T09:30:00
145```
146
147### 5. Complete Pipeline
148
149```python
150cleaner = ClinicalDataCleaner(
151 domain='DM',
152 missing_strategy='median',
153 outlier_method='iqr',
154 outlier_action='flag'
155)
156cleaned_data = cleaner.clean(data)
157cleaner.save_report('output.csv')
158```
159
160**Output Files:**
161- `output.csv` - Cleaned SDTM data
162- `output.report.json` - Audit trail for regulatory submission
163
164## CLI Usage
165
166```text
167
168# Clean demographics
169python scripts/main.py \
170 --input dm_raw.csv \
171 --domain DM \
172 --output dm_clean.csv \
173 --missing-strategy median \
174 --outlier-method iqr \
175 --outlier-action flag
176
177# Clean lab data with clinical thresholds
178python scripts/main.py \
179 --input lb_raw.csv \
180 --domain LB \
181 --output lb_clean.csv \
182 --outlier-method domain
183```
184
185## Common Patterns
186
187See [references/common-patterns.md](references/common-patterns.md) for detailed examples:
188- Regulatory Submission Preparation
189- Interim Analysis Data Preparation
190- Database Migration Cleanup
191- External Lab Data Integration
192
193## Troubleshooting
194
195See [references/troubleshooting.md](references/troubleshooting.md) for solutions to:
196- Validation failures
197- Date parsing errors
198- Memory errors with large datasets
199- Outlier detection issues
200
201## Quality Checklist
202
203**Pre-Cleaning:**
204- [ ] IACUC approval obtained (animal studies)
205- [ ] Sample size adequately powered
206- [ ] Randomization method documented
207
208**Post-Cleaning:**
209- [ ] Validate against CDISC SDTM IG
210- [ ] Review all cleaning actions in audit trail
211- [ ] Test import to analysis software
212
213## References
214
215- `references/sdtm_ig_guide.md` - CDISC SDTM Implementation Guide
216- `references/domain_specs.json` - Domain-specific field requirements
217- `references/outlier_thresholds.json` - Clinical outlier thresholds
218- `references/common-patterns.md` - Detailed usage patterns
219- `references/troubleshooting.md` - Problem-solving guide
220
221---
222
223**Skill ID**: 189 | **Version**: 2.0 | **License**: MIT
224
225## Output Requirements
226
227Every final response should make these items explicit when they are relevant:
228
229- Objective or requested deliverable
230- Inputs used and assumptions introduced
231- Workflow or decision path
232- Core result, recommendation, or artifact
233- Constraints, risks, caveats, or validation needs
234- Unresolved items and next-step checks
235
236## Error Handling
237
238- If required inputs are missing, state exactly which fields are missing and request only the minimum additional information.
239- If the task goes outside the documented scope, stop instead of guessing or silently widening the assignment.
240- If `scripts/main.py` fails, report the failure point, summarize what still can be completed safely, and provide a manual fallback.
241- Do not fabricate files, citations, data, search results, or execution outcomes.
242
243## Input Validation
244
245This skill accepts requests that match the documented purpose of `clinical-data-cleaner` and include enough context to complete the workflow safely.
246
247Do not continue the workflow when the request is out of scope, missing a critical input, or would require unsupported assumptions. Instead respond:
248
249> `clinical-data-cleaner` only handles its documented workflow. Please provide the missing required inputs or switch to a more suitable skill.
250
251## Response Template
252
253Use the following fixed structure for non-trivial requests:
254
2551. Objective
2562. Inputs Received
2573. Assumptions
2584. Workflow
2595. Deliverable
2606. Risks and Limits
2617. Next Checks
262
263If the request is simple, you may compress the structure, but still keep assumptions and limits explicit when they affect correctness.