Clinical Data Cleaner
Clean, validate, and standardize clinical trial data to meet CDISC SDTM standards for regulatory submissions to FDA or EMA.
When to Use
- Use this skill when the task needs Use when cleaning clinical trial data, preparing data for FDA/EMA submission, standardizing SDTM datasets, handling missing values in clinical studies, detecting outliers in lab results, or converting raw CRF data to CDISC format. Cleans and standardizes clinical trial data for regulatory compliance with audit trails.
- Use this skill for data analysis tasks that require explicit assumptions, bounded scope, and a reproducible output format.
- Use this skill when you need a documented fallback path for missing inputs, execution errors, or partial evidence.
Key Features
- Scope-focused workflow aligned to: Use when cleaning clinical trial data, preparing data for FDA/EMA submission, standardizing SDTM datasets, handling missing values in clinical studies, detecting outliers in lab results, or converting raw CRF data to CDISC format. Cleans and standardizes clinical trial data for regulatory compliance with audit trails.
- Packaged executable path(s):
scripts/main.py.
- Reference material available in
references/ for task-specific guidance.
- Structured execution path designed to keep outputs consistent and reviewable.
Dependencies
Python: 3.10+. Repository baseline for current packaged skills.
numpy: unspecified. Declared in requirements.txt.
pandas: unspecified. Declared in requirements.txt.
scipy: unspecified. Declared in requirements.txt.
Example Usage
cd "20260318/scientific-skills/Data Analytics/clinical-data-cleaner"
python -m py_compile scripts/main.py
python scripts/main.py --help
Example run plan:
- Confirm the user input, output path, and any required config values.
- Edit the in-file
CONFIG block or documented parameters if the script uses fixed settings.
- Run
python scripts/main.py with the validated inputs.
- Review the generated output and return the final artifact with any assumptions called out.
Implementation Details
See ## Workflow above for related details.
- Execution model: validate the request, choose the packaged workflow, and produce a bounded deliverable.
- Input controls: confirm the source files, scope limits, output format, and acceptance criteria before running any script.
- Primary implementation surface:
scripts/main.py.
- Reference guidance:
references/ contains supporting rules, prompts, or checklists.
- Parameters to clarify first: input path, output path, scope filters, thresholds, and any domain-specific constraints.
- Output discipline: keep results reproducible, identify assumptions explicitly, and avoid undocumented side effects.
Quick Check
Use this command to verify that the packaged script entry point can be parsed before deeper execution.
python -m py_compile scripts/main.py
Audit-Ready Commands
Use these concrete commands for validation. They are intentionally self-contained and avoid placeholder paths.
python -m py_compile scripts/main.py
python scripts/main.py --help
python scripts/main.py --input "Audit validation sample with explicit symptoms, history, assessment, and next-step plan."
Workflow
- Confirm the user objective, required inputs, and non-negotiable constraints before doing detailed work.
- Validate that the request matches the documented scope and stop early if the task would require unsupported assumptions.
- Use the packaged script path or the documented reasoning path with only the inputs that are actually available.
- Return a structured result that separates assumptions, deliverables, risks, and unresolved items.
- If execution fails or inputs are incomplete, switch to the fallback path and state exactly what blocked full completion.
Quick Start
from scripts.main import ClinicalDataCleaner
# Initialize for Demographics domain
cleaner = ClinicalDataCleaner(domain='DM')
# Clean data with default settings
cleaned = cleaner.clean(raw_data)
# Save with audit trail
cleaner.save_report('output.csv')
Core Capabilities
1. SDTM Domain Validation
cleaner = ClinicalDataCleaner(domain='DM') # or 'LB', 'VS'
is_valid, missing = cleaner.validate_domain(data)
Required Fields:
- DM: STUDYID, USUBJID, SUBJID, RFSTDTC, RFENDTC, SITEID, AGE, SEX, RACE
- LB: STUDYID, USUBJID, LBTESTCD, LBCAT, LBORRES, LBORRESU, LBSTRESC, LBDTC
- VS: STUDYID, USUBJID, VSTESTCD, VSORRES, VSORRESU, VSSTRESC, VSDTC
2. Missing Value Handling
cleaner = ClinicalDataCleaner(
domain='DM',
missing_strategy='median' # mean, median, mode, forward, drop
)
cleaned = cleaner.handle_missing_values(data)
3. Outlier Detection
cleaner = ClinicalDataCleaner(
domain='LB',
outlier_method='domain', # iqr, zscore, domain
outlier_action='flag' # flag, remove, cap
)
flagged = cleaner.detect_outliers(data)
Clinical Thresholds:
| Parameter |
Range |
Unit |
| Glucose |
50-500 |
mg/dL |
| Hemoglobin |
5-20 |
g/dL |
| Systolic BP |
70-220 |
mmHg |
4. Date Standardization
standardized = cleaner.standardize_dates(data)
# Converts to ISO 8601: 2023-01-15T09:30:00
5. Complete Pipeline
cleaner = ClinicalDataCleaner(
domain='DM',
missing_strategy='median',
outlier_method='iqr',
outlier_action='flag'
)
cleaned_data = cleaner.clean(data)
cleaner.save_report('output.csv')
Output Files:
output.csv - Cleaned SDTM data
output.report.json - Audit trail for regulatory submission
CLI Usage
# Clean demographics
python scripts/main.py \
--input dm_raw.csv \
--domain DM \
--output dm_clean.csv \
--missing-strategy median \
--outlier-method iqr \
--outlier-action flag
# Clean lab data with clinical thresholds
python scripts/main.py \
--input lb_raw.csv \
--domain LB \
--output lb_clean.csv \
--outlier-method domain
Common Patterns
See references/common-patterns.md for detailed examples:
- Regulatory Submission Preparation
- Interim Analysis Data Preparation
- Database Migration Cleanup
- External Lab Data Integration
Troubleshooting
See references/troubleshooting.md for solutions to:
- Validation failures
- Date parsing errors
- Memory errors with large datasets
- Outlier detection issues
Quality Checklist
Pre-Cleaning:
Post-Cleaning:
References
references/sdtm_ig_guide.md - CDISC SDTM Implementation Guide
references/domain_specs.json - Domain-specific field requirements
references/outlier_thresholds.json - Clinical outlier thresholds
references/common-patterns.md - Detailed usage patterns
references/troubleshooting.md - Problem-solving guide
Skill ID: 189 | Version: 2.0 | License: MIT
Output Requirements
Every final response should make these items explicit when they are relevant:
- Objective or requested deliverable
- Inputs used and assumptions introduced
- Workflow or decision path
- Core result, recommendation, or artifact
- Constraints, risks, caveats, or validation needs
- Unresolved items and next-step checks
Error Handling
- If required inputs are missing, state exactly which fields are missing and request only the minimum additional information.
- If the task goes outside the documented scope, stop instead of guessing or silently widening the assignment.
- If
scripts/main.py fails, report the failure point, summarize what still can be completed safely, and provide a manual fallback.
- Do not fabricate files, citations, data, search results, or execution outcomes.
Input Validation
This skill accepts requests that match the documented purpose of clinical-data-cleaner and include enough context to complete the workflow safely.
Do not continue the workflow when the request is out of scope, missing a critical input, or would require unsupported assumptions. Instead respond:
clinical-data-cleaner only handles its documented workflow. Please provide the missing required inputs or switch to a more suitable skill.
Response Template
Use the following fixed structure for non-trivial requests:
- Objective
- Inputs Received
- Assumptions
- Workflow
- Deliverable
- Risks and Limits
- Next Checks
If the request is simple, you may compress the structure, but still keep assumptions and limits explicit when they affect correctness.
1---2name: clinical-data-cleaner3description: Use when cleaning clinical trial data, preparing data for FDA/EMA submission, standardizing SDTM datasets, handling missing values in clinical studies, detecting outliers in lab results, or converting raw CRF data to CDISC format. Cleans and standardizes clinical trial data for regulatory compliance with audit trails.4license: MIT5---6# Clinical Data Cleaner78Clean, validate, and standardize clinical trial data to meet CDISC SDTM standards for regulatory submissions to FDA or EMA.910## When to Use1112- Use this skill when the task needs Use when cleaning clinical trial data, preparing data for FDA/EMA submission, standardizing SDTM datasets, handling missing values in clinical studies, detecting outliers in lab results, or converting raw CRF data to CDISC format. Cleans and standardizes clinical trial data for regulatory compliance with audit trails.13- Use this skill for data analysis tasks that require explicit assumptions, bounded scope, and a reproducible output format.14- Use this skill when you need a documented fallback path for missing inputs, execution errors, or partial evidence.1516## Key Features1718- Scope-focused workflow aligned to: Use when cleaning clinical trial data, preparing data for FDA/EMA submission, standardizing SDTM datasets, handling missing values in clinical studies, detecting outliers in lab results, or converting raw CRF data to CDISC format. Cleans and standardizes clinical trial data for regulatory compliance with audit trails.19- Packaged executable path(s): `scripts/main.py`.20- Reference material available in `references/` for task-specific guidance.21- Structured execution path designed to keep outputs consistent and reviewable.2223## Dependencies2425- `Python`: `3.10+`. Repository baseline for current packaged skills.26- `numpy`: `unspecified`. Declared in `requirements.txt`.27- `pandas`: `unspecified`. Declared in `requirements.txt`.28- `scipy`: `unspecified`. Declared in `requirements.txt`.2930## Example Usage3132```bash33cd "20260318/scientific-skills/Data Analytics/clinical-data-cleaner"34python -m py_compile scripts/main.py35python scripts/main.py --help36```3738Example run plan:391. Confirm the user input, output path, and any required config values.402. Edit the in-file `CONFIG` block or documented parameters if the script uses fixed settings.413. Run `python scripts/main.py` with the validated inputs.424. Review the generated output and return the final artifact with any assumptions called out.4344## Implementation Details4546See `## Workflow` above for related details.4748- Execution model: validate the request, choose the packaged workflow, and produce a bounded deliverable.49- Input controls: confirm the source files, scope limits, output format, and acceptance criteria before running any script.50- Primary implementation surface: `scripts/main.py`.51- Reference guidance: `references/` contains supporting rules, prompts, or checklists.52- Parameters to clarify first: input path, output path, scope filters, thresholds, and any domain-specific constraints.53- Output discipline: keep results reproducible, identify assumptions explicitly, and avoid undocumented side effects.5455## Quick Check5657Use this command to verify that the packaged script entry point can be parsed before deeper execution.5859```bash60python -m py_compile scripts/main.py61```6263## Audit-Ready Commands6465Use these concrete commands for validation. They are intentionally self-contained and avoid placeholder paths.6667```bash68python -m py_compile scripts/main.py69python scripts/main.py --help70python scripts/main.py --input "Audit validation sample with explicit symptoms, history, assessment, and next-step plan."71```7273## Workflow74751. Confirm the user objective, required inputs, and non-negotiable constraints before doing detailed work.762. Validate that the request matches the documented scope and stop early if the task would require unsupported assumptions.773. Use the packaged script path or the documented reasoning path with only the inputs that are actually available.784. Return a structured result that separates assumptions, deliverables, risks, and unresolved items.795. If execution fails or inputs are incomplete, switch to the fallback path and state exactly what blocked full completion.8081## Quick Start8283```python84from scripts.main import ClinicalDataCleaner8586# Initialize for Demographics domain87cleaner = ClinicalDataCleaner(domain='DM')8889# Clean data with default settings90cleaned = cleaner.clean(raw_data)9192# Save with audit trail93cleaner.save_report('output.csv')94```9596## Core Capabilities9798### 1. SDTM Domain Validation99100```python101cleaner = ClinicalDataCleaner(domain='DM') # or 'LB', 'VS'102is_valid, missing = cleaner.validate_domain(data)103```104105**Required Fields:**106- **DM**: STUDYID, USUBJID, SUBJID, RFSTDTC, RFENDTC, SITEID, AGE, SEX, RACE107- **LB**: STUDYID, USUBJID, LBTESTCD, LBCAT, LBORRES, LBORRESU, LBSTRESC, LBDTC108- **VS**: STUDYID, USUBJID, VSTESTCD, VSORRES, VSORRESU, VSSTRESC, VSDTC109110### 2. Missing Value Handling111112```python113cleaner = ClinicalDataCleaner(114 domain='DM',115 missing_strategy='median' # mean, median, mode, forward, drop116)117cleaned = cleaner.handle_missing_values(data)118```119120### 3. Outlier Detection121122```python123cleaner = ClinicalDataCleaner(124 domain='LB',125 outlier_method='domain', # iqr, zscore, domain126 outlier_action='flag' # flag, remove, cap127)128flagged = cleaner.detect_outliers(data)129```130131**Clinical Thresholds:**132| Parameter | Range | Unit |133|-----------|-------|------|134| Glucose | 50-500 | mg/dL |135| Hemoglobin | 5-20 | g/dL |136| Systolic BP | 70-220 | mmHg |137138### 4. Date Standardization139140```python141standardized = cleaner.standardize_dates(data)142143# Converts to ISO 8601: 2023-01-15T09:30:00144```145146### 5. Complete Pipeline147148```python149cleaner = ClinicalDataCleaner(150 domain='DM',151 missing_strategy='median',152 outlier_method='iqr',153 outlier_action='flag'154)155cleaned_data = cleaner.clean(data)156cleaner.save_report('output.csv')157```158159**Output Files:**160- `output.csv` - Cleaned SDTM data161- `output.report.json` - Audit trail for regulatory submission162163## CLI Usage164165```text166167# Clean demographics168python scripts/main.py \169 --input dm_raw.csv \170 --domain DM \171 --output dm_clean.csv \172 --missing-strategy median \173 --outlier-method iqr \174 --outlier-action flag175176# Clean lab data with clinical thresholds177python scripts/main.py \178 --input lb_raw.csv \179 --domain LB \180 --output lb_clean.csv \181 --outlier-method domain182```183184## Common Patterns185186See [references/common-patterns.md](references/common-patterns.md) for detailed examples:187- Regulatory Submission Preparation188- Interim Analysis Data Preparation189- Database Migration Cleanup190- External Lab Data Integration191192## Troubleshooting193194See [references/troubleshooting.md](references/troubleshooting.md) for solutions to:195- Validation failures196- Date parsing errors197- Memory errors with large datasets198- Outlier detection issues199200## Quality Checklist201202**Pre-Cleaning:**203- [ ] IACUC approval obtained (animal studies)204- [ ] Sample size adequately powered205- [ ] Randomization method documented206207**Post-Cleaning:**208- [ ] Validate against CDISC SDTM IG209- [ ] Review all cleaning actions in audit trail210- [ ] Test import to analysis software211212## References213214- `references/sdtm_ig_guide.md` - CDISC SDTM Implementation Guide215- `references/domain_specs.json` - Domain-specific field requirements216- `references/outlier_thresholds.json` - Clinical outlier thresholds217- `references/common-patterns.md` - Detailed usage patterns218- `references/troubleshooting.md` - Problem-solving guide219220---221222**Skill ID**: 189 | **Version**: 2.0 | **License**: MIT223224## Output Requirements225226Every final response should make these items explicit when they are relevant:227228- Objective or requested deliverable229- Inputs used and assumptions introduced230- Workflow or decision path231- Core result, recommendation, or artifact232- Constraints, risks, caveats, or validation needs233- Unresolved items and next-step checks234235## Error Handling236237- If required inputs are missing, state exactly which fields are missing and request only the minimum additional information.238- If the task goes outside the documented scope, stop instead of guessing or silently widening the assignment.239- If `scripts/main.py` fails, report the failure point, summarize what still can be completed safely, and provide a manual fallback.240- Do not fabricate files, citations, data, search results, or execution outcomes.241242## Input Validation243244This skill accepts requests that match the documented purpose of `clinical-data-cleaner` and include enough context to complete the workflow safely.245246Do not continue the workflow when the request is out of scope, missing a critical input, or would require unsupported assumptions. Instead respond:247248> `clinical-data-cleaner` only handles its documented workflow. Please provide the missing required inputs or switch to a more suitable skill.249250## Response Template251252Use the following fixed structure for non-trivial requests:2532541. Objective2552. Inputs Received2563. Assumptions2574. Workflow2585. Deliverable2596. Risks and Limits2607. Next Checks261262If the request is simple, you may compress the structure, but still keep assumptions and limits explicit when they affect correctness.