ClinicalTrials.gov Disease Landscape Scanner
When to Use This Skill
- Map competitive landscape across therapeutic mechanisms for any disease
- Track specific mechanism classes (e.g., anti-IL23, anti-TL1A, JAK inhibitors)
- Identify sponsors and their pipeline positions by phase
- Phase distribution analysis for business development diligence
- Pipeline monitoring for a specific sponsor's disease portfolio
- Pre-built disease configs available (IBD with 14 mechanism classes); generic mode for any other disease
Do NOT use for:
- Detailed single-trial protocol analysis
- Efficacy/safety comparisons (requires literature review skill)
Installation
| Software |
Version |
License |
Commercial Use |
Installation |
| pandas |
≥1.3 |
BSD-3 |
✅ Permitted |
pip install pandas |
| requests |
≥2.25 |
Apache-2.0 |
✅ Permitted |
pip install requests |
| numpy |
≥1.20 |
BSD-3 |
✅ Permitted |
pip install numpy |
| plotnine |
≥0.10 |
MIT |
✅ Permitted |
pip install plotnine |
| plotnine-prism |
≥0.3 |
MIT |
✅ Permitted |
pip install plotnine-prism |
| seaborn |
≥0.11 |
BSD-3 |
✅ Permitted |
pip install seaborn |
| matplotlib |
≥3.4 |
PSF |
✅ Permitted |
pip install matplotlib |
| reportlab |
≥3.6 |
BSD |
✅ Permitted |
pip install reportlab |
| pyyaml |
≥5.0 |
MIT |
✅ Permitted |
pip install pyyaml |
pip install pandas requests numpy plotnine plotnine-prism seaborn matplotlib reportlab pyyaml
System requirements: Internet connection for ClinicalTrials.gov API calls.
Inputs
Required:
- Disease / condition terms — list of conditions to search ClinicalTrials.gov
Optional:
- Disease config — pre-built config ID (e.g.,
"ibd") for mechanism taxonomy, or None for generic
- Mechanism filter — e.g., "Anti-IL-23 (p19)", "Anti-TL1A", "JAK Inhibitor"
- Sponsor filter — e.g., "Takeda", "AbbVie"
- Status filter — Default: all active (Recruiting + Active not recruiting + Not yet recruiting)
- Phase filter — Phase 1, 2, 3, 4
Outputs
Visualizations (PNG + SVG):
landscape_overview.png/.svg — 6-panel landscape figure (300 DPI)
- Mechanism × Phase heatmap, top sponsors, phase stacked bars, mechanism counts, timeline, sponsor type
landscape_supplementary.png/.svg — 4-panel supplementary figure
- Top 15 countries, study design by phase, enrollment distribution, phase transition funnel
Results (CSV):
trials_all.csv — All trials with 46 columns (mechanism, phase, sponsor, geography, study design, arms, endpoints, eligibility, regulatory)
trials_by_mechanism.csv — Mechanism × phase cross-tabulation
trials_by_sponsor.csv — Sponsor summary with trial counts
trials_filtered.csv — Filtered subset (if mechanism/sponsor filter applied)
Reports:
landscape_report.pdf — Publication-quality PDF with 24 sections: executive summary, mechanism deep-dives, geographic landscape, study design, phase transition funnel, endpoint comparison, combination therapies, biosimilar assessment, whitespace analysis, and more
landscape_report.md — Markdown version with identical 24-section structure
Analysis objects (Pickle):
analysis_object.pkl — Complete landscape for downstream use
- Load with:
import pickle; obj = pickle.load(open('analysis_object.pkl', 'rb'))
- Contains: trials_df (46 columns), mechanism/phase/sponsor distributions, geographic stats, design stats, parameters
Clarification Questions
Data Source (ASK THIS FIRST):
- This skill queries the ClinicalTrials.gov API v2 directly (free, no key needed).
- Use live API data? (recommended, ~30 seconds)
- Or use cached demo data? Pre-loaded IBD landscape snapshot for quick demo
Disease Area:
- Which disease area to analyze?
- a) IBD (Inflammatory Bowel Disease) — pre-built config with 14 mechanism classes
- b) Oncology (generic intervention-type classification)
- c) Autoimmune / Rheumatology (generic classification)
- d) Other (specify disease and condition terms)
Scope (if IBD selected):
- Which conditions?
- a) All IBD (Crohn's, UC, and IBD unspecified) — recommended
- b) Crohn's Disease only
- c) Ulcerative Colitis only
- (If other disease) — Provide list of condition search terms
Focus:
- Any mechanism or sponsor to highlight?
- (IBD) a) Anti-IL-23 — recommended for demo | b) Anti-TL1A | c) All mechanisms
- (Other) Specify or skip highlighting
Standard Workflow
Note: Run from the OmicsClaw root directory and add the workflow scripts to sys.path:
import sys; import os; sys.path.insert(0, os.path.abspath('knowledge_base/scripts/clinicaltrials-landscape'))
🚨 MANDATORY: USE SCRIPTS EXACTLY AS SHOWN - DO NOT WRITE INLINE CODE 🚨
Step 1 — Load config and query ClinicalTrials.gov:
from disease_config import load_disease_config, get_default_conditions
from query_clinicaltrials import query_trials
# Load disease config (use "ibd" for IBD, or None for generic)
config = load_disease_config("ibd")
# Get conditions from config or specify manually
conditions = get_default_conditions(config) or ["Crohn's Disease", "Ulcerative Colitis", "Inflammatory Bowel Disease"]
raw_trials = query_trials(
conditions=conditions,
statuses=["RECRUITING", "ACTIVE_NOT_RECRUITING", "ENROLLING_BY_INVITATION", "NOT_YET_RECRUITING"],
)
✅ VERIFICATION: "✓ Retrieved {N} trials from ClinicalTrials.gov"
Step 2 — Classify and compile:
from classify_mechanisms import classify_all
from compile_trials import compile_trials
classified = classify_all(raw_trials, config=config)
trials_df = compile_trials(classified, output_dir="landscape_results")
DO NOT write inline classification code. The script loads mechanism taxonomy from config.
✅ VERIFICATION: "✓ Trial data compiled successfully!"
Step 3 — Generate visualizations:
from generate_landscape_plots import generate_landscape_plots
generate_landscape_plots(
trials_df,
output_dir="landscape_results",
highlight_mechanism="Anti-IL-23 (p19)", # or None for no highlight
highlight_sponsor=None, # or "Takeda" to highlight
config=config,
)
🚨 DO NOT write inline plotting code. The script handles all 6 panels + PNG/SVG export. 🚨
✅ VERIFICATION: "✓ All landscape visualizations generated successfully!"
Step 4 — Export results:
from export_all import export_all
export_all(
trials_df,
parameters={
"conditions": conditions,
"statuses": ["RECRUITING", "ACTIVE_NOT_RECRUITING", "ENROLLING_BY_INVITATION", "NOT_YET_RECRUITING"],
"highlight_mechanism": "Anti-IL-23 (p19)",
},
output_dir="landscape_results",
config=config,
)
DO NOT write custom export code. Use export_all().
✅ VERIFICATION: "=== Export Complete ==="
⚠️ CRITICAL — DO NOT:
- ❌ Write inline classification code → STOP: Use
classify_all() from scripts
- ❌ Write inline plotting code (ggplot, plt, sns) → STOP: Use
generate_landscape_plots()
- ❌ Write custom export code → STOP: Use
export_all()
- ❌ Try to scrape ClinicalTrials.gov HTML → Use the API via
query_trials()
⚠️ IF SCRIPTS FAIL — Script Failure Hierarchy:
- Fix and Retry (90%) — Install missing package, re-run script
- Modify Script (5%) — Edit the script file itself, document changes
- Use as Reference (4%) — Read script, adapt approach, cite source
- Write from Scratch (1%) — Only if genuinely impossible, explain why
NEVER skip directly to writing inline code without trying the script first.
Common Issues
| Error |
Cause |
Solution |
| ConnectionError / Timeout |
ClinicalTrials.gov unreachable |
Check internet connection; retry after 30 seconds |
| HTTP 429 Too Many Requests |
Rate limit exceeded |
Increase RATE_LIMIT_DELAY in query_clinicaltrials.py |
| ModuleNotFoundError: plotnine |
Missing visualization package |
pip install plotnine plotnine-prism |
| Empty results (0 trials) |
Overly restrictive filters |
Broaden condition/status/phase filters |
| Many "Unclassified" mechanisms |
No disease config or new drugs |
Use a disease config (e.g., "ibd") or update disease_configs/*.yaml |
| SVG export failed |
Missing SVG backend |
Normal — PNG is always generated as fallback |
| Sponsor name variants |
Same company, different names |
Update SPONSOR_NORMALIZATION in compile_trials.py |
| ModuleNotFoundError: yaml |
Missing pyyaml |
pip install pyyaml |
Interpretation Guidelines
- Mechanism classification is based on intervention names and descriptions — some trials with vague descriptions (e.g., "Study Drug") will be classified as "Other Biologic" or "Unclassified"
- Phase 2/3 indicates a combined Phase 2/3 study design
- Sponsor normalization groups subsidiaries under parent company (e.g., Millennium → Takeda)
- Industry vs Academic based on ClinicalTrials.gov
leadSponsor.class field
- The landscape reflects registered trials, not all pipeline programs (pre-IND programs won't appear)
- Disease configs provide curated mechanism taxonomies; without config, classification uses generic intervention types
Suggested Next Steps
- Deep-dive a mechanism — Use
literature-preclinical to review mechanism biology
- Track a sponsor's full pipeline — Use
development-landscape for broader pipeline view
- Biomarker analysis — Use
lasso-biomarker-panel to identify response biomarkers from trial data
- Export to presentation — Use landscape_report.md and plots for stakeholder review
Related Skills
development-landscape — Broader, multi-source pipeline landscape for any target
literature-preclinical — Literature review for mechanism biology
lasso-biomarker-panel — Biomarker discovery from expression data
References
1---2name: clinicaltrials-gov-disease-landscape-scanner3description: ClinicalTrials.gov Disease Landscape Scanner4---56# ClinicalTrials.gov Disease Landscape Scanner78## When to Use This Skill910- **Map competitive landscape** across therapeutic mechanisms for any disease11- **Track specific mechanism classes** (e.g., anti-IL23, anti-TL1A, JAK inhibitors)12- **Identify sponsors** and their pipeline positions by phase13- **Phase distribution analysis** for business development diligence14- **Pipeline monitoring** for a specific sponsor's disease portfolio15- **Pre-built disease configs** available (IBD with 14 mechanism classes); generic mode for any other disease1617**Do NOT use for:**18- Detailed single-trial protocol analysis19- Efficacy/safety comparisons (requires literature review skill)2021---2223## Installation2425| Software | Version | License | Commercial Use | Installation |26|----------|---------|---------|----------------|-------------|27| pandas | ≥1.3 | BSD-3 | ✅ Permitted | `pip install pandas` |28| requests | ≥2.25 | Apache-2.0 | ✅ Permitted | `pip install requests` |29| numpy | ≥1.20 | BSD-3 | ✅ Permitted | `pip install numpy` |30| plotnine | ≥0.10 | MIT | ✅ Permitted | `pip install plotnine` |31| plotnine-prism | ≥0.3 | MIT | ✅ Permitted | `pip install plotnine-prism` |32| seaborn | ≥0.11 | BSD-3 | ✅ Permitted | `pip install seaborn` |33| matplotlib | ≥3.4 | PSF | ✅ Permitted | `pip install matplotlib` |34| reportlab | ≥3.6 | BSD | ✅ Permitted | `pip install reportlab` |35| pyyaml | ≥5.0 | MIT | ✅ Permitted | `pip install pyyaml` |3637```bash38pip install pandas requests numpy plotnine plotnine-prism seaborn matplotlib reportlab pyyaml39```4041**System requirements:** Internet connection for ClinicalTrials.gov API calls.4243---4445## Inputs4647**Required:**48- **Disease / condition terms** — list of conditions to search ClinicalTrials.gov4950**Optional:**51- **Disease config** — pre-built config ID (e.g., `"ibd"`) for mechanism taxonomy, or `None` for generic52- **Mechanism filter** — e.g., "Anti-IL-23 (p19)", "Anti-TL1A", "JAK Inhibitor"53- **Sponsor filter** — e.g., "Takeda", "AbbVie"54- **Status filter** — Default: all active (Recruiting + Active not recruiting + Not yet recruiting)55- **Phase filter** — Phase 1, 2, 3, 45657---5859## Outputs6061**Visualizations (PNG + SVG):**62- `landscape_overview.png/.svg` — 6-panel landscape figure (300 DPI)63 - Mechanism × Phase heatmap, top sponsors, phase stacked bars, mechanism counts, timeline, sponsor type64- `landscape_supplementary.png/.svg` — 4-panel supplementary figure65 - Top 15 countries, study design by phase, enrollment distribution, phase transition funnel6667**Results (CSV):**68- `trials_all.csv` — All trials with 46 columns (mechanism, phase, sponsor, geography, study design, arms, endpoints, eligibility, regulatory)69- `trials_by_mechanism.csv` — Mechanism × phase cross-tabulation70- `trials_by_sponsor.csv` — Sponsor summary with trial counts71- `trials_filtered.csv` — Filtered subset (if mechanism/sponsor filter applied)7273**Reports:**74- `landscape_report.pdf` — Publication-quality PDF with 24 sections: executive summary, mechanism deep-dives, geographic landscape, study design, phase transition funnel, endpoint comparison, combination therapies, biosimilar assessment, whitespace analysis, and more75- `landscape_report.md` — Markdown version with identical 24-section structure7677**Analysis objects (Pickle):**78- `analysis_object.pkl` — Complete landscape for downstream use79 - Load with: `import pickle; obj = pickle.load(open('analysis_object.pkl', 'rb'))`80 - Contains: trials_df (46 columns), mechanism/phase/sponsor distributions, geographic stats, design stats, parameters8182---8384## Clarification Questions85861. **Data Source** (ASK THIS FIRST):87 - This skill queries the ClinicalTrials.gov API v2 directly (free, no key needed).88 - **Use live API data?** (recommended, ~30 seconds)89 - **Or use cached demo data?** Pre-loaded IBD landscape snapshot for quick demo90912. **Disease Area:**92 - Which disease area to analyze?93 - a) IBD (Inflammatory Bowel Disease) — pre-built config with 14 mechanism classes94 - b) Oncology (generic intervention-type classification)95 - c) Autoimmune / Rheumatology (generic classification)96 - d) Other (specify disease and condition terms)97983. **Scope** *(if IBD selected)*:99 - Which conditions?100 - a) All IBD (Crohn's, UC, and IBD unspecified) — recommended101 - b) Crohn's Disease only102 - c) Ulcerative Colitis only103 - *(If other disease)* — Provide list of condition search terms1041054. **Focus:**106 - Any mechanism or sponsor to highlight?107 - *(IBD)* a) Anti-IL-23 — recommended for demo | b) Anti-TL1A | c) All mechanisms108 - *(Other)* Specify or skip highlighting109110---111112## Standard Workflow113114> **Note:** Run from the OmicsClaw root directory and add the workflow scripts to `sys.path`:115> ```python116> import sys; import os; sys.path.insert(0, os.path.abspath('knowledge_base/scripts/clinicaltrials-landscape'))117> ```118119🚨 **MANDATORY: USE SCRIPTS EXACTLY AS SHOWN - DO NOT WRITE INLINE CODE** 🚨120121**Step 1 — Load config and query ClinicalTrials.gov:**122```python123124from disease_config import load_disease_config, get_default_conditions125from query_clinicaltrials import query_trials126127# Load disease config (use "ibd" for IBD, or None for generic)128config = load_disease_config("ibd")129130# Get conditions from config or specify manually131conditions = get_default_conditions(config) or ["Crohn's Disease", "Ulcerative Colitis", "Inflammatory Bowel Disease"]132133raw_trials = query_trials(134 conditions=conditions,135 statuses=["RECRUITING", "ACTIVE_NOT_RECRUITING", "ENROLLING_BY_INVITATION", "NOT_YET_RECRUITING"],136)137```138**✅ VERIFICATION:** `"✓ Retrieved {N} trials from ClinicalTrials.gov"`139140**Step 2 — Classify and compile:**141```python142from classify_mechanisms import classify_all143from compile_trials import compile_trials144145classified = classify_all(raw_trials, config=config)146trials_df = compile_trials(classified, output_dir="landscape_results")147```148**DO NOT write inline classification code. The script loads mechanism taxonomy from config.**149150**✅ VERIFICATION:** `"✓ Trial data compiled successfully!"`151152**Step 3 — Generate visualizations:**153```python154from generate_landscape_plots import generate_landscape_plots155156generate_landscape_plots(157 trials_df,158 output_dir="landscape_results",159 highlight_mechanism="Anti-IL-23 (p19)", # or None for no highlight160 highlight_sponsor=None, # or "Takeda" to highlight161 config=config,162)163```164🚨 **DO NOT write inline plotting code. The script handles all 6 panels + PNG/SVG export.** 🚨165166**✅ VERIFICATION:** `"✓ All landscape visualizations generated successfully!"`167168**Step 4 — Export results:**169```python170from export_all import export_all171172export_all(173 trials_df,174 parameters={175 "conditions": conditions,176 "statuses": ["RECRUITING", "ACTIVE_NOT_RECRUITING", "ENROLLING_BY_INVITATION", "NOT_YET_RECRUITING"],177 "highlight_mechanism": "Anti-IL-23 (p19)",178 },179 output_dir="landscape_results",180 config=config,181)182```183**DO NOT write custom export code. Use export_all().**184185**✅ VERIFICATION:** `"=== Export Complete ==="`186187---188189## ⚠️ CRITICAL — DO NOT:190191- ❌ **Write inline classification code** → **STOP: Use `classify_all()` from scripts**192- ❌ **Write inline plotting code (ggplot, plt, sns)** → **STOP: Use `generate_landscape_plots()`**193- ❌ **Write custom export code** → **STOP: Use `export_all()`**194- ❌ **Try to scrape ClinicalTrials.gov HTML** → **Use the API via `query_trials()`**195196---197198## ⚠️ IF SCRIPTS FAIL — Script Failure Hierarchy:1992001. **Fix and Retry (90%)** — Install missing package, re-run script2012. **Modify Script (5%)** — Edit the script file itself, document changes2023. **Use as Reference (4%)** — Read script, adapt approach, cite source2034. **Write from Scratch (1%)** — Only if genuinely impossible, explain why204205**NEVER skip directly to writing inline code without trying the script first.**206207---208209## Common Issues210211| Error | Cause | Solution |212|-------|-------|----------|213| **ConnectionError / Timeout** | ClinicalTrials.gov unreachable | Check internet connection; retry after 30 seconds |214| **HTTP 429 Too Many Requests** | Rate limit exceeded | Increase `RATE_LIMIT_DELAY` in query_clinicaltrials.py |215| **ModuleNotFoundError: plotnine** | Missing visualization package | `pip install plotnine plotnine-prism` |216| **Empty results (0 trials)** | Overly restrictive filters | Broaden condition/status/phase filters |217| **Many "Unclassified" mechanisms** | No disease config or new drugs | Use a disease config (e.g., `"ibd"`) or update `disease_configs/*.yaml` |218| **SVG export failed** | Missing SVG backend | Normal — PNG is always generated as fallback |219| **Sponsor name variants** | Same company, different names | Update `SPONSOR_NORMALIZATION` in compile_trials.py |220| **ModuleNotFoundError: yaml** | Missing pyyaml | `pip install pyyaml` |221222---223224## Interpretation Guidelines225226- **Mechanism classification** is based on intervention names and descriptions — some trials with vague descriptions (e.g., "Study Drug") will be classified as "Other Biologic" or "Unclassified"227- **Phase 2/3** indicates a combined Phase 2/3 study design228- **Sponsor normalization** groups subsidiaries under parent company (e.g., Millennium → Takeda)229- **Industry vs Academic** based on ClinicalTrials.gov `leadSponsor.class` field230- The landscape reflects **registered trials**, not all pipeline programs (pre-IND programs won't appear)231- **Disease configs** provide curated mechanism taxonomies; without config, classification uses generic intervention types232233---234235## Suggested Next Steps2362371. **Deep-dive a mechanism** — Use `literature-preclinical` to review mechanism biology2382. **Track a sponsor's full pipeline** — Use `development-landscape` for broader pipeline view2393. **Biomarker analysis** — Use `lasso-biomarker-panel` to identify response biomarkers from trial data2404. **Export to presentation** — Use landscape_report.md and plots for stakeholder review241242---243244## Related Skills245246- `development-landscape` — Broader, multi-source pipeline landscape for any target247- `literature-preclinical` — Literature review for mechanism biology248- `lasso-biomarker-panel` — Biomarker discovery from expression data249250---251252## References253254- ClinicalTrials.gov API v2: https://clinicaltrials.gov/data-api/api255- ClinicalTrials.gov: https://clinicaltrials.gov/256- See `references/api-parameters.md` for full API parameter reference257- See `references/mechanisms.md` for mechanism taxonomy details258- See `references/output-schema.md` for output column definitions