name: tooluniverse-multiomic-disease-characterization
description: Comprehensive multi-omics disease characterization integrating genomics, transcriptomics, proteomics, pathway, and therapeutic layers for systems-level understanding. Produces a detailed multi-omics report with quantitative confidence scoring (0-100), cross-layer gene concordance analysis, biomarker candidates, therapeutic opportunities, and mechanistic hypotheses. Uses 80+ ToolUniverse tools across 8 analysis layers. Use when users ask about disease mechanisms, multi-omics analysis, systems biology of disease, biomarker discovery, or therapeutic target identification from a disease perspective.
Multi-Omics Disease Characterization Pipeline
Characterize diseases across multiple molecular layers (genomics, transcriptomics, proteomics, pathways) to provide systems-level understanding of disease mechanisms, identify therapeutic opportunities, and discover biomarker candidates.
KEY PRINCIPLES:
- Report-first approach - Create report file FIRST, then populate progressively
- Disease disambiguation FIRST - Resolve all identifiers before omics analysis
- Layer-by-layer analysis - Systematically cover all omics layers
- Cross-layer integration - Identify genes/targets appearing in multiple layers
- Evidence grading - Grade all evidence as T1 (human/clinical) to T4 (computational)
- Tissue context - Emphasize disease-relevant tissues/organs
- Quantitative scoring - Multi-Omics Confidence Score (0-100)
- Druggable focus - Prioritize targets with therapeutic potential
- Biomarker identification - Highlight diagnostic/prognostic markers
- Mechanistic synthesis - Generate testable hypotheses
- Source references - Every statement must cite tool/database
- Completeness checklist - Mandatory section showing analysis coverage
- English-first queries - Always use English terms in tool calls. Respond in user's language
When to Use This Skill
Apply when users:
- Ask about disease mechanisms across omics layers
- Need multi-omics characterization of a disease
- Want to understand disease at the systems biology level
- Ask "What pathways/genes/proteins are involved in [disease]?"
- Need biomarker discovery for a disease
- Want to identify druggable targets from disease profiling
- Ask for integrated genomics + transcriptomics + proteomics analysis
- Need cross-layer concordance analysis
- Ask about disease network biology / hub genes
NOT for (use other skills instead):
- Single gene/target validation -> Use
tooluniverse-drug-target-validation
- Drug safety profiling -> Use
tooluniverse-adverse-event-detection
- General disease overview -> Use
tooluniverse-disease-research
- Variant interpretation -> Use
tooluniverse-variant-interpretation
- GWAS-specific analysis -> Use
tooluniverse-gwas-* skills
- Pathway-only analysis -> Use
tooluniverse-systems-biology
Input Parameters
| Parameter |
Required |
Description |
Example |
| disease |
Yes |
Disease name, OMIM ID, EFO ID, or MONDO ID |
Alzheimer disease, MONDO_0004975 |
| tissue |
No |
Tissue/organ of interest |
brain, liver, blood |
| focus_layers |
No |
Specific omics layers to emphasize |
genomics, transcriptomics, pathways |
Multi-Omics Confidence Score (0-100)
Score Components
Data Availability (0-40 points):
- Genomics data available (GWAS or rare variants): 10 points
- Transcriptomics data available (DEGs or expression): 10 points
- Protein data available (PPI or expression): 5 points
- Pathway data available (enriched pathways): 10 points
- Clinical/drug data available (approved drugs or trials): 5 points
Evidence Concordance (0-40 points):
- Multi-layer genes (appear in 3+ layers): up to 20 points (2 per gene, max 10 genes)
- Consistent direction (genetics + expression concordant): 10 points
- Pathway-gene concordance (genes found in enriched pathways): 10 points
Evidence Quality (0-20 points):
- Strong genetic evidence (GWAS p < 5e-8): 10 points
- Clinical validation (approved drugs): 10 points
Score Interpretation
| Score |
Tier |
Interpretation |
| 80-100 |
Excellent |
Comprehensive multi-omics coverage, high confidence, strong cross-layer concordance |
| 60-79 |
Good |
Good coverage across most layers, some gaps |
| 40-59 |
Moderate |
Moderate coverage, limited cross-layer integration |
| 0-39 |
Limited |
Limited data, single-layer analysis dominates |
Evidence Grading System
| Tier |
Symbol |
Criteria |
Examples |
| T1 |
[T1] |
Direct human evidence, clinical proof |
FDA-approved drug, GWAS hit (p<5e-8), clinical trial result |
| T2 |
[T2] |
Experimental evidence |
Differential expression (validated), functional screen, mouse KO |
| T3 |
[T3] |
Computational/database evidence |
PPI network, pathway mapping, expression correlation |
| T4 |
[T4] |
Annotation/prediction only |
GO annotation, text-mined association, predicted interaction |
Report Template
Create this file structure at the start: {disease_name}_multiomic_report.md
# Multi-Omics Disease Characterization: {Disease Name}
**Report Generated**: {date}
**Disease Identifiers**: (to be filled)
**Multi-Omics Confidence Score**: (to be calculated)
---
## Executive Summary
(2-3 sentence disease mechanism synthesis - fill after all layers complete)
---
## 1. Disease Definition & Context
### Disease Identifiers
| System | ID | Source |
|--------|-----|--------|
### Description
### Synonyms
### Disease Hierarchy (parents/children)
### Affected Tissues/Organs
### Therapeutic Areas
**Sources**: (tools used)
---
## 2. Genomics Layer
### 2.1 GWAS Associations
| SNP | P-value | Effect | Gene | Study | Source |
|-----|---------|--------|------|-------|--------|
### 2.2 GWAS Studies Summary
| Study ID | Trait | Sample Size | Year | Source |
|----------|-------|-------------|------|--------|
### 2.3 Associated Genes (Genetic Evidence)
| Gene | Ensembl ID | Association Score | Evidence Type | Source |
|------|------------|-------------------|---------------|--------|
### 2.4 Rare Variants (ClinVar)
| Variant | Gene | Clinical Significance | Source |
|---------|------|-----------------------|--------|
### Genomics Layer Summary
- Total GWAS hits:
- Top genes by genetic evidence:
- Genetic architecture:
**Sources**: (tools used)
---
## 3. Transcriptomics Layer
### 3.1 Differential Expression Studies
| Experiment | Condition | Up-regulated | Down-regulated | Source |
|------------|-----------|--------------|----------------|--------|
### 3.2 Expression Atlas Disease Evidence
| Gene | Score | Source |
|------|-------|--------|
### 3.3 Tissue Expression Patterns (GTEx/HPA)
| Gene | Tissue | Expression Level | Source |
|------|--------|-----------------|--------|
### 3.4 Biomarker Candidates (Expression-Based)
| Gene | Tissue Specificity | Fold Change | Evidence | Source |
|------|-------------------|-------------|----------|--------|
### Transcriptomics Layer Summary
- Differential expression datasets:
- Top DEGs:
- Tissue-specific patterns:
**Sources**: (tools used)
---
## 4. Proteomics & Interaction Layer
### 4.1 Protein-Protein Interactions (STRING)
| Protein A | Protein B | Score | Source |
|-----------|-----------|-------|--------|
### 4.2 Hub Genes (Network Centrality)
| Gene | Degree | Betweenness | Role | Source |
|------|--------|-------------|------|--------|
### 4.3 Protein Complexes (IntAct)
| Complex | Members | Function | Source |
|---------|---------|----------|--------|
### 4.4 Tissue-Specific PPI Network
| Gene | Interaction Score | Tissue | Source |
|------|-------------------|--------|--------|
### Proteomics Layer Summary
- Total PPIs:
- Hub genes:
- Network modules:
**Sources**: (tools used)
---
## 5. Pathway & Network Layer
### 5.1 Enriched Pathways (Enrichr/Reactome)
| Pathway | Database | P-value | Genes | Source |
|---------|----------|---------|-------|--------|
### 5.2 Reactome Pathway Details
| Pathway ID | Name | Genes Involved | Source |
|------------|------|----------------|--------|
### 5.3 KEGG Pathways
| Pathway ID | Name | Description | Source |
|------------|------|-------------|--------|
### 5.4 WikiPathways
| Pathway ID | Name | Organism | Source |
|------------|------|----------|--------|
### Pathway Layer Summary
- Top enriched pathways:
- Key pathway nodes:
- Cross-pathway connections:
**Sources**: (tools used)
---
## 6. Gene Ontology & Functional Annotation
### 6.1 Biological Processes
| GO Term | Name | P-value | Genes | Source |
|---------|------|---------|-------|--------|
### 6.2 Molecular Functions
| GO Term | Name | P-value | Genes | Source |
|---------|------|---------|-------|--------|
### 6.3 Cellular Components
| GO Term | Name | P-value | Genes | Source |
|---------|------|---------|-------|--------|
**Sources**: (tools used)
---
## 7. Therapeutic Landscape
### 7.1 Approved Drugs
| Drug | ChEMBL ID | Mechanism | Target | Phase | Source |
|------|-----------|-----------|--------|-------|--------|
### 7.2 Druggable Targets
| Gene | Tractability | Modality | Clinical Precedent | Source |
|------|-------------|----------|-------------------|--------|
### 7.3 Drug Repurposing Candidates
| Drug | Original Indication | Mechanism | Target | Source |
|------|---------------------|-----------|--------|--------|
### 7.4 Clinical Trials
| NCT ID | Title | Phase | Status | Intervention | Source |
|--------|-------|-------|--------|--------------|--------|
### Therapeutic Summary
- Approved drugs:
- Clinical pipeline:
- Novel targets:
**Sources**: (tools used)
---
## 8. Multi-Omics Integration
### 8.1 Cross-Layer Gene Concordance
| Gene | Genomics | Transcriptomics | Proteomics | Pathways | Layers | Evidence Tier |
|------|----------|-----------------|------------|----------|--------|---------------|
### 8.2 Multi-Omics Hub Genes (Top 20)
| Rank | Gene | Layers Found | Key Evidence | Druggable | Source |
|------|------|-------------|--------------|-----------|--------|
### 8.3 Biomarker Candidates
| Biomarker | Type | Evidence Layers | Confidence | Source |
|-----------|------|-----------------|------------|--------|
### 8.4 Mechanistic Hypotheses
1. (Hypothesis with supporting evidence from multiple layers)
2. ...
### 8.5 Systems-Level Insights
- Key disrupted processes:
- Critical pathway nodes:
- Therapeutic intervention points:
- Testable hypotheses:
---
## Multi-Omics Confidence Score
| Component | Points | Max | Details |
|-----------|--------|-----|---------|
| Genomics data | | 10 | |
| Transcriptomics data | | 10 | |
| Protein data | | 5 | |
| Pathway data | | 10 | |
| Clinical data | | 5 | |
| Multi-layer genes | | 20 | |
| Direction concordance | | 10 | |
| Pathway-gene concordance | | 10 | |
| Genetic evidence quality | | 10 | |
| Clinical validation | | 10 | |
| **TOTAL** | | **100** | |
**Score**: XX/100 - [Tier]
---
## Data Availability Checklist
| Omics Layer | Data Available | Tools Used | Findings |
|-------------|---------------|------------|----------|
| Genomics (GWAS) | Yes/No | | |
| Genomics (Rare Variants) | Yes/No | | |
| Transcriptomics (DEGs) | Yes/No | | |
| Transcriptomics (Expression) | Yes/No | | |
| Proteomics (PPI) | Yes/No | | |
| Proteomics (Expression) | Yes/No | | |
| Pathways (Enrichment) | Yes/No | | |
| Pathways (KEGG/Reactome) | Yes/No | | |
| Gene Ontology | Yes/No | | |
| Drugs/Therapeutics | Yes/No | | |
| Clinical Trials | Yes/No | | |
| Literature | Yes/No | | |
---
## Completeness Checklist
- [ ] Disease disambiguation complete (IDs resolved)
- [ ] Genomics layer analyzed (GWAS + variants)
- [ ] Transcriptomics layer analyzed (DEGs + expression)
- [ ] Proteomics layer analyzed (PPI + interactions)
- [ ] Pathway layer analyzed (enrichment + mapping)
- [ ] Gene Ontology analyzed (BP + MF + CC)
- [ ] Therapeutic landscape analyzed (drugs + targets + trials)
- [ ] Cross-layer integration complete (concordance analysis)
- [ ] Multi-Omics Confidence Score calculated
- [ ] Biomarker candidates identified
- [ ] Hub genes identified
- [ ] Mechanistic hypotheses generated
- [ ] Executive summary written
- [ ] All sections have source citations
---
## References
### Data Sources Used
| # | Tool | Parameters | Section | Items Retrieved |
|---|------|------------|---------|-----------------|
### Database Versions
- OpenTargets: (current)
- GWAS Catalog: (current)
- STRING: (current)
- Reactome: (current)
Phase 0: Disease Disambiguation (ALWAYS FIRST)
Objective: Resolve disease to standard identifiers for all downstream queries.
Tools Used
OpenTargets_get_disease_id_description_by_name (primary):
- Input:
diseaseName (string) - Disease name
- Output:
{data: {search: {hits: [{id, name, description}]}}}
- Use: Get MONDO/EFO IDs and description
- CRITICAL: Disease IDs from OpenTargets use underscore format (e.g.,
MONDO_0004975), NOT colon format
OSL_get_efo_id_by_disease_name (secondary):
- Input:
disease (string) - Disease name
- Output:
{efo_id, name}
- Use: Get EFO/MONDO ID
OpenTargets_get_disease_description_by_efoId:
- Input:
efoId (string) - Disease ID (e.g., MONDO_0004975)
- Output:
{data: {disease: {id, name, description, dbXRefs}}}
- Use: Get full description, cross-references (OMIM, UMLS, DOID, etc.)
OpenTargets_get_disease_synonyms_by_efoId:
- Input:
efoId (string)
- Output:
{data: {disease: {id, name, synonyms: [{relation, terms}]}}}
OpenTargets_get_disease_therapeutic_areas_by_efoId:
- Input:
efoId (string)
- Output:
{data: {disease: {id, name, therapeuticAreas: [{id, name}]}}}
OpenTargets_get_disease_ancestors_parents_by_efoId:
- Input:
efoId (string)
- Output:
{data: {disease: {id, name, ancestors: [{id, name}]}}}
OpenTargets_get_disease_descendants_children_by_efoId:
- Input:
efoId (string)
- Output:
{data: {disease: {id, name, descendants: [{id, name}]}}}
OpenTargets_map_any_disease_id_to_all_other_ids:
- Input:
inputId (string) - Any known disease ID (e.g., OMIM:104300, UMLS:C0002395)
- Output:
{data: {disease: {id, name, dbXRefs: [str], ...}}}
- Use: Cross-map between OMIM, UMLS, ICD10, DOID, etc.
Workflow
- Search by disease name to get primary ID (OpenTargets)
- Get full description and cross-references
- Get synonyms for search term expansion
- Get therapeutic areas for context
- Get disease hierarchy (parents/children)
- If user provided OMIM/other ID, map to MONDO/EFO first
Collision-Aware Search
When disease name returns multiple hits:
- Check if user's input matches any hit exactly
- If ambiguous, present top 3-5 options and ask user to select
- Always prefer the most specific disease (not parent categories)
- For cancer, prefer the specific tumor type over generic "cancer"
Key Disease IDs to Track
After disambiguation, store these for all downstream queries:
efo_id - Primary ID for OpenTargets queries (e.g., MONDO_0004975)
disease_name - Canonical name (e.g., Alzheimer disease)
synonyms - For literature search expansion
therapeutic_areas - For context
dbXRefs - Cross-references (OMIM, UMLS, DOID, etc.)
Phase 1: Genomics Layer
Objective: Identify genetic variants, GWAS associations, and genetically implicated genes.
Tools Used
OpenTargets_get_associated_targets_by_disease_efoId (primary):
- Input:
efoId (string) - Disease EFO/MONDO ID
- Output:
{data: {disease: {id, name, associatedTargets: {count, rows: [{target: {id, approvedSymbol}, score}]}}}}
- Use: Get ALL disease-associated genes ranked by overall evidence score
- NOTE: Returns top 25 by default. For comprehensive analysis, note the total
count
OpenTargets_get_evidence_by_datasource:
- Input:
efoId (string), ensemblId (string), optional datasourceIds (array), size (int, default 50)
- Output:
{data: {disease: {evidences: {count, rows: [{...evidence details}]}}}}
- Use: Get specific evidence types. Key datasourceIds for genomics:
['ot_genetics_portal'] - GWAS/genetics
['gene2phenotype', 'genomics_england', 'orphanet'] - Rare variants
['eva'] - ClinVar variants
gwas_search_associations (GWAS Catalog):
- Input:
disease_trait (string), size (int, default 20)
- Output:
{data: [{association_id, p_value, or_per_copy_num, or_value, beta, risk_frequency, efo_traits: [{...}], ...}], metadata: {pagination: {totalElements}}}
- Use: Get genome-wide significant associations
- NOTE: Use disease name (e.g., "Alzheimer"), not ID. Returns paginated results
gwas_get_studies_for_trait:
- Input:
disease_trait (string), size (int)
- Output:
{data: [...studies], metadata: {pagination}}
- NOTE: May return empty if trait name does not match exactly. Try synonyms
gwas_get_variants_for_trait:
- Input:
disease_trait (string), size (int)
- Output:
{data: [...variants], metadata: {pagination}}
GWAS_search_associations_by_gene:
- Input:
gene_name (string)
- Output: Associations for a specific gene
OpenTargets_search_gwas_studies_by_disease:
- Input:
diseaseIds (array of strings), enableIndirect (bool, default true), size (int, default 10)
- Output:
{data: {studies: {count, rows: [{id, studyType, traitFromSource, publicationFirstAuthor, publicationDate, pubmedId, nSamples, nCases, nControls, ...}]}}}
- Use: Get GWAS studies from OpenTargets genetics portal
clinvar_search_variants:
- Input:
condition (string) or gene (string), optional max_results (int)
- Output: List of ClinVar variants with clinical significance
- Use: Rare variant / monogenic disease evidence
Workflow
- Get associated genes from OpenTargets (overall scores)
- For top 10-15 genes, get genetic evidence specifically via
OpenTargets_get_evidence_by_datasource
- Search GWAS Catalog for associations
- Search OpenTargets GWAS studies
- Search ClinVar for rare variants
- For top GWAS genes, check
GWAS_search_associations_by_gene
Gene Tracking
Maintain a dictionary of genes found in genomics layer:
genomics_genes = {
'PSEN1': {'score': 0.87, 'evidence': 'genetic', 'ensembl_id': 'ENSG00000080815', 'layer': 'genomics'},
'APP': {'score': 0.82, 'evidence': 'genetic', 'ensembl_id': 'ENSG00000142192', 'layer': 'genomics'},
# ...
}
Phase 2: Transcriptomics Layer
Objective: Identify differentially expressed genes, tissue-specific expression, and expression-based biomarkers.
Tools Used
ExpressionAtlas_search_differential:
- Input: optional
gene (string), condition (string), species (string, default 'homo sapiens')
- Output: Differential expression studies and results
- Use: Find studies where genes are differentially expressed in disease
ExpressionAtlas_search_experiments:
- Input: optional
gene (string), condition (string), species (string)
- Output: Expression experiments relevant to condition
- Use: Find all Expression Atlas experiments for the disease
expression_atlas_disease_target_score:
- Input:
efoId (string), pageSize (int, required)
- Output: Genes scored by expression evidence for the disease
- Use: Get expression-based disease-gene association scores
europepmc_disease_target_score:
- Input:
efoId (string), pageSize (int, required)
- Output: Genes scored by literature evidence for the disease
- Use: Complement expression evidence with literature-mined associations
HPA_get_rna_expression_by_source (Human Protein Atlas):
- Input:
gene_name (string), source_type (string: 'tissue', 'blood', 'brain'), source_name (string: e.g., 'brain', 'liver')
- Output:
{status, data: {gene_name, source_type, source_name, expression_value, expression_level, expression_unit}}
- NOTE: ALL 3 params required.
source_type options: 'tissue', 'blood', 'brain', 'cell_line', 'single_cell'
HPA_get_rna_expression_in_specific_tissues:
- Input:
gene_name (string), tissues (array of strings)
- Output: Expression across specified tissues
HPA_get_cancer_prognostics_by_gene:
- Input:
gene_name (string)
- Output: Cancer prognostic data (if cancer context)
HPA_get_subcellular_location:
- Input:
gene_name (string)
- Output: Subcellular localization data
HPA_search_genes_by_query:
- Input:
query (string)
- Output: Matching genes in HPA
Workflow
- Search Expression Atlas for differential expression studies
- Get expression-based disease scores
- Get literature-based disease scores (EuropePMC)
- For top 10-15 genes from genomics layer, check tissue expression via HPA
- Check disease-relevant tissue expression patterns
- For cancer: check prognostic biomarkers
Gene Tracking
Add transcriptomics genes to tracking:
transcriptomics_genes = {
'APOE': {'expression_score': 0.75, 'tissues': ['brain'], 'evidence': 'differential_expression', 'layer': 'transcriptomics'},
# ...
}
Phase 3: Proteomics & Interaction Layer
Objective: Map protein-protein interactions, identify hub genes, and characterize interaction networks.
Tools Used
STRING_get_interaction_partners (primary PPI):
- Input:
protein_ids (array of strings - gene names work), species (int, default 9606), confidence_score (float, default 0.4), limit (int, default 20)
- Output:
{status: 'success', data: [{stringId_A, stringId_B, preferredName_A, preferredName_B, ncbiTaxonId, score, nscore, fscore, pscore, ascore, escore, dscore, tscore}]}
- Use: Get interaction partners for disease genes
- NOTE:
protein_ids is an array, NOT string. Gene symbols like ['APOE'] work
STRING_get_network:
- Input:
protein_ids (array), species (int), confidence_score (float)
- Output: Network of interactions between input proteins
- Use: Build disease-specific PPI network
STRING_functional_enrichment:
- Input:
protein_ids (array), species (int)
- Output: Functional enrichment results (GO, KEGG, etc.)
- Use: Functional characterization of disease gene set
STRING_ppi_enrichment:
- Input:
protein_ids (array), species (int)
- Output: Statistical test for PPI enrichment (more interactions than expected)
- Use: Test if disease genes form a connected module
intact_get_interactions:
- Input:
identifier (string - UniProt ID or gene name)
- Output: Molecular interaction data from IntAct
intact_search_interactions:
- Input:
query (string), first (int, default 0), max (int, default 25)
- Output: Search results for interactions
HPA_get_protein_interactions_by_gene:
- Input:
gene_name (string)
- Output:
{gene, interactions, interactor_count, interactors: [...]}
humanbase_ppi_analysis:
- Input:
gene_list (array), tissue (string), max_node (int), interaction (string), string_mode (bool)
- Output: Tissue-specific PPI network
- NOTE: ALL params required.
interaction options: 'coexpression', 'interaction', 'coexpression_and_interaction'. string_mode: true/false
Workflow
- Take top 15-20 genes from genomics + transcriptomics layers
- Query STRING for interaction partners of each gene
- Build composite PPI network using STRING_get_network
- Test PPI enrichment (are genes more connected than random?)
- Get functional enrichment from STRING
- For disease-relevant tissue, get tissue-specific network (HumanBase)
- Identify hub genes (highest degree centrality)
- Check IntAct for experimentally validated interactions
Hub Gene Analysis
Calculate network centrality metrics:
- Degree: Number of interaction partners
- Betweenness: Number of shortest paths through node
- Hub score: Genes with degree > mean + 1 SD are hubs
Phase 4: Pathway & Network Layer
Objective: Identify enriched biological pathways and cross-pathway connections.
Tools Used
enrichr_gene_enrichment_analysis (primary enrichment):
- Input:
gene_list (array of gene symbols, min 2), libs (array of library names)
- Output:
{status: 'success', data: '{...JSON string with enrichment results...}'}
- Key libraries:
['KEGG_2021_Human'], ['Reactome_2022'], ['WikiPathway_2023_Human'], ['GO_Biological_Process_2023'], ['GO_Molecular_Function_2023'], ['GO_Cellular_Component_2023']
- NOTE:
data field is a JSON string, needs parsing. Contains connected_paths and per-library results
- NOTE:
libs is REQUIRED as array
ReactomeAnalysis_pathway_enrichment:
- Input:
identifiers (string - space-separated gene list), optional page_size (int, default 20), include_disease (bool), projection (bool)
- Output:
{data: {token, analysis_type, pathways_found, pathways: [{pathway_id, name, species, is_disease, is_lowest_level, entities_found, entities_total, entities_ratio, p_value, fdr, reactions_found, reactions_total}]}}
- Use: Reactome-specific pathway enrichment with statistical testing
Reactome_map_uniprot_to_pathways:
- Input:
id (string - UniProt accession)
- Output: List of Reactome pathways containing this protein
- Use: Map individual proteins to pathways
Reactome_get_pathway:
- Input:
stId (string - Reactome stable ID, e.g., 'R-HSA-73817')
- Output: Pathway details
Reactome_get_pathway_reactions:
- Input:
stId (string)
- Output: Reactions within pathway
kegg_search_pathway:
- Input:
keyword (string)
- Output: Array of KEGG pathway matches
kegg_get_pathway_info:
- Input:
pathway_id (string, e.g., 'hsa04930')
- Output: Detailed pathway information
WikiPathways_search:
- Input:
query (string), optional organism (string, e.g., 'Homo sapiens')
- Output: Matching community-curated pathways
Workflow
- Collect all genes from genomics + transcriptomics layers (top 20-30)
- Run Enrichr enrichment for KEGG, Reactome, WikiPathways
- Run ReactomeAnalysis for more detailed Reactome enrichment with p-values
- Search KEGG for disease-specific pathways
- Search WikiPathways for disease pathways
- For top Reactome pathways, get detailed reactions
- Identify cross-pathway connections (genes in multiple pathways)
Phase 5: Gene Ontology & Functional Annotation
Objective: Characterize biological processes, molecular functions, and cellular components.
Tools Used
enrichr_gene_enrichment_analysis (GO enrichment):
- Use with
libs=['GO_Biological_Process_2023'] for BP
- Use with
libs=['GO_Molecular_Function_2023'] for MF
- Use with
libs=['GO_Cellular_Component_2023'] for CC
GO_get_annotations_for_gene:
- Input:
gene_id (string - gene symbol or UniProt ID)
- Output: List of GO annotations with terms, aspects, evidence codes
GO_search_terms:
- Input:
query (string)
- Output: Matching GO terms
QuickGO_annotations_by_gene:
- Input:
gene_product_id (string - UniProt accession, e.g., 'UniProtKB:P02649'), optional aspect (string: 'biological_process', 'molecular_function', 'cellular_component'), taxon_id (int: 9606), limit (int: 25)
- Output: GO annotations with evidence codes
OpenTargets_get_target_gene_ontology_by_ensemblID:
- Input:
ensemblId (string)
- Output: GO terms associated with target
Workflow
- Run Enrichr GO enrichment for all 3 aspects using combined gene list
- For top 5 genes, get detailed GO annotations from QuickGO
- For top genes, get OpenTargets GO terms
- Summarize key biological processes, molecular functions, cellular components
Phase 6: Therapeutic Landscape
Objective: Map approved drugs, druggable targets, repurposing opportunities, and clinical trials.
Tools Used
OpenTargets_get_associated_drugs_by_disease_efoId (primary):
- Input:
efoId (string), size (int, REQUIRED - use 100)
- Output:
{data: {disease: {knownDrugs: {count, rows: [{drug: {id, name, tradeNames, maximumClinicalTrialPhase, isApproved, hasBeenWithdrawn}, phase, mechanismOfAction, target: {id, approvedSymbol}, disease: {id, name}, urls: [{url, name}]}]}}}}
- Use: All drugs associated with disease (approved + investigational)
OpenTargets_get_target_tractability_by_ensemblID:
- Input:
ensemblId (string)
- Output: Tractability assessment (small molecule, antibody, PROTAC, etc.)
OpenTargets_get_associated_drugs_by_target_ensemblID:
- Input:
ensemblId (string), size (int, REQUIRED)
- Output: Drugs targeting this gene/protein
search_clinical_trials:
- Input:
query_term (string, REQUIRED), optional condition (string), intervention (string), pageSize (int, default 10)
- Output: Clinical trial results
- NOTE:
query_term is REQUIRED even if condition is provided
OpenTargets_get_drug_mechanisms_of_action_by_chemblId:
- Input:
chemblId (string)
- Output: Mechanism of action details
Workflow
- Get all drugs for disease from OpenTargets
- For top disease-associated genes, check tractability
- For top genes with no approved drugs, identify repurposing candidates
- Search clinical trials for disease
- For top approved drugs, get mechanism of action
Drug Tracking
drug_targets = {
'PSEN1': {'drugs': ['Semagacestat'], 'tractability': 'small_molecule', 'clinical_phase': 3},
'ACHE': {'drugs': ['Donepezil', 'Galantamine'], 'tractability': 'small_molecule', 'clinical_phase': 4},
# ...
}
Phase 7: Multi-Omics Integration
Objective: Integrate findings across all layers to identify cross-layer genes, calculate concordance, and generate mechanistic hypotheses.
Cross-Layer Gene Concordance Analysis
This is the core integrative step. For each gene found in the analysis:
Count layers: In how many omics layers does this gene appear?
- Genomics (GWAS, rare variants, genetic association)
- Transcriptomics (DEGs, expression score)
- Proteomics (PPI hub, protein expression)
- Pathways (enriched pathway member)
- Therapeutics (drug target)
Score genes: Genes appearing in 3+ layers are "multi-omics hub genes"
Direction concordance: Do genetics and expression agree?
- Risk allele + upregulated = concordant gain-of-function
- Risk allele + downregulated = concordant loss-of-function
- Discordant = needs investigation
Biomarker Identification
For each multi-omics hub gene, assess biomarker potential:
- Diagnostic: Gene expression distinguishes disease vs healthy
- Prognostic: Expression/variant predicts outcome (cancer prognostics from HPA)
- Predictive: Variant/expression predicts treatment response (pharmacogenomics)
- Evidence level: Number of supporting omics layers
Mechanistic Hypothesis Generation
From the integrated data:
- Identify the most supported biological processes (GO + pathways)
- Map causal chain: genetic variant -> gene expression -> protein function -> pathway disruption -> disease
- Identify intervention points (druggable nodes in the causal chain)
- Generate testable hypotheses
Confidence Score Calculation
Calculate the Multi-Omics Confidence Score (0-100) based on:
- Data availability across layers
- Cross-layer concordance
- Evidence quality
- Clinical validation
Phase 8: Report Finalization
Executive Summary
Write a 2-3 sentence synthesis covering:
- Disease mechanism in systems terms
- Key genes/pathways identified
- Therapeutic opportunities
Final Report Quality Checklist
Before presenting to user, verify:
Tool Parameter Quick Reference
| Tool |
Key Parameters |
Notes |
OpenTargets_get_disease_id_description_by_name |
diseaseName |
Primary disambiguation |
OSL_get_efo_id_by_disease_name |
disease |
Secondary disambiguation |
OpenTargets_get_associated_targets_by_disease_efoId |
efoId |
Returns top 25 genes |
OpenTargets_get_evidence_by_datasource |
efoId, ensemblId, datasourceIds[], size |
Per-gene evidence |
OpenTargets_search_gwas_studies_by_disease |
diseaseIds[], size |
GWAS studies |
gwas_search_associations |
disease_trait, size |
GWAS Catalog |
clinvar_search_variants |
condition or gene, max_results |
Rare variants |
ExpressionAtlas_search_differential |
condition, species |
DEGs |
expression_atlas_disease_target_score |
efoId, pageSize (REQUIRED) |
Expression scores |
europepmc_disease_target_score |
efoId, pageSize (REQUIRED) |
Literature scores |
HPA_get_rna_expression_by_source |
gene_name, source_type, source_name (ALL REQUIRED) |
Tissue expression |
STRING_get_interaction_partners |
protein_ids[], species (9606), limit |
PPI partners |
STRING_get_network |
protein_ids[], species |
PPI network |
STRING_functional_enrichment |
protein_ids[], species |
Functional enrichment |
STRING_ppi_enrichment |
protein_ids[], species |
Network significance |
intact_search_interactions |
query, max |
Experimental PPIs |
humanbase_ppi_analysis |
gene_list[], tissue, max_node, interaction, string_mode (ALL REQ) |
Tissue PPI |
enrichr_gene_enrichment_analysis |
gene_list[], libs[] (BOTH REQUIRED) |
Pathway/GO enrichment |
ReactomeAnalysis_pathway_enrichment |
identifiers (space-sep string) |
Reactome enrichment |
Reactome_map_uniprot_to_pathways |
id (UniProt accession) |
Protein-pathway mapping |
kegg_search_pathway |
keyword |
KEGG pathway search |
WikiPathways_search |
query, organism |
WikiPathways search |
GO_get_annotations_for_gene |
gene_id |
GO annotations |
QuickGO_annotations_by_gene |
gene_product_id (e.g., 'UniProtKB:P02649') |
Detailed GO |
OpenTargets_get_associated_drugs_by_disease_efoId |
efoId, size (REQUIRED) |
Disease drugs |
OpenTargets_get_target_tractability_by_ensemblID |
ensemblId |
Druggability |
search_clinical_trials |
query_term (REQUIRED), condition, pageSize |
Clinical trials |
PubMed_search_articles |
query, limit |
Literature |
ensembl_lookup_gene |
gene_id, species ('homo_sapiens' REQUIRED) |
Gene lookup |
MyGene_query_genes |
query, species, fields, size |
Gene info |
OpenTargets_get_similar_entities_by_disease_efoId |
efoId, threshold, size (ALL REQUIRED) |
Similar diseases |
Response Format Notes (Verified)
OpenTargets Associated Targets
{
"data": {
"disease": {
"id": "MONDO_0004975",
"name": "Alzheimer disease",
"associatedTargets": {
"count": 2456,
"rows": [
{
"target": {"id": "ENSG00000080815", "approvedSymbol": "PSEN1"},
"score": 0.87
}
]
}
}
}
}
GWAS Catalog Associations
{
"data": [
{
"association_id": 216440893,
"p_value": 2e-09,
"or_per_copy_num": 0.94,
"or_value": "0.94",
"efo_traits": [{"..."}],
"risk_frequency": "NR"
}
],
"metadata": {"pagination": {"totalElements": 1061816}}
}
STRING Interactions
{
"status": "success",
"data": [
{
"stringId_A": "9606.ENSP00000252486",
"stringId_B": "9606.ENSP00000466775",
"preferredName_A": "APOE",
"preferredName_B": "APOC2",
"score": 0.999
}
]
}
Reactome Enrichment
{
"data": {
"token": "...",
"pathways_found": 154,
"pathways": [
{
"pathway_id": "R-HSA-1251985",
"name": "Nuclear signaling by ERBB4",
"species": "Homo sapiens",
"is_disease": false,
"is_lowest_level": true,
"entities_found": 3,
"entities_total": 47,
"entities_ratio": 0.00291,
"p_value": 4.0e-06,
"fdr": 0.00068,
"reactions_found": 3,
"reactions_total": 34
}
]
}
}
HPA RNA Expression
{
"status": "success",
"data": {
"gene_name": "APOE",
"source_type": "tissue",
"source_name": "brain",
"expression_value": "2714.9",
"expression_level": "very high",
"expression_unit": "nTPM"
}
}
Enrichr Results
{
"status": "success",
"data": "{\"connected_paths\": {\"Path: ...\": \"Total Weight: ...\"}}"
}
NOTE: The data field is a JSON string that needs parsing.
Common Use Patterns
1. Comprehensive Disease Profiling
User: "Characterize Alzheimer's disease across omics layers"
-> Run all 8 phases
-> Produce full multi-omics report
2. Therapeutic Target Discovery
User: "What are druggable targets for rheumatoid arthritis?"
-> Emphasize Phase 1 (genomics), Phase 6 (therapeutics), Phase 7 (integration)
-> Focus on tractability and clinical precedent
3. Biomarker Identification
User: "Find diagnostic biomarkers for pancreatic cancer"
-> Emphasize Phase 2 (transcriptomics), Phase 3 (proteomics), Phase 7 (biomarkers)
-> Focus on tissue-specific expression and diagnostic potential
4. Mechanism Elucidation
User: "What pathways are dysregulated in Crohn's disease?"
-> Emphasize Phase 4 (pathways), Phase 5 (GO), Phase 7 (mechanistic hypotheses)
-> Focus on pathway enrichment and cross-pathway connections
5. Drug Repurposing
User: "What existing drugs could be repurposed for ALS?"
-> Emphasize Phase 1 (genetics), Phase 6 (therapeutic landscape), Phase 7 (repurposing)
-> Focus on drugs targeting disease-associated genes
6. Systems Biology
User: "What are the hub genes and key pathways in type 2 diabetes?"
-> Emphasize Phase 3 (PPI network), Phase 4 (pathways), Phase 7 (network analysis)
-> Focus on hub genes and network modules
Edge Case Handling
Rare Diseases (limited data)
- Genomics layer may dominate (single gene)
- Limited GWAS data (monogenic)
- Focus on ClinVar variants, pathway consequences
- Confidence score will be lower (less cross-layer data)
Common Diseases (overwhelming data)
- Thousands of GWAS associations
- Prioritize by effect size and significance
- Focus on top 20-30 genes for downstream analysis
- Use strict significance thresholds (p < 5e-8)
Cancer
- Include somatic mutations (if CIViC/cBioPortal available)
- Check cancer prognostics via HPA
- Include tumor-specific expression patterns
- Clinical trial landscape may be extensive
Monogenic Diseases
- Single gene dominates
- ClinVar/OMIM evidence is primary
- Pathway analysis reveals downstream effects
- Therapeutic landscape may be limited (gene therapy, enzyme replacement)
Polygenic Diseases
- Many weak genetic signals
- GWAS provides the gene list
- Pathway enrichment reveals convergent biology
- Network analysis identifies hub genes
Tissue Ambiguity
- Diseases affecting multiple tissues
- Query HPA for all relevant tissues
- Compare tissue-specific
…(truncated)
1---2name: multiomic-disease-characterization3description: ToolUniverse workflow — Multiomic Disease Characterization4---56---7name: tooluniverse-multiomic-disease-characterization8description: Comprehensive multi-omics disease characterization integrating genomics, transcriptomics, proteomics, pathway, and therapeutic layers for systems-level understanding. Produces a detailed multi-omics report with quantitative confidence scoring (0-100), cross-layer gene concordance analysis, biomarker candidates, therapeutic opportunities, and mechanistic hypotheses. Uses 80+ ToolUniverse tools across 8 analysis layers. Use when users ask about disease mechanisms, multi-omics analysis, systems biology of disease, biomarker discovery, or therapeutic target identification from a disease perspective.9---1011# Multi-Omics Disease Characterization Pipeline1213Characterize diseases across multiple molecular layers (genomics, transcriptomics, proteomics, pathways) to provide systems-level understanding of disease mechanisms, identify therapeutic opportunities, and discover biomarker candidates.1415**KEY PRINCIPLES**:161. **Report-first approach** - Create report file FIRST, then populate progressively172. **Disease disambiguation FIRST** - Resolve all identifiers before omics analysis183. **Layer-by-layer analysis** - Systematically cover all omics layers194. **Cross-layer integration** - Identify genes/targets appearing in multiple layers205. **Evidence grading** - Grade all evidence as T1 (human/clinical) to T4 (computational)216. **Tissue context** - Emphasize disease-relevant tissues/organs227. **Quantitative scoring** - Multi-Omics Confidence Score (0-100)238. **Druggable focus** - Prioritize targets with therapeutic potential249. **Biomarker identification** - Highlight diagnostic/prognostic markers2510. **Mechanistic synthesis** - Generate testable hypotheses2611. **Source references** - Every statement must cite tool/database2712. **Completeness checklist** - Mandatory section showing analysis coverage2813. **English-first queries** - Always use English terms in tool calls. Respond in user's language2930---3132## When to Use This Skill3334Apply when users:35- Ask about disease mechanisms across omics layers36- Need multi-omics characterization of a disease37- Want to understand disease at the systems biology level38- Ask "What pathways/genes/proteins are involved in [disease]?"39- Need biomarker discovery for a disease40- Want to identify druggable targets from disease profiling41- Ask for integrated genomics + transcriptomics + proteomics analysis42- Need cross-layer concordance analysis43- Ask about disease network biology / hub genes4445**NOT for** (use other skills instead):46- Single gene/target validation -> Use `tooluniverse-drug-target-validation`47- Drug safety profiling -> Use `tooluniverse-adverse-event-detection`48- General disease overview -> Use `tooluniverse-disease-research`49- Variant interpretation -> Use `tooluniverse-variant-interpretation`50- GWAS-specific analysis -> Use `tooluniverse-gwas-*` skills51- Pathway-only analysis -> Use `tooluniverse-systems-biology`5253---5455## Input Parameters5657| Parameter | Required | Description | Example |58|-----------|----------|-------------|---------|59| **disease** | Yes | Disease name, OMIM ID, EFO ID, or MONDO ID | `Alzheimer disease`, `MONDO_0004975` |60| **tissue** | No | Tissue/organ of interest | `brain`, `liver`, `blood` |61| **focus_layers** | No | Specific omics layers to emphasize | `genomics`, `transcriptomics`, `pathways` |6263---6465## Multi-Omics Confidence Score (0-100)6667### Score Components6869**Data Availability (0-40 points)**:70- Genomics data available (GWAS or rare variants): 10 points71- Transcriptomics data available (DEGs or expression): 10 points72- Protein data available (PPI or expression): 5 points73- Pathway data available (enriched pathways): 10 points74- Clinical/drug data available (approved drugs or trials): 5 points7576**Evidence Concordance (0-40 points)**:77- Multi-layer genes (appear in 3+ layers): up to 20 points (2 per gene, max 10 genes)78- Consistent direction (genetics + expression concordant): 10 points79- Pathway-gene concordance (genes found in enriched pathways): 10 points8081**Evidence Quality (0-20 points)**:82- Strong genetic evidence (GWAS p < 5e-8): 10 points83- Clinical validation (approved drugs): 10 points8485### Score Interpretation8687| Score | Tier | Interpretation |88|-------|------|----------------|89| **80-100** | Excellent | Comprehensive multi-omics coverage, high confidence, strong cross-layer concordance |90| **60-79** | Good | Good coverage across most layers, some gaps |91| **40-59** | Moderate | Moderate coverage, limited cross-layer integration |92| **0-39** | Limited | Limited data, single-layer analysis dominates |9394### Evidence Grading System9596| Tier | Symbol | Criteria | Examples |97|------|--------|----------|----------|98| **T1** | [T1] | Direct human evidence, clinical proof | FDA-approved drug, GWAS hit (p<5e-8), clinical trial result |99| **T2** | [T2] | Experimental evidence | Differential expression (validated), functional screen, mouse KO |100| **T3** | [T3] | Computational/database evidence | PPI network, pathway mapping, expression correlation |101| **T4** | [T4] | Annotation/prediction only | GO annotation, text-mined association, predicted interaction |102103---104105## Report Template106107Create this file structure at the start: `{disease_name}_multiomic_report.md`108109```markdown110# Multi-Omics Disease Characterization: {Disease Name}111112**Report Generated**: {date}113**Disease Identifiers**: (to be filled)114**Multi-Omics Confidence Score**: (to be calculated)115116---117118## Executive Summary119120(2-3 sentence disease mechanism synthesis - fill after all layers complete)121122---123124## 1. Disease Definition & Context125126### Disease Identifiers127| System | ID | Source |128|--------|-----|--------|129130### Description131### Synonyms132### Disease Hierarchy (parents/children)133### Affected Tissues/Organs134### Therapeutic Areas135136**Sources**: (tools used)137138---139140## 2. Genomics Layer141142### 2.1 GWAS Associations143| SNP | P-value | Effect | Gene | Study | Source |144|-----|---------|--------|------|-------|--------|145146### 2.2 GWAS Studies Summary147| Study ID | Trait | Sample Size | Year | Source |148|----------|-------|-------------|------|--------|149150### 2.3 Associated Genes (Genetic Evidence)151| Gene | Ensembl ID | Association Score | Evidence Type | Source |152|------|------------|-------------------|---------------|--------|153154### 2.4 Rare Variants (ClinVar)155| Variant | Gene | Clinical Significance | Source |156|---------|------|-----------------------|--------|157158### Genomics Layer Summary159- Total GWAS hits:160- Top genes by genetic evidence:161- Genetic architecture:162163**Sources**: (tools used)164165---166167## 3. Transcriptomics Layer168169### 3.1 Differential Expression Studies170| Experiment | Condition | Up-regulated | Down-regulated | Source |171|------------|-----------|--------------|----------------|--------|172173### 3.2 Expression Atlas Disease Evidence174| Gene | Score | Source |175|------|-------|--------|176177### 3.3 Tissue Expression Patterns (GTEx/HPA)178| Gene | Tissue | Expression Level | Source |179|------|--------|-----------------|--------|180181### 3.4 Biomarker Candidates (Expression-Based)182| Gene | Tissue Specificity | Fold Change | Evidence | Source |183|------|-------------------|-------------|----------|--------|184185### Transcriptomics Layer Summary186- Differential expression datasets:187- Top DEGs:188- Tissue-specific patterns:189190**Sources**: (tools used)191192---193194## 4. Proteomics & Interaction Layer195196### 4.1 Protein-Protein Interactions (STRING)197| Protein A | Protein B | Score | Source |198|-----------|-----------|-------|--------|199200### 4.2 Hub Genes (Network Centrality)201| Gene | Degree | Betweenness | Role | Source |202|------|--------|-------------|------|--------|203204### 4.3 Protein Complexes (IntAct)205| Complex | Members | Function | Source |206|---------|---------|----------|--------|207208### 4.4 Tissue-Specific PPI Network209| Gene | Interaction Score | Tissue | Source |210|------|-------------------|--------|--------|211212### Proteomics Layer Summary213- Total PPIs:214- Hub genes:215- Network modules:216217**Sources**: (tools used)218219---220221## 5. Pathway & Network Layer222223### 5.1 Enriched Pathways (Enrichr/Reactome)224| Pathway | Database | P-value | Genes | Source |225|---------|----------|---------|-------|--------|226227### 5.2 Reactome Pathway Details228| Pathway ID | Name | Genes Involved | Source |229|------------|------|----------------|--------|230231### 5.3 KEGG Pathways232| Pathway ID | Name | Description | Source |233|------------|------|-------------|--------|234235### 5.4 WikiPathways236| Pathway ID | Name | Organism | Source |237|------------|------|----------|--------|238239### Pathway Layer Summary240- Top enriched pathways:241- Key pathway nodes:242- Cross-pathway connections:243244**Sources**: (tools used)245246---247248## 6. Gene Ontology & Functional Annotation249250### 6.1 Biological Processes251| GO Term | Name | P-value | Genes | Source |252|---------|------|---------|-------|--------|253254### 6.2 Molecular Functions255| GO Term | Name | P-value | Genes | Source |256|---------|------|---------|-------|--------|257258### 6.3 Cellular Components259| GO Term | Name | P-value | Genes | Source |260|---------|------|---------|-------|--------|261262**Sources**: (tools used)263264---265266## 7. Therapeutic Landscape267268### 7.1 Approved Drugs269| Drug | ChEMBL ID | Mechanism | Target | Phase | Source |270|------|-----------|-----------|--------|-------|--------|271272### 7.2 Druggable Targets273| Gene | Tractability | Modality | Clinical Precedent | Source |274|------|-------------|----------|-------------------|--------|275276### 7.3 Drug Repurposing Candidates277| Drug | Original Indication | Mechanism | Target | Source |278|------|---------------------|-----------|--------|--------|279280### 7.4 Clinical Trials281| NCT ID | Title | Phase | Status | Intervention | Source |282|--------|-------|-------|--------|--------------|--------|283284### Therapeutic Summary285- Approved drugs:286- Clinical pipeline:287- Novel targets:288289**Sources**: (tools used)290291---292293## 8. Multi-Omics Integration294295### 8.1 Cross-Layer Gene Concordance296| Gene | Genomics | Transcriptomics | Proteomics | Pathways | Layers | Evidence Tier |297|------|----------|-----------------|------------|----------|--------|---------------|298299### 8.2 Multi-Omics Hub Genes (Top 20)300| Rank | Gene | Layers Found | Key Evidence | Druggable | Source |301|------|------|-------------|--------------|-----------|--------|302303### 8.3 Biomarker Candidates304| Biomarker | Type | Evidence Layers | Confidence | Source |305|-----------|------|-----------------|------------|--------|306307### 8.4 Mechanistic Hypotheses3081. (Hypothesis with supporting evidence from multiple layers)3092. ...310311### 8.5 Systems-Level Insights312- Key disrupted processes:313- Critical pathway nodes:314- Therapeutic intervention points:315- Testable hypotheses:316317---318319## Multi-Omics Confidence Score320321| Component | Points | Max | Details |322|-----------|--------|-----|---------|323| Genomics data | | 10 | |324| Transcriptomics data | | 10 | |325| Protein data | | 5 | |326| Pathway data | | 10 | |327| Clinical data | | 5 | |328| Multi-layer genes | | 20 | |329| Direction concordance | | 10 | |330| Pathway-gene concordance | | 10 | |331| Genetic evidence quality | | 10 | |332| Clinical validation | | 10 | |333| **TOTAL** | | **100** | |334335**Score**: XX/100 - [Tier]336337---338339## Data Availability Checklist340341| Omics Layer | Data Available | Tools Used | Findings |342|-------------|---------------|------------|----------|343| Genomics (GWAS) | Yes/No | | |344| Genomics (Rare Variants) | Yes/No | | |345| Transcriptomics (DEGs) | Yes/No | | |346| Transcriptomics (Expression) | Yes/No | | |347| Proteomics (PPI) | Yes/No | | |348| Proteomics (Expression) | Yes/No | | |349| Pathways (Enrichment) | Yes/No | | |350| Pathways (KEGG/Reactome) | Yes/No | | |351| Gene Ontology | Yes/No | | |352| Drugs/Therapeutics | Yes/No | | |353| Clinical Trials | Yes/No | | |354| Literature | Yes/No | | |355356---357358## Completeness Checklist359360- [ ] Disease disambiguation complete (IDs resolved)361- [ ] Genomics layer analyzed (GWAS + variants)362- [ ] Transcriptomics layer analyzed (DEGs + expression)363- [ ] Proteomics layer analyzed (PPI + interactions)364- [ ] Pathway layer analyzed (enrichment + mapping)365- [ ] Gene Ontology analyzed (BP + MF + CC)366- [ ] Therapeutic landscape analyzed (drugs + targets + trials)367- [ ] Cross-layer integration complete (concordance analysis)368- [ ] Multi-Omics Confidence Score calculated369- [ ] Biomarker candidates identified370- [ ] Hub genes identified371- [ ] Mechanistic hypotheses generated372- [ ] Executive summary written373- [ ] All sections have source citations374375---376377## References378379### Data Sources Used380| # | Tool | Parameters | Section | Items Retrieved |381|---|------|------------|---------|-----------------|382383### Database Versions384- OpenTargets: (current)385- GWAS Catalog: (current)386- STRING: (current)387- Reactome: (current)388```389390---391392## Phase 0: Disease Disambiguation (ALWAYS FIRST)393394**Objective**: Resolve disease to standard identifiers for all downstream queries.395396### Tools Used397398**OpenTargets_get_disease_id_description_by_name** (primary):399- **Input**: `diseaseName` (string) - Disease name400- **Output**: `{data: {search: {hits: [{id, name, description}]}}}`401- **Use**: Get MONDO/EFO IDs and description402- **CRITICAL**: Disease IDs from OpenTargets use underscore format (e.g., `MONDO_0004975`), NOT colon format403404**OSL_get_efo_id_by_disease_name** (secondary):405- **Input**: `disease` (string) - Disease name406- **Output**: `{efo_id, name}`407- **Use**: Get EFO/MONDO ID408409**OpenTargets_get_disease_description_by_efoId**:410- **Input**: `efoId` (string) - Disease ID (e.g., `MONDO_0004975`)411- **Output**: `{data: {disease: {id, name, description, dbXRefs}}}`412- **Use**: Get full description, cross-references (OMIM, UMLS, DOID, etc.)413414**OpenTargets_get_disease_synonyms_by_efoId**:415- **Input**: `efoId` (string)416- **Output**: `{data: {disease: {id, name, synonyms: [{relation, terms}]}}}`417418**OpenTargets_get_disease_therapeutic_areas_by_efoId**:419- **Input**: `efoId` (string)420- **Output**: `{data: {disease: {id, name, therapeuticAreas: [{id, name}]}}}`421422**OpenTargets_get_disease_ancestors_parents_by_efoId**:423- **Input**: `efoId` (string)424- **Output**: `{data: {disease: {id, name, ancestors: [{id, name}]}}}`425426**OpenTargets_get_disease_descendants_children_by_efoId**:427- **Input**: `efoId` (string)428- **Output**: `{data: {disease: {id, name, descendants: [{id, name}]}}}`429430**OpenTargets_map_any_disease_id_to_all_other_ids**:431- **Input**: `inputId` (string) - Any known disease ID (e.g., `OMIM:104300`, `UMLS:C0002395`)432- **Output**: `{data: {disease: {id, name, dbXRefs: [str], ...}}}`433- **Use**: Cross-map between OMIM, UMLS, ICD10, DOID, etc.434435### Workflow4364371. Search by disease name to get primary ID (OpenTargets)4382. Get full description and cross-references4393. Get synonyms for search term expansion4404. Get therapeutic areas for context4415. Get disease hierarchy (parents/children)4426. If user provided OMIM/other ID, map to MONDO/EFO first443444### Collision-Aware Search445446When disease name returns multiple hits:447- Check if user's input matches any hit exactly448- If ambiguous, present top 3-5 options and ask user to select449- Always prefer the most specific disease (not parent categories)450- For cancer, prefer the specific tumor type over generic "cancer"451452### Key Disease IDs to Track453454After disambiguation, store these for all downstream queries:455- `efo_id` - Primary ID for OpenTargets queries (e.g., `MONDO_0004975`)456- `disease_name` - Canonical name (e.g., `Alzheimer disease`)457- `synonyms` - For literature search expansion458- `therapeutic_areas` - For context459- `dbXRefs` - Cross-references (OMIM, UMLS, DOID, etc.)460461---462463## Phase 1: Genomics Layer464465**Objective**: Identify genetic variants, GWAS associations, and genetically implicated genes.466467### Tools Used468469**OpenTargets_get_associated_targets_by_disease_efoId** (primary):470- **Input**: `efoId` (string) - Disease EFO/MONDO ID471- **Output**: `{data: {disease: {id, name, associatedTargets: {count, rows: [{target: {id, approvedSymbol}, score}]}}}}`472- **Use**: Get ALL disease-associated genes ranked by overall evidence score473- **NOTE**: Returns top 25 by default. For comprehensive analysis, note the total `count`474475**OpenTargets_get_evidence_by_datasource**:476- **Input**: `efoId` (string), `ensemblId` (string), optional `datasourceIds` (array), `size` (int, default 50)477- **Output**: `{data: {disease: {evidences: {count, rows: [{...evidence details}]}}}}`478- **Use**: Get specific evidence types. Key datasourceIds for genomics:479 - `['ot_genetics_portal']` - GWAS/genetics480 - `['gene2phenotype', 'genomics_england', 'orphanet']` - Rare variants481 - `['eva']` - ClinVar variants482483**gwas_search_associations** (GWAS Catalog):484- **Input**: `disease_trait` (string), `size` (int, default 20)485- **Output**: `{data: [{association_id, p_value, or_per_copy_num, or_value, beta, risk_frequency, efo_traits: [{...}], ...}], metadata: {pagination: {totalElements}}}`486- **Use**: Get genome-wide significant associations487- **NOTE**: Use disease name (e.g., "Alzheimer"), not ID. Returns paginated results488489**gwas_get_studies_for_trait**:490- **Input**: `disease_trait` (string), `size` (int)491- **Output**: `{data: [...studies], metadata: {pagination}}`492- **NOTE**: May return empty if trait name does not match exactly. Try synonyms493494**gwas_get_variants_for_trait**:495- **Input**: `disease_trait` (string), `size` (int)496- **Output**: `{data: [...variants], metadata: {pagination}}`497498**GWAS_search_associations_by_gene**:499- **Input**: `gene_name` (string)500- **Output**: Associations for a specific gene501502**OpenTargets_search_gwas_studies_by_disease**:503- **Input**: `diseaseIds` (array of strings), `enableIndirect` (bool, default true), `size` (int, default 10)504- **Output**: `{data: {studies: {count, rows: [{id, studyType, traitFromSource, publicationFirstAuthor, publicationDate, pubmedId, nSamples, nCases, nControls, ...}]}}}`505- **Use**: Get GWAS studies from OpenTargets genetics portal506507**clinvar_search_variants**:508- **Input**: `condition` (string) or `gene` (string), optional `max_results` (int)509- **Output**: List of ClinVar variants with clinical significance510- **Use**: Rare variant / monogenic disease evidence511512### Workflow5135141. Get associated genes from OpenTargets (overall scores)5152. For top 10-15 genes, get genetic evidence specifically via `OpenTargets_get_evidence_by_datasource`5163. Search GWAS Catalog for associations5174. Search OpenTargets GWAS studies5185. Search ClinVar for rare variants5196. For top GWAS genes, check `GWAS_search_associations_by_gene`520521### Gene Tracking522523Maintain a dictionary of genes found in genomics layer:524```python525genomics_genes = {526 'PSEN1': {'score': 0.87, 'evidence': 'genetic', 'ensembl_id': 'ENSG00000080815', 'layer': 'genomics'},527 'APP': {'score': 0.82, 'evidence': 'genetic', 'ensembl_id': 'ENSG00000142192', 'layer': 'genomics'},528 # ...529}530```531532---533534## Phase 2: Transcriptomics Layer535536**Objective**: Identify differentially expressed genes, tissue-specific expression, and expression-based biomarkers.537538### Tools Used539540**ExpressionAtlas_search_differential**:541- **Input**: optional `gene` (string), `condition` (string), `species` (string, default 'homo sapiens')542- **Output**: Differential expression studies and results543- **Use**: Find studies where genes are differentially expressed in disease544545**ExpressionAtlas_search_experiments**:546- **Input**: optional `gene` (string), `condition` (string), `species` (string)547- **Output**: Expression experiments relevant to condition548- **Use**: Find all Expression Atlas experiments for the disease549550**expression_atlas_disease_target_score**:551- **Input**: `efoId` (string), `pageSize` (int, required)552- **Output**: Genes scored by expression evidence for the disease553- **Use**: Get expression-based disease-gene association scores554555**europepmc_disease_target_score**:556- **Input**: `efoId` (string), `pageSize` (int, required)557- **Output**: Genes scored by literature evidence for the disease558- **Use**: Complement expression evidence with literature-mined associations559560**HPA_get_rna_expression_by_source** (Human Protein Atlas):561- **Input**: `gene_name` (string), `source_type` (string: 'tissue', 'blood', 'brain'), `source_name` (string: e.g., 'brain', 'liver')562- **Output**: `{status, data: {gene_name, source_type, source_name, expression_value, expression_level, expression_unit}}`563- **NOTE**: ALL 3 params required. `source_type` options: 'tissue', 'blood', 'brain', 'cell_line', 'single_cell'564565**HPA_get_rna_expression_in_specific_tissues**:566- **Input**: `gene_name` (string), `tissues` (array of strings)567- **Output**: Expression across specified tissues568569**HPA_get_cancer_prognostics_by_gene**:570- **Input**: `gene_name` (string)571- **Output**: Cancer prognostic data (if cancer context)572573**HPA_get_subcellular_location**:574- **Input**: `gene_name` (string)575- **Output**: Subcellular localization data576577**HPA_search_genes_by_query**:578- **Input**: `query` (string)579- **Output**: Matching genes in HPA580581### Workflow5825831. Search Expression Atlas for differential expression studies5842. Get expression-based disease scores5853. Get literature-based disease scores (EuropePMC)5864. For top 10-15 genes from genomics layer, check tissue expression via HPA5875. Check disease-relevant tissue expression patterns5886. For cancer: check prognostic biomarkers589590### Gene Tracking591592Add transcriptomics genes to tracking:593```python594transcriptomics_genes = {595 'APOE': {'expression_score': 0.75, 'tissues': ['brain'], 'evidence': 'differential_expression', 'layer': 'transcriptomics'},596 # ...597}598```599600---601602## Phase 3: Proteomics & Interaction Layer603604**Objective**: Map protein-protein interactions, identify hub genes, and characterize interaction networks.605606### Tools Used607608**STRING_get_interaction_partners** (primary PPI):609- **Input**: `protein_ids` (array of strings - gene names work), `species` (int, default 9606), `confidence_score` (float, default 0.4), `limit` (int, default 20)610- **Output**: `{status: 'success', data: [{stringId_A, stringId_B, preferredName_A, preferredName_B, ncbiTaxonId, score, nscore, fscore, pscore, ascore, escore, dscore, tscore}]}`611- **Use**: Get interaction partners for disease genes612- **NOTE**: `protein_ids` is an array, NOT string. Gene symbols like `['APOE']` work613614**STRING_get_network**:615- **Input**: `protein_ids` (array), `species` (int), `confidence_score` (float)616- **Output**: Network of interactions between input proteins617- **Use**: Build disease-specific PPI network618619**STRING_functional_enrichment**:620- **Input**: `protein_ids` (array), `species` (int)621- **Output**: Functional enrichment results (GO, KEGG, etc.)622- **Use**: Functional characterization of disease gene set623624**STRING_ppi_enrichment**:625- **Input**: `protein_ids` (array), `species` (int)626- **Output**: Statistical test for PPI enrichment (more interactions than expected)627- **Use**: Test if disease genes form a connected module628629**intact_get_interactions**:630- **Input**: `identifier` (string - UniProt ID or gene name)631- **Output**: Molecular interaction data from IntAct632633**intact_search_interactions**:634- **Input**: `query` (string), `first` (int, default 0), `max` (int, default 25)635- **Output**: Search results for interactions636637**HPA_get_protein_interactions_by_gene**:638- **Input**: `gene_name` (string)639- **Output**: `{gene, interactions, interactor_count, interactors: [...]}`640641**humanbase_ppi_analysis**:642- **Input**: `gene_list` (array), `tissue` (string), `max_node` (int), `interaction` (string), `string_mode` (bool)643- **Output**: Tissue-specific PPI network644- **NOTE**: ALL params required. `interaction` options: 'coexpression', 'interaction', 'coexpression_and_interaction'. `string_mode`: true/false645646### Workflow6476481. Take top 15-20 genes from genomics + transcriptomics layers6492. Query STRING for interaction partners of each gene6503. Build composite PPI network using STRING_get_network6514. Test PPI enrichment (are genes more connected than random?)6525. Get functional enrichment from STRING6536. For disease-relevant tissue, get tissue-specific network (HumanBase)6547. Identify hub genes (highest degree centrality)6558. Check IntAct for experimentally validated interactions656657### Hub Gene Analysis658659Calculate network centrality metrics:660- **Degree**: Number of interaction partners661- **Betweenness**: Number of shortest paths through node662- **Hub score**: Genes with degree > mean + 1 SD are hubs663664---665666## Phase 4: Pathway & Network Layer667668**Objective**: Identify enriched biological pathways and cross-pathway connections.669670### Tools Used671672**enrichr_gene_enrichment_analysis** (primary enrichment):673- **Input**: `gene_list` (array of gene symbols, min 2), `libs` (array of library names)674- **Output**: `{status: 'success', data: '{...JSON string with enrichment results...}'}`675- **Key libraries**: `['KEGG_2021_Human']`, `['Reactome_2022']`, `['WikiPathway_2023_Human']`, `['GO_Biological_Process_2023']`, `['GO_Molecular_Function_2023']`, `['GO_Cellular_Component_2023']`676- **NOTE**: `data` field is a JSON string, needs parsing. Contains `connected_paths` and per-library results677- **NOTE**: `libs` is REQUIRED as array678679**ReactomeAnalysis_pathway_enrichment**:680- **Input**: `identifiers` (string - space-separated gene list), optional `page_size` (int, default 20), `include_disease` (bool), `projection` (bool)681- **Output**: `{data: {token, analysis_type, pathways_found, pathways: [{pathway_id, name, species, is_disease, is_lowest_level, entities_found, entities_total, entities_ratio, p_value, fdr, reactions_found, reactions_total}]}}`682- **Use**: Reactome-specific pathway enrichment with statistical testing683684**Reactome_map_uniprot_to_pathways**:685- **Input**: `id` (string - UniProt accession)686- **Output**: List of Reactome pathways containing this protein687- **Use**: Map individual proteins to pathways688689**Reactome_get_pathway**:690- **Input**: `stId` (string - Reactome stable ID, e.g., 'R-HSA-73817')691- **Output**: Pathway details692693**Reactome_get_pathway_reactions**:694- **Input**: `stId` (string)695- **Output**: Reactions within pathway696697**kegg_search_pathway**:698- **Input**: `keyword` (string)699- **Output**: Array of KEGG pathway matches700701**kegg_get_pathway_info**:702- **Input**: `pathway_id` (string, e.g., 'hsa04930')703- **Output**: Detailed pathway information704705**WikiPathways_search**:706- **Input**: `query` (string), optional `organism` (string, e.g., 'Homo sapiens')707- **Output**: Matching community-curated pathways708709### Workflow7107111. Collect all genes from genomics + transcriptomics layers (top 20-30)7122. Run Enrichr enrichment for KEGG, Reactome, WikiPathways7133. Run ReactomeAnalysis for more detailed Reactome enrichment with p-values7144. Search KEGG for disease-specific pathways7155. Search WikiPathways for disease pathways7166. For top Reactome pathways, get detailed reactions7177. Identify cross-pathway connections (genes in multiple pathways)718719---720721## Phase 5: Gene Ontology & Functional Annotation722723**Objective**: Characterize biological processes, molecular functions, and cellular components.724725### Tools Used726727**enrichr_gene_enrichment_analysis** (GO enrichment):728- Use with `libs=['GO_Biological_Process_2023']` for BP729- Use with `libs=['GO_Molecular_Function_2023']` for MF730- Use with `libs=['GO_Cellular_Component_2023']` for CC731732**GO_get_annotations_for_gene**:733- **Input**: `gene_id` (string - gene symbol or UniProt ID)734- **Output**: List of GO annotations with terms, aspects, evidence codes735736**GO_search_terms**:737- **Input**: `query` (string)738- **Output**: Matching GO terms739740**QuickGO_annotations_by_gene**:741- **Input**: `gene_product_id` (string - UniProt accession, e.g., 'UniProtKB:P02649'), optional `aspect` (string: 'biological_process', 'molecular_function', 'cellular_component'), `taxon_id` (int: 9606), `limit` (int: 25)742- **Output**: GO annotations with evidence codes743744**OpenTargets_get_target_gene_ontology_by_ensemblID**:745- **Input**: `ensemblId` (string)746- **Output**: GO terms associated with target747748### Workflow7497501. Run Enrichr GO enrichment for all 3 aspects using combined gene list7512. For top 5 genes, get detailed GO annotations from QuickGO7523. For top genes, get OpenTargets GO terms7534. Summarize key biological processes, molecular functions, cellular components754755---756757## Phase 6: Therapeutic Landscape758759**Objective**: Map approved drugs, druggable targets, repurposing opportunities, and clinical trials.760761### Tools Used762763**OpenTargets_get_associated_drugs_by_disease_efoId** (primary):764- **Input**: `efoId` (string), `size` (int, REQUIRED - use 100)765- **Output**: `{data: {disease: {knownDrugs: {count, rows: [{drug: {id, name, tradeNames, maximumClinicalTrialPhase, isApproved, hasBeenWithdrawn}, phase, mechanismOfAction, target: {id, approvedSymbol}, disease: {id, name}, urls: [{url, name}]}]}}}}`766- **Use**: All drugs associated with disease (approved + investigational)767768**OpenTargets_get_target_tractability_by_ensemblID**:769- **Input**: `ensemblId` (string)770- **Output**: Tractability assessment (small molecule, antibody, PROTAC, etc.)771772**OpenTargets_get_associated_drugs_by_target_ensemblID**:773- **Input**: `ensemblId` (string), `size` (int, REQUIRED)774- **Output**: Drugs targeting this gene/protein775776**search_clinical_trials**:777- **Input**: `query_term` (string, REQUIRED), optional `condition` (string), `intervention` (string), `pageSize` (int, default 10)778- **Output**: Clinical trial results779- **NOTE**: `query_term` is REQUIRED even if `condition` is provided780781**OpenTargets_get_drug_mechanisms_of_action_by_chemblId**:782- **Input**: `chemblId` (string)783- **Output**: Mechanism of action details784785### Workflow7867871. Get all drugs for disease from OpenTargets7882. For top disease-associated genes, check tractability7893. For top genes with no approved drugs, identify repurposing candidates7904. Search clinical trials for disease7915. For top approved drugs, get mechanism of action792793### Drug Tracking794795```python796drug_targets = {797 'PSEN1': {'drugs': ['Semagacestat'], 'tractability': 'small_molecule', 'clinical_phase': 3},798 'ACHE': {'drugs': ['Donepezil', 'Galantamine'], 'tractability': 'small_molecule', 'clinical_phase': 4},799 # ...800}801```802803---804805## Phase 7: Multi-Omics Integration806807**Objective**: Integrate findings across all layers to identify cross-layer genes, calculate concordance, and generate mechanistic hypotheses.808809### Cross-Layer Gene Concordance Analysis810811This is the core integrative step. For each gene found in the analysis:8128131. **Count layers**: In how many omics layers does this gene appear?814 - Genomics (GWAS, rare variants, genetic association)815 - Transcriptomics (DEGs, expression score)816 - Proteomics (PPI hub, protein expression)817 - Pathways (enriched pathway member)818 - Therapeutics (drug target)8198202. **Score genes**: Genes appearing in 3+ layers are "multi-omics hub genes"8218223. **Direction concordance**: Do genetics and expression agree?823 - Risk allele + upregulated = concordant gain-of-function824 - Risk allele + downregulated = concordant loss-of-function825 - Discordant = needs investigation826827### Biomarker Identification828829For each multi-omics hub gene, assess biomarker potential:830- **Diagnostic**: Gene expression distinguishes disease vs healthy831- **Prognostic**: Expression/variant predicts outcome (cancer prognostics from HPA)832- **Predictive**: Variant/expression predicts treatment response (pharmacogenomics)833- **Evidence level**: Number of supporting omics layers834835### Mechanistic Hypothesis Generation836837From the integrated data:8381. Identify the most supported biological processes (GO + pathways)8392. Map causal chain: genetic variant -> gene expression -> protein function -> pathway disruption -> disease8403. Identify intervention points (druggable nodes in the causal chain)8414. Generate testable hypotheses842843### Confidence Score Calculation844845Calculate the Multi-Omics Confidence Score (0-100) based on:846- Data availability across layers847- Cross-layer concordance848- Evidence quality849- Clinical validation850851---852853## Phase 8: Report Finalization854855### Executive Summary856857Write a 2-3 sentence synthesis covering:858- Disease mechanism in systems terms859- Key genes/pathways identified860- Therapeutic opportunities861862### Final Report Quality Checklist863864Before presenting to user, verify:865- [ ] All 8 sections have content (or marked as "No data available")866- [ ] Every data point has a source citation867- [ ] Executive summary reflects key findings868- [ ] Multi-Omics Confidence Score calculated869- [ ] Top 20 genes ranked by multi-omics evidence870- [ ] Top 10 enriched pathways listed871- [ ] Biomarker candidates identified872- [ ] Cross-layer concordance table complete873- [ ] Therapeutic opportunities summarized874- [ ] Mechanistic hypotheses generated875- [ ] Data Availability Checklist complete876- [ ] Completeness Checklist complete877- [ ] References section lists all tools used878879---880881## Tool Parameter Quick Reference882883| Tool | Key Parameters | Notes |884|------|---------------|-------|885| `OpenTargets_get_disease_id_description_by_name` | `diseaseName` | Primary disambiguation |886| `OSL_get_efo_id_by_disease_name` | `disease` | Secondary disambiguation |887| `OpenTargets_get_associated_targets_by_disease_efoId` | `efoId` | Returns top 25 genes |888| `OpenTargets_get_evidence_by_datasource` | `efoId`, `ensemblId`, `datasourceIds[]`, `size` | Per-gene evidence |889| `OpenTargets_search_gwas_studies_by_disease` | `diseaseIds[]`, `size` | GWAS studies |890| `gwas_search_associations` | `disease_trait`, `size` | GWAS Catalog |891| `clinvar_search_variants` | `condition` or `gene`, `max_results` | Rare variants |892| `ExpressionAtlas_search_differential` | `condition`, `species` | DEGs |893| `expression_atlas_disease_target_score` | `efoId`, `pageSize` (REQUIRED) | Expression scores |894| `europepmc_disease_target_score` | `efoId`, `pageSize` (REQUIRED) | Literature scores |895| `HPA_get_rna_expression_by_source` | `gene_name`, `source_type`, `source_name` (ALL REQUIRED) | Tissue expression |896| `STRING_get_interaction_partners` | `protein_ids[]`, `species` (9606), `limit` | PPI partners |897| `STRING_get_network` | `protein_ids[]`, `species` | PPI network |898| `STRING_functional_enrichment` | `protein_ids[]`, `species` | Functional enrichment |899| `STRING_ppi_enrichment` | `protein_ids[]`, `species` | Network significance |900| `intact_search_interactions` | `query`, `max` | Experimental PPIs |901| `humanbase_ppi_analysis` | `gene_list[]`, `tissue`, `max_node`, `interaction`, `string_mode` (ALL REQ) | Tissue PPI |902| `enrichr_gene_enrichment_analysis` | `gene_list[]`, `libs[]` (BOTH REQUIRED) | Pathway/GO enrichment |903| `ReactomeAnalysis_pathway_enrichment` | `identifiers` (space-sep string) | Reactome enrichment |904| `Reactome_map_uniprot_to_pathways` | `id` (UniProt accession) | Protein-pathway mapping |905| `kegg_search_pathway` | `keyword` | KEGG pathway search |906| `WikiPathways_search` | `query`, `organism` | WikiPathways search |907| `GO_get_annotations_for_gene` | `gene_id` | GO annotations |908| `QuickGO_annotations_by_gene` | `gene_product_id` (e.g., 'UniProtKB:P02649') | Detailed GO |909| `OpenTargets_get_associated_drugs_by_disease_efoId` | `efoId`, `size` (REQUIRED) | Disease drugs |910| `OpenTargets_get_target_tractability_by_ensemblID` | `ensemblId` | Druggability |911| `search_clinical_trials` | `query_term` (REQUIRED), `condition`, `pageSize` | Clinical trials |912| `PubMed_search_articles` | `query`, `limit` | Literature |913| `ensembl_lookup_gene` | `gene_id`, `species` ('homo_sapiens' REQUIRED) | Gene lookup |914| `MyGene_query_genes` | `query`, `species`, `fields`, `size` | Gene info |915| `OpenTargets_get_similar_entities_by_disease_efoId` | `efoId`, `threshold`, `size` (ALL REQUIRED) | Similar diseases |916917---918919## Response Format Notes (Verified)920921### OpenTargets Associated Targets922```json923{924 "data": {925 "disease": {926 "id": "MONDO_0004975",927 "name": "Alzheimer disease",928 "associatedTargets": {929 "count": 2456,930 "rows": [931 {932 "target": {"id": "ENSG00000080815", "approvedSymbol": "PSEN1"},933 "score": 0.87934 }935 ]936 }937 }938 }939}940```941942### GWAS Catalog Associations943```json944{945 "data": [946 {947 "association_id": 216440893,948 "p_value": 2e-09,949 "or_per_copy_num": 0.94,950 "or_value": "0.94",951 "efo_traits": [{"..."}],952 "risk_frequency": "NR"953 }954 ],955 "metadata": {"pagination": {"totalElements": 1061816}}956}957```958959### STRING Interactions960```json961{962 "status": "success",963 "data": [964 {965 "stringId_A": "9606.ENSP00000252486",966 "stringId_B": "9606.ENSP00000466775",967 "preferredName_A": "APOE",968 "preferredName_B": "APOC2",969 "score": 0.999970 }971 ]972}973```974975### Reactome Enrichment976```json977{978 "data": {979 "token": "...",980 "pathways_found": 154,981 "pathways": [982 {983 "pathway_id": "R-HSA-1251985",984 "name": "Nuclear signaling by ERBB4",985 "species": "Homo sapiens",986 "is_disease": false,987 "is_lowest_level": true,988 "entities_found": 3,989 "entities_total": 47,990 "entities_ratio": 0.00291,991 "p_value": 4.0e-06,992 "fdr": 0.00068,993 "reactions_found": 3,994 "reactions_total": 34995 }996 ]997 }998}999```10001001### HPA RNA Expression1002```json1003{1004 "status": "success",1005 "data": {1006 "gene_name": "APOE",1007 "source_type": "tissue",1008 "source_name": "brain",1009 "expression_value": "2714.9",1010 "expression_level": "very high",1011 "expression_unit": "nTPM"1012 }1013}1014```10151016### Enrichr Results1017```json1018{1019 "status": "success",1020 "data": "{\"connected_paths\": {\"Path: ...\": \"Total Weight: ...\"}}"1021}1022```1023**NOTE**: The `data` field is a JSON string that needs parsing.10241025---10261027## Common Use Patterns10281029### 1. Comprehensive Disease Profiling1030```1031User: "Characterize Alzheimer's disease across omics layers"1032-> Run all 8 phases1033-> Produce full multi-omics report1034```10351036### 2. Therapeutic Target Discovery1037```1038User: "What are druggable targets for rheumatoid arthritis?"1039-> Emphasize Phase 1 (genomics), Phase 6 (therapeutics), Phase 7 (integration)1040-> Focus on tractability and clinical precedent1041```10421043### 3. Biomarker Identification1044```1045User: "Find diagnostic biomarkers for pancreatic cancer"1046-> Emphasize Phase 2 (transcriptomics), Phase 3 (proteomics), Phase 7 (biomarkers)1047-> Focus on tissue-specific expression and diagnostic potential1048```10491050### 4. Mechanism Elucidation1051```1052User: "What pathways are dysregulated in Crohn's disease?"1053-> Emphasize Phase 4 (pathways), Phase 5 (GO), Phase 7 (mechanistic hypotheses)1054-> Focus on pathway enrichment and cross-pathway connections1055```10561057### 5. Drug Repurposing1058```1059User: "What existing drugs could be repurposed for ALS?"1060-> Emphasize Phase 1 (genetics), Phase 6 (therapeutic landscape), Phase 7 (repurposing)1061-> Focus on drugs targeting disease-associated genes1062```10631064### 6. Systems Biology1065```1066User: "What are the hub genes and key pathways in type 2 diabetes?"1067-> Emphasize Phase 3 (PPI network), Phase 4 (pathways), Phase 7 (network analysis)1068-> Focus on hub genes and network modules1069```10701071---10721073## Edge Case Handling10741075### Rare Diseases (limited data)1076- Genomics layer may dominate (single gene)1077- Limited GWAS data (monogenic)1078- Focus on ClinVar variants, pathway consequences1079- Confidence score will be lower (less cross-layer data)10801081### Common Diseases (overwhelming data)1082- Thousands of GWAS associations1083- Prioritize by effect size and significance1084- Focus on top 20-30 genes for downstream analysis1085- Use strict significance thresholds (p < 5e-8)10861087### Cancer1088- Include somatic mutations (if CIViC/cBioPortal available)1089- Check cancer prognostics via HPA1090- Include tumor-specific expression patterns1091- Clinical trial landscape may be extensive10921093### Monogenic Diseases1094- Single gene dominates1095- ClinVar/OMIM evidence is primary1096- Pathway analysis reveals downstream effects1097- Therapeutic landscape may be limited (gene therapy, enzyme replacement)10981099### Polygenic Diseases1100- Many weak genetic signals1101- GWAS provides the gene list1102- Pathway enrichment reveals convergent biology1103- Network analysis identifies hub genes11041105### Tissue Ambiguity1106- Diseases affecting multiple tissues1107- Query HPA for all relevant tissues1108- Compare tissue-specific11091110…(truncated)