name: tooluniverse-chemical-safety
description: Comprehensive chemical safety and toxicology assessment integrating ADMET-AI predictions, CTD toxicogenomics, FDA label safety data, DrugBank safety profiles, and STITCH chemical-protein interactions. Performs predictive toxicology (AMES, DILI, LD50, carcinogenicity), organ/system toxicity profiling, chemical-gene-disease relationship mapping, regulatory safety extraction, and environmental hazard assessment. Use when asked about chemical toxicity, drug safety profiling, ADMET properties, environmental health risks, chemical hazard assessment, or toxicogenomic analysis.
Chemical Safety & Toxicology Assessment
Comprehensive chemical safety and toxicology analysis integrating predictive AI models, curated toxicogenomics databases, regulatory safety data, and chemical-biological interaction networks. Generates structured risk assessment reports with evidence grading.
When to Use This Skill
Triggers:
- "Is this chemical toxic?" / "What are the toxicity endpoints for [compound]?"
- "Assess the safety profile of [drug/chemical]"
- "What are the ADMET properties of [SMILES]?"
- "What genes does [chemical] interact with?"
- "What diseases are linked to [chemical] exposure?"
- "Predict toxicity for these molecules"
- "Drug safety assessment for [drug name]"
- "Environmental health risk of [chemical]"
- "Chemical hazard profiling"
- "Toxicogenomic analysis of [compound]"
Use Cases:
- Predictive Toxicology: AI-predicted toxicity endpoints (AMES mutagenicity, DILI, LD50, carcinogenicity, skin reactions) for novel compounds via SMILES
- ADMET Profiling: Full absorption, distribution, metabolism, excretion, toxicity characterization
- Toxicogenomics: Chemical-gene interaction mapping, gene-disease associations from CTD
- Regulatory Safety: FDA label warnings, boxed warnings, contraindications, adverse reactions
- Drug Safety Assessment: Combined DrugBank safety + FDA labels + adverse event data
- Chemical-Protein Interactions: STITCH-based chemical-protein binding and interaction networks
- Environmental Toxicology: Chemical-disease associations for environmental contaminants
KEY PRINCIPLES
- Report-first approach - Create report file FIRST, then populate progressively
- Tool parameter verification - Verify params via
get_tool_info before calling unfamiliar tools
- Evidence grading - Grade all safety claims by evidence strength (T1-T4)
- Citation requirements - Every toxicity finding must have inline source attribution
- Mandatory completeness - All sections must exist with data minimums or explicit "No data" notes
- Disambiguation first - Resolve compound identity (name -> SMILES, CID, ChEMBL ID) before analysis
- Negative results documented - "No toxicity signals found" is data; empty sections are failures
- Conservative risk assessment - When evidence is ambiguous, flag as "requires further investigation"
- English-first queries - Always use English chemical/drug names in tool calls
Evidence Grading System (MANDATORY)
Grade every toxicity claim by evidence strength:
| Tier |
Symbol |
Criteria |
Examples |
| T1 |
[T1] |
Direct human evidence, regulatory finding |
FDA boxed warning, clinical trial toxicity, human case reports |
| T2 |
[T2] |
Animal studies, validated in vitro |
Nonclinical toxicology, AMES positive, animal LD50 |
| T3 |
[T3] |
Computational prediction, association data |
ADMET-AI prediction, CTD association, QSAR model |
| T4 |
[T4] |
Database annotation, text-mined |
Literature mention, database entry without validation |
Required Evidence Grading Locations
Evidence grades MUST appear in:
- Executive Summary - Key toxicity findings graded
- Toxicity Predictions - Every ADMET-AI endpoint with confidence note
- Regulatory Safety - FDA findings marked [T1]
- Chemical-Gene Interactions - CTD data marked by curation status
- Risk Assessment - Final risk classification with supporting evidence tiers
Core Strategy: 8 Research Dimensions
Chemical/Drug Query
|
+-- PHASE 0: Compound Disambiguation (ALWAYS FIRST)
| +-- Resolve name -> SMILES, PubChem CID, ChEMBL ID
| +-- Get molecular formula, weight, canonical structure
|
+-- PHASE 1: Predictive Toxicology (ADMET-AI)
| +-- Mutagenicity (AMES)
| +-- Hepatotoxicity (DILI, ClinTox)
| +-- Carcinogenicity
| +-- Acute toxicity (LD50)
| +-- Skin reactions
| +-- Stress response pathways
| +-- Nuclear receptor activity
|
+-- PHASE 2: ADMET Properties
| +-- Absorption: BBB penetrance, bioavailability
| +-- Distribution: clearance, volume of distribution
| +-- Metabolism: CYP interactions (1A2, 2C9, 2C19, 2D6, 3A4)
| +-- Physicochemical: solubility, lipophilicity, pKa
|
+-- PHASE 3: Toxicogenomics (CTD)
| +-- Chemical-gene interactions
| +-- Chemical-disease associations
| +-- Affected biological pathways
|
+-- PHASE 4: Regulatory Safety (FDA Labels)
| +-- Boxed warnings (Black Box)
| +-- Contraindications
| +-- Adverse reactions
| +-- Warnings and precautions
| +-- Nonclinical toxicology
|
+-- PHASE 5: Drug Safety Profile (DrugBank)
| +-- Toxicity data
| +-- Contraindications
| +-- Drug interactions affecting safety
|
+-- PHASE 6: Chemical-Protein Interactions (STITCH)
| +-- Direct chemical-protein binding
| +-- Interaction confidence scores
| +-- Off-target effects
|
+-- PHASE 7: Structural Alerts (ChEMBL)
| +-- Known toxic substructures (PAINS, Brenk)
| +-- Structural alert flags
|
+-- SYNTHESIS: Integrated Risk Assessment
+-- Aggregate all evidence tiers
+-- Risk classification (Low/Medium/High/Critical)
+-- Data gaps and recommendations
Phase 0: Compound Disambiguation (ALWAYS FIRST)
CRITICAL: Resolve compound identity before any analysis.
Input Types Handled
| Input Format |
Resolution Strategy |
| Drug name (e.g., "Aspirin") |
PubChem_get_CID_by_compound_name -> get SMILES from properties |
| SMILES string |
Use directly for ADMET-AI; resolve to CID for other tools |
| PubChem CID |
PubChem_get_compound_properties_by_CID -> get SMILES + name |
| ChEMBL ID |
ChEMBL_get_molecule -> get SMILES + properties |
Resolution Steps
- Input detection: Determine if input is name, SMILES, CID, or ChEMBL ID
- SMILES: contains typical SMILES characters (=, #, [, ], (, ), c, n, o and no spaces in middle)
- CID: numeric only
- ChEMBL: starts with "CHEMBL"
- Otherwise: treat as compound name
- Name to CID:
PubChem_get_CID_by_compound_name(name=<compound_name>)
- CID to properties:
PubChem_get_compound_properties_by_CID(cid=<cid>)
- Extract SMILES: Get SMILES from PubChem properties (field:
ConnectivitySMILES, CanonicalSMILES, or IsomericSMILES depending on response format)
- Store resolved IDs: Maintain dict with
name, smiles, cid, formula, weight, inchi
Disambiguation Output
## Compound Identity
| Property | Value |
|----------|-------|
| **Name** | Acetaminophen |
| **PubChem CID** | 1983 |
| **SMILES** | CC(=O)Nc1ccc(O)cc1 |
| **Formula** | C8H9NO2 |
| **Molecular Weight** | 151.16 |
| **InChI** | InChI=1S/C8H9NO2/... |
Phase 1: Predictive Toxicology (ADMET-AI)
When: SMILES is available (from Phase 0 or provided directly)
Objective: Run comprehensive AI-predicted toxicity endpoints
Tools Used
All ADMET-AI tools take the same parameter format:
| Tool |
Predicted Endpoints |
Parameter |
ADMETAI_predict_toxicity |
AMES, Carcinogens_Lagunin, ClinTox, DILI, LD50_Zhu, Skin_Reaction, hERG |
smiles: list[str] |
ADMETAI_predict_stress_response |
Stress response pathway activation (ARE, ATAD5, HSE, MMP, p53) |
smiles: list[str] |
ADMETAI_predict_nuclear_receptor_activity |
AhR, AR, ER, PPARg, Aromatase nuclear receptor activity |
smiles: list[str] |
Workflow
- Call
ADMETAI_predict_toxicity(smiles=[resolved_smiles])
- Call
ADMETAI_predict_stress_response(smiles=[resolved_smiles])
- Call
ADMETAI_predict_nuclear_receptor_activity(smiles=[resolved_smiles])
- For each endpoint, interpret prediction:
- Classification endpoints: Active (1) = toxic signal, Inactive (0) = no signal
- Regression endpoints (LD50): Report numerical value with context
- All predictions graded [T3] (computational prediction)
Decision Logic
- Multiple SMILES: Can batch up to ~10 SMILES in single call
- Failed prediction: If ADMET-AI fails, note "prediction unavailable" (don't fail entire report)
- Confidence: Note that AI predictions are [T3] evidence, not definitive
- hERG flag: If hERG = Active, flag prominently (cardiac safety risk)
- AMES flag: If AMES = Active, flag prominently (mutagenicity concern)
- DILI flag: If DILI = Active, flag prominently (liver toxicity concern)
Output Table
### Toxicity Predictions [T3]
| Endpoint | Prediction | Interpretation | Concern Level |
|----------|-----------|---------------|---------------|
| AMES Mutagenicity | Inactive | No mutagenic signal | Low |
| Carcinogenicity | Inactive | No carcinogenic signal | Low |
| ClinTox | Active | Clinical toxicity signal | HIGH |
| DILI | Active | Drug-induced liver injury risk | HIGH |
| LD50 (Zhu) | 2.45 log(mg/kg) | ~282 mg/kg (moderate) | Medium |
| Skin Reaction | Inactive | No skin sensitization signal | Low |
| hERG Inhibition | Active | Cardiac arrhythmia risk | HIGH |
*All predictions from ADMET-AI. Evidence tier: [T3] (computational prediction)*
Phase 2: ADMET Properties
When: SMILES is available
Objective: Full ADMET characterization beyond toxicity
Tools Used
| Tool |
Properties Predicted |
Parameter |
ADMETAI_predict_BBB_penetrance |
Blood-brain barrier crossing probability |
smiles: list[str] |
ADMETAI_predict_bioavailability |
Oral bioavailability (F20%, F30%) |
smiles: list[str] |
ADMETAI_predict_clearance_distribution |
Clearance, VDss, half-life, PPB |
smiles: list[str] |
ADMETAI_predict_CYP_interactions |
CYP1A2, 2C9, 2C19, 2D6, 3A4 inhibition/substrate |
smiles: list[str] |
ADMETAI_predict_physicochemical_properties |
LogP, LogD, LogS, MW, pKa |
smiles: list[str] |
ADMETAI_predict_solubility_lipophilicity_hydration |
Aqueous solubility, lipophilicity, hydration free energy |
smiles: list[str] |
Workflow
- Call all 6 ADMET tools in parallel (independent calls)
- Compile results into Absorption / Distribution / Metabolism / Excretion sections
- Assess Lipinski Rule of 5 compliance from physicochemical properties
- Flag drug-drug interaction risks from CYP inhibition profiles
Decision Logic
- BBB penetrant + toxicity: If BBB = Yes and any CNS toxicity endpoint active, flag as neurotoxicity risk
- Low bioavailability: If F20% = Low, note absorption concerns
- CYP inhibitor: If CYP3A4 inhibitor = Yes, flag high DDI risk
- Lipinski violations: Count violations and report drug-likeness assessment
Output Format
### ADMET Profile [T3]
#### Absorption
| Property | Value | Interpretation |
|----------|-------|----------------|
| BBB Penetrance | Yes | Crosses blood-brain barrier |
| Bioavailability (F20%) | 85% | Good oral absorption |
#### Distribution
| Property | Value | Interpretation |
|----------|-------|----------------|
| VDss | 1.2 L/kg | Moderate tissue distribution |
| PPB | 92% | Highly protein bound |
#### Metabolism
| CYP Enzyme | Substrate | Inhibitor |
|------------|-----------|-----------|
| CYP1A2 | No | No |
| CYP2C9 | Yes | No |
| CYP2C19 | No | No |
| CYP2D6 | No | No |
| CYP3A4 | Yes | Yes (DDI risk) |
#### Excretion
| Property | Value | Interpretation |
|----------|-------|----------------|
| Clearance | 8.5 mL/min/kg | Moderate clearance |
| Half-life | 6.2 h | Moderate half-life |
Phase 3: Toxicogenomics (CTD)
When: Compound name is resolved
Objective: Map chemical-gene-disease relationships from curated CTD data
Tools Used
| Tool |
Function |
Parameter |
CTD_get_chemical_gene_interactions |
Genes affected by chemical |
input_terms: str (chemical name) |
CTD_get_chemical_diseases |
Diseases linked to chemical exposure |
input_terms: str (chemical name) |
Workflow
- Call
CTD_get_chemical_gene_interactions(input_terms=compound_name)
- Call
CTD_get_chemical_diseases(input_terms=compound_name)
- Parse gene interactions: extract gene symbols, interaction types (increases/decreases expression, binding, etc.)
- Parse disease associations: extract disease names, evidence types (marker/mechanism/therapeutic)
- Identify most affected biological processes from gene list
Decision Logic
- Direct evidence vs inferred: CTD separates curated direct evidence from inferred associations
- Therapeutic vs toxic: Disease associations can be therapeutic (drug treats disease) or adverse (chemical causes disease)
- Gene interaction types: Distinguish between expression changes, binding, and activity modulation
- Prioritize marker/mechanism: These indicate stronger causal evidence than simple associations
- Grade curated as [T2]: Direct curated CTD evidence from literature
- Grade inferred as [T3]: Computationally inferred associations
Output Format
### Toxicogenomics (CTD) [T2/T3]
#### Chemical-Gene Interactions (Top 20)
| Gene | Interaction | Type | Evidence |
|------|------------|------|----------|
| CYP1A2 | increases expression | mRNA | [T2] curated |
| TP53 | affects activity | protein | [T2] curated |
| ... | ... | ... | ... |
**Total interactions found**: 156
**Top affected pathways**: Xenobiotic metabolism, Apoptosis, DNA damage response
#### Chemical-Disease Associations (Top 10)
| Disease | Association Type | Evidence |
|---------|-----------------|----------|
| Liver Neoplasms | marker/mechanism | [T2] curated |
| Contact Dermatitis | therapeutic | [T2] curated |
| ... | ... | ... |
Phase 4: Regulatory Safety (FDA Labels)
When: Compound has an approved drug name
Objective: Extract regulatory safety information from FDA drug labels
Tools Used
| Tool |
Information Retrieved |
Parameter |
FDA_get_boxed_warning_info_by_drug_name |
Black box warnings (most serious) |
drug_name: str |
FDA_get_contraindications_by_drug_name |
Absolute contraindications |
drug_name: str |
FDA_get_adverse_reactions_by_drug_name |
Known adverse reactions |
drug_name: str |
FDA_get_warnings_by_drug_name |
Warnings and precautions |
drug_name: str |
FDA_get_nonclinical_toxicology_info_by_drug_name |
Animal toxicology data |
drug_name: str |
FDA_get_carcinogenic_mutagenic_fertility_by_drug_name |
Carcinogenicity/mutagenicity/fertility data |
drug_name: str |
Workflow
- Call all 6 FDA tools in parallel (independent queries by drug name)
- Parse and structure each response
- Prioritize: Boxed Warnings > Contraindications > Warnings > Adverse Reactions
- All FDA label data is [T1] evidence (regulatory finding based on human/animal data)
Decision Logic
- Boxed warning present: Flag as CRITICAL safety concern in executive summary
- No FDA data: Chemical may not be an approved drug; note "Not an FDA-approved drug" and continue with other phases
- Multiple warnings: Categorize by organ system (hepatic, cardiac, renal, CNS, etc.)
- Nonclinical toxicology: Grade as [T2] (animal data supporting human risk)
Output Format
### Regulatory Safety (FDA) [T1]
#### Boxed Warning
**PRESENT** - Hepatotoxicity risk with doses >4g/day. Liver failure reported. [T1]
#### Contraindications
- Severe hepatic impairment [T1]
- Known hypersensitivity [T1]
#### Adverse Reactions (by frequency)
| Reaction | Frequency | Severity |
|----------|-----------|----------|
| Nausea | Common (>1%) | Mild |
| Hepatotoxicity | Rare (<0.1%) | Severe |
| ... | ... | ... |
#### Nonclinical Toxicology [T2]
- **Carcinogenicity**: No carcinogenic potential in 2-year rat/mouse studies
- **Mutagenicity**: Negative in Ames assay and in vivo micronucleus test
- **Fertility**: No effects on fertility at doses up to 10x human dose
Phase 5: Drug Safety Profile (DrugBank)
When: Compound is a known drug
Objective: Retrieve curated drug safety data from DrugBank
Tools Used
| Tool |
Information |
Parameters |
drugbank_get_safety_by_drug_name_or_drugbank_id |
Toxicity, contraindications |
query: str, case_sensitive: bool, exact_match: bool, limit: int |
Workflow
- Call
drugbank_get_safety_by_drug_name_or_drugbank_id(query=drug_name, case_sensitive=False, exact_match=False, limit=5)
- Parse toxicity information, overdose data, contraindications
- Cross-reference with FDA data from Phase 4
Decision Logic
- Toxicity field: Contains LD50 values, overdose symptoms, organ toxicity data
- DrugBank ID: Note if found for cross-referencing
- Conflict with FDA: If DrugBank and FDA disagree, note discrepancy and defer to FDA [T1]
- Not found: Chemical may not be in DrugBank; continue with other phases
Phase 6: Chemical-Protein Interactions (STITCH)
When: Compound can be identified by name or SMILES
Objective: Map chemical-protein interaction network for off-target assessment
Tools Used
| Tool |
Function |
Parameters |
STITCH_resolve_identifier |
Resolve chemical name to STITCH ID |
identifier: str, species: int (9606=human) |
STITCH_get_chemical_protein_interactions |
Get chemical-protein interactions |
identifiers: list[str], species: int, required_score: int |
STITCH_get_interaction_partners |
Get interaction network |
identifiers: list[str], species: int, limit: int |
Workflow
- Resolve compound:
STITCH_resolve_identifier(identifier=compound_name, species=9606)
- Get interactions:
STITCH_get_chemical_protein_interactions(identifiers=[stitch_id], species=9606, required_score=700)
- Identify off-target proteins (not the intended drug target)
- Flag safety-relevant targets: hERG (cardiac), CYP enzymes (metabolism), nuclear receptors (endocrine)
Decision Logic
- High confidence (>900): Well-established interaction [T2]
- Medium confidence (700-900): Probable interaction [T3]
- Low confidence (400-700): Possible interaction, needs validation [T4]
- Safety-relevant targets: Flag interactions with known safety targets
- No STITCH data: Chemical may be too novel; note and continue
Phase 7: Structural Alerts (ChEMBL)
When: ChEMBL molecule ID is available (from Phase 0)
Objective: Check for known toxic substructures
Tools Used
| Tool |
Function |
Parameters |
ChEMBL_search_compound_structural_alerts |
Find structural alert matches |
molecule_chembl_id: str, limit: int |
Workflow
- If ChEMBL ID available:
ChEMBL_search_compound_structural_alerts(molecule_chembl_id=chembl_id, limit=20)
- Parse alert types: PAINS (pan-assay interference), Brenk (medicinal chemistry), Glaxo (GSK structural alerts)
- Categorize severity: Some alerts are informational, others indicate likely toxicity
Decision Logic
- PAINS alerts: May cause false positives in screening; note for medicinal chemistry
- Brenk alerts: Known problematic substructures; flag if present
- No alerts: Good sign but not definitive proof of safety
- No ChEMBL ID: Skip this phase gracefully; note "structural alert analysis not available"
Synthesis: Integrated Risk Assessment (MANDATORY)
Always the final section. Integrates all evidence into actionable risk classification.
Risk Classification Matrix
| Risk Level |
Criteria |
| CRITICAL |
FDA boxed warning present OR multiple [T1] toxicity findings OR active DILI + active hERG |
| HIGH |
FDA warnings present OR [T2] animal toxicity OR multiple active ADMET endpoints |
| MEDIUM |
Some [T3] predictions positive OR CTD disease associations OR structural alerts |
| LOW |
All ADMET endpoints negative AND no FDA/DrugBank safety flags AND no CTD concerns |
| INSUFFICIENT DATA |
Fewer than 3 phases returned data; cannot make confident assessment |
Synthesis Template
## Integrated Risk Assessment
### Overall Risk Classification: [HIGH]
### Evidence Summary
| Dimension | Finding | Evidence Tier | Concern |
|-----------|---------|--------------|---------|
| ADMET Toxicity | DILI active, hERG active | [T3] | HIGH |
| FDA Label | Boxed warning for hepatotoxicity | [T1] | CRITICAL |
| CTD Toxicogenomics | 156 gene interactions, liver neoplasms | [T2] | HIGH |
| DrugBank | Known hepatotoxicity at high doses | [T2] | HIGH |
| STITCH | Binds CYP3A4, hERG | [T3] | MEDIUM |
| Structural Alerts | 2 Brenk alerts | [T3] | MEDIUM |
### Key Safety Concerns
1. **Hepatotoxicity** [T1]: FDA boxed warning + ADMET-AI DILI prediction + CTD liver disease associations
2. **Cardiac Risk** [T3]: ADMET-AI hERG prediction + STITCH hERG interaction
3. **Drug Interactions** [T3]: CYP3A4 substrate/inhibitor, potential DDI risk
### Data Gaps
- [ ] No in vivo genotoxicity data available
- [ ] STITCH interaction scores moderate (700-900)
- [ ] No environmental exposure data
### Recommendations
1. Avoid doses >4g/day (hepatotoxicity threshold) [T1]
2. Monitor liver function in chronic use [T1]
3. Screen for CYP3A4 interactions before co-administration [T3]
4. Consider cardiac monitoring for at-risk patients [T3]
Mandatory Completeness Checklist
Before finalizing any report, verify:
Tool Parameter Reference
Critical Parameter Notes (verified from source code):
| Tool |
Parameter Name |
Type |
Notes |
| All ADMETAI tools |
smiles |
list[str] |
Always a list, even for single compound |
| All CTD tools |
input_terms |
str |
Chemical name, MeSH name, CAS RN, or MeSH ID |
| All FDA tools |
drug_name |
str |
Brand or generic drug name |
| drugbank_get_safety_* |
query, case_sensitive, exact_match, limit |
str, bool, bool, int |
All 4 required |
| STITCH_resolve_identifier |
identifier, species |
str, int |
species=9606 for human |
| STITCH_get_chemical_protein_interactions |
identifiers, species, required_score |
list[str], int, int |
required_score=400 default |
| PubChem_get_CID_by_compound_name |
name |
str |
Compound name (not SMILES) |
| PubChem_get_compound_properties_by_CID |
cid |
int |
Numeric CID |
| ChEMBL_search_compound_structural_alerts |
molecule_chembl_id |
str |
ChEMBL ID (e.g., "CHEMBL112") |
Response Format Notes
- ADMET-AI: Returns
{status: "success", data: {...}} with prediction values
- CTD: Returns list of interaction/association objects
- FDA: Returns
{status, data} with label text
- DrugBank: Returns
{data: [...]} with drug records
- STITCH: Returns list of interaction objects with scores
- PubChem CID lookup: Returns
{IdentifierList: {CID: [...]}} (may or may not have data wrapper)
- PubChem properties: Returns dict with
CID, MolecularWeight, ConnectivitySMILES, IUPACName
Fallback Strategies
Compound Resolution
- Primary: PubChem by name -> CID -> properties -> SMILES
- Fallback 1: ChEMBL search by name -> molecule -> SMILES
- Fallback 2: If SMILES provided directly, skip name resolution
Toxicity Prediction
- Primary: All 9 ADMET-AI endpoints
- Fallback: If ADMET-AI fails for a compound, note "prediction failed" and continue with database evidence
- Note: ADMET-AI may fail for very large or unusual SMILES
Regulatory Data
- Primary: FDA labels by drug name
- Fallback: If FDA returns no data, try alternative drug names (brand vs generic)
- Note: Non-drug chemicals (pesticides, industrial) will not have FDA labels
CTD Data
- Primary: Search by common chemical name
- Fallback: Try MeSH name if common name fails
- Note: Novel compounds may not be in CTD
Common Use Patterns
Pattern 1: Novel Compound Assessment
Input: SMILES string for new molecule
Workflow: Phase 0 (SMILES->CID) -> Phase 1 (toxicity) -> Phase 2 (ADMET) -> Phase 7 (structural alerts) -> Synthesis
Output: Predictive safety profile for novel compound
Pattern 2: Approved Drug Safety Review
Input: Drug name (e.g., "Acetaminophen")
Workflow: All phases (0-7 + Synthesis)
Output: Complete safety dossier with regulatory + predictive + database evidence
Pattern 3: Environmental Chemical Risk
Input: Chemical name (e.g., "Bisphenol A")
Workflow: Phase 0 -> Phase 1 -> Phase 2 -> Phase 3 (CTD, key for env chemicals) -> Phase 6 -> Synthesis
Output: Environmental health risk assessment focused on gene-disease associations
Pattern 4: Batch Toxicity Screening
Input: Multiple SMILES strings
Workflow: Phase 0 -> Phase 1 (batch) -> Phase 2 (batch) -> Comparative table -> Synthesis
Output: Comparative toxicity table ranking compounds by safety
Pattern 5: Toxicogenomic Deep-Dive
Input: Chemical name + specific gene or disease interest
Workflow: Phase 0 -> Phase 3 (CTD expanded) -> Literature search -> Synthesis
Output: Detailed chemical-gene-disease mechanistic analysis
Output Report Structure
All analyses generate a structured markdown report with progressive sections:
# Chemical Safety & Toxicology Report: [Compound Name]
**Generated**: YYYY-MM-DD HH:MM
**Compound**: [Name] | SMILES: [SMILES] | CID: [CID]
## Executive Summary
[2-3 sentence overview with risk classification and key findings, all graded]
## 1. Compound Identity
[Phase 0 results - disambiguation table]
## 2. Predictive Toxicology
[Phase 1 results - ADMET-AI toxicity endpoints]
## 3. ADMET Profile
[Phase 2 results - absorption, distribution, metabolism, excretion]
## 4. Toxicogenomics
[Phase 3 results - CTD chemical-gene-disease relationships]
## 5. Regulatory Safety
[Phase 4 results - FDA label information]
## 6. Drug Safety Profile
[Phase 5 results - DrugBank data]
## 7. Chemical-Protein Interactions
[Phase 6 results - STITCH network]
## 8. Structural Alerts
[Phase 7 results - ChEMBL alerts]
## 9. Integrated Risk Assessment
[Synthesis - risk classification, evidence summary, data gaps, recommendations]
## Appendix: Methods and Data Sources
[Tool versions, databases queried, date of access]
Limitations & Known Issues
Tool-Specific
- ADMET-AI: Predictions are computational [T3]; should not replace experimental testing
- CTD: Curated but may lag behind latest literature by 6-12 months
- FDA: Only covers FDA-approved drugs; not applicable to environmental chemicals or supplements
- DrugBank: Primarily drugs; limited coverage of industrial chemicals
- STITCH: Score thresholds affect sensitivity; lower scores increase false positives
- ChEMBL: Structural alerts require ChEMBL ID; not all compounds have one
Analysis
- Novel compounds: May only have ADMET-AI predictions (no database evidence)
- Environmental chemicals: FDA/DrugBank phases will be empty; rely on CTD and ADMET-AI
- Batch mode: ADMET-AI can handle batches; other tools require individual queries
- Species specificity: Most data is human-centric; animal data noted where applicable
Technical
- SMILES validity: Invalid SMILES will cause ADMET-AI failures
- Name ambiguity: Chemical names can be ambiguous; always verify with CID
- Rate limits: Some FDA endpoints may rate-limit for rapid queries
Summary
Chemical Safety & Toxicology Assessment Skill provides comprehensive safety evaluation by integrating:
- Predictive toxicology (ADMET-AI) - 9 tools covering toxicity, ADMET, physicochemical properties
- Toxicogenomics (CTD) - Chemical-gene-disease relationship mapping
- Regulatory safety (FDA) - 6 tools for label-based safety extraction
- Drug safety (DrugBank) - Curated toxicity and contraindication data
- Chemical interactions (STITCH) - Chemical-protein interaction networks
- Structural alerts (ChEMBL) - Known toxic substructure detection
Outputs: Structured markdown report with risk classification, evidence grading, and actionable recommendations
Best for: Drug safety assessment, chemical hazard profiling, environmental toxicology, ADMET characterization, toxicogenomic analysis
Total tools integrated: 25+ tools across 6 databases
1---2name: chemical-safety3description: ToolUniverse workflow — Chemical Safety4---56---7name: tooluniverse-chemical-safety8description: Comprehensive chemical safety and toxicology assessment integrating ADMET-AI predictions, CTD toxicogenomics, FDA label safety data, DrugBank safety profiles, and STITCH chemical-protein interactions. Performs predictive toxicology (AMES, DILI, LD50, carcinogenicity), organ/system toxicity profiling, chemical-gene-disease relationship mapping, regulatory safety extraction, and environmental hazard assessment. Use when asked about chemical toxicity, drug safety profiling, ADMET properties, environmental health risks, chemical hazard assessment, or toxicogenomic analysis.9---1011# Chemical Safety & Toxicology Assessment1213Comprehensive chemical safety and toxicology analysis integrating predictive AI models, curated toxicogenomics databases, regulatory safety data, and chemical-biological interaction networks. Generates structured risk assessment reports with evidence grading.1415## When to Use This Skill1617**Triggers**:18- "Is this chemical toxic?" / "What are the toxicity endpoints for [compound]?"19- "Assess the safety profile of [drug/chemical]"20- "What are the ADMET properties of [SMILES]?"21- "What genes does [chemical] interact with?"22- "What diseases are linked to [chemical] exposure?"23- "Predict toxicity for these molecules"24- "Drug safety assessment for [drug name]"25- "Environmental health risk of [chemical]"26- "Chemical hazard profiling"27- "Toxicogenomic analysis of [compound]"2829**Use Cases**:301. **Predictive Toxicology**: AI-predicted toxicity endpoints (AMES mutagenicity, DILI, LD50, carcinogenicity, skin reactions) for novel compounds via SMILES312. **ADMET Profiling**: Full absorption, distribution, metabolism, excretion, toxicity characterization323. **Toxicogenomics**: Chemical-gene interaction mapping, gene-disease associations from CTD334. **Regulatory Safety**: FDA label warnings, boxed warnings, contraindications, adverse reactions345. **Drug Safety Assessment**: Combined DrugBank safety + FDA labels + adverse event data356. **Chemical-Protein Interactions**: STITCH-based chemical-protein binding and interaction networks367. **Environmental Toxicology**: Chemical-disease associations for environmental contaminants3738---3940## KEY PRINCIPLES41421. **Report-first approach** - Create report file FIRST, then populate progressively432. **Tool parameter verification** - Verify params via `get_tool_info` before calling unfamiliar tools443. **Evidence grading** - Grade all safety claims by evidence strength (T1-T4)454. **Citation requirements** - Every toxicity finding must have inline source attribution465. **Mandatory completeness** - All sections must exist with data minimums or explicit "No data" notes476. **Disambiguation first** - Resolve compound identity (name -> SMILES, CID, ChEMBL ID) before analysis487. **Negative results documented** - "No toxicity signals found" is data; empty sections are failures498. **Conservative risk assessment** - When evidence is ambiguous, flag as "requires further investigation"509. **English-first queries** - Always use English chemical/drug names in tool calls5152---5354## Evidence Grading System (MANDATORY)5556Grade every toxicity claim by evidence strength:5758| Tier | Symbol | Criteria | Examples |59|------|--------|----------|----------|60| **T1** | [T1] | Direct human evidence, regulatory finding | FDA boxed warning, clinical trial toxicity, human case reports |61| **T2** | [T2] | Animal studies, validated in vitro | Nonclinical toxicology, AMES positive, animal LD50 |62| **T3** | [T3] | Computational prediction, association data | ADMET-AI prediction, CTD association, QSAR model |63| **T4** | [T4] | Database annotation, text-mined | Literature mention, database entry without validation |6465### Required Evidence Grading Locations6667Evidence grades MUST appear in:681. **Executive Summary** - Key toxicity findings graded692. **Toxicity Predictions** - Every ADMET-AI endpoint with confidence note703. **Regulatory Safety** - FDA findings marked [T1]714. **Chemical-Gene Interactions** - CTD data marked by curation status725. **Risk Assessment** - Final risk classification with supporting evidence tiers7374---7576## Core Strategy: 8 Research Dimensions7778```79Chemical/Drug Query80|81+-- PHASE 0: Compound Disambiguation (ALWAYS FIRST)82| +-- Resolve name -> SMILES, PubChem CID, ChEMBL ID83| +-- Get molecular formula, weight, canonical structure84|85+-- PHASE 1: Predictive Toxicology (ADMET-AI)86| +-- Mutagenicity (AMES)87| +-- Hepatotoxicity (DILI, ClinTox)88| +-- Carcinogenicity89| +-- Acute toxicity (LD50)90| +-- Skin reactions91| +-- Stress response pathways92| +-- Nuclear receptor activity93|94+-- PHASE 2: ADMET Properties95| +-- Absorption: BBB penetrance, bioavailability96| +-- Distribution: clearance, volume of distribution97| +-- Metabolism: CYP interactions (1A2, 2C9, 2C19, 2D6, 3A4)98| +-- Physicochemical: solubility, lipophilicity, pKa99|100+-- PHASE 3: Toxicogenomics (CTD)101| +-- Chemical-gene interactions102| +-- Chemical-disease associations103| +-- Affected biological pathways104|105+-- PHASE 4: Regulatory Safety (FDA Labels)106| +-- Boxed warnings (Black Box)107| +-- Contraindications108| +-- Adverse reactions109| +-- Warnings and precautions110| +-- Nonclinical toxicology111|112+-- PHASE 5: Drug Safety Profile (DrugBank)113| +-- Toxicity data114| +-- Contraindications115| +-- Drug interactions affecting safety116|117+-- PHASE 6: Chemical-Protein Interactions (STITCH)118| +-- Direct chemical-protein binding119| +-- Interaction confidence scores120| +-- Off-target effects121|122+-- PHASE 7: Structural Alerts (ChEMBL)123| +-- Known toxic substructures (PAINS, Brenk)124| +-- Structural alert flags125|126+-- SYNTHESIS: Integrated Risk Assessment127 +-- Aggregate all evidence tiers128 +-- Risk classification (Low/Medium/High/Critical)129 +-- Data gaps and recommendations130```131132---133134## Phase 0: Compound Disambiguation (ALWAYS FIRST)135136**CRITICAL**: Resolve compound identity before any analysis.137138### Input Types Handled139140| Input Format | Resolution Strategy |141|-------------|---------------------|142| Drug name (e.g., "Aspirin") | PubChem_get_CID_by_compound_name -> get SMILES from properties |143| SMILES string | Use directly for ADMET-AI; resolve to CID for other tools |144| PubChem CID | PubChem_get_compound_properties_by_CID -> get SMILES + name |145| ChEMBL ID | ChEMBL_get_molecule -> get SMILES + properties |146147### Resolution Steps1481491. **Input detection**: Determine if input is name, SMILES, CID, or ChEMBL ID150 - SMILES: contains typical SMILES characters (=, #, [, ], (, ), c, n, o and no spaces in middle)151 - CID: numeric only152 - ChEMBL: starts with "CHEMBL"153 - Otherwise: treat as compound name1542. **Name to CID**: `PubChem_get_CID_by_compound_name(name=<compound_name>)`1553. **CID to properties**: `PubChem_get_compound_properties_by_CID(cid=<cid>)`1564. **Extract SMILES**: Get SMILES from PubChem properties (field: `ConnectivitySMILES`, `CanonicalSMILES`, or `IsomericSMILES` depending on response format)1575. **Store resolved IDs**: Maintain dict with `name`, `smiles`, `cid`, `formula`, `weight`, `inchi`158159### Disambiguation Output160161```markdown162## Compound Identity163164| Property | Value |165|----------|-------|166| **Name** | Acetaminophen |167| **PubChem CID** | 1983 |168| **SMILES** | CC(=O)Nc1ccc(O)cc1 |169| **Formula** | C8H9NO2 |170| **Molecular Weight** | 151.16 |171| **InChI** | InChI=1S/C8H9NO2/... |172```173174---175176## Phase 1: Predictive Toxicology (ADMET-AI)177178**When**: SMILES is available (from Phase 0 or provided directly)179180**Objective**: Run comprehensive AI-predicted toxicity endpoints181182### Tools Used183184All ADMET-AI tools take the same parameter format:185186| Tool | Predicted Endpoints | Parameter |187|------|---------------------|-----------|188| `ADMETAI_predict_toxicity` | AMES, Carcinogens_Lagunin, ClinTox, DILI, LD50_Zhu, Skin_Reaction, hERG | `smiles`: list[str] |189| `ADMETAI_predict_stress_response` | Stress response pathway activation (ARE, ATAD5, HSE, MMP, p53) | `smiles`: list[str] |190| `ADMETAI_predict_nuclear_receptor_activity` | AhR, AR, ER, PPARg, Aromatase nuclear receptor activity | `smiles`: list[str] |191192### Workflow1931941. Call `ADMETAI_predict_toxicity(smiles=[resolved_smiles])`1952. Call `ADMETAI_predict_stress_response(smiles=[resolved_smiles])`1963. Call `ADMETAI_predict_nuclear_receptor_activity(smiles=[resolved_smiles])`1974. For each endpoint, interpret prediction:198 - Classification endpoints: Active (1) = toxic signal, Inactive (0) = no signal199 - Regression endpoints (LD50): Report numerical value with context200 - All predictions graded [T3] (computational prediction)201202### Decision Logic203204- **Multiple SMILES**: Can batch up to ~10 SMILES in single call205- **Failed prediction**: If ADMET-AI fails, note "prediction unavailable" (don't fail entire report)206- **Confidence**: Note that AI predictions are [T3] evidence, not definitive207- **hERG flag**: If hERG = Active, flag prominently (cardiac safety risk)208- **AMES flag**: If AMES = Active, flag prominently (mutagenicity concern)209- **DILI flag**: If DILI = Active, flag prominently (liver toxicity concern)210211### Output Table212213```markdown214### Toxicity Predictions [T3]215216| Endpoint | Prediction | Interpretation | Concern Level |217|----------|-----------|---------------|---------------|218| AMES Mutagenicity | Inactive | No mutagenic signal | Low |219| Carcinogenicity | Inactive | No carcinogenic signal | Low |220| ClinTox | Active | Clinical toxicity signal | HIGH |221| DILI | Active | Drug-induced liver injury risk | HIGH |222| LD50 (Zhu) | 2.45 log(mg/kg) | ~282 mg/kg (moderate) | Medium |223| Skin Reaction | Inactive | No skin sensitization signal | Low |224| hERG Inhibition | Active | Cardiac arrhythmia risk | HIGH |225226*All predictions from ADMET-AI. Evidence tier: [T3] (computational prediction)*227```228229---230231## Phase 2: ADMET Properties232233**When**: SMILES is available234235**Objective**: Full ADMET characterization beyond toxicity236237### Tools Used238239| Tool | Properties Predicted | Parameter |240|------|---------------------|-----------|241| `ADMETAI_predict_BBB_penetrance` | Blood-brain barrier crossing probability | `smiles`: list[str] |242| `ADMETAI_predict_bioavailability` | Oral bioavailability (F20%, F30%) | `smiles`: list[str] |243| `ADMETAI_predict_clearance_distribution` | Clearance, VDss, half-life, PPB | `smiles`: list[str] |244| `ADMETAI_predict_CYP_interactions` | CYP1A2, 2C9, 2C19, 2D6, 3A4 inhibition/substrate | `smiles`: list[str] |245| `ADMETAI_predict_physicochemical_properties` | LogP, LogD, LogS, MW, pKa | `smiles`: list[str] |246| `ADMETAI_predict_solubility_lipophilicity_hydration` | Aqueous solubility, lipophilicity, hydration free energy | `smiles`: list[str] |247248### Workflow2492501. Call all 6 ADMET tools in parallel (independent calls)2512. Compile results into Absorption / Distribution / Metabolism / Excretion sections2523. Assess Lipinski Rule of 5 compliance from physicochemical properties2534. Flag drug-drug interaction risks from CYP inhibition profiles254255### Decision Logic256257- **BBB penetrant + toxicity**: If BBB = Yes and any CNS toxicity endpoint active, flag as neurotoxicity risk258- **Low bioavailability**: If F20% = Low, note absorption concerns259- **CYP inhibitor**: If CYP3A4 inhibitor = Yes, flag high DDI risk260- **Lipinski violations**: Count violations and report drug-likeness assessment261262### Output Format263264```markdown265### ADMET Profile [T3]266267#### Absorption268| Property | Value | Interpretation |269|----------|-------|----------------|270| BBB Penetrance | Yes | Crosses blood-brain barrier |271| Bioavailability (F20%) | 85% | Good oral absorption |272273#### Distribution274| Property | Value | Interpretation |275|----------|-------|----------------|276| VDss | 1.2 L/kg | Moderate tissue distribution |277| PPB | 92% | Highly protein bound |278279#### Metabolism280| CYP Enzyme | Substrate | Inhibitor |281|------------|-----------|-----------|282| CYP1A2 | No | No |283| CYP2C9 | Yes | No |284| CYP2C19 | No | No |285| CYP2D6 | No | No |286| CYP3A4 | Yes | Yes (DDI risk) |287288#### Excretion289| Property | Value | Interpretation |290|----------|-------|----------------|291| Clearance | 8.5 mL/min/kg | Moderate clearance |292| Half-life | 6.2 h | Moderate half-life |293```294295---296297## Phase 3: Toxicogenomics (CTD)298299**When**: Compound name is resolved300301**Objective**: Map chemical-gene-disease relationships from curated CTD data302303### Tools Used304305| Tool | Function | Parameter |306|------|----------|-----------|307| `CTD_get_chemical_gene_interactions` | Genes affected by chemical | `input_terms`: str (chemical name) |308| `CTD_get_chemical_diseases` | Diseases linked to chemical exposure | `input_terms`: str (chemical name) |309310### Workflow3113121. Call `CTD_get_chemical_gene_interactions(input_terms=compound_name)`3132. Call `CTD_get_chemical_diseases(input_terms=compound_name)`3143. Parse gene interactions: extract gene symbols, interaction types (increases/decreases expression, binding, etc.)3154. Parse disease associations: extract disease names, evidence types (marker/mechanism/therapeutic)3165. Identify most affected biological processes from gene list317318### Decision Logic319320- **Direct evidence vs inferred**: CTD separates curated direct evidence from inferred associations321- **Therapeutic vs toxic**: Disease associations can be therapeutic (drug treats disease) or adverse (chemical causes disease)322- **Gene interaction types**: Distinguish between expression changes, binding, and activity modulation323- **Prioritize marker/mechanism**: These indicate stronger causal evidence than simple associations324- **Grade curated as [T2]**: Direct curated CTD evidence from literature325- **Grade inferred as [T3]**: Computationally inferred associations326327### Output Format328329```markdown330### Toxicogenomics (CTD) [T2/T3]331332#### Chemical-Gene Interactions (Top 20)333| Gene | Interaction | Type | Evidence |334|------|------------|------|----------|335| CYP1A2 | increases expression | mRNA | [T2] curated |336| TP53 | affects activity | protein | [T2] curated |337| ... | ... | ... | ... |338339**Total interactions found**: 156340**Top affected pathways**: Xenobiotic metabolism, Apoptosis, DNA damage response341342#### Chemical-Disease Associations (Top 10)343| Disease | Association Type | Evidence |344|---------|-----------------|----------|345| Liver Neoplasms | marker/mechanism | [T2] curated |346| Contact Dermatitis | therapeutic | [T2] curated |347| ... | ... | ... |348```349350---351352## Phase 4: Regulatory Safety (FDA Labels)353354**When**: Compound has an approved drug name355356**Objective**: Extract regulatory safety information from FDA drug labels357358### Tools Used359360| Tool | Information Retrieved | Parameter |361|------|---------------------|-----------|362| `FDA_get_boxed_warning_info_by_drug_name` | Black box warnings (most serious) | `drug_name`: str |363| `FDA_get_contraindications_by_drug_name` | Absolute contraindications | `drug_name`: str |364| `FDA_get_adverse_reactions_by_drug_name` | Known adverse reactions | `drug_name`: str |365| `FDA_get_warnings_by_drug_name` | Warnings and precautions | `drug_name`: str |366| `FDA_get_nonclinical_toxicology_info_by_drug_name` | Animal toxicology data | `drug_name`: str |367| `FDA_get_carcinogenic_mutagenic_fertility_by_drug_name` | Carcinogenicity/mutagenicity/fertility data | `drug_name`: str |368369### Workflow3703711. Call all 6 FDA tools in parallel (independent queries by drug name)3722. Parse and structure each response3733. Prioritize: Boxed Warnings > Contraindications > Warnings > Adverse Reactions3744. All FDA label data is [T1] evidence (regulatory finding based on human/animal data)375376### Decision Logic377378- **Boxed warning present**: Flag as CRITICAL safety concern in executive summary379- **No FDA data**: Chemical may not be an approved drug; note "Not an FDA-approved drug" and continue with other phases380- **Multiple warnings**: Categorize by organ system (hepatic, cardiac, renal, CNS, etc.)381- **Nonclinical toxicology**: Grade as [T2] (animal data supporting human risk)382383### Output Format384385```markdown386### Regulatory Safety (FDA) [T1]387388#### Boxed Warning389**PRESENT** - Hepatotoxicity risk with doses >4g/day. Liver failure reported. [T1]390391#### Contraindications392- Severe hepatic impairment [T1]393- Known hypersensitivity [T1]394395#### Adverse Reactions (by frequency)396| Reaction | Frequency | Severity |397|----------|-----------|----------|398| Nausea | Common (>1%) | Mild |399| Hepatotoxicity | Rare (<0.1%) | Severe |400| ... | ... | ... |401402#### Nonclinical Toxicology [T2]403- **Carcinogenicity**: No carcinogenic potential in 2-year rat/mouse studies404- **Mutagenicity**: Negative in Ames assay and in vivo micronucleus test405- **Fertility**: No effects on fertility at doses up to 10x human dose406```407408---409410## Phase 5: Drug Safety Profile (DrugBank)411412**When**: Compound is a known drug413414**Objective**: Retrieve curated drug safety data from DrugBank415416### Tools Used417418| Tool | Information | Parameters |419|------|------------|------------|420| `drugbank_get_safety_by_drug_name_or_drugbank_id` | Toxicity, contraindications | `query`: str, `case_sensitive`: bool, `exact_match`: bool, `limit`: int |421422### Workflow4234241. Call `drugbank_get_safety_by_drug_name_or_drugbank_id(query=drug_name, case_sensitive=False, exact_match=False, limit=5)`4252. Parse toxicity information, overdose data, contraindications4263. Cross-reference with FDA data from Phase 4427428### Decision Logic429430- **Toxicity field**: Contains LD50 values, overdose symptoms, organ toxicity data431- **DrugBank ID**: Note if found for cross-referencing432- **Conflict with FDA**: If DrugBank and FDA disagree, note discrepancy and defer to FDA [T1]433- **Not found**: Chemical may not be in DrugBank; continue with other phases434435---436437## Phase 6: Chemical-Protein Interactions (STITCH)438439**When**: Compound can be identified by name or SMILES440441**Objective**: Map chemical-protein interaction network for off-target assessment442443### Tools Used444445| Tool | Function | Parameters |446|------|----------|------------|447| `STITCH_resolve_identifier` | Resolve chemical name to STITCH ID | `identifier`: str, `species`: int (9606=human) |448| `STITCH_get_chemical_protein_interactions` | Get chemical-protein interactions | `identifiers`: list[str], `species`: int, `required_score`: int |449| `STITCH_get_interaction_partners` | Get interaction network | `identifiers`: list[str], `species`: int, `limit`: int |450451### Workflow4524531. Resolve compound: `STITCH_resolve_identifier(identifier=compound_name, species=9606)`4542. Get interactions: `STITCH_get_chemical_protein_interactions(identifiers=[stitch_id], species=9606, required_score=700)`4553. Identify off-target proteins (not the intended drug target)4564. Flag safety-relevant targets: hERG (cardiac), CYP enzymes (metabolism), nuclear receptors (endocrine)457458### Decision Logic459460- **High confidence (>900)**: Well-established interaction [T2]461- **Medium confidence (700-900)**: Probable interaction [T3]462- **Low confidence (400-700)**: Possible interaction, needs validation [T4]463- **Safety-relevant targets**: Flag interactions with known safety targets464- **No STITCH data**: Chemical may be too novel; note and continue465466---467468## Phase 7: Structural Alerts (ChEMBL)469470**When**: ChEMBL molecule ID is available (from Phase 0)471472**Objective**: Check for known toxic substructures473474### Tools Used475476| Tool | Function | Parameters |477|------|----------|------------|478| `ChEMBL_search_compound_structural_alerts` | Find structural alert matches | `molecule_chembl_id`: str, `limit`: int |479480### Workflow4814821. If ChEMBL ID available: `ChEMBL_search_compound_structural_alerts(molecule_chembl_id=chembl_id, limit=20)`4832. Parse alert types: PAINS (pan-assay interference), Brenk (medicinal chemistry), Glaxo (GSK structural alerts)4843. Categorize severity: Some alerts are informational, others indicate likely toxicity485486### Decision Logic487488- **PAINS alerts**: May cause false positives in screening; note for medicinal chemistry489- **Brenk alerts**: Known problematic substructures; flag if present490- **No alerts**: Good sign but not definitive proof of safety491- **No ChEMBL ID**: Skip this phase gracefully; note "structural alert analysis not available"492493---494495## Synthesis: Integrated Risk Assessment (MANDATORY)496497**Always the final section**. Integrates all evidence into actionable risk classification.498499### Risk Classification Matrix500501| Risk Level | Criteria |502|-----------|----------|503| **CRITICAL** | FDA boxed warning present OR multiple [T1] toxicity findings OR active DILI + active hERG |504| **HIGH** | FDA warnings present OR [T2] animal toxicity OR multiple active ADMET endpoints |505| **MEDIUM** | Some [T3] predictions positive OR CTD disease associations OR structural alerts |506| **LOW** | All ADMET endpoints negative AND no FDA/DrugBank safety flags AND no CTD concerns |507| **INSUFFICIENT DATA** | Fewer than 3 phases returned data; cannot make confident assessment |508509### Synthesis Template510511```markdown512## Integrated Risk Assessment513514### Overall Risk Classification: [HIGH]515516### Evidence Summary517| Dimension | Finding | Evidence Tier | Concern |518|-----------|---------|--------------|---------|519| ADMET Toxicity | DILI active, hERG active | [T3] | HIGH |520| FDA Label | Boxed warning for hepatotoxicity | [T1] | CRITICAL |521| CTD Toxicogenomics | 156 gene interactions, liver neoplasms | [T2] | HIGH |522| DrugBank | Known hepatotoxicity at high doses | [T2] | HIGH |523| STITCH | Binds CYP3A4, hERG | [T3] | MEDIUM |524| Structural Alerts | 2 Brenk alerts | [T3] | MEDIUM |525526### Key Safety Concerns5271. **Hepatotoxicity** [T1]: FDA boxed warning + ADMET-AI DILI prediction + CTD liver disease associations5282. **Cardiac Risk** [T3]: ADMET-AI hERG prediction + STITCH hERG interaction5293. **Drug Interactions** [T3]: CYP3A4 substrate/inhibitor, potential DDI risk530531### Data Gaps532- [ ] No in vivo genotoxicity data available533- [ ] STITCH interaction scores moderate (700-900)534- [ ] No environmental exposure data535536### Recommendations5371. Avoid doses >4g/day (hepatotoxicity threshold) [T1]5382. Monitor liver function in chronic use [T1]5393. Screen for CYP3A4 interactions before co-administration [T3]5404. Consider cardiac monitoring for at-risk patients [T3]541```542543---544545## Mandatory Completeness Checklist546547Before finalizing any report, verify:548549- [ ] **Phase 0**: Compound fully disambiguated (SMILES + CID at minimum)550- [ ] **Phase 1**: At least 5 toxicity endpoints reported or "prediction unavailable" noted551- [ ] **Phase 2**: ADMET profile with A/D/M/E sections or "not available" noted552- [ ] **Phase 3**: CTD queried; gene interactions and disease associations reported or "no data in CTD"553- [ ] **Phase 4**: FDA labels queried; results or "not an FDA-approved drug" noted554- [ ] **Phase 5**: DrugBank queried; results or "not found in DrugBank" noted555- [ ] **Phase 6**: STITCH queried; results or "no STITCH data available" noted556- [ ] **Phase 7**: Structural alerts checked or "ChEMBL ID not available" noted557- [ ] **Synthesis**: Risk classification provided with evidence summary558- [ ] **Evidence Grading**: All findings have [T1]-[T4] annotations559- [ ] **Data Gaps**: Explicitly listed in synthesis section560561---562563## Tool Parameter Reference564565**Critical Parameter Notes** (verified from source code):566567| Tool | Parameter Name | Type | Notes |568|------|---------------|------|-------|569| All ADMETAI tools | `smiles` | `list[str]` | Always a list, even for single compound |570| All CTD tools | `input_terms` | `str` | Chemical name, MeSH name, CAS RN, or MeSH ID |571| All FDA tools | `drug_name` | `str` | Brand or generic drug name |572| drugbank_get_safety_* | `query`, `case_sensitive`, `exact_match`, `limit` | str, bool, bool, int | All 4 required |573| STITCH_resolve_identifier | `identifier`, `species` | str, int | species=9606 for human |574| STITCH_get_chemical_protein_interactions | `identifiers`, `species`, `required_score` | list[str], int, int | required_score=400 default |575| PubChem_get_CID_by_compound_name | `name` | `str` | Compound name (not SMILES) |576| PubChem_get_compound_properties_by_CID | `cid` | `int` | Numeric CID |577| ChEMBL_search_compound_structural_alerts | `molecule_chembl_id` | `str` | ChEMBL ID (e.g., "CHEMBL112") |578579### Response Format Notes580581- **ADMET-AI**: Returns `{status: "success", data: {...}}` with prediction values582- **CTD**: Returns list of interaction/association objects583- **FDA**: Returns `{status, data}` with label text584- **DrugBank**: Returns `{data: [...]}` with drug records585- **STITCH**: Returns list of interaction objects with scores586- **PubChem CID lookup**: Returns `{IdentifierList: {CID: [...]}}` (may or may not have `data` wrapper)587- **PubChem properties**: Returns dict with `CID`, `MolecularWeight`, `ConnectivitySMILES`, `IUPACName`588589---590591## Fallback Strategies592593### Compound Resolution594- **Primary**: PubChem by name -> CID -> properties -> SMILES595- **Fallback 1**: ChEMBL search by name -> molecule -> SMILES596- **Fallback 2**: If SMILES provided directly, skip name resolution597598### Toxicity Prediction599- **Primary**: All 9 ADMET-AI endpoints600- **Fallback**: If ADMET-AI fails for a compound, note "prediction failed" and continue with database evidence601- **Note**: ADMET-AI may fail for very large or unusual SMILES602603### Regulatory Data604- **Primary**: FDA labels by drug name605- **Fallback**: If FDA returns no data, try alternative drug names (brand vs generic)606- **Note**: Non-drug chemicals (pesticides, industrial) will not have FDA labels607608### CTD Data609- **Primary**: Search by common chemical name610- **Fallback**: Try MeSH name if common name fails611- **Note**: Novel compounds may not be in CTD612613---614615## Common Use Patterns616617### Pattern 1: Novel Compound Assessment618```619Input: SMILES string for new molecule620Workflow: Phase 0 (SMILES->CID) -> Phase 1 (toxicity) -> Phase 2 (ADMET) -> Phase 7 (structural alerts) -> Synthesis621Output: Predictive safety profile for novel compound622```623624### Pattern 2: Approved Drug Safety Review625```626Input: Drug name (e.g., "Acetaminophen")627Workflow: All phases (0-7 + Synthesis)628Output: Complete safety dossier with regulatory + predictive + database evidence629```630631### Pattern 3: Environmental Chemical Risk632```633Input: Chemical name (e.g., "Bisphenol A")634Workflow: Phase 0 -> Phase 1 -> Phase 2 -> Phase 3 (CTD, key for env chemicals) -> Phase 6 -> Synthesis635Output: Environmental health risk assessment focused on gene-disease associations636```637638### Pattern 4: Batch Toxicity Screening639```640Input: Multiple SMILES strings641Workflow: Phase 0 -> Phase 1 (batch) -> Phase 2 (batch) -> Comparative table -> Synthesis642Output: Comparative toxicity table ranking compounds by safety643```644645### Pattern 5: Toxicogenomic Deep-Dive646```647Input: Chemical name + specific gene or disease interest648Workflow: Phase 0 -> Phase 3 (CTD expanded) -> Literature search -> Synthesis649Output: Detailed chemical-gene-disease mechanistic analysis650```651652---653654## Output Report Structure655656All analyses generate a structured markdown report with progressive sections:657658```markdown659# Chemical Safety & Toxicology Report: [Compound Name]660661**Generated**: YYYY-MM-DD HH:MM662**Compound**: [Name] | SMILES: [SMILES] | CID: [CID]663664## Executive Summary665[2-3 sentence overview with risk classification and key findings, all graded]666667## 1. Compound Identity668[Phase 0 results - disambiguation table]669670## 2. Predictive Toxicology671[Phase 1 results - ADMET-AI toxicity endpoints]672673## 3. ADMET Profile674[Phase 2 results - absorption, distribution, metabolism, excretion]675676## 4. Toxicogenomics677[Phase 3 results - CTD chemical-gene-disease relationships]678679## 5. Regulatory Safety680[Phase 4 results - FDA label information]681682## 6. Drug Safety Profile683[Phase 5 results - DrugBank data]684685## 7. Chemical-Protein Interactions686[Phase 6 results - STITCH network]687688## 8. Structural Alerts689[Phase 7 results - ChEMBL alerts]690691## 9. Integrated Risk Assessment692[Synthesis - risk classification, evidence summary, data gaps, recommendations]693694## Appendix: Methods and Data Sources695[Tool versions, databases queried, date of access]696```697698---699700## Limitations & Known Issues701702### Tool-Specific703- **ADMET-AI**: Predictions are computational [T3]; should not replace experimental testing704- **CTD**: Curated but may lag behind latest literature by 6-12 months705- **FDA**: Only covers FDA-approved drugs; not applicable to environmental chemicals or supplements706- **DrugBank**: Primarily drugs; limited coverage of industrial chemicals707- **STITCH**: Score thresholds affect sensitivity; lower scores increase false positives708- **ChEMBL**: Structural alerts require ChEMBL ID; not all compounds have one709710### Analysis711- **Novel compounds**: May only have ADMET-AI predictions (no database evidence)712- **Environmental chemicals**: FDA/DrugBank phases will be empty; rely on CTD and ADMET-AI713- **Batch mode**: ADMET-AI can handle batches; other tools require individual queries714- **Species specificity**: Most data is human-centric; animal data noted where applicable715716### Technical717- **SMILES validity**: Invalid SMILES will cause ADMET-AI failures718- **Name ambiguity**: Chemical names can be ambiguous; always verify with CID719- **Rate limits**: Some FDA endpoints may rate-limit for rapid queries720721---722723## Summary724725**Chemical Safety & Toxicology Assessment Skill** provides comprehensive safety evaluation by integrating:7267271. **Predictive toxicology** (ADMET-AI) - 9 tools covering toxicity, ADMET, physicochemical properties7282. **Toxicogenomics** (CTD) - Chemical-gene-disease relationship mapping7293. **Regulatory safety** (FDA) - 6 tools for label-based safety extraction7304. **Drug safety** (DrugBank) - Curated toxicity and contraindication data7315. **Chemical interactions** (STITCH) - Chemical-protein interaction networks7326. **Structural alerts** (ChEMBL) - Known toxic substructure detection733734**Outputs**: Structured markdown report with risk classification, evidence grading, and actionable recommendations735736**Best for**: Drug safety assessment, chemical hazard profiling, environmental toxicology, ADMET characterization, toxicogenomic analysis737738**Total tools integrated**: 25+ tools across 6 databases