ToolUniverse
Overview
ToolUniverse is a unified ecosystem that enables AI agents to function as research scientists by providing standardized access to 600+ scientific resources. Use this skill to discover, execute, and compose scientific tools across multiple research domains including bioinformatics, cheminformatics, genomics, structural biology, proteomics, and drug discovery.
Key Capabilities:
- Access 600+ scientific tools, models, datasets, and APIs
- Discover tools using natural language, semantic search, or keywords
- Execute tools through standardized AI-Tool Interaction Protocol
- Compose multi-step workflows for complex research problems
- Integration with Claude Desktop/Code via Model Context Protocol (MCP)
When to Use This Skill
Use this skill when:
- Searching for scientific tools by function or domain (e.g., "find protein structure prediction tools")
- Executing computational biology workflows (e.g., disease target identification, drug discovery, genomics analysis)
- Accessing scientific databases (OpenTargets, PubChem, UniProt, PDB, ChEMBL, KEGG, etc.)
- Composing multi-step research pipelines (e.g., target discovery → structure prediction → virtual screening)
- Working with bioinformatics, cheminformatics, or structural biology tasks
- Analyzing gene expression, protein sequences, molecular structures, or clinical data
- Performing literature searches, pathway enrichment, or variant annotation
- Building automated scientific research workflows
Quick Start
Basic Setup
from tooluniverse import ToolUniverse
# Initialize and load tools
tu = ToolUniverse()
tu.load_tools() # Loads 600+ scientific tools
# Discover tools
tools = tu.run({
"name": "Tool_Finder_Keyword",
"arguments": {
"description": "disease target associations",
"limit": 10
}
})
# Execute a tool
result = tu.run({
"name": "OpenTargets_get_associated_targets_by_disease_efoId",
"arguments": {"efoId": "EFO_0000537"} # Hypertension
})
Model Context Protocol (MCP)
For Claude Desktop/Code integration:
tooluniverse-smcp
Core Workflows
1. Tool Discovery
Find relevant tools for your research task:
Three discovery methods:
Tool_Finder - Embedding-based semantic search (requires GPU)
Tool_Finder_LLM - LLM-based semantic search (no GPU required)
Tool_Finder_Keyword - Fast keyword search
Example:
# Search by natural language description
tools = tu.run({
"name": "Tool_Finder_LLM",
"arguments": {
"description": "Find tools for RNA sequencing differential expression analysis",
"limit": 10
}
})
# Review available tools
for tool in tools:
print(f"{tool['name']}: {tool['description']}")
See references/tool-discovery.md for:
- Detailed discovery methods and search strategies
- Domain-specific keyword suggestions
- Best practices for finding tools
2. Tool Execution
Execute individual tools through the standardized interface:
Example:
# Execute disease-target lookup
targets = tu.run({
"name": "OpenTargets_get_associated_targets_by_disease_efoId",
"arguments": {"efoId": "EFO_0000616"} # Breast cancer
})
# Get protein structure
structure = tu.run({
"name": "AlphaFold_get_structure",
"arguments": {"uniprot_id": "P12345"}
})
# Calculate molecular properties
properties = tu.run({
"name": "RDKit_calculate_descriptors",
"arguments": {"smiles": "CCO"} # Ethanol
})
See references/tool-execution.md for:
- Real-world execution examples across domains
- Tool parameter handling and validation
- Result processing and error handling
- Best practices for production use
3. Tool Composition and Workflows
Compose multiple tools for complex research workflows:
Drug Discovery Example:
# 1. Find disease targets
targets = tu.run({
"name": "OpenTargets_get_associated_targets_by_disease_efoId",
"arguments": {"efoId": "EFO_0000616"}
})
# 2. Get protein structures
structures = []
for target in targets[:5]:
structure = tu.run({
"name": "AlphaFold_get_structure",
"arguments": {"uniprot_id": target['uniprot_id']}
})
structures.append(structure)
# 3. Screen compounds
hits = []
for structure in structures:
compounds = tu.run({
"name": "ZINC_virtual_screening",
"arguments": {
"structure": structure,
"library": "lead-like",
"top_n": 100
}
})
hits.extend(compounds)
# 4. Evaluate drug-likeness
drug_candidates = []
for compound in hits:
props = tu.run({
"name": "RDKit_calculate_drug_properties",
"arguments": {"smiles": compound['smiles']}
})
if props['lipinski_pass']:
drug_candidates.append(compound)
See references/tool-composition.md for:
- Complete workflow examples (drug discovery, genomics, clinical)
- Sequential and parallel tool composition patterns
- Output processing hooks
- Workflow best practices
Scientific Domains
ToolUniverse supports 600+ tools across major scientific domains:
Bioinformatics:
- Sequence analysis, alignment, BLAST
- Gene expression (RNA-seq, DESeq2)
- Pathway enrichment (KEGG, Reactome, GO)
- Variant annotation (VEP, ClinVar)
Cheminformatics:
- Molecular descriptors and fingerprints
- Drug discovery and virtual screening
- ADMET prediction and drug-likeness
- Chemical databases (PubChem, ChEMBL, ZINC)
Structural Biology:
- Protein structure prediction (AlphaFold)
- Structure retrieval (PDB)
- Binding site detection
- Protein-protein interactions
Proteomics:
- Mass spectrometry analysis
- Protein databases (UniProt, STRING)
- Post-translational modifications
Genomics:
- Genome assembly and annotation
- Copy number variation
- Clinical genomics workflows
Medical/Clinical:
- Disease databases (OpenTargets, OMIM)
- Clinical trials and FDA data
- Variant classification
See references/domains.md for:
- Complete domain categorization
- Tool examples by discipline
- Cross-domain applications
- Search strategies by domain
Reference Documentation
This skill includes comprehensive reference files that provide detailed information for specific aspects:
references/installation.md - Installation, setup, MCP configuration, platform integration
references/tool-discovery.md - Discovery methods, search strategies, listing tools
references/tool-execution.md - Execution patterns, real-world examples, error handling
references/tool-composition.md - Workflow composition, complex pipelines, parallel execution
references/domains.md - Tool categorization by domain, use case examples
references/api_reference.md - Python API documentation, hooks, protocols
Workflow: When helping with specific tasks, reference the appropriate file for detailed instructions. For example, if searching for tools, consult references/tool-discovery.md for search strategies.
Example Scripts
Two executable example scripts demonstrate common use cases:
scripts/example_tool_search.py - Demonstrates all three discovery methods:
- Keyword-based search
- LLM-based search
- Domain-specific searches
- Getting detailed tool information
scripts/example_workflow.py - Complete workflow examples:
- Drug discovery pipeline (disease → targets → structures → screening → candidates)
- Genomics analysis (expression data → differential analysis → pathways)
Run examples to understand typical usage patterns and workflow composition.
Best Practices
Tool Discovery:
- Start with broad searches, then refine based on results
- Use
Tool_Finder_Keyword for fast searches with known terms
- Use
Tool_Finder_LLM for complex semantic queries
- Set appropriate
limit parameter (default: 10)
Tool Execution:
- Always verify tool parameters before execution
- Implement error handling for production workflows
- Validate input data formats (SMILES, UniProt IDs, gene symbols)
- Check result types and structures
Workflow Composition:
- Test each step individually before composing full workflows
- Implement checkpointing for long workflows
- Consider rate limits for remote APIs
- Use parallel execution when tools are independent
Integration:
- Initialize ToolUniverse once and reuse the instance
- Call
load_tools() once at startup
- Cache frequently used tool information
- Enable logging for debugging
Key Terminology
- Tool: A scientific resource (model, dataset, API, package) accessible through ToolUniverse
- Tool Discovery: Finding relevant tools using search methods (Finder, LLM, Keyword)
- Tool Execution: Running a tool with specific arguments via
tu.run()
- Tool Composition: Chaining multiple tools for multi-step workflows
- MCP: Model Context Protocol for integration with Claude Desktop/Code
- AI-Tool Interaction Protocol: Standardized interface for LLM-tool communication
Resources
Converted and distributed by TomeVault — claim your Tome and manage your conversions.
1---2name: tooluniverse3description: Use this skill when working with scientific research tools and workflows across bioinformatics, cheminformatics, genomics, structural biology, proteomics, and drug discovery. This skill provides access to 600+ scientific tools including machine learning models, datasets, APIs, and analysis packages. Use when searching for scientific tools, executing computational biology workflows, composing multi-step research pipelines, accessing databases like OpenTargets/PubChem/UniProt/PDB/ChEMBL, performing tool discovery for research tasks, or integrating scientific computational resources into LLM workflows.4---56# ToolUniverse78## Overview910ToolUniverse is a unified ecosystem that enables AI agents to function as research scientists by providing standardized access to 600+ scientific resources. Use this skill to discover, execute, and compose scientific tools across multiple research domains including bioinformatics, cheminformatics, genomics, structural biology, proteomics, and drug discovery.1112**Key Capabilities:**13- Access 600+ scientific tools, models, datasets, and APIs14- Discover tools using natural language, semantic search, or keywords15- Execute tools through standardized AI-Tool Interaction Protocol16- Compose multi-step workflows for complex research problems17- Integration with Claude Desktop/Code via Model Context Protocol (MCP)1819## When to Use This Skill2021Use this skill when:22- Searching for scientific tools by function or domain (e.g., "find protein structure prediction tools")23- Executing computational biology workflows (e.g., disease target identification, drug discovery, genomics analysis)24- Accessing scientific databases (OpenTargets, PubChem, UniProt, PDB, ChEMBL, KEGG, etc.)25- Composing multi-step research pipelines (e.g., target discovery → structure prediction → virtual screening)26- Working with bioinformatics, cheminformatics, or structural biology tasks27- Analyzing gene expression, protein sequences, molecular structures, or clinical data28- Performing literature searches, pathway enrichment, or variant annotation29- Building automated scientific research workflows3031## Quick Start3233### Basic Setup34```python35from tooluniverse import ToolUniverse3637# Initialize and load tools38tu = ToolUniverse()39tu.load_tools() # Loads 600+ scientific tools4041# Discover tools42tools = tu.run({43 "name": "Tool_Finder_Keyword",44 "arguments": {45 "description": "disease target associations",46 "limit": 1047 }48})4950# Execute a tool51result = tu.run({52 "name": "OpenTargets_get_associated_targets_by_disease_efoId",53 "arguments": {"efoId": "EFO_0000537"} # Hypertension54})55```5657### Model Context Protocol (MCP)58For Claude Desktop/Code integration:59```bash60tooluniverse-smcp61```6263## Core Workflows6465### 1. Tool Discovery6667Find relevant tools for your research task:6869**Three discovery methods:**70- `Tool_Finder` - Embedding-based semantic search (requires GPU)71- `Tool_Finder_LLM` - LLM-based semantic search (no GPU required)72- `Tool_Finder_Keyword` - Fast keyword search7374**Example:**75```python76# Search by natural language description77tools = tu.run({78 "name": "Tool_Finder_LLM",79 "arguments": {80 "description": "Find tools for RNA sequencing differential expression analysis",81 "limit": 1082 }83})8485# Review available tools86for tool in tools:87 print(f"{tool['name']}: {tool['description']}")88```8990**See `references/tool-discovery.md` for:**91- Detailed discovery methods and search strategies92- Domain-specific keyword suggestions93- Best practices for finding tools9495### 2. Tool Execution9697Execute individual tools through the standardized interface:9899**Example:**100```python101# Execute disease-target lookup102targets = tu.run({103 "name": "OpenTargets_get_associated_targets_by_disease_efoId",104 "arguments": {"efoId": "EFO_0000616"} # Breast cancer105})106107# Get protein structure108structure = tu.run({109 "name": "AlphaFold_get_structure",110 "arguments": {"uniprot_id": "P12345"}111})112113# Calculate molecular properties114properties = tu.run({115 "name": "RDKit_calculate_descriptors",116 "arguments": {"smiles": "CCO"} # Ethanol117})118```119120**See `references/tool-execution.md` for:**121- Real-world execution examples across domains122- Tool parameter handling and validation123- Result processing and error handling124- Best practices for production use125126### 3. Tool Composition and Workflows127128Compose multiple tools for complex research workflows:129130**Drug Discovery Example:**131```python132# 1. Find disease targets133targets = tu.run({134 "name": "OpenTargets_get_associated_targets_by_disease_efoId",135 "arguments": {"efoId": "EFO_0000616"}136})137138# 2. Get protein structures139structures = []140for target in targets[:5]:141 structure = tu.run({142 "name": "AlphaFold_get_structure",143 "arguments": {"uniprot_id": target['uniprot_id']}144 })145 structures.append(structure)146147# 3. Screen compounds148hits = []149for structure in structures:150 compounds = tu.run({151 "name": "ZINC_virtual_screening",152 "arguments": {153 "structure": structure,154 "library": "lead-like",155 "top_n": 100156 }157 })158 hits.extend(compounds)159160# 4. Evaluate drug-likeness161drug_candidates = []162for compound in hits:163 props = tu.run({164 "name": "RDKit_calculate_drug_properties",165 "arguments": {"smiles": compound['smiles']}166 })167 if props['lipinski_pass']:168 drug_candidates.append(compound)169```170171**See `references/tool-composition.md` for:**172- Complete workflow examples (drug discovery, genomics, clinical)173- Sequential and parallel tool composition patterns174- Output processing hooks175- Workflow best practices176177## Scientific Domains178179ToolUniverse supports 600+ tools across major scientific domains:180181**Bioinformatics:**182- Sequence analysis, alignment, BLAST183- Gene expression (RNA-seq, DESeq2)184- Pathway enrichment (KEGG, Reactome, GO)185- Variant annotation (VEP, ClinVar)186187**Cheminformatics:**188- Molecular descriptors and fingerprints189- Drug discovery and virtual screening190- ADMET prediction and drug-likeness191- Chemical databases (PubChem, ChEMBL, ZINC)192193**Structural Biology:**194- Protein structure prediction (AlphaFold)195- Structure retrieval (PDB)196- Binding site detection197- Protein-protein interactions198199**Proteomics:**200- Mass spectrometry analysis201- Protein databases (UniProt, STRING)202- Post-translational modifications203204**Genomics:**205- Genome assembly and annotation206- Copy number variation207- Clinical genomics workflows208209**Medical/Clinical:**210- Disease databases (OpenTargets, OMIM)211- Clinical trials and FDA data212- Variant classification213214**See `references/domains.md` for:**215- Complete domain categorization216- Tool examples by discipline217- Cross-domain applications218- Search strategies by domain219220## Reference Documentation221222This skill includes comprehensive reference files that provide detailed information for specific aspects:223224- **`references/installation.md`** - Installation, setup, MCP configuration, platform integration225- **`references/tool-discovery.md`** - Discovery methods, search strategies, listing tools226- **`references/tool-execution.md`** - Execution patterns, real-world examples, error handling227- **`references/tool-composition.md`** - Workflow composition, complex pipelines, parallel execution228- **`references/domains.md`** - Tool categorization by domain, use case examples229- **`references/api_reference.md`** - Python API documentation, hooks, protocols230231**Workflow:** When helping with specific tasks, reference the appropriate file for detailed instructions. For example, if searching for tools, consult `references/tool-discovery.md` for search strategies.232233## Example Scripts234235Two executable example scripts demonstrate common use cases:236237**`scripts/example_tool_search.py`** - Demonstrates all three discovery methods:238- Keyword-based search239- LLM-based search240- Domain-specific searches241- Getting detailed tool information242243**`scripts/example_workflow.py`** - Complete workflow examples:244- Drug discovery pipeline (disease → targets → structures → screening → candidates)245- Genomics analysis (expression data → differential analysis → pathways)246247Run examples to understand typical usage patterns and workflow composition.248249## Best Practices2502511. **Tool Discovery:**252 - Start with broad searches, then refine based on results253 - Use `Tool_Finder_Keyword` for fast searches with known terms254 - Use `Tool_Finder_LLM` for complex semantic queries255 - Set appropriate `limit` parameter (default: 10)2562572. **Tool Execution:**258 - Always verify tool parameters before execution259 - Implement error handling for production workflows260 - Validate input data formats (SMILES, UniProt IDs, gene symbols)261 - Check result types and structures2622633. **Workflow Composition:**264 - Test each step individually before composing full workflows265 - Implement checkpointing for long workflows266 - Consider rate limits for remote APIs267 - Use parallel execution when tools are independent2682694. **Integration:**270 - Initialize ToolUniverse once and reuse the instance271 - Call `load_tools()` once at startup272 - Cache frequently used tool information273 - Enable logging for debugging274275## Key Terminology276277- **Tool**: A scientific resource (model, dataset, API, package) accessible through ToolUniverse278- **Tool Discovery**: Finding relevant tools using search methods (Finder, LLM, Keyword)279- **Tool Execution**: Running a tool with specific arguments via `tu.run()`280- **Tool Composition**: Chaining multiple tools for multi-step workflows281- **MCP**: Model Context Protocol for integration with Claude Desktop/Code282- **AI-Tool Interaction Protocol**: Standardized interface for LLM-tool communication283284## Resources285286- **Official Website**: https://aiscientist.tools287- **GitHub**: https://github.com/mims-harvard/ToolUniverse288- **Documentation**: https://zitniklab.hms.harvard.edu/ToolUniverse/289- **Installation**: `uv pip install tooluniverse`290- **MCP Server**: `tooluniverse-smcp`291292---293> Converted and distributed by [TomeVault](https://tomevault.io/claim/lifangda) — claim your Tome and manage your conversions.294<!-- tomevault:4.0:skill_md:2026-04-11 -->