Comprehensive Protein Analysis
Usage
1. MCP Server Definition
Use the same BioInfoToolsClient class as defined in the protein-blast-search skill.
2. Comprehensive Protein Analysis Workflow
This workflow combines InterProScan domain analysis with BLAST similarity search to provide a complete functional and evolutionary annotation of a protein sequence.
Workflow Steps:
- Validate Input - Check protein sequence format
- Run InterProScan - Identify functional domains and GO terms
- Run BLAST Search - Find similar sequences and homologs
- Integrate Results - Combine domain and homology information for comprehensive annotation
Implementation:
from datetime import timedelta
## Initialize client
client = BioInfoToolsClient(
"https://scp.intern-ai.org.cn/api/v1/mcp/17/BioInfo-Tools",
"<your-api-key>"
)
if not await client.connect():
print("connection failed")
exit()
## Input: Protein sequence to analyze
protein_sequence = """
MALWMRLLPLLALLALWGPDPAAAFVNQHLCGSHLVEALYLVCGERGFFYTPKTRREAEDLQVGQVELGGGPGAGSLQPLALEGSLQKRGIVEQCCTSICSLYQLENYCN
"""
sequence_id = "INS_HUMAN"
## Step 1, 2 & 3: Run comprehensive analysis (InterProScan + BLAST)
result = await client.session.call_tool(
"analyze_protein",
arguments={
"sequence": protein_sequence.strip(),
"sequence_id": sequence_id,
"databases": ["Pfam"], # InterProScan databases
"evalue": 1e-5, # BLAST E-value threshold (more stringent)
"max_hits": 10 # BLAST max hits
},
read_timeout_seconds=timedelta(seconds=1200) # Allow up to 20 minutes
)
## Step 4: Parse and display comprehensive results
result_data = client.parse_result(result)
print(f"{'='*80}")
print(f"Comprehensive Protein Analysis: {sequence_id}")
print(f"{'='*80}\n")
# InterProScan Results
ips_result = result_data.get("interproscan", {})
if ips_result.get("success"):
ips_data = ips_result.get("results", {})
domains = ips_data.get('domains', [])
go_terms = ips_data.get('go_terms', [])
print("=== DOMAIN ANALYSIS (InterProScan) ===")
print(f"Execution time: {ips_result.get('time_seconds', '?')} seconds")
print(f"Domains found: {len(domains)}")
print(f"GO annotations: {len(go_terms)}\n")
if domains:
print("Functional Domains:")
for domain in domains:
print(f" • {domain.get('name', 'N/A')} ({domain.get('database', 'N/A')})")
if domain.get('description'):
print(f" Description: {domain.get('description')}")
locations = domain.get('locations', [])
if locations:
loc = locations[0]
print(f" Position: {loc.get('start')}-{loc.get('end')} aa")
print()
if go_terms:
print("Gene Ontology Annotations:")
for go in go_terms[:5]: # Show top 5
print(f" • {go.get('id', 'N/A')}: {go.get('name', 'N/A')}")
print(f" Category: {go.get('category', 'N/A')}")
if len(go_terms) > 5:
print(f" ... and {len(go_terms) - 5} more")
print()
else:
print(f"❌ InterProScan failed: {ips_result.get('error', 'Unknown')}\n")
# BLAST Results
blast_result = result_data.get("blast", {})
if blast_result.get("success"):
hits = blast_result.get('hits', [])
print("=== HOMOLOGY SEARCH (BLAST) ===")
print(f"Execution time: {blast_result.get('time_seconds', '?')} seconds")
print(f"Similar sequences found: {blast_result.get('total_hits', 0)}")
print(f"E-value threshold: {1e-5}\n")
if hits:
print("Top Homologous Proteins:")
for i, hit in enumerate(hits[:5], 1):
print(f" {i}. {hit['uniprot_id']} - {hit.get('organism', 'N/A')}")
print(f" Description: {hit['description']}")
print(f" Identity: {hit['identity_percent']:.1f}%, E-value: {hit['evalue']:.2e}")
if len(hits) > 5:
print(f" ... and {len(hits) - 5} more matches")
print()
else:
print("No significant homologs found (E-value threshold may be too stringent)\n")
else:
print(f"❌ BLAST failed: {blast_result.get('error', 'Unknown')}\n")
# Summary
print("=== FUNCTIONAL SUMMARY ===")
if domains:
print(f"Protein Family: {domains[0].get('name', 'Unknown')}")
if hits:
most_similar = hits[0]
print(f"Most Similar Protein: {most_similar['uniprot_id']} ({most_similar['identity_percent']:.1f}% identity)")
print(f"Organism: {most_similar.get('organism', 'Unknown')}")
print(f"{'='*80}")
await client.disconnect()
Tool Descriptions
BioInfo-Tools Server:
analyze_protein: Comprehensive protein analysis combining InterProScan and BLAST
- Args:
sequence (str): Protein sequence in amino acid single-letter code
sequence_id (str, optional): Identifier for the query sequence
databases (list, optional): InterProScan databases (default: ["Pfam"])
evalue (float, optional): BLAST E-value threshold (default: 0.01)
max_hits (int, optional): Maximum BLAST hits (default: 10)
- Returns:
interproscan (dict): InterProScan analysis results
success (bool): Whether InterProScan completed
results (dict): Domains and GO terms
time_seconds (float): Execution time
blast (dict): BLAST search results
success (bool): Whether BLAST completed
hits (list): Similar proteins
total_hits (int): Number of matches
time_seconds (float): Execution time
Input/Output
Input:
sequence: Protein sequence (amino acid single-letter code)
sequence_id: Optional identifier for the query
databases: List of InterProScan databases to query
evalue: BLAST E-value threshold (lower = more stringent)
max_hits: Maximum number of BLAST hits to return
Output:
- InterProScan Results:
- Functional domains with positions
- Protein family classifications
- Gene Ontology annotations
- BLAST Results:
- Homologous proteins across species
- Sequence identity and alignment statistics
- Evolutionary relationships
Analysis Strategy
This comprehensive approach provides:
Structural Information (InterProScan):
- Domain architecture and organization
- Functional motifs and active sites
- Protein family membership
Evolutionary Context (BLAST):
- Homologs in other species
- Sequence conservation patterns
- Potential orthologs and paralogs
Functional Prediction:
- Combining domain and homology information
- GO term annotations for molecular function
- Biological process involvement
Performance Notes
- Total execution time: 2-20 minutes depending on sequence length
- InterProScan: 30 seconds to 15 minutes
- BLAST: 10-90 seconds
- Both run sequentially in this workflow
- Timeout recommendation: Set to at least 1200 seconds (20 minutes)
- E-value tuning: Use lower E-values (e.g., 1e-10) for highly conserved proteins, higher (e.g., 0.01) for divergent families
Use Cases
- Complete functional annotation of unknown proteins
- Validate predicted protein functions
- Study protein evolution and conservation
- Identify potential drug targets
- Annotate proteomes and genome sequences
- Compare protein function across species
Interpretation Tips
- High domain coverage + high homology: Well-characterized protein with known function
- Domains but no homologs: Novel protein with conserved domains, function can be inferred from domains
- Homologs but no domains: May need more sensitive domain detection or represents a novel fold
- Neither domains nor homologs: Potentially novel protein, may require experimental characterization
1---2name: comprehensive-protein-analysis3description: Comprehensive protein analysis combining InterProScan domain identification with BLAST similarity search to provide complete functional and evolutionary annotation.4license: MIT license5---6
7# Comprehensive Protein Analysis
8
9## Usage
10
11### 1. MCP Server Definition
12
13Use the same `BioInfoToolsClient` class as defined in the protein-blast-search skill.
14
15### 2. Comprehensive Protein Analysis Workflow
16
17This workflow combines InterProScan domain analysis with BLAST similarity search to provide a complete functional and evolutionary annotation of a protein sequence.
18
19**Workflow Steps:**
20
211. **Validate Input** - Check protein sequence format
222. **Run InterProScan** - Identify functional domains and GO terms
233. **Run BLAST Search** - Find similar sequences and homologs
244. **Integrate Results** - Combine domain and homology information for comprehensive annotation
25
26**Implementation:**
27
28```python
29from datetime import timedelta
30
31## Initialize client
32client = BioInfoToolsClient(
33 "https://scp.intern-ai.org.cn/api/v1/mcp/17/BioInfo-Tools",
34 "<your-api-key>"
35)
36
37if not await client.connect():
38 print("connection failed")
39 exit()
40
41## Input: Protein sequence to analyze
42protein_sequence = """
43MALWMRLLPLLALLALWGPDPAAAFVNQHLCGSHLVEALYLVCGERGFFYTPKTRREAEDLQVGQVELGGGPGAGSLQPLALEGSLQKRGIVEQCCTSICSLYQLENYCN
44"""
45
46sequence_id = "INS_HUMAN"
47
48## Step 1, 2 & 3: Run comprehensive analysis (InterProScan + BLAST)
49result = await client.session.call_tool(
50 "analyze_protein",
51 arguments={
52 "sequence": protein_sequence.strip(),
53 "sequence_id": sequence_id,
54 "databases": ["Pfam"], # InterProScan databases
55 "evalue": 1e-5, # BLAST E-value threshold (more stringent)
56 "max_hits": 10 # BLAST max hits
57 },
58 read_timeout_seconds=timedelta(seconds=1200) # Allow up to 20 minutes
59)
60
61## Step 4: Parse and display comprehensive results
62result_data = client.parse_result(result)
63
64print(f"{'='*80}")
65print(f"Comprehensive Protein Analysis: {sequence_id}")
66print(f"{'='*80}\n")
67
68# InterProScan Results
69ips_result = result_data.get("interproscan", {})
70if ips_result.get("success"):
71 ips_data = ips_result.get("results", {})
72 domains = ips_data.get('domains', [])
73 go_terms = ips_data.get('go_terms', [])
74
75 print("=== DOMAIN ANALYSIS (InterProScan) ===")
76 print(f"Execution time: {ips_result.get('time_seconds', '?')} seconds")
77 print(f"Domains found: {len(domains)}")
78 print(f"GO annotations: {len(go_terms)}\n")
79
80 if domains:
81 print("Functional Domains:")
82 for domain in domains:
83 print(f" • {domain.get('name', 'N/A')} ({domain.get('database', 'N/A')})")
84 if domain.get('description'):
85 print(f" Description: {domain.get('description')}")
86 locations = domain.get('locations', [])
87 if locations:
88 loc = locations[0]
89 print(f" Position: {loc.get('start')}-{loc.get('end')} aa")
90 print()
91
92 if go_terms:
93 print("Gene Ontology Annotations:")
94 for go in go_terms[:5]: # Show top 5
95 print(f" • {go.get('id', 'N/A')}: {go.get('name', 'N/A')}")
96 print(f" Category: {go.get('category', 'N/A')}")
97 if len(go_terms) > 5:
98 print(f" ... and {len(go_terms) - 5} more")
99 print()
100else:
101 print(f"❌ InterProScan failed: {ips_result.get('error', 'Unknown')}\n")
102
103# BLAST Results
104blast_result = result_data.get("blast", {})
105if blast_result.get("success"):
106 hits = blast_result.get('hits', [])
107
108 print("=== HOMOLOGY SEARCH (BLAST) ===")
109 print(f"Execution time: {blast_result.get('time_seconds', '?')} seconds")
110 print(f"Similar sequences found: {blast_result.get('total_hits', 0)}")
111 print(f"E-value threshold: {1e-5}\n")
112
113 if hits:
114 print("Top Homologous Proteins:")
115 for i, hit in enumerate(hits[:5], 1):
116 print(f" {i}. {hit['uniprot_id']} - {hit.get('organism', 'N/A')}")
117 print(f" Description: {hit['description']}")
118 print(f" Identity: {hit['identity_percent']:.1f}%, E-value: {hit['evalue']:.2e}")
119 if len(hits) > 5:
120 print(f" ... and {len(hits) - 5} more matches")
121 print()
122 else:
123 print("No significant homologs found (E-value threshold may be too stringent)\n")
124else:
125 print(f"❌ BLAST failed: {blast_result.get('error', 'Unknown')}\n")
126
127# Summary
128print("=== FUNCTIONAL SUMMARY ===")
129if domains:
130 print(f"Protein Family: {domains[0].get('name', 'Unknown')}")
131if hits:
132 most_similar = hits[0]
133 print(f"Most Similar Protein: {most_similar['uniprot_id']} ({most_similar['identity_percent']:.1f}% identity)")
134 print(f"Organism: {most_similar.get('organism', 'Unknown')}")
135print(f"{'='*80}")
136
137await client.disconnect()
138```
139
140### Tool Descriptions
141
142**BioInfo-Tools Server:**
143- `analyze_protein`: Comprehensive protein analysis combining InterProScan and BLAST
144 - Args:
145 - `sequence` (str): Protein sequence in amino acid single-letter code
146 - `sequence_id` (str, optional): Identifier for the query sequence
147 - `databases` (list, optional): InterProScan databases (default: ["Pfam"])
148 - `evalue` (float, optional): BLAST E-value threshold (default: 0.01)
149 - `max_hits` (int, optional): Maximum BLAST hits (default: 10)
150 - Returns:
151 - `interproscan` (dict): InterProScan analysis results
152 - `success` (bool): Whether InterProScan completed
153 - `results` (dict): Domains and GO terms
154 - `time_seconds` (float): Execution time
155 - `blast` (dict): BLAST search results
156 - `success` (bool): Whether BLAST completed
157 - `hits` (list): Similar proteins
158 - `total_hits` (int): Number of matches
159 - `time_seconds` (float): Execution time
160
161### Input/Output
162
163**Input:**
164- `sequence`: Protein sequence (amino acid single-letter code)
165- `sequence_id`: Optional identifier for the query
166- `databases`: List of InterProScan databases to query
167- `evalue`: BLAST E-value threshold (lower = more stringent)
168- `max_hits`: Maximum number of BLAST hits to return
169
170**Output:**
171- **InterProScan Results**:
172 - Functional domains with positions
173 - Protein family classifications
174 - Gene Ontology annotations
175- **BLAST Results**:
176 - Homologous proteins across species
177 - Sequence identity and alignment statistics
178 - Evolutionary relationships
179
180### Analysis Strategy
181
182This comprehensive approach provides:
183
1841. **Structural Information** (InterProScan):
185 - Domain architecture and organization
186 - Functional motifs and active sites
187 - Protein family membership
188
1892. **Evolutionary Context** (BLAST):
190 - Homologs in other species
191 - Sequence conservation patterns
192 - Potential orthologs and paralogs
193
1943. **Functional Prediction**:
195 - Combining domain and homology information
196 - GO term annotations for molecular function
197 - Biological process involvement
198
199### Performance Notes
200
201- **Total execution time**: 2-20 minutes depending on sequence length
202 - InterProScan: 30 seconds to 15 minutes
203 - BLAST: 10-90 seconds
204 - Both run sequentially in this workflow
205- **Timeout recommendation**: Set to at least 1200 seconds (20 minutes)
206- **E-value tuning**: Use lower E-values (e.g., 1e-10) for highly conserved proteins, higher (e.g., 0.01) for divergent families
207
208### Use Cases
209
210- Complete functional annotation of unknown proteins
211- Validate predicted protein functions
212- Study protein evolution and conservation
213- Identify potential drug targets
214- Annotate proteomes and genome sequences
215- Compare protein function across species
216
217### Interpretation Tips
218
219- **High domain coverage + high homology**: Well-characterized protein with known function
220- **Domains but no homologs**: Novel protein with conserved domains, function can be inferred from domains
221- **Homologs but no domains**: May need more sensitive domain detection or represents a novel fold
222- **Neither domains nor homologs**: Potentially novel protein, may require experimental characterization