Therapeutic Protein Designer
AI-guided de novo protein design using RFdiffusion backbone generation, ProteinMPNN sequence optimization, and structure validation for therapeutic protein development.
KEY PRINCIPLES:
- Structure-first - Generate backbone geometry before sequence
- Target-guided - Design binders with target structure in mind
- Iterative validation - Predict structure to validate designs
- Developability-aware - Consider aggregation, immunogenicity, expression
- Evidence-graded - Grade designs by confidence metrics
- Actionable output - Provide sequences ready for experimental testing
- English-first queries - Always use English terms in tool calls
Therapeutic protein design starts with the target interaction. What binding surface do you need to cover? A small pocket = nanobody or peptide. A large flat surface = designed protein. Stability, immunogenicity, and manufacturability constrain the design space.
LOOK UP, DON'T GUESS
When uncertain about any scientific fact, SEARCH databases first rather than reasoning from memory. A database-verified answer is always more reliable than a guess.
COMPUTE, DON'T DESCRIBE
When analysis requires computation (statistics, data processing, scoring, enrichment), write and run Python code via Bash. Don't describe what you would do — execute it and report actual results. Use ToolUniverse tools to retrieve data, then Python (pandas, scipy, statsmodels, matplotlib) to analyze it.
When to Use
Apply when user asks to:
- Design a protein binder, therapeutic protein, or scaffold
- Optimize a protein sequence for function
- Design a de novo enzyme
- Generate protein variants for target binding
Workflow Overview
Phase 1: Target Characterization
Get structure (PDB, EMDB cryo-EM, AlphaFold), identify binding epitope
Phase 2: Backbone Generation (RFdiffusion)
Define constraints, generate >= 5 backbones, filter by geometry
Phase 3: Sequence Design (ProteinMPNN)
Design >= 8 sequences per backbone, sample with temperature control
Phase 4: Structure Validation (ESMFold/AlphaFold2)
Predict structure, compare to backbone, assess pLDDT/pTM
Phase 5: Developability Assessment
Aggregation, pI, expression prediction
Phase 6: Report Synthesis
Ranked candidates, FASTA, experimental recommendations
Critical Requirements
Report-First Approach (MANDATORY)
- Create
[TARGET]_protein_design_report.md first with section headers
- Progressively update as designs are generated
- Output
[TARGET]_designed_sequences.fasta and [TARGET]_top_candidates.csv
Design Documentation (MANDATORY)
Every design MUST include: Sequence, Length, Target, Method, and Quality Metrics (pLDDT, pTM, MPNN score, binding prediction).
NVIDIA NIM Tools
| Tool |
Purpose |
Key Parameter |
NvidiaNIM_rfdiffusion (requires NVIDIA_API_KEY env var; free key at build.nvidia.com) |
Backbone generation |
diffusion_steps (NOT num_steps) |
NvidiaNIM_proteinmpnn (requires NVIDIA_API_KEY env var; free key at build.nvidia.com) |
Sequence design |
pdb_string (NOT pdb) |
ESMFold_predict_structure |
Fast validation |
sequence (NOT seq) |
NvidiaNIM_alphafold2 (requires NVIDIA_API_KEY env var; free key at build.nvidia.com) |
High-accuracy structure inference from sequence |
sequence, algorithm |
NvidiaNIM_esm2_650m (requires NVIDIA_API_KEY env var; free key at build.nvidia.com) |
Sequence embeddings |
sequences, format |
Common Parameter Mistakes
| Tool |
Wrong |
Correct |
NvidiaNIM_rfdiffusion (requires NVIDIA_API_KEY) |
num_steps=50 |
diffusion_steps=50 |
NvidiaNIM_proteinmpnn (requires NVIDIA_API_KEY) |
pdb=content |
pdb_string=content |
ESMFold_predict_structure |
seq="MVLS..." |
sequence="MVLS..." |
NvidiaNIM_alphafold2 (requires NVIDIA_API_KEY) |
seq="MVLS..." |
sequence="MVLS..." |
NVIDIA NIM Requirements
- API Key:
NVIDIA_API_KEY environment variable required
- Rate limits: 40 RPM (1.5 second minimum between calls)
- AlphaFold2 may return 202 (polling required); RFdiffusion and ESMFold are synchronous
Supporting Tools
| Tool |
Purpose |
Key Parameters |
PDBe_get_uniprot_mappings |
Find PDB structures |
uniprot_id |
RCSBData_get_entry |
Download PDB file |
pdb_id |
alphafold_get_prediction |
Get AlphaFold DB structure |
accession |
EMDB_search_structures |
Search cryo-EM maps |
query |
EMDB_get_structure |
Get entry details |
entry_id |
UniProt_get_entry_by_accession |
Get target sequence |
accession |
InterPro_get_protein_domains |
Get domains |
accession |
Evidence Grading
| Tier |
Criteria |
| T1 (best) |
pLDDT >85, pTM >0.8, low aggregation, neutral pI |
| T2 |
pLDDT >75, pTM >0.7, acceptable developability |
| T3 |
pLDDT >70, pTM >0.65, developability concerns |
| T4 |
Failed validation or major developability issues |
Completeness Checklist
Reference Files
- DESIGN_PROCEDURES.md - Phase-by-phase code examples, sampling parameters, fallback chains
- TOOLS_REFERENCE.md - Complete tool documentation with code examples
- EXAMPLES.md - Sample design workflows and outputs
- CHECKLIST.md - Detailed phase checklists and quality metrics
- design_templates.md - Report templates and output format examples
1---2name: tooluniverse-protein-therapeutic-design3description: AI-guided de novo protein design — RFdiffusion backbone generation, ProteinMPNN sequence design, structure validation (pLDDT, pTM, MPNN scores). Use for designing therapeutic protein binders, novel scaffolds, enzyme variants, and miniprotein/protein-interface design before experimental validation.4---5
6# Therapeutic Protein Designer
7
8AI-guided de novo protein design using RFdiffusion backbone generation, ProteinMPNN sequence optimization, and structure validation for therapeutic protein development.
9
10**KEY PRINCIPLES**:
111. **Structure-first** - Generate backbone geometry before sequence
122. **Target-guided** - Design binders with target structure in mind
133. **Iterative validation** - Predict structure to validate designs
144. **Developability-aware** - Consider aggregation, immunogenicity, expression
155. **Evidence-graded** - Grade designs by confidence metrics
166. **Actionable output** - Provide sequences ready for experimental testing
177. **English-first queries** - Always use English terms in tool calls
18
19Therapeutic protein design starts with the target interaction. What binding surface do you need to cover? A small pocket = nanobody or peptide. A large flat surface = designed protein. Stability, immunogenicity, and manufacturability constrain the design space.
20
21## LOOK UP, DON'T GUESS
22When uncertain about any scientific fact, SEARCH databases first rather than reasoning from memory. A database-verified answer is always more reliable than a guess.
23
24---
25
26## COMPUTE, DON'T DESCRIBE
27When analysis requires computation (statistics, data processing, scoring, enrichment), write and run Python code via Bash. Don't describe what you would do — execute it and report actual results. Use ToolUniverse tools to retrieve data, then Python (pandas, scipy, statsmodels, matplotlib) to analyze it.
28
29## When to Use
30
31Apply when user asks to:
32- Design a protein binder, therapeutic protein, or scaffold
33- Optimize a protein sequence for function
34- Design a de novo enzyme
35- Generate protein variants for target binding
36
37---
38
39## Workflow Overview
40
41```
42Phase 1: Target Characterization
43 Get structure (PDB, EMDB cryo-EM, AlphaFold), identify binding epitope
44
45Phase 2: Backbone Generation (RFdiffusion)
46 Define constraints, generate >= 5 backbones, filter by geometry
47
48Phase 3: Sequence Design (ProteinMPNN)
49 Design >= 8 sequences per backbone, sample with temperature control
50
51Phase 4: Structure Validation (ESMFold/AlphaFold2)
52 Predict structure, compare to backbone, assess pLDDT/pTM
53
54Phase 5: Developability Assessment
55 Aggregation, pI, expression prediction
56
57Phase 6: Report Synthesis
58 Ranked candidates, FASTA, experimental recommendations
59```
60
61---
62
63## Critical Requirements
64
65### Report-First Approach (MANDATORY)
661. Create `[TARGET]_protein_design_report.md` first with section headers
672. Progressively update as designs are generated
683. Output `[TARGET]_designed_sequences.fasta` and `[TARGET]_top_candidates.csv`
69
70### Design Documentation (MANDATORY)
71Every design MUST include: Sequence, Length, Target, Method, and Quality Metrics (pLDDT, pTM, MPNN score, binding prediction).
72
73---
74
75## NVIDIA NIM Tools
76
77| Tool | Purpose | Key Parameter |
78|------|---------|---------------|
79| `NvidiaNIM_rfdiffusion` *(requires NVIDIA_API_KEY env var; free key at build.nvidia.com)* | Backbone generation | `diffusion_steps` (NOT `num_steps`) |
80| `NvidiaNIM_proteinmpnn` *(requires NVIDIA_API_KEY env var; free key at build.nvidia.com)* | Sequence design | `pdb_string` (NOT `pdb`) |
81| `ESMFold_predict_structure` | Fast validation | `sequence` (NOT `seq`) |
82| `NvidiaNIM_alphafold2` *(requires NVIDIA_API_KEY env var; free key at build.nvidia.com)* | High-accuracy structure inference from sequence | `sequence`, `algorithm` |
83| `NvidiaNIM_esm2_650m` *(requires NVIDIA_API_KEY env var; free key at build.nvidia.com)* | Sequence embeddings | `sequences`, `format` |
84
85### Common Parameter Mistakes
86
87| Tool | Wrong | Correct |
88|------|-------|---------|
89| `NvidiaNIM_rfdiffusion` *(requires NVIDIA_API_KEY)* | `num_steps=50` | `diffusion_steps=50` |
90| `NvidiaNIM_proteinmpnn` *(requires NVIDIA_API_KEY)* | `pdb=content` | `pdb_string=content` |
91| `ESMFold_predict_structure` | `seq="MVLS..."` | `sequence="MVLS..."` |
92| `NvidiaNIM_alphafold2` *(requires NVIDIA_API_KEY)* | `seq="MVLS..."` | `sequence="MVLS..."` |
93
94### NVIDIA NIM Requirements
95- **API Key**: `NVIDIA_API_KEY` environment variable required
96- **Rate limits**: 40 RPM (1.5 second minimum between calls)
97- AlphaFold2 may return 202 (polling required); RFdiffusion and ESMFold are synchronous
98
99---
100
101## Supporting Tools
102
103| Tool | Purpose | Key Parameters |
104|------|---------|----------------|
105| `PDBe_get_uniprot_mappings` | Find PDB structures | `uniprot_id` |
106| `RCSBData_get_entry` | Download PDB file | `pdb_id` |
107| `alphafold_get_prediction` | Get AlphaFold DB structure | `accession` |
108| `EMDB_search_structures` | Search cryo-EM maps | `query` |
109| `EMDB_get_structure` | Get entry details | `entry_id` |
110| `UniProt_get_entry_by_accession` | Get target sequence | `accession` |
111| `InterPro_get_protein_domains` | Get domains | `accession` |
112
113---
114
115## Evidence Grading
116
117| Tier | Criteria |
118|------|----------|
119| T1 (best) | pLDDT >85, pTM >0.8, low aggregation, neutral pI |
120| T2 | pLDDT >75, pTM >0.7, acceptable developability |
121| T3 | pLDDT >70, pTM >0.65, developability concerns |
122| T4 | Failed validation or major developability issues |
123
124---
125
126## Completeness Checklist
127
128- [ ] Target structure obtained (PDB or predicted)
129- [ ] Binding epitope identified
130- [ ] >= 5 backbones generated, top 3-5 selected
131- [ ] >= 8 sequences per backbone, MPNN scores reported
132- [ ] All sequences validated (ESMFold), pLDDT/pTM reported, >= 3 passing
133- [ ] Developability assessed (aggregation, pI, expression)
134- [ ] Ranked candidate list, FASTA file, experimental recommendations
135
136---
137
138## Reference Files
139
140- **DESIGN_PROCEDURES.md** - Phase-by-phase code examples, sampling parameters, fallback chains
141- **TOOLS_REFERENCE.md** - Complete tool documentation with code examples
142- **EXAMPLES.md** - Sample design workflows and outputs
143- **CHECKLIST.md** - Detailed phase checklists and quality metrics
144- **design_templates.md** - Report templates and output format examples