LatchBio Integration
Overview
Latch is a Python framework for building and deploying bioinformatics workflows as serverless pipelines. Built on Flyte, create workflows with @workflow/@task decorators, manage cloud data with LatchFile/LatchDir, configure resources, and integrate Nextflow/Snakemake pipelines.
Core Capabilities
The Latch platform provides four main areas of functionality:
1. Workflow Creation and Deployment
- Define serverless workflows using Python decorators
- Support for native Python, Nextflow, and Snakemake pipelines
- Automatic containerization with Docker
- Auto-generated no-code user interfaces
- Version control and reproducibility
2. Data Management
- Cloud storage abstractions (LatchFile, LatchDir)
- Structured data organization with Registry (Projects → Tables → Records)
- Type-safe data operations with links and enums
- Automatic file transfer between local and cloud
- Glob pattern matching for file selection
3. Resource Configuration
- Pre-configured task decorators (@small_task, @large_task, @small_gpu_task, @large_gpu_task)
- Custom resource specifications (CPU, memory, GPU, storage)
- GPU support (K80, V100, A100)
- Timeout and storage configuration
- Cost optimization strategies
4. Verified Workflows
- Production-ready pre-built pipelines
- Bulk RNA-seq, DESeq2, pathway analysis
- AlphaFold and ColabFold for protein structure prediction
- Single-cell tools (ArchR, scVelo, emptyDropsR)
- CRISPR analysis, phylogenetics, and more
Quick Start
Installation and Setup
# Install Latch SDK
python3 -m uv pip install latch
# Login to Latch
latch login
# Initialize a new workflow
latch init my-workflow
# Register workflow to platform
latch register my-workflow
Prerequisites:
- Docker installed and running
- Latch account credentials
- Python 3.8+
Basic Workflow Example
from latch import workflow, small_task
from latch.types import LatchFile
@small_task
def process_file(input_file: LatchFile) -> LatchFile:
"""Process a single file"""
# Processing logic
return output_file
@workflow
def my_workflow(input_file: LatchFile) -> LatchFile:
"""
My bioinformatics workflow
Args:
input_file: Input data file
"""
return process_file(input_file=input_file)
When to Use This Skill
This skill should be used when encountering any of the following scenarios:
Workflow Development:
- "Create a Latch workflow for RNA-seq analysis"
- "Deploy my pipeline to Latch"
- "Convert my Nextflow pipeline to Latch"
- "Add GPU support to my workflow"
- Working with
@workflow, @task decorators
Data Management:
- "Organize my sequencing data in Latch Registry"
- "How do I use LatchFile and LatchDir?"
- "Set up sample tracking in Latch"
- Working with
latch:/// paths
Resource Configuration:
- "Configure GPU for AlphaFold on Latch"
- "My task is running out of memory"
- "How do I optimize workflow costs?"
- Working with task decorators
Verified Workflows:
- "Run AlphaFold on Latch"
- "Use DESeq2 for differential expression"
- "Available pre-built workflows"
- Using
latch.verified module
Detailed Documentation
This skill includes comprehensive reference documentation organized by capability:
references/workflow-creation.md
Read this for:
- Creating and registering workflows
- Task definition and decorators
- Supporting Python, Nextflow, Snakemake
- Launch plans and conditional sections
- Workflow execution (CLI and programmatic)
- Multi-step and parallel pipelines
- Troubleshooting registration issues
Key topics:
latch init and latch register commands
@workflow and @task decorators
- LatchFile and LatchDir basics
- Type annotations and docstrings
- Launch plans with preset parameters
- Conditional UI sections
references/data-management.md
Read this for:
- Cloud storage with LatchFile and LatchDir
- Registry system (Projects, Tables, Records)
- Linked records and relationships
- Enum and typed columns
- Bulk operations and transactions
- Integration with workflows
- Account and workspace management
Key topics:
latch:/// path format
- File transfer and glob patterns
- Creating and querying Registry tables
- Column types (string, number, file, link, enum)
- Record CRUD operations
- Workflow-Registry integration
references/resource-configuration.md
Read this for:
- Task resource decorators
- Custom CPU, memory, GPU configuration
- GPU types (K80, V100, A100)
- Timeout and storage settings
- Resource optimization strategies
- Cost-effective workflow design
- Monitoring and debugging
Key topics:
@small_task, @large_task, @small_gpu_task, @large_gpu_task
@custom_task with precise specifications
- Multi-GPU configuration
- Resource selection by workload type
- Platform limits and quotas
references/verified-workflows.md
Read this for:
- Pre-built production workflows
- Bulk RNA-seq and DESeq2
- AlphaFold and ColabFold
- Single-cell analysis (ArchR, scVelo)
- CRISPR editing analysis
- Pathway enrichment
- Integration with custom workflows
Key topics:
latch.verified module imports
- Available verified workflows
- Workflow parameters and options
- Combining verified and custom steps
- Version management
Common Workflow Patterns
Complete RNA-seq Pipeline
from latch import workflow, small_task, large_task
from latch.types import LatchFile, LatchDir
@small_task
def quality_control(fastq: LatchFile) -> LatchFile:
"""Run FastQC"""
return qc_output
@large_task
def alignment(fastq: LatchFile, genome: str) -> LatchFile:
"""STAR alignment"""
return bam_output
@small_task
def quantification(bam: LatchFile) -> LatchFile:
"""featureCounts"""
return counts
@workflow
def rnaseq_pipeline(
input_fastq: LatchFile,
genome: str,
output_dir: LatchDir
) -> LatchFile:
"""RNA-seq analysis pipeline"""
qc = quality_control(fastq=input_fastq)
aligned = alignment(fastq=qc, genome=genome)
return quantification(bam=aligned)
GPU-Accelerated Workflow
from latch import workflow, small_task, large_gpu_task
from latch.types import LatchFile
@small_task
def preprocess(input_file: LatchFile) -> LatchFile:
"""Prepare data"""
return processed
@large_gpu_task
def gpu_computation(data: LatchFile) -> LatchFile:
"""GPU-accelerated analysis"""
return results
@workflow
def gpu_pipeline(input_file: LatchFile) -> LatchFile:
"""Pipeline with GPU tasks"""
preprocessed = preprocess(input_file=input_file)
return gpu_computation(data=preprocessed)
Registry-Integrated Workflow
from latch import workflow, small_task
from latch.registry.table import Table
from latch.registry.record import Record
from latch.types import LatchFile
@small_task
def process_and_track(sample_id: str, table_id: str) -> str:
"""Process sample and update Registry"""
# Get sample from registry
table = Table.get(table_id=table_id)
records = Record.list(table_id=table_id, filter={"sample_id": sample_id})
sample = records[0]
# Process
input_file = sample.values["fastq_file"]
output = process(input_file)
# Update registry
sample.update(values={"status": "completed", "result": output})
return "Success"
@workflow
def registry_workflow(sample_id: str, table_id: str):
"""Workflow integrated with Registry"""
return process_and_track(sample_id=sample_id, table_id=table_id)
Best Practices
Workflow Design
- Use type annotations for all parameters
- Write clear docstrings (appear in UI)
- Start with standard task decorators, scale up if needed
- Break complex workflows into modular tasks
- Implement proper error handling
Data Management
- Use consistent folder structures
- Define Registry schemas before bulk entry
- Use linked records for relationships
- Store metadata in Registry for traceability
Resource Configuration
- Right-size resources (don't over-allocate)
- Use GPU only when algorithms support it
- Monitor execution metrics and optimize
- Design for parallel execution when possible
Development Workflow
- Test locally with Docker before registration
- Use version control for workflow code
- Document resource requirements
- Profile workflows to determine actual needs
Troubleshooting
Common Issues
Registration Failures:
- Ensure Docker is running
- Check authentication with
latch login
- Verify all dependencies in Dockerfile
- Use
--verbose flag for detailed logs
Resource Problems:
- Out of memory: Increase memory in task decorator
- Timeouts: Increase timeout parameter
- Storage issues: Increase ephemeral storage_gib
Data Access:
- Use correct
latch:/// path format
- Verify file exists in workspace
- Check permissions for shared workspaces
Type Errors:
- Add type annotations to all parameters
- Use LatchFile/LatchDir for file/directory parameters
- Ensure workflow return type matches actual return
Additional Resources
Support
For issues or questions:
- Check documentation links above
- Search GitHub issues
- Ask in Slack community
- Contact support@latch.bio
1---2name: latchbio-integration3description: Latch platform for bioinformatics workflows. Build pipelines with Latch SDK, @workflow/@task decorators, deploy serverless workflows, LatchFile/LatchDir, Nextflow/Snakemake integration.4---5
6# LatchBio Integration
7
8## Overview
9
10Latch is a Python framework for building and deploying bioinformatics workflows as serverless pipelines. Built on Flyte, create workflows with @workflow/@task decorators, manage cloud data with LatchFile/LatchDir, configure resources, and integrate Nextflow/Snakemake pipelines.
11
12## Core Capabilities
13
14The Latch platform provides four main areas of functionality:
15
16### 1. Workflow Creation and Deployment
17- Define serverless workflows using Python decorators
18- Support for native Python, Nextflow, and Snakemake pipelines
19- Automatic containerization with Docker
20- Auto-generated no-code user interfaces
21- Version control and reproducibility
22
23### 2. Data Management
24- Cloud storage abstractions (LatchFile, LatchDir)
25- Structured data organization with Registry (Projects → Tables → Records)
26- Type-safe data operations with links and enums
27- Automatic file transfer between local and cloud
28- Glob pattern matching for file selection
29
30### 3. Resource Configuration
31- Pre-configured task decorators (@small_task, @large_task, @small_gpu_task, @large_gpu_task)
32- Custom resource specifications (CPU, memory, GPU, storage)
33- GPU support (K80, V100, A100)
34- Timeout and storage configuration
35- Cost optimization strategies
36
37### 4. Verified Workflows
38- Production-ready pre-built pipelines
39- Bulk RNA-seq, DESeq2, pathway analysis
40- AlphaFold and ColabFold for protein structure prediction
41- Single-cell tools (ArchR, scVelo, emptyDropsR)
42- CRISPR analysis, phylogenetics, and more
43
44## Quick Start
45
46### Installation and Setup
47
48```bash
49# Install Latch SDK
50python3 -m uv pip install latch
51
52# Login to Latch
53latch login
54
55# Initialize a new workflow
56latch init my-workflow
57
58# Register workflow to platform
59latch register my-workflow
60```
61
62**Prerequisites:**
63- Docker installed and running
64- Latch account credentials
65- Python 3.8+
66
67### Basic Workflow Example
68
69```python
70from latch import workflow, small_task
71from latch.types import LatchFile
72
73@small_task
74def process_file(input_file: LatchFile) -> LatchFile:
75 """Process a single file"""
76 # Processing logic
77 return output_file
78
79@workflow
80def my_workflow(input_file: LatchFile) -> LatchFile:
81 """
82 My bioinformatics workflow
83
84 Args:
85 input_file: Input data file
86 """
87 return process_file(input_file=input_file)
88```
89
90## When to Use This Skill
91
92This skill should be used when encountering any of the following scenarios:
93
94**Workflow Development:**
95- "Create a Latch workflow for RNA-seq analysis"
96- "Deploy my pipeline to Latch"
97- "Convert my Nextflow pipeline to Latch"
98- "Add GPU support to my workflow"
99- Working with `@workflow`, `@task` decorators
100
101**Data Management:**
102- "Organize my sequencing data in Latch Registry"
103- "How do I use LatchFile and LatchDir?"
104- "Set up sample tracking in Latch"
105- Working with `latch:///` paths
106
107**Resource Configuration:**
108- "Configure GPU for AlphaFold on Latch"
109- "My task is running out of memory"
110- "How do I optimize workflow costs?"
111- Working with task decorators
112
113**Verified Workflows:**
114- "Run AlphaFold on Latch"
115- "Use DESeq2 for differential expression"
116- "Available pre-built workflows"
117- Using `latch.verified` module
118
119## Detailed Documentation
120
121This skill includes comprehensive reference documentation organized by capability:
122
123### references/workflow-creation.md
124**Read this for:**
125- Creating and registering workflows
126- Task definition and decorators
127- Supporting Python, Nextflow, Snakemake
128- Launch plans and conditional sections
129- Workflow execution (CLI and programmatic)
130- Multi-step and parallel pipelines
131- Troubleshooting registration issues
132
133**Key topics:**
134- `latch init` and `latch register` commands
135- `@workflow` and `@task` decorators
136- LatchFile and LatchDir basics
137- Type annotations and docstrings
138- Launch plans with preset parameters
139- Conditional UI sections
140
141### references/data-management.md
142**Read this for:**
143- Cloud storage with LatchFile and LatchDir
144- Registry system (Projects, Tables, Records)
145- Linked records and relationships
146- Enum and typed columns
147- Bulk operations and transactions
148- Integration with workflows
149- Account and workspace management
150
151**Key topics:**
152- `latch:///` path format
153- File transfer and glob patterns
154- Creating and querying Registry tables
155- Column types (string, number, file, link, enum)
156- Record CRUD operations
157- Workflow-Registry integration
158
159### references/resource-configuration.md
160**Read this for:**
161- Task resource decorators
162- Custom CPU, memory, GPU configuration
163- GPU types (K80, V100, A100)
164- Timeout and storage settings
165- Resource optimization strategies
166- Cost-effective workflow design
167- Monitoring and debugging
168
169**Key topics:**
170- `@small_task`, `@large_task`, `@small_gpu_task`, `@large_gpu_task`
171- `@custom_task` with precise specifications
172- Multi-GPU configuration
173- Resource selection by workload type
174- Platform limits and quotas
175
176### references/verified-workflows.md
177**Read this for:**
178- Pre-built production workflows
179- Bulk RNA-seq and DESeq2
180- AlphaFold and ColabFold
181- Single-cell analysis (ArchR, scVelo)
182- CRISPR editing analysis
183- Pathway enrichment
184- Integration with custom workflows
185
186**Key topics:**
187- `latch.verified` module imports
188- Available verified workflows
189- Workflow parameters and options
190- Combining verified and custom steps
191- Version management
192
193## Common Workflow Patterns
194
195### Complete RNA-seq Pipeline
196
197```python
198from latch import workflow, small_task, large_task
199from latch.types import LatchFile, LatchDir
200
201@small_task
202def quality_control(fastq: LatchFile) -> LatchFile:
203 """Run FastQC"""
204 return qc_output
205
206@large_task
207def alignment(fastq: LatchFile, genome: str) -> LatchFile:
208 """STAR alignment"""
209 return bam_output
210
211@small_task
212def quantification(bam: LatchFile) -> LatchFile:
213 """featureCounts"""
214 return counts
215
216@workflow
217def rnaseq_pipeline(
218 input_fastq: LatchFile,
219 genome: str,
220 output_dir: LatchDir
221) -> LatchFile:
222 """RNA-seq analysis pipeline"""
223 qc = quality_control(fastq=input_fastq)
224 aligned = alignment(fastq=qc, genome=genome)
225 return quantification(bam=aligned)
226```
227
228### GPU-Accelerated Workflow
229
230```python
231from latch import workflow, small_task, large_gpu_task
232from latch.types import LatchFile
233
234@small_task
235def preprocess(input_file: LatchFile) -> LatchFile:
236 """Prepare data"""
237 return processed
238
239@large_gpu_task
240def gpu_computation(data: LatchFile) -> LatchFile:
241 """GPU-accelerated analysis"""
242 return results
243
244@workflow
245def gpu_pipeline(input_file: LatchFile) -> LatchFile:
246 """Pipeline with GPU tasks"""
247 preprocessed = preprocess(input_file=input_file)
248 return gpu_computation(data=preprocessed)
249```
250
251### Registry-Integrated Workflow
252
253```python
254from latch import workflow, small_task
255from latch.registry.table import Table
256from latch.registry.record import Record
257from latch.types import LatchFile
258
259@small_task
260def process_and_track(sample_id: str, table_id: str) -> str:
261 """Process sample and update Registry"""
262 # Get sample from registry
263 table = Table.get(table_id=table_id)
264 records = Record.list(table_id=table_id, filter={"sample_id": sample_id})
265 sample = records[0]
266
267 # Process
268 input_file = sample.values["fastq_file"]
269 output = process(input_file)
270
271 # Update registry
272 sample.update(values={"status": "completed", "result": output})
273 return "Success"
274
275@workflow
276def registry_workflow(sample_id: str, table_id: str):
277 """Workflow integrated with Registry"""
278 return process_and_track(sample_id=sample_id, table_id=table_id)
279```
280
281## Best Practices
282
283### Workflow Design
2841. Use type annotations for all parameters
2852. Write clear docstrings (appear in UI)
2863. Start with standard task decorators, scale up if needed
2874. Break complex workflows into modular tasks
2885. Implement proper error handling
289
290### Data Management
2916. Use consistent folder structures
2927. Define Registry schemas before bulk entry
2938. Use linked records for relationships
2949. Store metadata in Registry for traceability
295
296### Resource Configuration
29710. Right-size resources (don't over-allocate)
29811. Use GPU only when algorithms support it
29912. Monitor execution metrics and optimize
30013. Design for parallel execution when possible
301
302### Development Workflow
30314. Test locally with Docker before registration
30415. Use version control for workflow code
30516. Document resource requirements
30617. Profile workflows to determine actual needs
307
308## Troubleshooting
309
310### Common Issues
311
312**Registration Failures:**
313- Ensure Docker is running
314- Check authentication with `latch login`
315- Verify all dependencies in Dockerfile
316- Use `--verbose` flag for detailed logs
317
318**Resource Problems:**
319- Out of memory: Increase memory in task decorator
320- Timeouts: Increase timeout parameter
321- Storage issues: Increase ephemeral storage_gib
322
323**Data Access:**
324- Use correct `latch:///` path format
325- Verify file exists in workspace
326- Check permissions for shared workspaces
327
328**Type Errors:**
329- Add type annotations to all parameters
330- Use LatchFile/LatchDir for file/directory parameters
331- Ensure workflow return type matches actual return
332
333## Additional Resources
334
335- **Official Documentation**: https://docs.latch.bio
336- **GitHub Repository**: https://github.com/latchbio/latch
337- **Slack Community**: Join Latch SDK workspace
338- **API Reference**: https://docs.latch.bio/api/latch.html
339- **Blog**: https://blog.latch.bio
340
341## Support
342
343For issues or questions:
3441. Check documentation links above
3452. Search GitHub issues
3463. Ask in Slack community
3474. Contact support@latch.bio