Bioinformatics Sequences Analysis
Overview
Systematically process biological sequences (DNA, RNA, proteins) using industry-standard tools. Covers FASTA/FASTQ/SAM/BAM/VCF formats, quality control, alignment, assembly, annotation, and statistical analysis. Produces reproducible analysis pipelines.
When to Use
- "Analyze DNA/RNA/protein sequences"
- "Run BLAST search and interpret results"
- "Perform genome assembly from raw reads"
- "Do phylogenetic tree construction"
File Formats
| Format | Content | Tools |
|---|---|---|
| FASTA | Sequences with headers | biopython, samtools |
| FASTQ | Sequences + quality scores | fastp, fastqc |
| SAM/BAM | Aligned reads | samtools, picard |
| VCF | Variants | bcftools, gatk |
Core Pipeline
- Quality check (FastQC)
- Trimming (fastp)
- Alignment (BWA/STAR/minimap2)
- Post-processing (samtools/picard)
- Quantification/Annotation
Common Pitfalls
- Wrong aligner for data type — use splice-aware for RNA-seq
- Skipping QC — wastes compute on bad data
- Not indexing BAM files — samtools fails
- Ignoring reference genome version mismatch
Verification Checklist
- FASTQ validated with BioPython
- Quality passes thresholds
- Alignment rate >80%
- Output properly indexed
- Results reproducible