# Bioinformatics Sequences Analysis

> Use when analyzing bio sequences. Genomics, alignment.

- Skill: `loopyluci/bioinformatics-sequences-analysis` (Agent Skill)
- Install (CLI): `npx skillmds@latest add loopyluci/bioinformatics-sequences-analysis`
- Raw SKILL.md: https://api.skillmd.com/api/skills/loopyluci/bioinformatics-sequences-analysis/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- License: MIT
- Author: LoopyLuci (https://skillmd.com/u/loopyluci)
- Updated: 2026-09-08
- Page: https://skillmd.com/skills/loopyluci/bioinformatics-sequences-analysis

---


# Bioinformatics Sequences Analysis

## Overview
Systematically process biological sequences (DNA, RNA, proteins) using industry-standard tools. Covers FASTA/FASTQ/SAM/BAM/VCF formats, quality control, alignment, assembly, annotation, and statistical analysis. Produces reproducible analysis pipelines.

## When to Use
- "Analyze DNA/RNA/protein sequences"
- "Run BLAST search and interpret results"
- "Perform genome assembly from raw reads"
- "Do phylogenetic tree construction"

## File Formats
| Format | Content | Tools |
|--------|---------|-------|
| FASTA | Sequences with headers | biopython, samtools |
| FASTQ | Sequences + quality scores | fastp, fastqc |
| SAM/BAM | Aligned reads | samtools, picard |
| VCF | Variants | bcftools, gatk |

## Core Pipeline
1. Quality check (FastQC)
2. Trimming (fastp)
3. Alignment (BWA/STAR/minimap2)
4. Post-processing (samtools/picard)
5. Quantification/Annotation

## Common Pitfalls
1. Wrong aligner for data type — use splice-aware for RNA-seq
2. Skipping QC — wastes compute on bad data
3. Not indexing BAM files — samtools fails
4. Ignoring reference genome version mismatch

## Verification Checklist
- [ ] FASTQ validated with BioPython
- [ ] Quality passes thresholds
- [ ] Alignment rate >80%
- [ ] Output properly indexed
- [ ] Results reproducible
