seqkit
Quick Start
- Command:
seqkit <subcommand> [options] <input> - Local executable:
/home/vimalinx/miniforge3/envs/bio/bin/seqkit - Version: 2.13.0
- Full reference: See references/help.md for complete subcommand options
When To Use This Tool
- General-purpose FASTA/FASTQ inspection, filtering, conversion, and manipulation.
- Quick sequence operations without writing custom scripts.
- Use it when you need one clean subcommand instead of an ad hoc awk or Python one-liner.
- Especially useful for stats, grep-like selection, subsequences, and format conversion.
Common Patterns
# 1) Basic FASTQ summary statistics
seqkit stats reads.fastq.gz
# 2) Filter sequences by minimum length and keep only IDs
seqkit seq -m 200 input.fa.gz
# 3) Search by ID or pattern
seqkit grep -p TP53 proteins.fa.gz
# 4) Convert FASTA/Q to tabular form
seqkit fx2tab reads.fastq.gz
# 5) Remove duplicates
seqkit rmdup input.fa.gz -o dedup.fa.gz
Recommended Workflow
- Start with
seqkit statsto understand the file you are about to manipulate. - Choose the narrowest subcommand that matches the operation instead of overloading shell text processing.
- Keep outputs compressed when appropriate; SeqKit handles compressed I/O directly.
- Use subcommand-specific help for anything beyond the common workflows.
Guardrails
seqkit versionis the version command;seqkit --versionis not the right interface here.- Use
seqkit <subcommand> --helpfor real usage details, because the top-level command is just a dispatcher. - Output compression is handled automatically from the filename suffix.
- SeqKit can parse IDs with its own regex rules, so be explicit if FASTA headers use unusual formats.