seqtk
Quick Start
- Command:
seqtk <subcommand> ... - Local executable:
/home/vimalinx/miniforge3/envs/bio/bin/seqtk - Version: 1.5-r133
- Full reference: See
references/help.md
When To Use This Tool
- Convert FASTQ to FASTA or normalize sequence formatting quickly.
- Subsample reads, extract named intervals or records, and perform quick read trimming.
- Run lightweight QC summaries without pulling in a larger workflow engine.
- Prefer
seqtkfor small, surgical sequence manipulations rather than full-blown alignment or variant tooling.
Common Patterns
# 1) Convert FASTQ to FASTA
seqtk seq -A reads.fq.gz > reads.fa
# 2) Subsample reads reproducibly
seqtk sample -s100 reads.fq.gz 100000 > reads.subset.fq
# 3) Extract named regions or sequences
seqtk subseq genome.fa regions.bed > subset.fa
# 4) Trim low-quality FASTQ ends
seqtk trimfq reads.fq.gz > reads.trimmed.fq
Recommended Workflow
- Pick the exact subcommand first:
seqfor format conversion,samplefor downsampling,subseqfor extraction,trimfqfor simple trimming, orfqchkfor quick QC. - Treat
seqtkas a stream-oriented filter and redirect output explicitly. - Use a fixed sampling seed when you need reproducible subsets.
- Validate output record counts or names before feeding the result into heavier downstream tools.
Guardrails
seqtkdoes not use global--helpor--versionflags; run it without arguments or use the specific subcommand help patterns instead.- Most subcommands write to stdout by default, so forgetting redirection can make pipelines look like they succeeded without leaving files behind.
subseqis convenient, but if you only need a few genomic intervals from a large indexed FASTA,samtools faidxmay be a better fit.- For paired-end subsampling, keep mate handling and random seed strategy consistent across both files.