BLAST+ — Command-Line Sequence Search
Run local NCBI BLAST+ 2.17.0 searches end-to-end: build a database with
makeblastdb, search it with blastn / blastp / blastx / tblastn, emit
machine-parseable tabular output, and scope by taxonomy. For very large protein
searches, hand off to DIAMOND blastp --ultra-sensitive (100x–10,000x the
speed of BLAST, per the DIAMOND project). This is the CLI / local-database
skill; it is deliberately distinct from the Biopython web API and the gget
one-liner (see routing table below).
Bulk DB builds and large searches are CPU/IO-heavy and fully offline — good
candidates to run on local compute rather than burning API calls.
When to Use This Skill
Use this skill when the request involves any of:
- "BLAST these sequences", "run blastn/blastp/blastx/tblastn", "command-line BLAST"
- "build a local BLAST database", "makeblastdb", "index this FASTA for BLAST"
- "search my reads against a local nt/nr database", "get tabular BLAST hits I can parse"
- "scope the BLAST search to a taxon" (
-taxids / -negative_taxids)
- "BLAST is too slow on millions of proteins" → DIAMOND
blastp
- retrieving sequences out of a BLAST DB (
blastdbcmd, requires -parse_seqids)
Does NOT Trigger
Route adjacent requests to the right sibling skill instead of forcing BLAST+:
| The request is really about… |
Route to |
The web BLAST API (Bio.Blast.NCBIWWW.qblast), or scripting BLAST inside a Python pipeline with Bio.Blast parsing |
alterlab-biopython |
A quick one-liner BLAST/database lookup (gget blast, gene/structure/enrichment lookups) |
alterlab-gget |
| Unified programmatic access to many bio web services (UniProt, KEGG, Ensembl REST, NCBI eUtils) |
alterlab-bioservices |
| Building/searching a phylogenetic tree from sequences, not a similarity search |
alterlab-phylogenetics |
| Read alignment to a reference genome (BWA/minimap2 → BAM) and SAM/BAM handling |
alterlab-pysam |
| FASTQ→VCF variant calling pipeline |
alterlab-nf-core-sarek |
| Transcript-level RNA-seq quantification (salmon/kallisto) |
alterlab-rnaseq-quant |
| 16S/ITS amplicon classification (QIIME 2) |
alterlab-qiime2-amplicon |
| Protein structure prediction / embeddings (ESM, AlphaFold) |
alterlab-esm |
If the user explicitly says "web BLAST", "NCBIWWW", or "without installing
anything", they want alterlab-biopython, not this skill.
Quick Start
# 1. Build a protein DB (‑parse_seqids enables blastdbcmd retrieval + DIAMOND reuse)
makeblastdb -in proteins.fasta -dbtype prot -parse_seqids -out mydb -title "my proteins"
# 2. Search, tabular output you can parse, std 12 columns
blastp -query query.faa -db mydb -outfmt 6 -evalue 1e-5 -out hits.tsv
# 3. QC / summarize the tabular output (stdlib only)
uv run python scripts/parse_blast_tab.py hits.tsv --best-hit
-outfmt 6 is the canonical machine-readable format; its default columns are
the std set: qseqid sseqid pident length mismatch gapopen qstart qend sstart send evalue bitscore. Use -outfmt 7 for the same columns plus comment lines.
Choosing the Right Program
| Query |
Subject DB |
Program |
| nucleotide |
nucleotide |
blastn |
| protein |
protein |
blastp |
| nucleotide (translated) |
protein |
blastx |
| protein |
nucleotide (translated) |
tblastn |
-dbtype for makeblastdb is nucl for nucleotide subjects, prot for protein.
The Five Things People Get Wrong
-max_target_seqs is NOT a "top N best hits" filter. It is the number of
aligned sequences to keep, applied during the search as a heuristic cutoff;
ties are broken "by order of sequences in the database", not by score. Setting
-max_target_seqs 1 does not reliably return the single best hit. To get
the best hit, keep a generous value and pick the top row after sorting by
bitscore (see scripts/parse_blast_tab.py --best-hit). Default is 500.
- Wrong
-task for blastn. megablast (default) is for highly similar
sequences; use blastn for cross-species / more divergent hits and
blastn-short for queries < ~30 nt (primers, sgRNAs). dc-megablast is the
discontiguous option for inter-species comparison.
- Forgetting
-parse_seqids at DB-build time. Without it you cannot pull
sequences back out with blastdbcmd -entry, and DIAMOND cannot reuse the
sequence IDs cleanly. You cannot add it later without rebuilding.
- Quoting the
-outfmt custom column list for DIAMOND. BLAST+ wants the
spec quoted (-outfmt '6 qseqid sseqid pident evalue'); DIAMOND wants it
unquoted (--outfmt 6 qseqid sseqid pident evalue). Mixing these up is a
common silent error.
- Multithreading. Use
-num_threads N. For many small queries, set
-mt_mode 1 (split by query) so all threads stay busy; -mt_mode 0 (default,
split by database volume) suits few large queries. BLAST+ 2.15+ can choose
automatically, but set it explicitly when in doubt.
Full option reference, taxonomy scoping, and DB-prep details:
references/blast_cli.md.
Taxonomic Scoping
Restrict a search to (or away from) clades by NCBI taxid:
blastn -query q.fna -db nt -taxids 9606 -outfmt 6 -out human_only.tsv
blastp -query q.faa -db nr -negative_taxids 2 -outfmt 6 -out no_bacteria.tsv
Scoping by taxid requires a taxonomy-aware database (one built/downloaded with
its *.taxid mapping, e.g. NCBI's pre-formatted nt / nr). See
references/blast_cli.md.
DIAMOND — Fast Path for Large Protein Searches
When blastp / blastx against millions of proteins is too slow, DIAMOND is a
drop-in for protein-space search:
diamond makedb --in nr.faa -d nr_diamond
diamond blastp -d nr_diamond -q query.faa -o hits.tsv \
--ultra-sensitive --outfmt 6 qseqid sseqid pident length evalue bitscore
Sensitivity ladder (fast → most sensitive): --fast, --mid-sensitive,
--sensitive, --more-sensitive, --very-sensitive, --ultra-sensitive.
Use --ultra-sensitive when you need BLAST-comparable recall; default fast mode
trades sensitivity for speed. DIAMOND's --outfmt 6 is compatible with the
BLAST+ tabular parser below. Details and tradeoffs:
references/diamond.md.
Recommended Workflow
- Pick the program from the query/subject table above.
- Build the DB with
makeblastdb -parse_seqids (or download a pre-formatted
NCBI DB). For >~1M proteins, build a DIAMOND DB instead.
- Search with
-outfmt 6, an explicit -evalue threshold, the right
-task (blastn), and -num_threads. Add -taxids if scoping.
- Parse & QC with
scripts/parse_blast_tab.py — it sorts by bitscore,
extracts best-hit-per-query, applies identity/coverage/e-value filters, and
flags the -max_target_seqs pitfall if the column count looks truncated.
- Retrieve any hit sequence with
blastdbcmd -db mydb -entry <id> (needs -parse_seqids).
Verify Before Reporting
- Confirm
blastn -version / diamond version actually ran — never report hits
you did not produce.
- State the program,
-task, -evalue, and DB used; results are meaningless
without them.
- If you used
-max_target_seqs, confirm best-hit selection was done by
post-hoc bitscore sort, not by trusting the keep-count as a top-N.
- For DIAMOND results, note the sensitivity level used.
References
references/blast_cli.md — full BLAST+ 2.17.0 option
reference: programs, makeblastdb, -outfmt columns, -task, taxonomy
scoping, -mt_mode, blastdbcmd retrieval, and the -max_target_seqs caveat.
references/diamond.md — DIAMOND DB build, sensitivity
modes, output formats, and when to choose it over BLAST+.
- NCBI BLAST+ manual: https://www.ncbi.nlm.nih.gov/books/NBK569856/
- DIAMOND: https://github.com/bbuchfink/diamond
Part of the AlterLab Academic Skills suite.
1---2name: alterlab-blast3description: Runs NCBI BLAST+ 2.17.0 sequence searches from the command line: makeblastdb (with -parse_seqids), blastn/blastp/blastx/tblastn with tabular -outfmt 6/7 for parsing, correct -task choice (megablast vs blastn vs blastn-short), -taxids/-negative_taxids taxonomic scoping, and -mt_mode multithreading; plus a DIAMOND blastp --ultra-sensitive path for large protein searches. Warns that -max_target_seqs is a heuristic keep-count, not a top-N best-hits filter. Use when the user wants command-line BLAST, makeblastdb, a local BLAST database, blastn/blastp/blastx/tblastn searches, or DIAMOND protein search. For the Bio.Blast web NCBIWWW API prefer alterlab-biopython; for quick one-liner database lookups prefer alterlab-gget. Part of the AlterLab Academic Skills suite.4license: MIT5---67# BLAST+ — Command-Line Sequence Search89Run local NCBI **BLAST+ 2.17.0** searches end-to-end: build a database with10`makeblastdb`, search it with `blastn` / `blastp` / `blastx` / `tblastn`, emit11machine-parseable tabular output, and scope by taxonomy. For very large protein12searches, hand off to **DIAMOND** `blastp --ultra-sensitive` (100x–10,000x the13speed of BLAST, per the DIAMOND project). This is the **CLI / local-database**14skill; it is deliberately distinct from the Biopython web API and the gget15one-liner (see routing table below).1617> Bulk DB builds and large searches are CPU/IO-heavy and fully offline — good18> candidates to run on local compute rather than burning API calls.1920## When to Use This Skill2122Use this skill when the request involves any of:2324- "BLAST these sequences", "run blastn/blastp/blastx/tblastn", "command-line BLAST"25- "build a local BLAST database", "makeblastdb", "index this FASTA for BLAST"26- "search my reads against a local nt/nr database", "get tabular BLAST hits I can parse"27- "scope the BLAST search to a taxon" (`-taxids` / `-negative_taxids`)28- "BLAST is too slow on millions of proteins" → DIAMOND `blastp`29- retrieving sequences out of a BLAST DB (`blastdbcmd`, requires `-parse_seqids`)3031### Does NOT Trigger3233Route adjacent requests to the right sibling skill instead of forcing BLAST+:3435| The request is really about… | Route to |36|------------------------------|----------|37| The **web** BLAST API (`Bio.Blast.NCBIWWW.qblast`), or scripting BLAST inside a Python pipeline with `Bio.Blast` parsing | `alterlab-biopython` |38| A **quick one-liner** BLAST/database lookup (`gget blast`, gene/structure/enrichment lookups) | `alterlab-gget` |39| Unified programmatic access to many bio web services (UniProt, KEGG, Ensembl REST, NCBI eUtils) | `alterlab-bioservices` |40| Building/searching a **phylogenetic tree** from sequences, not a similarity search | `alterlab-phylogenetics` |41| Read alignment to a reference genome (BWA/minimap2 → BAM) and SAM/BAM handling | `alterlab-pysam` |42| FASTQ→VCF variant calling pipeline | `alterlab-nf-core-sarek` |43| Transcript-level RNA-seq quantification (salmon/kallisto) | `alterlab-rnaseq-quant` |44| 16S/ITS amplicon classification (QIIME 2) | `alterlab-qiime2-amplicon` |45| Protein **structure** prediction / embeddings (ESM, AlphaFold) | `alterlab-esm` |4647If the user explicitly says "web BLAST", "NCBIWWW", or "without installing48anything", they want `alterlab-biopython`, not this skill.4950## Quick Start5152```bash53# 1. Build a protein DB (‑parse_seqids enables blastdbcmd retrieval + DIAMOND reuse)54makeblastdb -in proteins.fasta -dbtype prot -parse_seqids -out mydb -title "my proteins"5556# 2. Search, tabular output you can parse, std 12 columns57blastp -query query.faa -db mydb -outfmt 6 -evalue 1e-5 -out hits.tsv5859# 3. QC / summarize the tabular output (stdlib only)60uv run python scripts/parse_blast_tab.py hits.tsv --best-hit61```6263`-outfmt 6` is the canonical machine-readable format; its default columns are64the `std` set: `qseqid sseqid pident length mismatch gapopen qstart qend sstart65send evalue bitscore`. Use `-outfmt 7` for the same columns plus comment lines.6667## Choosing the Right Program6869| Query | Subject DB | Program |70|-------|-----------|---------|71| nucleotide | nucleotide | `blastn` |72| protein | protein | `blastp` |73| nucleotide (translated) | protein | `blastx` |74| protein | nucleotide (translated) | `tblastn` |7576`-dbtype` for `makeblastdb` is `nucl` for nucleotide subjects, `prot` for protein.7778## The Five Things People Get Wrong79801. **`-max_target_seqs` is NOT a "top N best hits" filter.** It is the number of81 aligned sequences to *keep*, applied during the search as a heuristic cutoff;82 ties are broken "by order of sequences in the database", not by score. Setting83 `-max_target_seqs 1` does **not** reliably return the single best hit. To get84 the best hit, keep a generous value and pick the top row *after* sorting by85 bitscore (see `scripts/parse_blast_tab.py --best-hit`). Default is 500.862. **Wrong `-task` for `blastn`.** `megablast` (default) is for highly similar87 sequences; use `blastn` for cross-species / more divergent hits and88 `blastn-short` for queries < ~30 nt (primers, sgRNAs). `dc-megablast` is the89 discontiguous option for inter-species comparison.903. **Forgetting `-parse_seqids` at DB-build time.** Without it you cannot pull91 sequences back out with `blastdbcmd -entry`, and DIAMOND cannot reuse the92 sequence IDs cleanly. You cannot add it later without rebuilding.934. **Quoting the `-outfmt` custom column list for DIAMOND.** BLAST+ wants the94 spec quoted (`-outfmt '6 qseqid sseqid pident evalue'`); **DIAMOND wants it95 unquoted** (`--outfmt 6 qseqid sseqid pident evalue`). Mixing these up is a96 common silent error.975. **Multithreading.** Use `-num_threads N`. For *many small queries*, set98 `-mt_mode 1` (split by query) so all threads stay busy; `-mt_mode 0` (default,99 split by database volume) suits few large queries. BLAST+ 2.15+ can choose100 automatically, but set it explicitly when in doubt.101102Full option reference, taxonomy scoping, and DB-prep details:103[`references/blast_cli.md`](references/blast_cli.md).104105## Taxonomic Scoping106107Restrict a search to (or away from) clades by NCBI taxid:108109```bash110blastn -query q.fna -db nt -taxids 9606 -outfmt 6 -out human_only.tsv111blastp -query q.faa -db nr -negative_taxids 2 -outfmt 6 -out no_bacteria.tsv112```113114Scoping by taxid requires a taxonomy-aware database (one built/downloaded with115its `*.taxid` mapping, e.g. NCBI's pre-formatted `nt` / `nr`). See116[`references/blast_cli.md`](references/blast_cli.md#taxonomy).117118## DIAMOND — Fast Path for Large Protein Searches119120When `blastp` / `blastx` against millions of proteins is too slow, DIAMOND is a121drop-in for protein-space search:122123```bash124diamond makedb --in nr.faa -d nr_diamond125diamond blastp -d nr_diamond -q query.faa -o hits.tsv \126 --ultra-sensitive --outfmt 6 qseqid sseqid pident length evalue bitscore127```128129Sensitivity ladder (fast → most sensitive): `--fast`, `--mid-sensitive`,130`--sensitive`, `--more-sensitive`, `--very-sensitive`, `--ultra-sensitive`.131Use `--ultra-sensitive` when you need BLAST-comparable recall; default fast mode132trades sensitivity for speed. DIAMOND's `--outfmt 6` is compatible with the133BLAST+ tabular parser below. Details and tradeoffs:134[`references/diamond.md`](references/diamond.md).135136## Recommended Workflow1371381. **Pick the program** from the query/subject table above.1392. **Build the DB** with `makeblastdb -parse_seqids` (or download a pre-formatted140 NCBI DB). For >~1M proteins, build a DIAMOND DB instead.1413. **Search** with `-outfmt 6`, an explicit `-evalue` threshold, the right142 `-task` (blastn), and `-num_threads`. Add `-taxids` if scoping.1434. **Parse & QC** with `scripts/parse_blast_tab.py` — it sorts by bitscore,144 extracts best-hit-per-query, applies identity/coverage/e-value filters, and145 flags the `-max_target_seqs` pitfall if the column count looks truncated.1465. **Retrieve** any hit sequence with147 `blastdbcmd -db mydb -entry <id>` (needs `-parse_seqids`).148149## Verify Before Reporting150151- Confirm `blastn -version` / `diamond version` actually ran — never report hits152 you did not produce.153- State the program, `-task`, `-evalue`, and DB used; results are meaningless154 without them.155- If you used `-max_target_seqs`, confirm best-hit selection was done by156 *post-hoc bitscore sort*, not by trusting the keep-count as a top-N.157- For DIAMOND results, note the sensitivity level used.158159## References160161- [`references/blast_cli.md`](references/blast_cli.md) — full BLAST+ 2.17.0 option162 reference: programs, `makeblastdb`, `-outfmt` columns, `-task`, taxonomy163 scoping, `-mt_mode`, `blastdbcmd` retrieval, and the `-max_target_seqs` caveat.164- [`references/diamond.md`](references/diamond.md) — DIAMOND DB build, sensitivity165 modes, output formats, and when to choose it over BLAST+.166- NCBI BLAST+ manual: https://www.ncbi.nlm.nih.gov/books/NBK569856/167- DIAMOND: https://github.com/bbuchfink/diamond168169Part of the AlterLab Academic Skills suite.