efetch
Quick Start
- Command:
efetch - Local executable:
/home/vimalinx/miniforge3/envs/bio/bin/efetch - Full reference: references/help.md
When To Use This Tool
- Retrieve full records or sequence data after
esearchorepost. - Fetch records by ID or accession from PubMed, Nucleotide, Protein, Gene, SRA, Assembly, and related databases.
- Export in FASTA, abstract, docsum, XML, JSON, or native record formats.
- Pull sequence subranges when you only need a locus, not the entire accession.
Common Patterns
# 1) Fetch PubMed abstracts from a search result
esearch -db pubmed -query 'ebola virus[Title/Abstract]' | efetch -format abstract
# 2) Fetch a nucleotide record in FASTA
efetch -db nucleotide -id NM_000546.6 -format fasta
# 3) Fetch only a sequence subrange
efetch \
-db nucleotide \
-id NC_000001.11 \
-format fasta \
-seq_start 100000 \
-seq_stop 101000
# 4) Fetch summaries in JSON-compatible mode
efetch -db assembly -id GCF_000001405.40 -format docsum -mode json
Recommended Workflow
- Get stable IDs first with
esearchorepost. - Choose
-formatdeliberately, because downstream parsing depends on it. - Save raw outputs for reproducibility when the fetched record is part of an analysis result.
- Pipe XML-style outputs into
xtractonly after verifying the record shape.
Guardrails
- Always specify
-dband usually specify-format; implicit defaults are easy to misread. - Large fetches should be chunked with
-startand-stop, or handled through batched EDirect pipelines. -seq_start,-seq_stop, and strand options only apply to sequence-style retrieval.- Invalid or mismatched IDs often yield empty output, so validate the upstream search first.