xml2fsa
Quick Start
- Command:
efetch ... -format gbc | xml2fsa > records.fasta - Local executable:
/home/vimalinx/miniforge3/envs/bio/bin/xml2fsa - Full reference: See references/help.md for detailed usage
When To Use This Tool
- Convert INSDSeq-style NCBI XML sequence records into FASTA.
- Flatten
efetchXML output into headers and plain sequence strings for downstream sequence tools. - Extract accession/definition-based FASTA entries without writing a custom
xtractcommand.
Common Patterns
# 1) Convert streamed INSDSeq XML into FASTA
cat records.xml | xml2fsa > records.fasta
# 2) Pipe efetch-style XML straight into FASTA output
efetch -db nuccore -id ABC123.1 -format gbc | xml2fsa > ABC123.fasta
Recommended Workflow
- Start from INSDSeq-style XML, usually streamed from
efetchor another NCBI XML source. - Pipe that XML into
xml2fsa; the wrapper itself is stdin-driven. - Inspect the first FASTA header to confirm the accession / definition mapping looks right.
- Send the resulting FASTA into downstream alignment, indexing, or QC tools.
Guardrails
- The wrapper is a fixed
xtractcommand overINSDSeq; it is not a general XML-to-FASTA converter. - It does not pass through positional filenames, so stdin / pipes are the reliable invocation path.
xtractmust be available onPATH, and--help/--versionjust fall through to xtract-style input errors.- FASTA headers are constructed from the first available accession/id/locus field plus the definition line, so spot-check headers before batch use.