HMDB Database
Overview
The Human Metabolome Database (HMDB) is a comprehensive, freely available resource containing detailed information about small molecule metabolites found in the human body.
Scripts
scripts/query_hmdb.py — fetch and parse an HMDB metabolite XML record by accession (stdlib only, JSON to stdout):
python scripts/query_hmdb.py HMDB0000001 # full accession
python scripts/query_hmdb.py 1 # bare number (zero-padded automatically)
Note: HMDB serves no public REST API and may rate-limit/block automated fetches; for bulk work download the XML/SDF dumps from https://www.hmdb.ca/downloads.
When to Use This Skill
This skill should be used when performing metabolomics research, clinical chemistry, biomarker discovery, or metabolite identification tasks.
Database Contents
HMDB version 5.0 (released 2022; latest major version as of mid-2026) contains:
- 220,945 metabolite entries covering both water-soluble and lipid-soluble compounds (the v5.0 paper reported 217,920; the live site count grows with curation)
- ~8,600 protein sequences for enzymes and transporters involved in metabolism
- 130+ data fields per metabolite including:
- Chemical properties (structure, formula, molecular weight, InChI, SMILES)
- Clinical data (biomarker associations, diseases, normal/abnormal concentrations)
- Biological information (pathways, reactions, locations)
- Spectroscopic data (NMR, MS, MS-MS spectra)
- External database links (KEGG, PubChem, MetaCyc, ChEBI, PDB, UniProt, GenBank)
Core Capabilities
1. Web-Based Metabolite Searches
Access HMDB through the web interface at https://www.hmdb.ca/ for:
Text Searches:
- Search by metabolite name, synonym, or identifier (HMDB ID)
- Example HMDB IDs: HMDB0000001, HMDB0001234
- Search by disease associations or pathway involvement
- Query by biological specimen type (urine, serum, CSF, saliva, feces, sweat)
Structure-Based Searches:
- Use ChemQuery for structure and substructure searches
- Search by molecular weight or molecular weight range
- Use SMILES or InChI strings to find compounds
Spectral Searches:
- LC-MS spectral matching
- GC-MS spectral matching
- NMR spectral searches for metabolite identification
Advanced Searches:
- Combine multiple criteria (name, properties, concentration ranges)
- Filter by biological locations or specimen types
- Search by protein/enzyme associations
2. Accessing Metabolite Information
When retrieving metabolite data, HMDB provides:
Chemical Information:
- Systematic name, traditional names, and synonyms
- Chemical formula and molecular weight
- Structure representations (2D/3D, SMILES, InChI, MOL file)
- Chemical taxonomy and classification
Biological Context:
- Metabolic pathways and reactions
- Associated enzymes and transporters
- Subcellular locations
- Biological roles and functions
Clinical Relevance:
- Normal concentration ranges in biological fluids
- Biomarker associations with diseases
- Clinical significance
- Toxicity information when applicable
Analytical Data:
- Experimental and predicted NMR spectra
- MS and MS-MS spectra
- Retention times and chromatographic data
- Reference peaks for identification
3. Downloadable Datasets
HMDB offers bulk data downloads at https://www.hmdb.ca/downloads in multiple formats:
Available Formats:
- XML: Complete metabolite, protein, and spectra data
- SDF: Metabolite structure files for cheminformatics
- FASTA: Protein and gene sequences
- TXT: Raw spectra peak lists
- CSV/TSV: Tabular data exports
Dataset Categories:
- All metabolites or filtered by specimen type
- Protein/enzyme sequences
- Experimental and predicted spectra (NMR, GC-MS, MS-MS)
- Pathway information
Best Practices:
- Download XML format for comprehensive data including all fields
- Use SDF format for structure-based analysis and cheminformatics workflows
- Parse CSV/TSV formats for integration with data analysis pipelines
- Check version dates to ensure up-to-date data (current major version: v5.0)
Usage Requirements:
- Free for academic and non-commercial research
- Commercial use requires explicit permission (contact samackay@ualberta.ca)
- Cite HMDB publication when using data
4. Programmatic Access
HMDB publishes no documented public REST API. Practical programmatic routes, in order of preference:
- Bulk downloads (preferred for any volume): Parse the XML/SDF/CSV dumps from https://www.hmdb.ca/downloads locally. This is the only route that scales and won't get rate-limited.
- Per-record XML endpoint: Each entry is served as XML at
https://www.hmdb.ca/metabolites/<ID>.xml (used by scripts/query_hmdb.py). Undocumented and aggressively rate-limited/blocked (HTTP 403/429) for automated clients — fine for a handful of ad-hoc lookups, not for batch jobs. Send a descriptive User-Agent and back off on failures.
- R/Bioconductor
hmdbQuery: BiocManager::install("hmdbQuery") wraps the same web endpoints for R workflows.
- Custom API: For sanctioned bulk/commercial API access, contact the HMDB team (see Usage Requirements above for the listed address).
5. Common Research Workflows
Metabolite Identification in Untargeted Metabolomics:
- Obtain experimental MS or NMR spectra from samples
- Use HMDB spectral search tools to match against reference spectra
- Verify candidates by checking molecular weight, retention time, and MS-MS fragmentation
- Review biological plausibility (expected in specimen type, known pathways)
Biomarker Discovery:
- Search HMDB for metabolites associated with disease of interest
- Review concentration ranges in normal vs. disease states
- Identify metabolites with strong differential abundance
- Examine pathway context and biological mechanisms
- Cross-reference with literature via PubMed links
Pathway Analysis:
- Identify metabolites of interest from experimental data
- Look up HMDB entries for each metabolite
- Extract pathway associations and enzymatic reactions
- Use linked SMPDB (Small Molecule Pathway Database) for pathway diagrams
- Identify pathway enrichment for biological interpretation
Database Integration:
- Download HMDB data in XML or CSV format
- Parse and extract relevant fields for local database
- Link with external IDs (KEGG, PubChem, ChEBI) for cross-database queries
- Build local tools or pipelines incorporating HMDB reference data
Related HMDB Resources
The HMDB ecosystem includes related databases:
- DrugBank: ~2,832 drug compounds with pharmaceutical information
- T3DB (Toxin and Toxin Target Database): ~3,670 toxic compounds
- SMPDB (Small Molecule Pathway Database): Pathway diagrams and maps
- FooDB: ~70,000 food component compounds
These databases share similar structure and identifiers, enabling integrated queries across human metabolome, drug, toxin, and food databases.
Best Practices
Data Quality:
- Verify metabolite identifications with multiple evidence types (spectra, structure, properties)
- Check experimental vs. predicted data quality indicators
- Review citations and evidence for biomarker associations
Version Tracking:
- Note HMDB version used in research (current: v5.0)
- Databases are updated periodically with new entries and corrections
- Re-query for updates when publishing to ensure current information
Citation:
- Always cite HMDB in publications using the database
- Reference specific HMDB IDs when discussing metabolites
- Acknowledge data sources for downloaded datasets
Performance:
- For large-scale analysis, download complete datasets rather than repeated web queries
- Use appropriate file formats (XML for comprehensive data, CSV for tabular analysis)
- Consider local caching of frequently accessed metabolite information
Reference Documentation
See references/hmdb_data_fields.md for detailed information about available data fields and their meanings.
1---2name: alterlab-hmdb3description: Access the Human Metabolome Database (HMDB, 220K+ metabolites), searching by name, HMDB ID, or structure to retrieve chemical properties, biomarker data, NMR/MS reference spectra, and associated pathways. Use when identifying a human metabolite, looking up its biomarker or disease associations, matching NMR/MS spectra, or running metabolomics annotation. Part of the AlterLab Academic Skills suite.4license: MIT5---67# HMDB Database89## Overview1011The Human Metabolome Database (HMDB) is a comprehensive, freely available resource containing detailed information about small molecule metabolites found in the human body.1213## Scripts1415`scripts/query_hmdb.py` — fetch and parse an HMDB metabolite XML record by accession (stdlib only, JSON to stdout):1617```bash18python scripts/query_hmdb.py HMDB0000001 # full accession19python scripts/query_hmdb.py 1 # bare number (zero-padded automatically)20```2122Note: HMDB serves no public REST API and may rate-limit/block automated fetches; for bulk work download the XML/SDF dumps from https://www.hmdb.ca/downloads.2324## When to Use This Skill2526This skill should be used when performing metabolomics research, clinical chemistry, biomarker discovery, or metabolite identification tasks.2728## Database Contents2930HMDB version 5.0 (released 2022; latest major version as of mid-2026) contains:3132- **220,945 metabolite entries** covering both water-soluble and lipid-soluble compounds (the v5.0 paper reported 217,920; the live site count grows with curation)33- **~8,600 protein sequences** for enzymes and transporters involved in metabolism34- **130+ data fields per metabolite** including:35 - Chemical properties (structure, formula, molecular weight, InChI, SMILES)36 - Clinical data (biomarker associations, diseases, normal/abnormal concentrations)37 - Biological information (pathways, reactions, locations)38 - Spectroscopic data (NMR, MS, MS-MS spectra)39 - External database links (KEGG, PubChem, MetaCyc, ChEBI, PDB, UniProt, GenBank)4041## Core Capabilities4243### 1. Web-Based Metabolite Searches4445Access HMDB through the web interface at https://www.hmdb.ca/ for:4647**Text Searches:**48- Search by metabolite name, synonym, or identifier (HMDB ID)49- Example HMDB IDs: HMDB0000001, HMDB000123450- Search by disease associations or pathway involvement51- Query by biological specimen type (urine, serum, CSF, saliva, feces, sweat)5253**Structure-Based Searches:**54- Use ChemQuery for structure and substructure searches55- Search by molecular weight or molecular weight range56- Use SMILES or InChI strings to find compounds5758**Spectral Searches:**59- LC-MS spectral matching60- GC-MS spectral matching61- NMR spectral searches for metabolite identification6263**Advanced Searches:**64- Combine multiple criteria (name, properties, concentration ranges)65- Filter by biological locations or specimen types66- Search by protein/enzyme associations6768### 2. Accessing Metabolite Information6970When retrieving metabolite data, HMDB provides:7172**Chemical Information:**73- Systematic name, traditional names, and synonyms74- Chemical formula and molecular weight75- Structure representations (2D/3D, SMILES, InChI, MOL file)76- Chemical taxonomy and classification7778**Biological Context:**79- Metabolic pathways and reactions80- Associated enzymes and transporters81- Subcellular locations82- Biological roles and functions8384**Clinical Relevance:**85- Normal concentration ranges in biological fluids86- Biomarker associations with diseases87- Clinical significance88- Toxicity information when applicable8990**Analytical Data:**91- Experimental and predicted NMR spectra92- MS and MS-MS spectra93- Retention times and chromatographic data94- Reference peaks for identification9596### 3. Downloadable Datasets9798HMDB offers bulk data downloads at https://www.hmdb.ca/downloads in multiple formats:99100**Available Formats:**101- **XML**: Complete metabolite, protein, and spectra data102- **SDF**: Metabolite structure files for cheminformatics103- **FASTA**: Protein and gene sequences104- **TXT**: Raw spectra peak lists105- **CSV/TSV**: Tabular data exports106107**Dataset Categories:**108- All metabolites or filtered by specimen type109- Protein/enzyme sequences110- Experimental and predicted spectra (NMR, GC-MS, MS-MS)111- Pathway information112113**Best Practices:**114- Download XML format for comprehensive data including all fields115- Use SDF format for structure-based analysis and cheminformatics workflows116- Parse CSV/TSV formats for integration with data analysis pipelines117- Check version dates to ensure up-to-date data (current major version: v5.0)118119**Usage Requirements:**120- Free for academic and non-commercial research121- Commercial use requires explicit permission (contact samackay@ualberta.ca)122- Cite HMDB publication when using data123124### 4. Programmatic Access125126HMDB publishes **no documented public REST API**. Practical programmatic routes, in order of preference:127128- **Bulk downloads (preferred for any volume):** Parse the XML/SDF/CSV dumps from https://www.hmdb.ca/downloads locally. This is the only route that scales and won't get rate-limited.129- **Per-record XML endpoint:** Each entry is served as XML at `https://www.hmdb.ca/metabolites/<ID>.xml` (used by `scripts/query_hmdb.py`). Undocumented and aggressively rate-limited/blocked (HTTP 403/429) for automated clients — fine for a handful of ad-hoc lookups, not for batch jobs. Send a descriptive User-Agent and back off on failures.130- **R/Bioconductor `hmdbQuery`:** `BiocManager::install("hmdbQuery")` wraps the same web endpoints for R workflows.131- **Custom API:** For sanctioned bulk/commercial API access, contact the HMDB team (see Usage Requirements above for the listed address).132133### 5. Common Research Workflows134135**Metabolite Identification in Untargeted Metabolomics:**1361. Obtain experimental MS or NMR spectra from samples1372. Use HMDB spectral search tools to match against reference spectra1383. Verify candidates by checking molecular weight, retention time, and MS-MS fragmentation1394. Review biological plausibility (expected in specimen type, known pathways)140141**Biomarker Discovery:**1421. Search HMDB for metabolites associated with disease of interest1432. Review concentration ranges in normal vs. disease states1443. Identify metabolites with strong differential abundance1454. Examine pathway context and biological mechanisms1465. Cross-reference with literature via PubMed links147148**Pathway Analysis:**1491. Identify metabolites of interest from experimental data1502. Look up HMDB entries for each metabolite1513. Extract pathway associations and enzymatic reactions1524. Use linked SMPDB (Small Molecule Pathway Database) for pathway diagrams1535. Identify pathway enrichment for biological interpretation154155**Database Integration:**1561. Download HMDB data in XML or CSV format1572. Parse and extract relevant fields for local database1583. Link with external IDs (KEGG, PubChem, ChEBI) for cross-database queries1594. Build local tools or pipelines incorporating HMDB reference data160161## Related HMDB Resources162163The HMDB ecosystem includes related databases:164165- **DrugBank**: ~2,832 drug compounds with pharmaceutical information166- **T3DB (Toxin and Toxin Target Database)**: ~3,670 toxic compounds167- **SMPDB (Small Molecule Pathway Database)**: Pathway diagrams and maps168- **FooDB**: ~70,000 food component compounds169170These databases share similar structure and identifiers, enabling integrated queries across human metabolome, drug, toxin, and food databases.171172## Best Practices173174**Data Quality:**175- Verify metabolite identifications with multiple evidence types (spectra, structure, properties)176- Check experimental vs. predicted data quality indicators177- Review citations and evidence for biomarker associations178179**Version Tracking:**180- Note HMDB version used in research (current: v5.0)181- Databases are updated periodically with new entries and corrections182- Re-query for updates when publishing to ensure current information183184**Citation:**185- Always cite HMDB in publications using the database186- Reference specific HMDB IDs when discussing metabolites187- Acknowledge data sources for downloaded datasets188189**Performance:**190- For large-scale analysis, download complete datasets rather than repeated web queries191- Use appropriate file formats (XML for comprehensive data, CSV for tabular analysis)192- Consider local caching of frequently accessed metabolite information193194## Reference Documentation195196See `references/hmdb_data_fields.md` for detailed information about available data fields and their meanings.197