GEO Database
Overview
The Gene Expression Omnibus (GEO) is NCBI's public repository for high-throughput gene expression and functional genomics data. It holds hundreds of thousands of studies (Series) and millions of samples from both array-based and sequence-based experiments. For current totals see the GEO browser at https://www.ncbi.nlm.nih.gov/geo/.
When to Use This Skill
Use this skill when searching for gene expression datasets, retrieving experimental data, downloading raw and processed files, querying expression profiles, or integrating GEO data into computational analysis workflows.
Core Workflow
- Search for relevant studies (E-utilities /
Bio.Entrez) → see references/searching.md.
- Retrieve a series with GEOparse:
GEOparse.get_GEO(geo="GSE...", destdir="./data") → see references/geoparse_usage.md.
- Extract the expression matrix via
gse.pivot_samples('VALUE') (genes × samples).
- Analyze — QC, differential expression, clustering, meta-analysis → see
references/expression_analysis.md.
For bulk downloads, skip GEOparse and pull files directly over FTP → see references/data_access.md.
GEO Data Organization
GEO organizes data hierarchically using different accession types:
- Series (GSE) — a complete experiment with related samples (e.g.
GSE123456).
Largest organizational unit. Holds experimental design, samples, study info.
- Sample (GSM) — a single experimental sample / biological replicate
(e.g.
GSM987654). Linked to platforms and series.
- Platform (GPL) — the microarray or sequencing platform (e.g.
GPL570,
Affymetrix Human Genome U133 Plus 2.0 Array). Shared across experiments.
- DataSet (GDS) — curated, consistently-formatted collections (e.g.
GDS5678)
processed for differential analysis. Ideal for quick comparative analyses, but
the curated GDS set is legacy and no longer growing — most recent studies exist
only as GSE Series, so prefer GSE for new work.
- Profiles — gene-specific expression data linked to sequence features,
queryable by gene name, cross-referenced to Entrez Gene.
Access Methods (routing)
- GEOparse (recommended) — easiest series-level access; download, parse
metadata, extract expression matrices, supplementary files, filter samples.
→
references/geoparse_usage.md
- E-utilities (
Bio.Entrez) — metadata searching and batch queries across
gds, geoprofiles. Always set Entrez.email. → references/searching.md
- Direct FTP / wget / curl — bulk downloads of series matrix, SOFT, MINiML,
and supplementary files; no rate limits. →
references/data_access.md
- GEO2R web tool — no-code differential expression analysis in the browser at
https://www.ncbi.nlm.nih.gov/geo/geo2r/?acc=GSExxxxx; generates reproducible
R scripts. Useful for exploratory analysis before downloading.
Installation and Setup
uv pip install GEOparse # primary GEO access (recommended)
uv pip install biopython # E-utilities / programmatic NCBI access
uv pip install pandas numpy scipy statsmodels scikit-learn # analysis
uv pip install matplotlib seaborn # visualization
Configure NCBI E-utilities access (email is required by NCBI; an API key raises
the rate limit from 3 to 10 requests/second):
from Bio import Entrez
Entrez.email = "your.email@example.com" # required
# Optional API key — https://www.ncbi.nlm.nih.gov/account/
Entrez.api_key = "your_api_key_here"
Key Concepts
- SOFT (Simple Omnibus Format in Text): GEO's primary text format with
metadata and data tables. Easily parsed by GEOparse.
- MINiML (MIAME Notation in Markup Language): XML format for programmatic
access and data exchange.
- Series Matrix: tab-delimited expression matrix (samples as columns,
genes/probes as rows). Fastest format for getting expression data.
- MIAME Compliance: minimum standardized annotation GEO enforces on
submissions.
- Expression Value Types: raw signal, normalized, or log-transformed — always
check platform and processing methods.
- Platform Annotation: maps probe/feature IDs to genes; essential for
biological interpretation.
Common Use Cases
- Transcriptomics research — download expression data, compare profiles
across studies, identify DEGs, run meta-analyses.
- Drug response studies — analyze post-treatment expression changes, identify
response biomarkers, build sensitivity models.
- Disease biology — disease vs. normal expression, disease signatures,
patient subgroup/stage comparisons, expression–outcome correlation.
- Biomarker discovery — screen diagnostic/prognostic markers, validate across
cohorts, integrate with clinical data.
Rate Limiting and Best Practices
- E-utilities limits: 3 req/s without an API key, 10 req/s with one. Insert
delays:
time.sleep(0.34) (no key) or time.sleep(0.1) (with key).
- FTP: no rate limits — preferred for bulk downloads (
wget -r for whole
directories).
- GEOparse caching: files cached in
destdir; subsequent calls reuse them.
- Method selection: GEOparse for series-level access, E-utilities for search /
batch metadata, FTP for direct/bulk file pulls. Cache locally; always set
Entrez.email with Biopython.
Important Notes
- Data quality: GEO accepts user-submitted data of varying quality. Check
platform annotation and processing methods, verify metadata and design, watch
for batch effects, and consider reprocessing raw data.
- File sizes: series matrix files can exceed 1 GB; supplementary files (e.g.
CEL) can be very large. Plan disk space; download incrementally.
- Citation: GEO data is free for research. Cite original studies plus the GEO
database (Barrett et al. 2013, Nucleic Acids Research). Check per-dataset
usage restrictions and follow NCBI guidelines.
- Common pitfalls: platforms use different probe IDs (need annotation
mapping); values may be raw/normalized/log-transformed (check metadata); sample
metadata is inconsistently formatted; older submissions may lack series matrix
files; platform annotations may be outdated.
Reference Index
references/searching.md — E-utilities search (datasets, profiles, advanced
queries, search→summary→fetch workflow, batch metadata).
references/geoparse_usage.md — GEOparse install/usage, expression matrices,
supplementary files, sample filtering, caching.
references/data_access.md — direct FTP access (ftplib, wget, curl) and the
accession-to-path rule.
references/expression_analysis.md — QC/preprocessing, differential expression,
correlation/clustering, batch processing, cross-study meta-analysis.
references/geo_reference.md — in-depth technical reference: E-utilities API
endpoints, SOFT/MINiML format docs, FTP directory structure, normalization
pipelines, platform quirks, and error-handling troubleshooting.
Additional Resources
Scripts
scripts/query_geo.py — runnable helper for GEO DataSets via NCBI E-utilities (no key; NCBI_API_KEY lifts the rate limit). Has a PEP 723 inline-dependency header, so uv run installs requests automatically:
uv run scripts/query_geo.py search "breast cancer AND Homo sapiens[ORGN]" --retmax 5
uv run scripts/query_geo.py summary 200000001,200000002
1---2name: alterlab-geo3description: Access NCBI GEO (Gene Expression Omnibus) for gene expression and functional genomics data — search and download microarray and RNA-seq datasets by GSE, GSM, GPL, or GDS accession and retrieve SOFT, MINiML, and series matrix files. Use when locating public expression datasets, fetching processed expression matrices, downloading a study's supplementary files, or sourcing per-study transcriptomics data for differential-expression analysis. For raw FASTQ sequencing reads by SRA/ENA run accession use alterlab-ena; for reference tissue-expression baselines (median TPM across human tissues) use alterlab-gtex; for cancer cohort somatic mutations and copy-number use alterlab-cbioportal. Part of the AlterLab Academic Skills suite.4license: MIT5---67# GEO Database89## Overview1011The Gene Expression Omnibus (GEO) is NCBI's public repository for high-throughput gene expression and functional genomics data. It holds hundreds of thousands of studies (Series) and millions of samples from both array-based and sequence-based experiments. For current totals see the GEO browser at https://www.ncbi.nlm.nih.gov/geo/.1213## When to Use This Skill1415Use this skill when searching for gene expression datasets, retrieving experimental data, downloading raw and processed files, querying expression profiles, or integrating GEO data into computational analysis workflows.1617## Core Workflow18191. **Search** for relevant studies (E-utilities / `Bio.Entrez`) → see `references/searching.md`.202. **Retrieve** a series with GEOparse: `GEOparse.get_GEO(geo="GSE...", destdir="./data")` → see `references/geoparse_usage.md`.213. **Extract** the expression matrix via `gse.pivot_samples('VALUE')` (genes × samples).224. **Analyze** — QC, differential expression, clustering, meta-analysis → see `references/expression_analysis.md`.2324For bulk downloads, skip GEOparse and pull files directly over FTP → see `references/data_access.md`.2526## GEO Data Organization2728GEO organizes data hierarchically using different accession types:2930- **Series (GSE)** — a complete experiment with related samples (e.g. `GSE123456`).31 Largest organizational unit. Holds experimental design, samples, study info.32- **Sample (GSM)** — a single experimental sample / biological replicate33 (e.g. `GSM987654`). Linked to platforms and series.34- **Platform (GPL)** — the microarray or sequencing platform (e.g. `GPL570`,35 Affymetrix Human Genome U133 Plus 2.0 Array). Shared across experiments.36- **DataSet (GDS)** — curated, consistently-formatted collections (e.g. `GDS5678`)37 processed for differential analysis. Ideal for quick comparative analyses, but38 the curated GDS set is legacy and no longer growing — most recent studies exist39 only as GSE Series, so prefer GSE for new work.40- **Profiles** — gene-specific expression data linked to sequence features,41 queryable by gene name, cross-referenced to Entrez Gene.4243## Access Methods (routing)4445- **GEOparse (recommended)** — easiest series-level access; download, parse46 metadata, extract expression matrices, supplementary files, filter samples.47 → `references/geoparse_usage.md`48- **E-utilities (`Bio.Entrez`)** — metadata searching and batch queries across49 `gds`, `geoprofiles`. Always set `Entrez.email`. → `references/searching.md`50- **Direct FTP / wget / curl** — bulk downloads of series matrix, SOFT, MINiML,51 and supplementary files; no rate limits. → `references/data_access.md`52- **GEO2R web tool** — no-code differential expression analysis in the browser at53 `https://www.ncbi.nlm.nih.gov/geo/geo2r/?acc=GSExxxxx`; generates reproducible54 R scripts. Useful for exploratory analysis before downloading.5556## Installation and Setup5758```bash59uv pip install GEOparse # primary GEO access (recommended)60uv pip install biopython # E-utilities / programmatic NCBI access61uv pip install pandas numpy scipy statsmodels scikit-learn # analysis62uv pip install matplotlib seaborn # visualization63```6465Configure NCBI E-utilities access (email is required by NCBI; an API key raises66the rate limit from 3 to 10 requests/second):6768```python69from Bio import Entrez7071Entrez.email = "your.email@example.com" # required72# Optional API key — https://www.ncbi.nlm.nih.gov/account/73Entrez.api_key = "your_api_key_here"74```7576## Key Concepts7778- **SOFT (Simple Omnibus Format in Text):** GEO's primary text format with79 metadata and data tables. Easily parsed by GEOparse.80- **MINiML (MIAME Notation in Markup Language):** XML format for programmatic81 access and data exchange.82- **Series Matrix:** tab-delimited expression matrix (samples as columns,83 genes/probes as rows). Fastest format for getting expression data.84- **MIAME Compliance:** minimum standardized annotation GEO enforces on85 submissions.86- **Expression Value Types:** raw signal, normalized, or log-transformed — always87 check platform and processing methods.88- **Platform Annotation:** maps probe/feature IDs to genes; essential for89 biological interpretation.9091## Common Use Cases9293- **Transcriptomics research** — download expression data, compare profiles94 across studies, identify DEGs, run meta-analyses.95- **Drug response studies** — analyze post-treatment expression changes, identify96 response biomarkers, build sensitivity models.97- **Disease biology** — disease vs. normal expression, disease signatures,98 patient subgroup/stage comparisons, expression–outcome correlation.99- **Biomarker discovery** — screen diagnostic/prognostic markers, validate across100 cohorts, integrate with clinical data.101102## Rate Limiting and Best Practices103104- **E-utilities limits:** 3 req/s without an API key, 10 req/s with one. Insert105 delays: `time.sleep(0.34)` (no key) or `time.sleep(0.1)` (with key).106- **FTP:** no rate limits — preferred for bulk downloads (`wget -r` for whole107 directories).108- **GEOparse caching:** files cached in `destdir`; subsequent calls reuse them.109- **Method selection:** GEOparse for series-level access, E-utilities for search /110 batch metadata, FTP for direct/bulk file pulls. Cache locally; always set111 `Entrez.email` with Biopython.112113## Important Notes114115- **Data quality:** GEO accepts user-submitted data of varying quality. Check116 platform annotation and processing methods, verify metadata and design, watch117 for batch effects, and consider reprocessing raw data.118- **File sizes:** series matrix files can exceed 1 GB; supplementary files (e.g.119 CEL) can be very large. Plan disk space; download incrementally.120- **Citation:** GEO data is free for research. Cite original studies plus the GEO121 database (Barrett et al. 2013, *Nucleic Acids Research*). Check per-dataset122 usage restrictions and follow NCBI guidelines.123- **Common pitfalls:** platforms use different probe IDs (need annotation124 mapping); values may be raw/normalized/log-transformed (check metadata); sample125 metadata is inconsistently formatted; older submissions may lack series matrix126 files; platform annotations may be outdated.127128## Reference Index129130- `references/searching.md` — E-utilities search (datasets, profiles, advanced131 queries, search→summary→fetch workflow, batch metadata).132- `references/geoparse_usage.md` — GEOparse install/usage, expression matrices,133 supplementary files, sample filtering, caching.134- `references/data_access.md` — direct FTP access (ftplib, wget, curl) and the135 accession-to-path rule.136- `references/expression_analysis.md` — QC/preprocessing, differential expression,137 correlation/clustering, batch processing, cross-study meta-analysis.138- `references/geo_reference.md` — in-depth technical reference: E-utilities API139 endpoints, SOFT/MINiML format docs, FTP directory structure, normalization140 pipelines, platform quirks, and error-handling troubleshooting.141142## Additional Resources143144- **GEO Website:** https://www.ncbi.nlm.nih.gov/geo/145- **GEO Submission Guidelines:** https://www.ncbi.nlm.nih.gov/geo/info/submission.html146- **GEOparse Documentation:** https://geoparse.readthedocs.io/147- **E-utilities Documentation:** https://www.ncbi.nlm.nih.gov/books/NBK25501/148- **GEO FTP Site:** ftp://ftp.ncbi.nlm.nih.gov/geo/149- **GEO2R Tool:** https://www.ncbi.nlm.nih.gov/geo/geo2r/150- **NCBI API Keys:** https://ncbiinsights.ncbi.nlm.nih.gov/2017/11/02/new-api-keys-for-the-e-utilities/151- **Biopython Tutorial:** https://biopython.org/DIST/docs/tutorial/Tutorial.html152153## Scripts154155`scripts/query_geo.py` — runnable helper for GEO DataSets via NCBI E-utilities (no key; `NCBI_API_KEY` lifts the rate limit). Has a PEP 723 inline-dependency header, so `uv run` installs `requests` automatically:156157```bash158uv run scripts/query_geo.py search "breast cancer AND Homo sapiens[ORGN]" --retmax 5159uv run scripts/query_geo.py summary 200000001,200000002160```