Datasets Reference
Overview
TorchDrug provides 40+ curated datasets across multiple domains: molecular property prediction, protein modeling, knowledge graph reasoning, and retrosynthesis. All datasets support lazy loading, automatic downloading, and customizable feature extraction.
Molecular Property Prediction Datasets
Drug Discovery Classification
| Dataset |
Size |
Task |
Classes |
Description |
| BACE |
1,513 |
Binary |
2 |
β-secretase inhibition for Alzheimer's |
| BBBP |
2,039 |
Binary |
2 |
Blood-brain barrier penetration |
| HIV |
41,127 |
Binary |
2 |
Inhibition of HIV replication |
| ClinTox |
1,478 |
Multi-label |
2 |
Clinical trial toxicity |
| SIDER |
1,427 |
Multi-label |
27 |
Side effects by system organ class |
| Tox21 |
7,831 |
Multi-label |
12 |
Toxicity across 12 targets |
| ToxCast |
8,576 |
Multi-label |
617 |
High-throughput toxicology |
| MUV |
93,087 |
Multi-label |
17 |
Unbiased validation for screening |
Key Features:
- All use scaffold splits for realistic evaluation
- Binary classification metrics: AUROC, AUPRC
- Multi-label handles missing values
Use Cases:
- Drug safety prediction
- Virtual screening
- ADMET property prediction
Drug Discovery Regression
| Dataset |
Size |
Property |
Units |
Description |
| ESOL |
1,128 |
Solubility |
log(mol/L) |
Water solubility |
| FreeSolv |
642 |
Hydration |
kcal/mol |
Hydration free energy |
| Lipophilicity |
4,200 |
LogD |
- |
Octanol/water distribution |
| SAMPL |
643 |
Solvation |
kcal/mol |
Solvation free energies |
Metrics: MAE, RMSE, R²
Use Cases: ADME optimization, lead optimization
Quantum Chemistry
| Dataset |
Size |
Properties |
Description |
| QM7 |
7,165 |
1 |
Atomization energy |
| QM8 |
21,786 |
12 |
Electronic spectra, excited states |
| QM9 |
133,885 |
12 |
Geometric, energetic, electronic, thermodynamic |
| PCQM4M |
3.8M |
1 |
Large-scale HOMO-LUMO gap |
Properties (QM9):
- Dipole moment
- Isotropic polarizability
- HOMO/LUMO energies
- Internal energy, enthalpy, free energy
- Heat capacity
- Electronic spatial extent
Use Cases:
- Quantum property prediction
- Method development benchmarking
- Pre-training molecular models
Large Molecule Databases
| Dataset |
Size |
Description |
Use Case |
| ZINC250k |
250,000 |
Drug-like molecules |
Generative model training |
| ZINC2M |
2,000,000 |
Drug-like molecules |
Large-scale pre-training |
| ChEMBL |
Millions |
Bioactive molecules |
Property prediction, generation |
Protein Datasets
Function Prediction
| Dataset |
Size |
Task |
Classes |
Description |
| EnzymeCommission |
17,562 |
Multi-class |
7 levels |
EC number classification |
| GeneOntology |
46,796 |
Multi-label |
489 |
GO term prediction (BP/MF/CC) |
| BetaLactamase |
5,864 |
Regression |
- |
Enzyme activity levels |
| Fluorescence |
54,025 |
Regression |
- |
GFP fluorescence intensity |
| Stability |
53,614 |
Regression |
- |
Thermostability (ΔΔG) |
Features:
- Sequence and/or structure input
- Evolutionary information available
- Multiple train/test splits
Use Cases:
- Protein engineering
- Function annotation
- Enzyme design
Localization and Solubility
| Dataset |
Size |
Task |
Classes |
Description |
| Solubility |
62,478 |
Binary |
2 |
Protein solubility |
| BinaryLocalization |
22,168 |
Binary |
2 |
Membrane vs soluble |
| SubcellularLocalization |
8,943 |
Multi-class |
10 |
Subcellular compartment |
Use Cases:
- Protein expression optimization
- Target identification
- Cell biology
Structure Prediction
| Dataset |
Size |
Task |
Description |
| Fold |
16,712 |
Multi-class (1,195) |
Structural fold recognition |
| SecondaryStructure |
8,678 |
Sequence labeling |
3-state or 8-state prediction |
| ProteinNet |
Varied |
Contact prediction |
Residue-residue contacts |
Use Cases:
- Structure prediction pipelines
- Fold recognition
- Contact map generation
Protein Interactions
| Dataset |
Size |
Positives |
Negatives |
Description |
| HumanPPI |
1,412 proteins |
6,584 |
- |
Human protein interactions |
| YeastPPI |
2,018 proteins |
6,451 |
- |
Yeast protein interactions |
| PPIAffinity |
2,156 pairs |
- |
- |
Binding affinity values |
Use Cases:
- PPI prediction
- Network biology
- Drug target identification
Protein-Ligand Binding
| Dataset |
Size |
Type |
Description |
| BindingDB |
~1.5M |
Affinity |
Comprehensive binding data |
| PDBBind |
20,000+ |
3D complexes |
Structure-based binding |
| - Refined Set |
5,316 |
High quality |
Curated crystal structures |
| - Core Set |
285 |
Benchmark |
Diverse test set |
Use Cases:
- Binding affinity prediction
- Structure-based drug design
- Scoring function development
Large Protein Databases
| Dataset |
Size |
Description |
| AlphaFoldDB |
200M+ |
Predicted structures for most known proteins |
| UniProt |
Integration |
Sequence and annotation data |
Knowledge Graph Datasets
General Knowledge
| Dataset |
Entities |
Relations |
Triples |
Domain |
| FB15k |
14,951 |
1,345 |
592,213 |
Freebase (general knowledge) |
| FB15k-237 |
14,541 |
237 |
310,116 |
Filtered Freebase |
| WN18 |
40,943 |
18 |
151,442 |
WordNet (lexical) |
| WN18RR |
40,943 |
11 |
93,003 |
Filtered WordNet |
Relation Types (FB15k-237):
/people/person/nationality
/film/film/genre
/location/location/contains
/business/company/founders
- Many more...
Use Cases:
- Link prediction
- Relation extraction
- Knowledge base completion
Biomedical Knowledge
| Dataset |
Entities |
Relations |
Triples |
Description |
| Hetionet |
45,158 |
24 |
2,250,197 |
Integrates 29 biomedical databases |
Entity Types in Hetionet:
- Genes (20,945)
- Compounds (1,552)
- Diseases (137)
- Anatomy (400)
- Pathways (1,822)
- Pharmacologic classes
- Side effects
- Symptoms
- Molecular functions
- Biological processes
- Cellular components
Relation Types:
- Compound-binds-Gene
- Gene-associates-Disease
- Disease-presents-Symptom
- Compound-treats-Disease
- Compound-causes-Side effect
- Gene-participates-Pathway
- And 18 more...
Use Cases:
- Drug repurposing
- Disease mechanism discovery
- Target identification
- Multi-hop reasoning in biomedicine
Citation Network Datasets
| Dataset |
Nodes |
Edges |
Classes |
Description |
| Cora |
2,708 |
5,429 |
7 |
Machine learning papers |
| CiteSeer |
3,327 |
4,732 |
6 |
Computer science papers |
| PubMed |
19,717 |
44,338 |
3 |
Biomedical papers |
Use Cases:
- Node classification
- GNN baseline comparisons
- Method development
Retrosynthesis Datasets
| Dataset |
Size |
Description |
| USPTO-50k |
50,017 |
Curated patent reactions, single-step |
Features:
- Product → Reactants mapping
- Atom mapping for reaction centers
- Canonicalized SMILES
- Balanced across reaction types
Splits:
- Train: ~40,000
- Validation: ~5,000
- Test: ~5,000
Use Cases:
- Retrosynthesis prediction
- Reaction type classification
- Synthetic route planning
Dataset Usage Patterns
Loading Datasets
from torchdrug import datasets
# Basic loading
dataset = datasets.BBBP("~/molecule-datasets/")
# With transforms
from torchdrug import transforms
transform = transforms.VirtualNode()
dataset = datasets.BBBP("~/molecule-datasets/", transform=transform)
# Protein dataset
dataset = datasets.EnzymeCommission("~/protein-datasets/")
# Knowledge graph
dataset = datasets.FB15k237("~/kg-datasets/")
Data Splitting
# Random split
train, valid, test = dataset.split([0.8, 0.1, 0.1])
# Scaffold split (for molecules)
from torchdrug import utils
train, valid, test = dataset.split(
utils.scaffold_split(dataset, [0.8, 0.1, 0.1])
)
# Predefined splits (some datasets)
train, valid, test = dataset.split()
Feature Extraction
Node Features (Molecules):
- Atom type (one-hot or embedding)
- Formal charge
- Hybridization
- Aromaticity
- Number of hydrogens
- Chirality
Edge Features (Molecules):
- Bond type (single, double, triple, aromatic)
- Stereochemistry
- Conjugation
- Ring membership
Node Features (Proteins):
- Amino acid type (one-hot)
- Physicochemical properties
- Position in sequence
- Secondary structure
- Solvent accessibility
Edge Features (Proteins):
- Edge type (sequential, spatial, contact)
- Distance
- Angles and dihedrals
Choosing Datasets
By Task
Molecular Property Prediction:
- Start with BBBP or HIV (medium size, clear task)
- Use QM9 for quantum properties
- ESOL/FreeSolv for regression
Protein Function:
- EnzymeCommission (well-defined classes)
- GeneOntology (comprehensive annotations)
Drug Safety:
- Tox21 (standard benchmark)
- ClinTox (clinical relevance)
Structure-Based:
- PDBBind (protein-ligand)
- ProteinNet (structure prediction)
Knowledge Graph:
- FB15k-237 (standard benchmark)
- Hetionet (biomedical applications)
Generation:
- ZINC250k (training)
- QM9 (with properties)
Retrosynthesis:
By Size and Resources
Small (<5k, for testing):
- BACE, FreeSolv, ClinTox
- Core set of PDBBind
Medium (5k-100k):
- BBBP, HIV, ESOL, Tox21
- EnzymeCommission, Fold
- FB15k-237, WN18RR
Large (>100k):
- QM9, MUV, PCQM4M
- GeneOntology, AlphaFoldDB
- ZINC2M, BindingDB
By Domain
Drug Discovery: BBBP, HIV, Tox21, ESOL, ZINC
Quantum Chemistry: QM7, QM8, QM9, PCQM4M
Protein Engineering: Fluorescence, Stability, Solubility
Structural Biology: Fold, PDBBind, ProteinNet, AlphaFoldDB
Biomedical: Hetionet, GeneOntology, EnzymeCommission
Retrosynthesis: USPTO-50k
Best Practices
- Start Small: Test on small datasets before scaling
- Scaffold Split: Use for realistic drug discovery evaluation
- Balanced Metrics: Use AUROC + AUPRC for imbalanced data
- Multiple Runs: Report mean ± std over multiple random seeds
- Data Leakage: Be careful with pre-trained models
- Domain Knowledge: Understand what you're predicting
- Validation: Always use held-out test set
- Preprocessing: Standardize features, handle missing values
1---2name: 033-torchdrug-6851951d3description: Datasets Reference4---5# Datasets Reference67## Overview89TorchDrug provides 40+ curated datasets across multiple domains: molecular property prediction, protein modeling, knowledge graph reasoning, and retrosynthesis. All datasets support lazy loading, automatic downloading, and customizable feature extraction.1011## Molecular Property Prediction Datasets1213### Drug Discovery Classification1415| Dataset | Size | Task | Classes | Description |16|---------|------|------|---------|-------------|17| **BACE** | 1,513 | Binary | 2 | β-secretase inhibition for Alzheimer's |18| **BBBP** | 2,039 | Binary | 2 | Blood-brain barrier penetration |19| **HIV** | 41,127 | Binary | 2 | Inhibition of HIV replication |20| **ClinTox** | 1,478 | Multi-label | 2 | Clinical trial toxicity |21| **SIDER** | 1,427 | Multi-label | 27 | Side effects by system organ class |22| **Tox21** | 7,831 | Multi-label | 12 | Toxicity across 12 targets |23| **ToxCast** | 8,576 | Multi-label | 617 | High-throughput toxicology |24| **MUV** | 93,087 | Multi-label | 17 | Unbiased validation for screening |2526**Key Features:**27- All use scaffold splits for realistic evaluation28- Binary classification metrics: AUROC, AUPRC29- Multi-label handles missing values3031**Use Cases:**32- Drug safety prediction33- Virtual screening34- ADMET property prediction3536### Drug Discovery Regression3738| Dataset | Size | Property | Units | Description |39|---------|------|----------|-------|-------------|40| **ESOL** | 1,128 | Solubility | log(mol/L) | Water solubility |41| **FreeSolv** | 642 | Hydration | kcal/mol | Hydration free energy |42| **Lipophilicity** | 4,200 | LogD | - | Octanol/water distribution |43| **SAMPL** | 643 | Solvation | kcal/mol | Solvation free energies |4445**Metrics:** MAE, RMSE, R²46**Use Cases:** ADME optimization, lead optimization4748### Quantum Chemistry4950| Dataset | Size | Properties | Description |51|---------|------|------------|-------------|52| **QM7** | 7,165 | 1 | Atomization energy |53| **QM8** | 21,786 | 12 | Electronic spectra, excited states |54| **QM9** | 133,885 | 12 | Geometric, energetic, electronic, thermodynamic |55| **PCQM4M** | 3.8M | 1 | Large-scale HOMO-LUMO gap |5657**Properties (QM9):**58- Dipole moment59- Isotropic polarizability60- HOMO/LUMO energies61- Internal energy, enthalpy, free energy62- Heat capacity63- Electronic spatial extent6465**Use Cases:**66- Quantum property prediction67- Method development benchmarking68- Pre-training molecular models6970### Large Molecule Databases7172| Dataset | Size | Description | Use Case |73|---------|------|-------------|----------|74| **ZINC250k** | 250,000 | Drug-like molecules | Generative model training |75| **ZINC2M** | 2,000,000 | Drug-like molecules | Large-scale pre-training |76| **ChEMBL** | Millions | Bioactive molecules | Property prediction, generation |7778## Protein Datasets7980### Function Prediction8182| Dataset | Size | Task | Classes | Description |83|---------|------|------|---------|-------------|84| **EnzymeCommission** | 17,562 | Multi-class | 7 levels | EC number classification |85| **GeneOntology** | 46,796 | Multi-label | 489 | GO term prediction (BP/MF/CC) |86| **BetaLactamase** | 5,864 | Regression | - | Enzyme activity levels |87| **Fluorescence** | 54,025 | Regression | - | GFP fluorescence intensity |88| **Stability** | 53,614 | Regression | - | Thermostability (ΔΔG) |8990**Features:**91- Sequence and/or structure input92- Evolutionary information available93- Multiple train/test splits9495**Use Cases:**96- Protein engineering97- Function annotation98- Enzyme design99100### Localization and Solubility101102| Dataset | Size | Task | Classes | Description |103|---------|------|------|---------|-------------|104| **Solubility** | 62,478 | Binary | 2 | Protein solubility |105| **BinaryLocalization** | 22,168 | Binary | 2 | Membrane vs soluble |106| **SubcellularLocalization** | 8,943 | Multi-class | 10 | Subcellular compartment |107108**Use Cases:**109- Protein expression optimization110- Target identification111- Cell biology112113### Structure Prediction114115| Dataset | Size | Task | Description |116|---------|------|------|-------------|117| **Fold** | 16,712 | Multi-class (1,195) | Structural fold recognition |118| **SecondaryStructure** | 8,678 | Sequence labeling | 3-state or 8-state prediction |119| **ProteinNet** | Varied | Contact prediction | Residue-residue contacts |120121**Use Cases:**122- Structure prediction pipelines123- Fold recognition124- Contact map generation125126### Protein Interactions127128| Dataset | Size | Positives | Negatives | Description |129|---------|------|-----------|-----------|-------------|130| **HumanPPI** | 1,412 proteins | 6,584 | - | Human protein interactions |131| **YeastPPI** | 2,018 proteins | 6,451 | - | Yeast protein interactions |132| **PPIAffinity** | 2,156 pairs | - | - | Binding affinity values |133134**Use Cases:**135- PPI prediction136- Network biology137- Drug target identification138139### Protein-Ligand Binding140141| Dataset | Size | Type | Description |142|---------|------|------|-------------|143| **BindingDB** | ~1.5M | Affinity | Comprehensive binding data |144| **PDBBind** | 20,000+ | 3D complexes | Structure-based binding |145| - Refined Set | 5,316 | High quality | Curated crystal structures |146| - Core Set | 285 | Benchmark | Diverse test set |147148**Use Cases:**149- Binding affinity prediction150- Structure-based drug design151- Scoring function development152153### Large Protein Databases154155| Dataset | Size | Description |156|---------|------|-------------|157| **AlphaFoldDB** | 200M+ | Predicted structures for most known proteins |158| **UniProt** | Integration | Sequence and annotation data |159160## Knowledge Graph Datasets161162### General Knowledge163164| Dataset | Entities | Relations | Triples | Domain |165|---------|----------|-----------|---------|--------|166| **FB15k** | 14,951 | 1,345 | 592,213 | Freebase (general knowledge) |167| **FB15k-237** | 14,541 | 237 | 310,116 | Filtered Freebase |168| **WN18** | 40,943 | 18 | 151,442 | WordNet (lexical) |169| **WN18RR** | 40,943 | 11 | 93,003 | Filtered WordNet |170171**Relation Types (FB15k-237):**172- `/people/person/nationality`173- `/film/film/genre`174- `/location/location/contains`175- `/business/company/founders`176- Many more...177178**Use Cases:**179- Link prediction180- Relation extraction181- Knowledge base completion182183### Biomedical Knowledge184185| Dataset | Entities | Relations | Triples | Description |186|---------|----------|-----------|---------|-------------|187| **Hetionet** | 45,158 | 24 | 2,250,197 | Integrates 29 biomedical databases |188189**Entity Types in Hetionet:**190- Genes (20,945)191- Compounds (1,552)192- Diseases (137)193- Anatomy (400)194- Pathways (1,822)195- Pharmacologic classes196- Side effects197- Symptoms198- Molecular functions199- Biological processes200- Cellular components201202**Relation Types:**203- Compound-binds-Gene204- Gene-associates-Disease205- Disease-presents-Symptom206- Compound-treats-Disease207- Compound-causes-Side effect208- Gene-participates-Pathway209- And 18 more...210211**Use Cases:**212- Drug repurposing213- Disease mechanism discovery214- Target identification215- Multi-hop reasoning in biomedicine216217## Citation Network Datasets218219| Dataset | Nodes | Edges | Classes | Description |220|---------|-------|-------|---------|-------------|221| **Cora** | 2,708 | 5,429 | 7 | Machine learning papers |222| **CiteSeer** | 3,327 | 4,732 | 6 | Computer science papers |223| **PubMed** | 19,717 | 44,338 | 3 | Biomedical papers |224225**Use Cases:**226- Node classification227- GNN baseline comparisons228- Method development229230## Retrosynthesis Datasets231232| Dataset | Size | Description |233|---------|------|-------------|234| **USPTO-50k** | 50,017 | Curated patent reactions, single-step |235236**Features:**237- Product → Reactants mapping238- Atom mapping for reaction centers239- Canonicalized SMILES240- Balanced across reaction types241242**Splits:**243- Train: ~40,000244- Validation: ~5,000245- Test: ~5,000246247**Use Cases:**248- Retrosynthesis prediction249- Reaction type classification250- Synthetic route planning251252## Dataset Usage Patterns253254### Loading Datasets255256```python257from torchdrug import datasets258259# Basic loading260dataset = datasets.BBBP("~/molecule-datasets/")261262# With transforms263from torchdrug import transforms264transform = transforms.VirtualNode()265dataset = datasets.BBBP("~/molecule-datasets/", transform=transform)266267# Protein dataset268dataset = datasets.EnzymeCommission("~/protein-datasets/")269270# Knowledge graph271dataset = datasets.FB15k237("~/kg-datasets/")272```273274### Data Splitting275276```python277# Random split278train, valid, test = dataset.split([0.8, 0.1, 0.1])279280# Scaffold split (for molecules)281from torchdrug import utils282train, valid, test = dataset.split(283 utils.scaffold_split(dataset, [0.8, 0.1, 0.1])284)285286# Predefined splits (some datasets)287train, valid, test = dataset.split()288```289290### Feature Extraction291292**Node Features (Molecules):**293- Atom type (one-hot or embedding)294- Formal charge295- Hybridization296- Aromaticity297- Number of hydrogens298- Chirality299300**Edge Features (Molecules):**301- Bond type (single, double, triple, aromatic)302- Stereochemistry303- Conjugation304- Ring membership305306**Node Features (Proteins):**307- Amino acid type (one-hot)308- Physicochemical properties309- Position in sequence310- Secondary structure311- Solvent accessibility312313**Edge Features (Proteins):**314- Edge type (sequential, spatial, contact)315- Distance316- Angles and dihedrals317318## Choosing Datasets319320### By Task321322**Molecular Property Prediction:**323- Start with BBBP or HIV (medium size, clear task)324- Use QM9 for quantum properties325- ESOL/FreeSolv for regression326327**Protein Function:**328- EnzymeCommission (well-defined classes)329- GeneOntology (comprehensive annotations)330331**Drug Safety:**332- Tox21 (standard benchmark)333- ClinTox (clinical relevance)334335**Structure-Based:**336- PDBBind (protein-ligand)337- ProteinNet (structure prediction)338339**Knowledge Graph:**340- FB15k-237 (standard benchmark)341- Hetionet (biomedical applications)342343**Generation:**344- ZINC250k (training)345- QM9 (with properties)346347**Retrosynthesis:**348- USPTO-50k (only choice)349350### By Size and Resources351352**Small (<5k, for testing):**353- BACE, FreeSolv, ClinTox354- Core set of PDBBind355356**Medium (5k-100k):**357- BBBP, HIV, ESOL, Tox21358- EnzymeCommission, Fold359- FB15k-237, WN18RR360361**Large (>100k):**362- QM9, MUV, PCQM4M363- GeneOntology, AlphaFoldDB364- ZINC2M, BindingDB365366### By Domain367368**Drug Discovery:** BBBP, HIV, Tox21, ESOL, ZINC369**Quantum Chemistry:** QM7, QM8, QM9, PCQM4M370**Protein Engineering:** Fluorescence, Stability, Solubility371**Structural Biology:** Fold, PDBBind, ProteinNet, AlphaFoldDB372**Biomedical:** Hetionet, GeneOntology, EnzymeCommission373**Retrosynthesis:** USPTO-50k374375## Best Practices3763771. **Start Small**: Test on small datasets before scaling3782. **Scaffold Split**: Use for realistic drug discovery evaluation3793. **Balanced Metrics**: Use AUROC + AUPRC for imbalanced data3804. **Multiple Runs**: Report mean ± std over multiple random seeds3815. **Data Leakage**: Be careful with pre-trained models3826. **Domain Knowledge**: Understand what you're predicting3837. **Validation**: Always use held-out test set3848. **Preprocessing**: Standardize features, handle missing values