Source: https://github.com/aipoch/medical-research-skills
When to Use
- You need curated, AI-ready datasets for drug discovery tasks (e.g., ADME, toxicity, bioactivity, DTI/DDI).
- You want standardized benchmarks with consistent evaluation protocols (including multi-seed benchmark groups).
- You require meaningful data splits such as scaffold splits (chemical diversity) or cold-start splits (unseen drugs/targets).
- You are building models for single-instance prediction (molecular/protein properties) or multi-instance prediction (interactions).
- You are doing goal-directed molecular generation and need property Oracles for scoring/optimization.
Key Features
- Dataset access by task type
- Single-instance prediction: ADME, Toxicity (Tox), HTS, QM, and more.
- Multi-instance prediction: DTI, DDI, PPI, and more.
- Generation: MolGen, RetroSyn, PairMolGen.
- Standardized splitting utilities
random, scaffold, and cold-start variants such as cold_drug, cold_target, cold_drug_target, plus temporal where applicable.
- Unified evaluation
- Built-in
Evaluator with common metrics (ROC-AUC, PR-AUC, RMSE, MAE, Spearman, etc.).
- Benchmark groups
- Curated collections (e.g., ADMET group) with recommended evaluation protocols (commonly 5 seeds).
- Chem/data utilities
- Molecular format conversion (e.g., SMILES → PyG), filtering, balancing, negative sampling, entity retrieval (CID→SMILES, UniProt→sequence).
- Molecular Oracles
- Property scoring functions usable for goal-directed generation workflows (see
references/oracles.md).
Dependencies
Install (recommended):
uv pip install PyTDC
Upgrade:
uv pip install PyTDC --upgrade
Core runtime dependencies (installed automatically; versions depend on the PyTDC release you install):
PyTDC (latest from PyPI)
numpy
pandas
scikit-learn
tqdm
seaborn
fuzzywuzzy
Optional dependencies may be pulled in automatically depending on which submodules you use (e.g., graph backends or chemistry toolchains).
Example Usage
A complete runnable example that:
- loads an ADME dataset,
- performs a scaffold split,
- trains a simple baseline model,
- evaluates with a standard metric,
- queries an Oracle score for a SMILES.
# pip install PyTDC scikit-learn
from tdc.single_pred import ADME
from tdc import Evaluator, Oracle
from sklearn.pipeline import Pipeline
from sklearn.feature_extraction.text import CountVectorizer
from sklearn.linear_model import Ridge
def main():
# 1) Load a single-instance prediction dataset (ADME)
data = ADME(name="Caco2_Wang")
# 2) Create a scaffold split (train/valid/test)
split = data.get_split(method="scaffold", seed=42, frac=[0.7, 0.1, 0.2])
train, valid, test = split["train"], split["valid"], split["test"]
# 3) Train a simple baseline model on SMILES strings
# (character n-gram + ridge regression; replace with your own model)
model = Pipeline(
steps=[
("featurizer", CountVectorizer(analyzer="char", ngram_range=(2, 5))),
("regressor", Ridge(alpha=1.0)),
]
)
model.fit(train["Drug"], train["Y"])
# 4) Evaluate on the test set using a TDC Evaluator
y_pred = model.predict(test["Drug"])
evaluator = Evaluator(name="MAE")
mae = evaluator(test["Y"], y_pred)
print(f"Test MAE: {mae:.4f}")
# 5) Oracle scoring example (property scoring for a SMILES)
oracle = Oracle(name="DRD2")
score = oracle("CC(C)Cc1ccc(cc1)C(C)C(O)=O")
print(f"DRD2 Oracle score: {score}")
if __name__ == "__main__":
main()
Related references and templates (if present in this skill package):
- Oracle catalog and usage:
references/oracles.md
- Utility functions (splits, processing, retrieval):
references/utilities.md
- Dataset catalog:
references/datasets.md
- Workflow templates:
scripts/load_and_split_data.py, scripts/benchmark_evaluation.py, scripts/molecular_generation.py
Implementation Details
1) Dataset Access Pattern
PyTDC datasets follow a consistent interface:
from tdc.<problem> import <Task>
data = <Task>(name="<DatasetName>")
df = data.get_data(format="df")
split = data.get_split(method="scaffold", seed=1, frac=[0.7, 0.1, 0.2])
<problem> is typically one of:
single_pred (single-entity property prediction)
multi_pred (pairwise/multi-entity interaction prediction)
generation (molecule/reaction generation tasks)
2) Splitting Strategies (Key Parameters)
Use get_split(...) to obtain {"train": ..., "valid": ..., "test": ...}.
Common parameters:
method: split strategy
seed: random seed for reproducibility
frac: [train, valid, test] fractions (when supported)
Typical methods:
random: random shuffling split
scaffold: Bemis–Murcko scaffold-based split to reduce scaffold leakage and improve chemical generalization
- Cold-start (commonly for interaction tasks like DTI/DDI):
cold_drug: test contains unseen drugs
cold_target: test contains unseen targets
cold_drug_target: test contains unseen drugs and targets
temporal: time-based split for datasets with timestamps (when available)
Example:
split = data.get_split(method="cold_target", seed=1)
3) Standardized Evaluation
TDC provides a unified evaluator:
from tdc import Evaluator
evaluator = Evaluator(name="ROC-AUC") # classification
score = evaluator(y_true, y_pred)
Choose metrics appropriate to the task type:
- Classification:
ROC-AUC, PR-AUC, F1, Accuracy, etc.
- Regression:
RMSE, MAE, R2, Spearman, Pearson, etc.
4) Data Schemas (What Columns to Expect)
While schemas vary by task, common conventions include:
Single-instance prediction (e.g., ADME/Tox):
Drug (often SMILES) and label Y
- sometimes
Drug_ID / Compound_ID
Multi-instance prediction (e.g., DTI):
Drug (SMILES), Target (protein sequence), label Y
- plus identifiers such as
Drug_ID, Target_ID
5) Oracles for Molecular Optimization
Oracles provide a callable scoring interface:
from tdc import Oracle
oracle = Oracle(name="GSK3B")
score = oracle("CCO...")
scores = oracle(["SMILES1", "SMILES2"])
Use Oracles to:
- score candidate molecules during generation,
- define optimization objectives,
- compare molecules under consistent property predictors.
For the full list of Oracles and their expected inputs/outputs, see references/oracles.md.
1---2name: pytdc3description: Therapeutics Data Commons (PyTDC) for AI-ready therapeutic ML datasets and benchmarks; use it when you need standardized dataset loading, meaningful splits (e.g., scaffold/cold-start), and consistent evaluation for ADME/Toxicity/DTI/DDI or molecular optimization.4license: MIT5---6> **Source**: [https://github.com/aipoch/medical-research-skills](https://github.com/aipoch/medical-research-skills)78## When to Use910- You need **curated, AI-ready datasets** for drug discovery tasks (e.g., ADME, toxicity, bioactivity, DTI/DDI).11- You want **standardized benchmarks** with consistent evaluation protocols (including multi-seed benchmark groups).12- You require **meaningful data splits** such as **scaffold splits** (chemical diversity) or **cold-start splits** (unseen drugs/targets).13- You are building models for **single-instance prediction** (molecular/protein properties) or **multi-instance prediction** (interactions).14- You are doing **goal-directed molecular generation** and need **property Oracles** for scoring/optimization.1516## Key Features1718- **Dataset access by task type**19 - Single-instance prediction: ADME, Toxicity (Tox), HTS, QM, and more.20 - Multi-instance prediction: DTI, DDI, PPI, and more.21 - Generation: MolGen, RetroSyn, PairMolGen.22- **Standardized splitting utilities**23 - `random`, `scaffold`, and cold-start variants such as `cold_drug`, `cold_target`, `cold_drug_target`, plus `temporal` where applicable.24- **Unified evaluation**25 - Built-in `Evaluator` with common metrics (ROC-AUC, PR-AUC, RMSE, MAE, Spearman, etc.).26- **Benchmark groups**27 - Curated collections (e.g., ADMET group) with recommended evaluation protocols (commonly 5 seeds).28- **Chem/data utilities**29 - Molecular format conversion (e.g., SMILES → PyG), filtering, balancing, negative sampling, entity retrieval (CID→SMILES, UniProt→sequence).30- **Molecular Oracles**31 - Property scoring functions usable for goal-directed generation workflows (see `references/oracles.md`).3233## Dependencies3435Install (recommended):3637```bash38uv pip install PyTDC39```4041Upgrade:4243```bash44uv pip install PyTDC --upgrade45```4647Core runtime dependencies (installed automatically; versions depend on the PyTDC release you install):4849- `PyTDC` (latest from PyPI)50- `numpy`51- `pandas`52- `scikit-learn`53- `tqdm`54- `seaborn`55- `fuzzywuzzy`5657Optional dependencies may be pulled in automatically depending on which submodules you use (e.g., graph backends or chemistry toolchains).5859## Example Usage6061A complete runnable example that:621) loads an ADME dataset, 632) performs a scaffold split, 643) trains a simple baseline model, 654) evaluates with a standard metric, 665) queries an Oracle score for a SMILES.6768```python69# pip install PyTDC scikit-learn7071from tdc.single_pred import ADME72from tdc import Evaluator, Oracle7374from sklearn.pipeline import Pipeline75from sklearn.feature_extraction.text import CountVectorizer76from sklearn.linear_model import Ridge7778def main():79 # 1) Load a single-instance prediction dataset (ADME)80 data = ADME(name="Caco2_Wang")8182 # 2) Create a scaffold split (train/valid/test)83 split = data.get_split(method="scaffold", seed=42, frac=[0.7, 0.1, 0.2])84 train, valid, test = split["train"], split["valid"], split["test"]8586 # 3) Train a simple baseline model on SMILES strings87 # (character n-gram + ridge regression; replace with your own model)88 model = Pipeline(89 steps=[90 ("featurizer", CountVectorizer(analyzer="char", ngram_range=(2, 5))),91 ("regressor", Ridge(alpha=1.0)),92 ]93 )94 model.fit(train["Drug"], train["Y"])9596 # 4) Evaluate on the test set using a TDC Evaluator97 y_pred = model.predict(test["Drug"])98 evaluator = Evaluator(name="MAE")99 mae = evaluator(test["Y"], y_pred)100 print(f"Test MAE: {mae:.4f}")101102 # 5) Oracle scoring example (property scoring for a SMILES)103 oracle = Oracle(name="DRD2")104 score = oracle("CC(C)Cc1ccc(cc1)C(C)C(O)=O")105 print(f"DRD2 Oracle score: {score}")106107if __name__ == "__main__":108 main()109```110111Related references and templates (if present in this skill package):112- Oracle catalog and usage: `references/oracles.md`113- Utility functions (splits, processing, retrieval): `references/utilities.md`114- Dataset catalog: `references/datasets.md`115- Workflow templates: `scripts/load_and_split_data.py`, `scripts/benchmark_evaluation.py`, `scripts/molecular_generation.py`116117## Implementation Details118119### 1) Dataset Access Pattern120121PyTDC datasets follow a consistent interface:122123```python124from tdc.<problem> import <Task>125126data = <Task>(name="<DatasetName>")127df = data.get_data(format="df")128split = data.get_split(method="scaffold", seed=1, frac=[0.7, 0.1, 0.2])129```130131- `<problem>` is typically one of:132 - `single_pred` (single-entity property prediction)133 - `multi_pred` (pairwise/multi-entity interaction prediction)134 - `generation` (molecule/reaction generation tasks)135136### 2) Splitting Strategies (Key Parameters)137138Use `get_split(...)` to obtain `{"train": ..., "valid": ..., "test": ...}`.139140Common parameters:141- `method`: split strategy142- `seed`: random seed for reproducibility143- `frac`: `[train, valid, test]` fractions (when supported)144145Typical methods:146- `random`: random shuffling split147- `scaffold`: Bemis–Murcko scaffold-based split to reduce scaffold leakage and improve chemical generalization148- Cold-start (commonly for interaction tasks like DTI/DDI):149 - `cold_drug`: test contains unseen drugs150 - `cold_target`: test contains unseen targets151 - `cold_drug_target`: test contains unseen drugs and targets152- `temporal`: time-based split for datasets with timestamps (when available)153154Example:155156```python157split = data.get_split(method="cold_target", seed=1)158```159160### 3) Standardized Evaluation161162TDC provides a unified evaluator:163164```python165from tdc import Evaluator166167evaluator = Evaluator(name="ROC-AUC") # classification168score = evaluator(y_true, y_pred)169```170171Choose metrics appropriate to the task type:172- Classification: `ROC-AUC`, `PR-AUC`, `F1`, `Accuracy`, etc.173- Regression: `RMSE`, `MAE`, `R2`, `Spearman`, `Pearson`, etc.174175### 4) Data Schemas (What Columns to Expect)176177While schemas vary by task, common conventions include:178179- Single-instance prediction (e.g., ADME/Tox):180 - `Drug` (often SMILES) and label `Y`181 - sometimes `Drug_ID` / `Compound_ID`182183- Multi-instance prediction (e.g., DTI):184 - `Drug` (SMILES), `Target` (protein sequence), label `Y`185 - plus identifiers such as `Drug_ID`, `Target_ID`186187### 5) Oracles for Molecular Optimization188189Oracles provide a callable scoring interface:190191```python192from tdc import Oracle193194oracle = Oracle(name="GSK3B")195score = oracle("CCO...")196scores = oracle(["SMILES1", "SMILES2"])197```198199Use Oracles to:200- score candidate molecules during generation,201- define optimization objectives,202- compare molecules under consistent property predictors.203204For the full list of Oracles and their expected inputs/outputs, see `references/oracles.md`.