Drug Repurposing Hub Query Skill
Search the Broad Institute Drug Repurposing Hub by any entity.
Auto-detects input type by pattern:
| Input Pattern |
Detected As |
Match Logic |
BRD-A12345678 |
Broad compound ID |
prefix on broad_id |
ABCDEFGHIJKLMN-OPQRSTUVWX-Y |
InChIKey |
exact on InChIKey |
EGFR, BRAF, TOP1 |
Gene / target |
exact token in target (pipe-separated) |
| anything else |
free text |
substring on pert_iname, moa, indication, disease_area |
API
| Function |
Input |
Returns |
load_drugs(path) |
drug TSV path |
list[dict] |
load_samples(path) |
sample TSV path |
list[dict] |
load_merged() |
— |
list[dict] (drugs + chemical IDs from samples) |
search(entity) |
single entity string |
list[dict] |
search_batch(entities) |
list of entity strings |
dict[str, list[dict]] |
summarize(hits, entity) |
hits + label |
compact LLM-readable text |
to_json(hits) |
list[dict] |
list[dict] (JSON-serialisable) |
Usage
See if __name__ == "__main__" block in 29_Drug_Repurposing_Hub.py for
runnable examples covering: drug name, gene target, MOA keyword, disease
area, batch search, and JSON output.
from importlib.machinery import SourceFileLoader
hub = SourceFileLoader("hub", "29_Drug_Repurposing_Hub.py").load_module()
# Single drug lookup
hits = hub.search("imatinib")
print(hub.summarize(hits, "imatinib"))
# Target-based search
hits = hub.search("EGFR")
# Batch
results = hub.search_batch(["metformin", "aspirin", "BRAF"])
Data
- Source: Broad Institute Drug Repurposing Hub (https://repo-hub.broadinstitute.org/repurposing)
- Drug file:
repo-drug-annotation-20200324.txt — tab-delimited, !-prefixed comment lines
- Columns:
pert_iname, clinical_phase, moa, target, disease_area, indication
- Sample file:
repo-sample-annotation-20240610.txt — tab-delimited, !-prefixed comment lines
- Columns include:
broad_id, pert_iname, InChIKey, pubchem_cid, smiles, vendor, purity, etc.
- Merge: on
pert_iname; first sample with non-empty InChIKey is kept per drug
- Path:
DATA_DIR variable in 29_Drug_Repurposing_Hub.py
- Citation: Corsello SM et al. Nature Medicine 23, 405–408 (2017). doi:10.1038/nm.4306
1---2name: drug-repurposing-hub3description: Query the Broad Institute Drug Repurposing Hub (~6,800 compounds). Look up drugs by name, gene target, MOA, disease area, Broad ID, or InChIKey. Returns clinical phase, mechanism of action, targets, disease area, indication, and chemical identifiers.4---5
6# Drug Repurposing Hub Query Skill
7
8Search the Broad Institute Drug Repurposing Hub by any entity.
9Auto-detects input type by pattern:
10
11| Input Pattern | Detected As | Match Logic |
12|---|---|---|
13| `BRD-A12345678` | Broad compound ID | prefix on `broad_id` |
14| `ABCDEFGHIJKLMN-OPQRSTUVWX-Y` | InChIKey | exact on `InChIKey` |
15| `EGFR`, `BRAF`, `TOP1` | Gene / target | exact token in `target` (pipe-separated) |
16| anything else | free text | substring on `pert_iname`, `moa`, `indication`, `disease_area` |
17
18## API
19
20| Function | Input | Returns |
21|---|---|---|
22| `load_drugs(path)` | drug TSV path | list[dict] |
23| `load_samples(path)` | sample TSV path | list[dict] |
24| `load_merged()` | — | list[dict] (drugs + chemical IDs from samples) |
25| `search(entity)` | single entity string | list[dict] |
26| `search_batch(entities)` | list of entity strings | dict[str, list[dict]] |
27| `summarize(hits, entity)` | hits + label | compact LLM-readable text |
28| `to_json(hits)` | list[dict] | list[dict] (JSON-serialisable) |
29
30## Usage
31
32See `if __name__ == "__main__"` block in `29_Drug_Repurposing_Hub.py` for
33runnable examples covering: drug name, gene target, MOA keyword, disease
34area, batch search, and JSON output.
35
36```python
37from importlib.machinery import SourceFileLoader
38hub = SourceFileLoader("hub", "29_Drug_Repurposing_Hub.py").load_module()
39
40# Single drug lookup
41hits = hub.search("imatinib")
42print(hub.summarize(hits, "imatinib"))
43
44# Target-based search
45hits = hub.search("EGFR")
46
47# Batch
48results = hub.search_batch(["metformin", "aspirin", "BRAF"])
49```
50
51## Data
52
53- **Source**: Broad Institute Drug Repurposing Hub (https://repo-hub.broadinstitute.org/repurposing)
54- **Drug file**: `repo-drug-annotation-20200324.txt` — tab-delimited, `!`-prefixed comment lines
55 - Columns: `pert_iname`, `clinical_phase`, `moa`, `target`, `disease_area`, `indication`
56- **Sample file**: `repo-sample-annotation-20240610.txt` — tab-delimited, `!`-prefixed comment lines
57 - Columns include: `broad_id`, `pert_iname`, `InChIKey`, `pubchem_cid`, `smiles`, `vendor`, `purity`, etc.
58- **Merge**: on `pert_iname`; first sample with non-empty `InChIKey` is kept per drug
59- **Path**: `DATA_DIR` variable in `29_Drug_Repurposing_Hub.py`
60- **Citation**: Corsello SM et al. *Nature Medicine* 23, 405–408 (2017). doi:10.1038/nm.4306