TAC 2017 ADR Query Skill
Search 200 FDA drug labels annotated with adverse reactions, severity,
and MedDRA normalization from the TAC 2017 shared task.
Entity Auto-detection
| Input Pattern |
Detected As |
Match Logic |
10019211 (8 digits) |
MedDRA ID |
exact on meddra_pt_id or meddra_llt_id |
ACTEMRA (known drug) |
Drug name |
exact (case-insensitive) on drug label name |
headache (known ADR) |
ADR string |
exact on ADR reaction string |
| anything else |
Free text |
substring on drug names, ADR strings, MedDRA PT/LLT names |
API
| Function |
Input |
Returns |
search(entity) |
single entity string |
list[dict] — matching label hit(s) |
search_batch(entities) |
list of entity strings |
dict[str, list[dict]] |
summarize(hits, entity) |
hit list + query label |
compact LLM-readable text |
to_json(hits) |
hit list |
list[dict] (JSON-serializable) |
list_drugs() |
— |
sorted list of all drug names |
stats() |
— |
dataset-level statistics dict |
Hit Dict Structure
Each hit returned by search() contains:
| Field |
Type |
Description |
drug |
str |
Drug label name |
source_file |
str |
XML filename |
sections |
list[str] |
Annotated section names (e.g. "adverse reactions") |
mention_counts |
dict |
Count per mention type (AdverseReaction, Severity, …) |
num_reactions |
int |
Total unique reactions in this label |
positive_adrs |
list[str] |
Positive (non-negated, non-hypothetical) ADR strings |
reactions |
list[dict] |
Each with adr, meddra_pt, meddra_pt_id, optional meddra_llt, meddra_llt_id, flag |
Usage
See if __name__ == "__main__" block in 37_TAC_2017_ADR.py for runnable
examples covering: drug name lookup, ADR string search, MedDRA ID search,
batch search, and JSON output.
from importlib.machinery import SourceFileLoader
tac = SourceFileLoader("tac2017", "/path/to/37_TAC_2017_ADR.py").load_module()
# Single drug
hits = tac.search("ACTEMRA")
print(tac.summarize(hits, "ACTEMRA"))
# ADR across all labels
hits = tac.search("headache")
print(tac.summarize(hits, "headache"))
# MedDRA PT ID
hits = tac.search("10019211")
# Batch
results = tac.search_batch(["ENBREL", "nausea", "10002198"])
Data
- Source: TAC 2017 ADR shared task (NLM / FDA)
- Files:
gold_xml/ (99 test labels) + train_xml/ (101 training labels), each annotated XML
- Annotations: Mentions (AdverseReaction, Severity, Factor, DrugClass, Negation, Animal), Relations (Negated, Hypothetical, Effect), Reactions (unique ADRs with MedDRA PT/LLT normalization)
- MedDRA version: 18.1
- Path:
DATA_DIR variable in 37_TAC_2017_ADR.py
- Reference: https://bionlp.nlm.nih.gov/tac2017adversereactions/
1---2name: tac2017-adr3description: Query TAC 2017 ADR annotated drug labels for adverse drug reactions. Use whenever the user asks about ADRs extracted from FDA drug labels, MedDRA-normalized adverse reactions, or wants to look up a drug name, ADR string, or MedDRA code in the TAC 2017 ADR corpus.4---5
6# TAC 2017 ADR Query Skill
7
8Search 200 FDA drug labels annotated with adverse reactions, severity,
9and MedDRA normalization from the TAC 2017 shared task.
10
11## Entity Auto-detection
12
13| Input Pattern | Detected As | Match Logic |
14|---|---|---|
15| `10019211` (8 digits) | MedDRA ID | exact on `meddra_pt_id` or `meddra_llt_id` |
16| `ACTEMRA` (known drug) | Drug name | exact (case-insensitive) on drug label name |
17| `headache` (known ADR) | ADR string | exact on ADR reaction string |
18| anything else | Free text | substring on drug names, ADR strings, MedDRA PT/LLT names |
19
20## API
21
22| Function | Input | Returns |
23|---|---|---|
24| `search(entity)` | single entity string | `list[dict]` — matching label hit(s) |
25| `search_batch(entities)` | list of entity strings | `dict[str, list[dict]]` |
26| `summarize(hits, entity)` | hit list + query label | compact LLM-readable text |
27| `to_json(hits)` | hit list | `list[dict]` (JSON-serializable) |
28| `list_drugs()` | — | sorted list of all drug names |
29| `stats()` | — | dataset-level statistics dict |
30
31## Hit Dict Structure
32
33Each hit returned by `search()` contains:
34
35| Field | Type | Description |
36|---|---|---|
37| `drug` | str | Drug label name |
38| `source_file` | str | XML filename |
39| `sections` | list[str] | Annotated section names (e.g. "adverse reactions") |
40| `mention_counts` | dict | Count per mention type (AdverseReaction, Severity, …) |
41| `num_reactions` | int | Total unique reactions in this label |
42| `positive_adrs` | list[str] | Positive (non-negated, non-hypothetical) ADR strings |
43| `reactions` | list[dict] | Each with `adr`, `meddra_pt`, `meddra_pt_id`, optional `meddra_llt`, `meddra_llt_id`, `flag` |
44
45## Usage
46
47See `if __name__ == "__main__"` block in `37_TAC_2017_ADR.py` for runnable
48examples covering: drug name lookup, ADR string search, MedDRA ID search,
49batch search, and JSON output.
50
51```python
52from importlib.machinery import SourceFileLoader
53tac = SourceFileLoader("tac2017", "/path/to/37_TAC_2017_ADR.py").load_module()
54
55# Single drug
56hits = tac.search("ACTEMRA")
57print(tac.summarize(hits, "ACTEMRA"))
58
59# ADR across all labels
60hits = tac.search("headache")
61print(tac.summarize(hits, "headache"))
62
63# MedDRA PT ID
64hits = tac.search("10019211")
65
66# Batch
67results = tac.search_batch(["ENBREL", "nausea", "10002198"])
68```
69
70## Data
71
72- **Source**: TAC 2017 ADR shared task (NLM / FDA)
73- **Files**: `gold_xml/` (99 test labels) + `train_xml/` (101 training labels), each annotated XML
74- **Annotations**: Mentions (AdverseReaction, Severity, Factor, DrugClass, Negation, Animal), Relations (Negated, Hypothetical, Effect), Reactions (unique ADRs with MedDRA PT/LLT normalization)
75- **MedDRA version**: 18.1
76- **Path**: `DATA_DIR` variable in `37_TAC_2017_ADR.py`
77- **Reference**: https://bionlp.nlm.nih.gov/tac2017adversereactions/