Source: https://github.com/aipoch/medical-research-skills
Baseline Extraction (RCT)
This skill extracts 10 key baseline characteristics from clinical trial articles. It implements a hybrid workflow:
- PubMed Lookup: Checks PubMed API using the PMID to verify existence and get basic metadata.
- LLM Extraction: Analyzes the article text to extract detailed baseline data (since PubMed metadata is limited).
When to Use
- Use this skill when you need extracts clinical trial baseline data (study, region, participants, etc.) from article text or pmid. checks pubmed for metadata; always falls back to llm extraction for full details in a reproducible workflow.
- Use this skill when a data analytics task needs a packaged method instead of ad-hoc freeform output.
- Use this skill when the user expects a concrete deliverable, validation step, or file-based result.
- Use this skill when
scripts/extract_pdf.py is the most direct path to complete the request.
- Use this skill when you need the
baseline-extraction for clinical trials package behavior rather than a generic answer.
Key Features
- Scope-focused workflow aligned to: Extracts clinical trial baseline data (study, region, participants, etc.) from article text or PMID. Checks PubMed for metadata; always falls back to LLM extraction for full details.
- Packaged executable path(s):
scripts/extract_pdf.py.
- Reference material available in
references/ for task-specific guidance.
- Structured execution path designed to keep outputs consistent and reviewable.
Dependencies
Python: 3.10+. Repository baseline for current packaged skills.
Third-party packages: not explicitly version-pinned in this skill package. Add pinned versions if this skill needs stricter environment control.
Example Usage
cd "20260316/scientific-skills/Data Analytics/baseline-extraction-for-clinical-trials"
python -m py_compile scripts/extract_pdf.py
python scripts/extract_pdf.py --help
Example run plan:
- Confirm the user input, output path, and any required config values.
- Edit the in-file
CONFIG block or documented parameters if the script uses fixed settings.
- Run
python scripts/extract_pdf.py with the validated inputs.
- Review the generated output and return the final artifact with any assumptions called out.
Implementation Details
See ## Workflow above for related details.
- Execution model: validate the request, choose the packaged workflow, and produce a bounded deliverable.
- Input controls: confirm the source files, scope limits, output format, and acceptance criteria before running any script.
- Primary implementation surface:
scripts/extract_pdf.py.
- Reference guidance:
references/ contains supporting rules, prompts, or checklists.
- Parameters to clarify first: input path, output path, scope filters, thresholds, and any domain-specific constraints.
- Output discipline: keep results reproducible, identify assumptions explicitly, and avoid undocumented side effects.
Workflow
Step 1: Check PubMed (Deterministic)
If the user provides a PMID, use the baseline_extractor.py script to check PubMed.
import subprocess
import json
# Replace <PMID> with actual PMID
result = subprocess.run(["python", "scripts/baseline_extractor.py", "<PMID>"], capture_output=True, text=True)
print(result.stdout)
Analyze the Script Output:
- If
status is "success": Stop here. Return the data JSON to the user.
- If
status is "not_found", "incomplete", or "error" (or if no PMID was provided): Proceed to Step 2.
Step 2: LLM Extraction (Fallback)
If Step 1 did not yield a complete result, use the LLM to extract the information from the full article text.
Input:
- Full article text provided by the user.
Instructions:
- Read the Extraction Schema carefully.
- Analyze the text to identify all 10 required fields.
- Ensure the output is strictly in the JSON format defined in the schema.
- Constraint: Do not hallucinate. If a field is not mentioned in the text, set it to
null or an empty string.
Output
Return the final result as a Markdown code block containing the JSON object.
{
"study": "...",
"region": "...",
...
}
Helper Scripts
PDF Text Extraction
When the user provides a PDF file path, use extract_pdf.py to extract the text content before assessment:
1---2name: baseline-extraction-for-clinical-trials3description: Extracts clinical trial baseline data (study, region, participants, etc.) from article text or PMID. Checks PubMed for metadata; always falls back to LLM extraction for full details.4license: MIT5---6> **Source**: [https://github.com/aipoch/medical-research-skills](https://github.com/aipoch/medical-research-skills)
7
8# Baseline Extraction (RCT)
9
10This skill extracts 10 key baseline characteristics from clinical trial articles. It implements a hybrid workflow:
111. **PubMed Lookup**: Checks PubMed API using the PMID to verify existence and get basic metadata.
122. **LLM Extraction**: Analyzes the article text to extract detailed baseline data (since PubMed metadata is limited).
13
14## When to Use
15
16- Use this skill when you need extracts clinical trial baseline data (study, region, participants, etc.) from article text or pmid. checks pubmed for metadata; always falls back to llm extraction for full details in a reproducible workflow.
17- Use this skill when a data analytics task needs a packaged method instead of ad-hoc freeform output.
18- Use this skill when the user expects a concrete deliverable, validation step, or file-based result.
19- Use this skill when `scripts/extract_pdf.py` is the most direct path to complete the request.
20- Use this skill when you need the `baseline-extraction for clinical trials` package behavior rather than a generic answer.
21
22## Key Features
23
24- Scope-focused workflow aligned to: Extracts clinical trial baseline data (study, region, participants, etc.) from article text or PMID. Checks PubMed for metadata; always falls back to LLM extraction for full details.
25- Packaged executable path(s): `scripts/extract_pdf.py`.
26- Reference material available in `references/` for task-specific guidance.
27- Structured execution path designed to keep outputs consistent and reviewable.
28
29## Dependencies
30
31- `Python`: `3.10+`. Repository baseline for current packaged skills.
32- `Third-party packages`: `not explicitly version-pinned in this skill package`. Add pinned versions if this skill needs stricter environment control.
33
34## Example Usage
35
36```bash
37cd "20260316/scientific-skills/Data Analytics/baseline-extraction-for-clinical-trials"
38python -m py_compile scripts/extract_pdf.py
39python scripts/extract_pdf.py --help
40```
41
42Example run plan:
431. Confirm the user input, output path, and any required config values.
442. Edit the in-file `CONFIG` block or documented parameters if the script uses fixed settings.
453. Run `python scripts/extract_pdf.py` with the validated inputs.
464. Review the generated output and return the final artifact with any assumptions called out.
47
48## Implementation Details
49
50See `## Workflow` above for related details.
51
52- Execution model: validate the request, choose the packaged workflow, and produce a bounded deliverable.
53- Input controls: confirm the source files, scope limits, output format, and acceptance criteria before running any script.
54- Primary implementation surface: `scripts/extract_pdf.py`.
55- Reference guidance: `references/` contains supporting rules, prompts, or checklists.
56- Parameters to clarify first: input path, output path, scope filters, thresholds, and any domain-specific constraints.
57- Output discipline: keep results reproducible, identify assumptions explicitly, and avoid undocumented side effects.
58
59## Workflow
60
61### Step 1: Check PubMed (Deterministic)
62
63If the user provides a PMID, use the `baseline_extractor.py` script to check PubMed.
64
65```python
66import subprocess
67import json
68
69# Replace <PMID> with actual PMID
70result = subprocess.run(["python", "scripts/baseline_extractor.py", "<PMID>"], capture_output=True, text=True)
71print(result.stdout)
72```
73
74**Analyze the Script Output:**
75* If `status` is `"success"`: **Stop here.** Return the `data` JSON to the user.
76* If `status` is `"not_found"`, `"incomplete"`, or `"error"` (or if no PMID was provided): Proceed to Step 2.
77
78### Step 2: LLM Extraction (Fallback)
79
80If Step 1 did not yield a complete result, use the LLM to extract the information from the **full article text**.
81
82**Input:**
83* Full article text provided by the user.
84
85**Instructions:**
861. Read the [Extraction Schema](references/extraction_schema.md) carefully.
872. Analyze the text to identify all 10 required fields.
883. Ensure the output is strictly in the JSON format defined in the schema.
894. **Constraint**: Do not hallucinate. If a field is not mentioned in the text, set it to `null` or an empty string.
90
91## Output
92
93Return the final result as a Markdown code block containing the JSON object.
94
95```json
96{
97 "study": "...",
98 "region": "...",
99 ...
100}
101```
102
103## Helper Scripts
104
105### PDF Text Extraction
106
107When the user provides a PDF file path, use `extract_pdf.py` to extract the text content before assessment: