Source: https://github.com/aipoch/medical-research-skills
Clinical Outcome Extraction
Extract structured outcome data from clinical research papers for meta-analysis.
When to Use
- Use this skill when you need clinical research outcome extraction for meta-analysis. use when users need to extract outcome measures (binary, continuous, or survival data) from clinical research papers for systematic review and meta-analysis. handles both database lookup by pmid and real-time llm extraction in a reproducible workflow.
- Use this skill when a data analytics task needs a packaged method instead of ad-hoc freeform output.
- Use this skill when the user expects a concrete deliverable, validation step, or file-based result.
- Use this skill when
scripts/extract_pdf.py is the most direct path to complete the request.
- Use this skill when you need the
outcome-extraction for clinical trials package behavior rather than a generic answer.
Key Features
- Scope-focused workflow aligned to: Clinical research outcome extraction for meta-analysis. Use when users need to extract outcome measures (binary, continuous, or survival data) from clinical research papers for systematic review and meta-analysis. Handles both database lookup by PMID and real-time LLM extraction.
- Packaged executable path(s):
scripts/extract_pdf.py.
- Reference material available in
references/ for task-specific guidance.
- Structured execution path designed to keep outputs consistent and reviewable.
Dependencies
Python: 3.10+. Repository baseline for current packaged skills.
Third-party packages: not explicitly version-pinned in this skill package. Add pinned versions if this skill needs stricter environment control.
Example Usage
cd "20260316/scientific-skills/Data Analytics/outcome-extraction-for-clinical-trials"
python -m py_compile scripts/extract_pdf.py
python scripts/extract_pdf.py --help
Example run plan:
- Confirm the user input, output path, and any required config values.
- Edit the in-file
CONFIG block or documented parameters if the script uses fixed settings.
- Run
python scripts/extract_pdf.py with the validated inputs.
- Review the generated output and return the final artifact with any assumptions called out.
Implementation Details
See ## Workflow above for related details.
- Execution model: validate the request, choose the packaged workflow, and produce a bounded deliverable.
- Input controls: confirm the source files, scope limits, output format, and acceptance criteria before running any script.
- Primary implementation surface:
scripts/extract_pdf.py.
- Reference guidance:
references/ contains supporting rules, prompts, or checklists.
- Parameters to clarify first: input path, output path, scope filters, thresholds, and any domain-specific constraints.
- Output discipline: keep results reproducible, identify assumptions explicitly, and avoid undocumented side effects.
Workflow
Input Processing
- User provides: full paper text + optional PMID
- If PMID provided: query database first for existing results
- If no PMID or no database match: proceed to LLM extraction
Outcome Identification (LLM)
- Extract all outcome measures from the paper
- Determine outcome types: binary, continuous, or survival
- Identify measurement time points
- Output JSON format with outcome classification
Data Classification (Code)
- Separate outcomes into three categories:
bi_outcomes: Binary/dichotomous outcomes
con_outcomes: Continuous outcomes
sur_outcomes: Survival outcomes
Data Extraction by Type
Binary Outcomes
Extract for each intervention group:
- Sample size (n)
- Number of events (event)
Continuous Outcomes
Extract for each intervention group:
- Sample size (n)
- Mean (mean)
- Standard deviation (sd)
Survival Outcomes
Extract for each intervention group:
- Sample size (n)
- Hazard ratio (HR)
- 95% Lower CI
- 95% Upper CI
- Output Formatting
- Combine all extracted data
- Ensure consistent JSON structure
- Convert values to strings
Output Format
[
{
"outcome_name": "PFS",
"detection_time_point": "12 months",
"groups": [
{
"group_name": "Treatment A",
"sample_size": "100",
"outcome_type": "Binary|Continuous|Survival",
"data": [
{"value_type": "Events|Mean|SD|HR|95%Lower CI|95%Upper CI", "value": "25"}
]
}
]
}
]
‼️‼️‼️See references (extraction-promots.md) for detailed JSON structures for each outcome type (binary, continuous, survival)‼️‼️‼️
Requirements
- Extract from full text, not just abstract
- Consider ALL intervention groups in the paper
- Include ALL outcome measures of interest
- Report all data regardless of statistical significance
- Use specific group names (intervention names in English), not generic terms like "treatment group"
- Output in JSON format
- Output language: English for all field values
- If data not found: output blank space ""
1---2name: outcome-extraction-for-clinical-trials3description: Clinical research outcome extraction for meta-analysis. Use when users need to extract outcome measures (binary, continuous, or survival data) from clinical research papers for systematic review and meta-analysis. Handles both database lookup by PMID and real-time LLM extraction.4license: MIT5---6> **Source**: [https://github.com/aipoch/medical-research-skills](https://github.com/aipoch/medical-research-skills)
7
8# Clinical Outcome Extraction
9
10Extract structured outcome data from clinical research papers for meta-analysis.
11
12## When to Use
13
14- Use this skill when you need clinical research outcome extraction for meta-analysis. use when users need to extract outcome measures (binary, continuous, or survival data) from clinical research papers for systematic review and meta-analysis. handles both database lookup by pmid and real-time llm extraction in a reproducible workflow.
15- Use this skill when a data analytics task needs a packaged method instead of ad-hoc freeform output.
16- Use this skill when the user expects a concrete deliverable, validation step, or file-based result.
17- Use this skill when `scripts/extract_pdf.py` is the most direct path to complete the request.
18- Use this skill when you need the `outcome-extraction for clinical trials` package behavior rather than a generic answer.
19
20## Key Features
21
22- Scope-focused workflow aligned to: Clinical research outcome extraction for meta-analysis. Use when users need to extract outcome measures (binary, continuous, or survival data) from clinical research papers for systematic review and meta-analysis. Handles both database lookup by PMID and real-time LLM extraction.
23- Packaged executable path(s): `scripts/extract_pdf.py`.
24- Reference material available in `references/` for task-specific guidance.
25- Structured execution path designed to keep outputs consistent and reviewable.
26
27## Dependencies
28
29- `Python`: `3.10+`. Repository baseline for current packaged skills.
30- `Third-party packages`: `not explicitly version-pinned in this skill package`. Add pinned versions if this skill needs stricter environment control.
31
32## Example Usage
33
34```bash
35cd "20260316/scientific-skills/Data Analytics/outcome-extraction-for-clinical-trials"
36python -m py_compile scripts/extract_pdf.py
37python scripts/extract_pdf.py --help
38```
39
40Example run plan:
411. Confirm the user input, output path, and any required config values.
422. Edit the in-file `CONFIG` block or documented parameters if the script uses fixed settings.
433. Run `python scripts/extract_pdf.py` with the validated inputs.
444. Review the generated output and return the final artifact with any assumptions called out.
45
46## Implementation Details
47
48See `## Workflow` above for related details.
49
50- Execution model: validate the request, choose the packaged workflow, and produce a bounded deliverable.
51- Input controls: confirm the source files, scope limits, output format, and acceptance criteria before running any script.
52- Primary implementation surface: `scripts/extract_pdf.py`.
53- Reference guidance: `references/` contains supporting rules, prompts, or checklists.
54- Parameters to clarify first: input path, output path, scope filters, thresholds, and any domain-specific constraints.
55- Output discipline: keep results reproducible, identify assumptions explicitly, and avoid undocumented side effects.
56
57## Workflow
58
591. **Input Processing**
60 - User provides: full paper text + optional PMID
61 - If PMID provided: query database first for existing results
62 - If no PMID or no database match: proceed to LLM extraction
63
642. **Outcome Identification** (LLM)
65 - Extract all outcome measures from the paper
66 - Determine outcome types: binary, continuous, or survival
67 - Identify measurement time points
68 - Output JSON format with outcome classification
69
703. **Data Classification** (Code)
71 - Separate outcomes into three categories:
72 - `bi_outcomes`: Binary/dichotomous outcomes
73 - `con_outcomes`: Continuous outcomes
74 - `sur_outcomes`: Survival outcomes
75
764. **Data Extraction by Type**
77
78### Binary Outcomes
79Extract for each intervention group:
80- Sample size (n)
81- Number of events (event)
82
83### Continuous Outcomes
84Extract for each intervention group:
85- Sample size (n)
86- Mean (mean)
87- Standard deviation (sd)
88
89### Survival Outcomes
90Extract for each intervention group:
91- Sample size (n)
92- Hazard ratio (HR)
93- 95% Lower CI
94- 95% Upper CI
95
965. **Output Formatting**
97 - Combine all extracted data
98 - Ensure consistent JSON structure
99 - Convert values to strings
100
101## Output Format
102
103```json
104[
105 {
106 "outcome_name": "PFS",
107 "detection_time_point": "12 months",
108 "groups": [
109 {
110 "group_name": "Treatment A",
111 "sample_size": "100",
112 "outcome_type": "Binary|Continuous|Survival",
113 "data": [
114 {"value_type": "Events|Mean|SD|HR|95%Lower CI|95%Upper CI", "value": "25"}
115 ]
116 }
117 ]
118 }
119]
120```
121
122## ‼️‼️‼️See references (extraction-promots.md) for detailed JSON structures for each outcome type (binary, continuous, survival)‼️‼️‼️
123
124## Requirements
125
126- Extract from full text, not just abstract
127- Consider ALL intervention groups in the paper
128- Include ALL outcome measures of interest
129- Report all data regardless of statistical significance
130- Use specific group names (intervention names in English), not generic terms like "treatment group"
131- Output in JSON format
132- Output language: English for all field values
133- If data not found: output blank space ""