Source: https://github.com/aipoch/medical-research-skills
Paper Reading Tweet Generator
This skill analyzes an academic paper (PDF, Word, or Text) and generates a structured reading tweet including basic info, background, results, and conclusion. It can highlight specific product/drug advantages and ensures standardized terminology.
When to Use
- Use this skill when the request matches its documented task boundary.
- Use it when the user can provide the required inputs and expects a structured deliverable.
- Prefer this skill for repeatable, checklist-driven execution rather than open-ended brainstorming.
Key Features
- Scope-focused workflow aligned to: Generates a structured reading tweet from an academic paper (PDF, Word, or Text), highlighting specific product advantages. Use when the user wants to turn a document into a social media post or reading summary.
- Packaged executable path(s):
scripts/extract_pdf.py plus 1 additional script(s).
- Reference material available in
references/ for task-specific guidance.
- Structured execution path designed to keep outputs consistent and reviewable.
Dependencies
Python: 3.10+. Repository baseline for current packaged skills.
Third-party packages: not explicitly version-pinned in this skill package. Add pinned versions if this skill needs stricter environment control.
Example Usage
cd "20260316/scientific-skills/Others/paper-tweet-generator"
python -m py_compile scripts/extract_pdf.py
python scripts/extract_pdf.py --help
Example run plan:
- Confirm the user input, output path, and any required config values.
- Edit the in-file
CONFIG block or documented parameters if the script uses fixed settings.
- Run
python scripts/extract_pdf.py with the validated inputs.
- Review the generated output and return the final artifact with any assumptions called out.
Implementation Details
See ## Workflow above for related details.
- Execution model: validate the request, choose the packaged workflow, and produce a bounded deliverable.
- Input controls: confirm the source files, scope limits, output format, and acceptance criteria before running any script.
- Primary implementation surface:
scripts/extract_pdf.py with additional helper scripts under scripts/.
- Reference guidance:
references/ contains supporting rules, prompts, or checklists.
- Parameters to clarify first: input path, output path, scope filters, thresholds, and any domain-specific constraints.
- Output discipline: keep results reproducible, identify assumptions explicitly, and avoid undocumented side effects.
Workflow
To generate a tweet, follow these steps sequentially:
1. Locate and Extract Content
First, locate the file and extract its text content.
- Locate File: If the user provides a file path, use it. If not (e.g., "uploaded file"), use
Glob to search for .pdf, .docx, or .txt files in the entire workspace (pattern: **/*.pdf). Select the most relevant file (e.g., recently added).
- Extract Text:
- Recommend using an output file to avoid console buffer limits.
- Run:
python scripts/extract_text.py <file_path> extracted_content.txt
- Read the content:
Read extracted_content.txt
- Handle Output:
- If the extraction fails or returns empty text (check stderr logs), inform the user.
- If "Warning: No text extracted" is logged, the PDF is likely a scanned image.
- Fallback: If the script fails, try reading the file directly with built-in tools (only for text files).
2. Generate Tweet Sections
Use the extracted text to generate the following sections using the prompts in references/prompt_templates.md.
Note: If the extracted text is very long (> 50k chars), focus on the Abstract, Introduction, Results, and Conclusion sections.
- Basic Info: Extract title, authors, journal, DOI.
- Background: Summarize the research background (< 500 words).
- Results: Summarize key findings highlighting the product (< 800 words).
- Conclusion: Summarize the main conclusion.
3. Final Assembly
- Title: Generate a catchy title based on the extracted info.
- Assembly: Assemble the final tweet in Markdown including all sections.
Requirements
- Python environment with
pypdf and python-docx installed.
- Access to an LLM for content extraction.
Scripts
scripts/extract_text.py: Extracts raw text from PDF, Word, or Text files. Supports output to file for large documents.
References
references/prompt_templates.md: Prompts for extracting and summarizing each section.
When Not to Use
- Do not use this skill when the required source data, identifiers, files, or credentials are missing.
- Do not use this skill when the user asks for fabricated results, unsupported claims, or out-of-scope conclusions.
- Do not use this skill when a simpler direct answer is more appropriate than the documented workflow.
Required Inputs
- A clearly specified task goal aligned with the documented scope.
- All required files, identifiers, parameters, or environment variables before execution.
- Any domain constraints, formatting requirements, and expected output destination if applicable.
Output Contract
- Return a structured deliverable that is directly usable without reformatting.
- If a file is produced, prefer a deterministic output name such as
paper_tweet_generator_result.md unless the skill documentation defines a better convention.
- Include a short validation summary describing what was checked, what assumptions were made, and any remaining limitations.
Validation and Safety Rules
- Validate required inputs before execution and stop early when mandatory fields or files are missing.
- Do not fabricate measurements, references, findings, or conclusions that are not supported by the provided source material.
- Emit a clear warning when credentials, privacy constraints, safety boundaries, or unsupported requests affect the result.
- Keep the output safe, reproducible, and within the documented scope at all times.
Failure Handling
- If validation fails, explain the exact missing field, file, or parameter and show the minimum fix required.
- If an external dependency or script fails, surface the command path, likely cause, and the next recovery step.
- If partial output is returned, label it clearly and identify which checks could not be completed.
Quick Validation
Run this minimal verification path before full execution when possible:
python scripts/extract_pdf.py --help
Expected output format:
Result file: paper_tweet_generator_result.md
Validation summary: PASS/FAIL with brief notes
Assumptions: explicit list if any
1---2name: paper-tweet-generator3description: Generates a structured reading tweet from an academic paper (PDF, Word, or Text), highlighting specific product advantages. Use when the user wants to turn a document into a social media post or reading summary.4license: MIT5---6> **Source**: [https://github.com/aipoch/medical-research-skills](https://github.com/aipoch/medical-research-skills)
7
8# Paper Reading Tweet Generator
9
10This skill analyzes an academic paper (PDF, Word, or Text) and generates a structured reading tweet including basic info, background, results, and conclusion. It can highlight specific product/drug advantages and ensures standardized terminology.
11
12## When to Use
13
14- Use this skill when the request matches its documented task boundary.
15- Use it when the user can provide the required inputs and expects a structured deliverable.
16- Prefer this skill for repeatable, checklist-driven execution rather than open-ended brainstorming.
17
18## Key Features
19
20- Scope-focused workflow aligned to: Generates a structured reading tweet from an academic paper (PDF, Word, or Text), highlighting specific product advantages. Use when the user wants to turn a document into a social media post or reading summary.
21- Packaged executable path(s): `scripts/extract_pdf.py` plus 1 additional script(s).
22- Reference material available in `references/` for task-specific guidance.
23- Structured execution path designed to keep outputs consistent and reviewable.
24
25## Dependencies
26
27- `Python`: `3.10+`. Repository baseline for current packaged skills.
28- `Third-party packages`: `not explicitly version-pinned in this skill package`. Add pinned versions if this skill needs stricter environment control.
29
30## Example Usage
31
32```bash
33cd "20260316/scientific-skills/Others/paper-tweet-generator"
34python -m py_compile scripts/extract_pdf.py
35python scripts/extract_pdf.py --help
36```
37
38Example run plan:
391. Confirm the user input, output path, and any required config values.
402. Edit the in-file `CONFIG` block or documented parameters if the script uses fixed settings.
413. Run `python scripts/extract_pdf.py` with the validated inputs.
424. Review the generated output and return the final artifact with any assumptions called out.
43
44## Implementation Details
45
46See `## Workflow` above for related details.
47
48- Execution model: validate the request, choose the packaged workflow, and produce a bounded deliverable.
49- Input controls: confirm the source files, scope limits, output format, and acceptance criteria before running any script.
50- Primary implementation surface: `scripts/extract_pdf.py` with additional helper scripts under `scripts/`.
51- Reference guidance: `references/` contains supporting rules, prompts, or checklists.
52- Parameters to clarify first: input path, output path, scope filters, thresholds, and any domain-specific constraints.
53- Output discipline: keep results reproducible, identify assumptions explicitly, and avoid undocumented side effects.
54
55## Workflow
56
57To generate a tweet, follow these steps sequentially:
58
59### 1. Locate and Extract Content
60First, locate the file and extract its text content.
61- **Locate File**: If the user provides a file path, use it. If not (e.g., "uploaded file"), use `Glob` to search for `.pdf`, `.docx`, or `.txt` files in the **entire workspace** (pattern: `**/*.pdf`). Select the most relevant file (e.g., recently added).
62- **Extract Text**:
63 - Recommend using an output file to avoid console buffer limits.
64 - Run: `python scripts/extract_text.py <file_path> extracted_content.txt`
65 - Read the content: `Read extracted_content.txt`
66- **Handle Output**:
67 - If the extraction fails or returns empty text (check stderr logs), inform the user.
68 - If "Warning: No text extracted" is logged, the PDF is likely a scanned image.
69- **Fallback**: If the script fails, try reading the file directly with built-in tools (only for text files).
70
71### 2. Generate Tweet Sections
72Use the extracted text to generate the following sections using the prompts in `references/prompt_templates.md`.
73*Note: If the extracted text is very long (> 50k chars), focus on the Abstract, Introduction, Results, and Conclusion sections.*
74
75- **Basic Info**: Extract title, authors, journal, DOI.
76- **Background**: Summarize the research background (< 500 words).
77- **Results**: Summarize key findings highlighting the product (< 800 words).
78- **Conclusion**: Summarize the main conclusion.
79
80### 3. Final Assembly
81- **Title**: Generate a catchy title based on the extracted info.
82- **Assembly**: Assemble the final tweet in Markdown including all sections.
83
84## Requirements
85
86* Python environment with `pypdf` and `python-docx` installed.
87* Access to an LLM for content extraction.
88
89## Scripts
90
91* `scripts/extract_text.py`: Extracts raw text from PDF, Word, or Text files. Supports output to file for large documents.
92
93## References
94
95* `references/prompt_templates.md`: Prompts for extracting and summarizing each section.
96
97## When Not to Use
98
99- Do not use this skill when the required source data, identifiers, files, or credentials are missing.
100- Do not use this skill when the user asks for fabricated results, unsupported claims, or out-of-scope conclusions.
101- Do not use this skill when a simpler direct answer is more appropriate than the documented workflow.
102
103## Required Inputs
104
105- A clearly specified task goal aligned with the documented scope.
106- All required files, identifiers, parameters, or environment variables before execution.
107- Any domain constraints, formatting requirements, and expected output destination if applicable.
108
109## Output Contract
110
111- Return a structured deliverable that is directly usable without reformatting.
112- If a file is produced, prefer a deterministic output name such as `paper_tweet_generator_result.md` unless the skill documentation defines a better convention.
113- Include a short validation summary describing what was checked, what assumptions were made, and any remaining limitations.
114
115## Validation and Safety Rules
116
117- Validate required inputs before execution and stop early when mandatory fields or files are missing.
118- Do not fabricate measurements, references, findings, or conclusions that are not supported by the provided source material.
119- Emit a clear warning when credentials, privacy constraints, safety boundaries, or unsupported requests affect the result.
120- Keep the output safe, reproducible, and within the documented scope at all times.
121
122## Failure Handling
123
124- If validation fails, explain the exact missing field, file, or parameter and show the minimum fix required.
125- If an external dependency or script fails, surface the command path, likely cause, and the next recovery step.
126- If partial output is returned, label it clearly and identify which checks could not be completed.
127
128## Quick Validation
129
130Run this minimal verification path before full execution when possible:
131
132```bash
133python scripts/extract_pdf.py --help
134```
135
136Expected output format:
137
138```text
139Result file: paper_tweet_generator_result.md
140Validation summary: PASS/FAIL with brief notes
141Assumptions: explicit list if any
142```