Source: https://github.com/aipoch/medical-research-skills
Bibliography
When to Use
- You are conducting a literature review and need consistent summaries plus structured metadata (theme/method/conclusion) across many papers.
- You have a mixed-format reading folder (
.pdf, .md, .docx, .txt) and want a single CSV for downstream analysis (e.g., Excel, R, Python).
- You need to organize annotations by keywords (theme), experimental methods (method), and key conclusions (conclusion).
- You want a two-step pipeline: first generate a human-readable summary Markdown, then generate a machine-friendly CSV from that Markdown.
- You need robust handling of PDFs by converting them to Markdown first (via
pdf-extract) and then using only Markdown content for extraction.
Key Features
- Batch scans an input directory for
.pdf, .md, .docx, and .txt literature files.
- Converts PDFs to Markdown via
pdf-extract, then ignores non-Markdown artifacts (e.g., image folders).
- Extracts and normalizes, per document:
- Title
- Summary (prefer original abstract)
- Keywords (theme)
- Experimental Methods (method names only)
- Key Conclusions (single sentence)
- Commentary (one-sentence, tactful evaluation)
- Produces exactly two outputs:
- A consolidated Summary Markdown saved under
outputs/
- A single CSV generated from that Summary Markdown
- Enforces UTF-8 output to prevent garbled characters; fills missing fields with
"Not recognized" instead of leaving blanks.
- Uses the CSV field order and headers defined in
assets/bibliography_template.csv.
Dependencies
pdf-extract (version: not specified; required when PDFs are present)
- Input formats supported (no external version constraints specified):
- PDF
- Markdown (
.md)
- DOCX (
.docx)
- Plain text (
.txt)
Example Usage
Goal
Read all literature files in a folder, generate a consolidated summary Markdown, then generate a CSV following assets/bibliography_template.csv.
Inputs
- Input directory (example):
./inputs/literature/
- Output directory:
./outputs/
- Output CSV path (example):
./outputs/bibliography.csv
Expected Outputs (exactly two files)
./outputs/bibliography_summary.md
./outputs/bibliography.csv
Example Summary Markdown Structure (generated first)
# Bibliography Summary
## Document 1
- Title: <Title>
- Summary: <Prefer the original Abstract; if missing, use the closest equivalent section>
- Keywords: keyword1 | keyword2 | keyword3
- Experimental Methods: <method1; method2; ... (names only)>
- Key Conclusions: <one sentence covering all main points>
- Commentary: <one tactful sentence>
## Document 2
...
Example CSV (generated from the Summary Markdown)
The CSV must follow the header order defined in:
assets/bibliography_template.csv
Rules:
- One row per document.
- No empty cells; use
Not recognized when extraction fails.
- Save as UTF-8.
Implementation Details
1) Input Reading and Normalization
- Traverse the input directory and process files with extensions:
PDF handling
- If PDFs exist, convert them to Markdown using
pdf-extract.
- Use only the generated
.md content; ignore image directories or other byproducts.
- Locate
pdf-extract as follows:
- First, look for a sibling skill directory containing
SKILL.md at the same level as this skill’s parent directory.
- If not found, ask the user to confirm the actual
pdf-extract path.
DOCX handling
- Extract body text while preserving title/paragraph order as much as possible.
MD/TXT handling
- Read text directly.
- If garbled characters appear or key fields cannot be recognized, attempt to detect and read using the original encoding (commonly
GB18030 / GBK) before extraction.
2) Generate the Summary Markdown First (Single Source of Truth)
Before producing the CSV, generate a consolidated Summary Markdown containing, for each document:
- Title
- Summary
- Prefer the original Abstract.
- If no “Abstract” exists, use the closest equivalent section (e.g., “Summary”, “Highlights”, or an “Objective–Method–Result–Conclusion” style segment).
- Keywords
- Experimental Methods
- Key Conclusions
- Commentary
- Exactly one sentence.
- Avoid harsh criticism; if the work has low value, use tactful phrasing.
This Summary Markdown must be saved with UTF-8 encoding and stored under outputs/. The CSV must be generated only from this Markdown (not directly from raw files).
3) Field Extraction Rules (Theme / Method / Conclusion)
- Keywords (theme)
- Prefer the original keywords from the document.
- Separate multiple keywords with
|.
- If no keywords are found, generate 3–5 keyword phrases based on the abstract and append:
(generated based on abstract)
- Experimental Methods (method)
- Output method names only (no long descriptions).
- Key Conclusions (conclusion)
- One sentence that covers all main points.
4) CSV Output Constraints
- Output exactly one CSV file at the end.
- CSV field order and headers must match
assets/bibliography_template.csv.
- Encoding must be UTF-8 to avoid garbled characters.
- If any field cannot be extracted, write
Not recognized (never leave empty).
- Only two files may be generated in total:
- Summary Markdown
- CSV
- No temporary/intermediate/auxiliary files may be left behind (including extracted text dumps, caches, logs, images, backups). If conversion/extraction requires intermediate artifacts, keep them in memory or ensure all non-target files are deleted before final output.
- Do not use PowerShell to directly write/manipulate CSV/Markdown to avoid encoding/newline issues; always generate and save using UTF-8.
Reference
- Detailed rules and field descriptions:
references/guide.md
1---2name: bibliography3description: Classifies and organizes literature by theme, method, and conclusion; use when you need to batch-read a folder of PDF/MD/DOCX/TXT files and output a structured CSV for literature reviews and annotation management.4license: MIT5---6> **Source**: [https://github.com/aipoch/medical-research-skills](https://github.com/aipoch/medical-research-skills)
7
8# Bibliography
9
10## When to Use
11
12- You are conducting a literature review and need consistent summaries plus structured metadata (theme/method/conclusion) across many papers.
13- You have a mixed-format reading folder (`.pdf`, `.md`, `.docx`, `.txt`) and want a single CSV for downstream analysis (e.g., Excel, R, Python).
14- You need to organize annotations by **keywords (theme)**, **experimental methods (method)**, and **key conclusions (conclusion)**.
15- You want a two-step pipeline: first generate a human-readable summary Markdown, then generate a machine-friendly CSV from that Markdown.
16- You need robust handling of PDFs by converting them to Markdown first (via `pdf-extract`) and then using **only Markdown content** for extraction.
17
18## Key Features
19
20- Batch scans an input directory for `.pdf`, `.md`, `.docx`, and `.txt` literature files.
21- Converts PDFs to Markdown via `pdf-extract`, then **ignores non-Markdown artifacts** (e.g., image folders).
22- Extracts and normalizes, per document:
23 - Title
24 - Summary (prefer original abstract)
25 - Keywords (theme)
26 - Experimental Methods (method names only)
27 - Key Conclusions (single sentence)
28 - Commentary (one-sentence, tactful evaluation)
29- Produces exactly **two outputs**:
30 1. A consolidated **Summary Markdown** saved under `outputs/`
31 2. A single **CSV** generated from that Summary Markdown
32- Enforces UTF-8 output to prevent garbled characters; fills missing fields with `"Not recognized"` instead of leaving blanks.
33- Uses the CSV field order and headers defined in `assets/bibliography_template.csv`.
34
35## Dependencies
36
37- `pdf-extract` (version: not specified; required when PDFs are present)
38- Input formats supported (no external version constraints specified):
39 - PDF
40 - Markdown (`.md`)
41 - DOCX (`.docx`)
42 - Plain text (`.txt`)
43
44## Example Usage
45
46### Goal
47Read all literature files in a folder, generate a consolidated summary Markdown, then generate a CSV following `assets/bibliography_template.csv`.
48
49### Inputs
50- Input directory (example): `./inputs/literature/`
51- Output directory: `./outputs/`
52- Output CSV path (example): `./outputs/bibliography.csv`
53
54### Expected Outputs (exactly two files)
55- `./outputs/bibliography_summary.md`
56- `./outputs/bibliography.csv`
57
58### Example Summary Markdown Structure (generated first)
59
60```md
61# Bibliography Summary
62
63## Document 1
64- Title: <Title>
65- Summary: <Prefer the original Abstract; if missing, use the closest equivalent section>
66- Keywords: keyword1 | keyword2 | keyword3
67- Experimental Methods: <method1; method2; ... (names only)>
68- Key Conclusions: <one sentence covering all main points>
69- Commentary: <one tactful sentence>
70
71## Document 2
72...
73```
74
75### Example CSV (generated from the Summary Markdown)
76
77The CSV must follow the header order defined in:
78
79- `assets/bibliography_template.csv`
80
81Rules:
82- One row per document.
83- No empty cells; use `Not recognized` when extraction fails.
84- Save as UTF-8.
85
86## Implementation Details
87
88### 1) Input Reading and Normalization
89
90- Traverse the input directory and process files with extensions:
91 - `.pdf`, `.md`, `.docx`, `.txt`
92
93**PDF handling**
94- If PDFs exist, convert them to Markdown using `pdf-extract`.
95- Use only the generated `.md` content; ignore image directories or other byproducts.
96- Locate `pdf-extract` as follows:
97 1. First, look for a sibling skill directory containing `SKILL.md` at the same level as this skill’s parent directory.
98 2. If not found, ask the user to confirm the actual `pdf-extract` path.
99
100**DOCX handling**
101- Extract body text while preserving title/paragraph order as much as possible.
102
103**MD/TXT handling**
104- Read text directly.
105- If garbled characters appear or key fields cannot be recognized, attempt to detect and read using the original encoding (commonly `GB18030` / `GBK`) before extraction.
106
107### 2) Generate the Summary Markdown First (Single Source of Truth)
108
109Before producing the CSV, generate a consolidated Summary Markdown containing, for each document:
110
111- **Title**
112- **Summary**
113 - Prefer the original **Abstract**.
114 - If no “Abstract” exists, use the closest equivalent section (e.g., “Summary”, “Highlights”, or an “Objective–Method–Result–Conclusion” style segment).
115- **Keywords**
116- **Experimental Methods**
117- **Key Conclusions**
118- **Commentary**
119 - Exactly one sentence.
120 - Avoid harsh criticism; if the work has low value, use tactful phrasing.
121
122This Summary Markdown must be saved with **UTF-8** encoding and stored under `outputs/`. The CSV must be generated **only** from this Markdown (not directly from raw files).
123
124### 3) Field Extraction Rules (Theme / Method / Conclusion)
125
126- **Keywords (theme)**
127 - Prefer the original keywords from the document.
128 - Separate multiple keywords with `|`.
129 - If no keywords are found, generate **3–5** keyword phrases based on the abstract and append:
130 - `(generated based on abstract)`
131- **Experimental Methods (method)**
132 - Output method names only (no long descriptions).
133- **Key Conclusions (conclusion)**
134 - One sentence that covers all main points.
135
136### 4) CSV Output Constraints
137
138- Output exactly **one** CSV file at the end.
139- CSV field order and headers must match `assets/bibliography_template.csv`.
140- Encoding must be **UTF-8** to avoid garbled characters.
141- If any field cannot be extracted, write `Not recognized` (never leave empty).
142- Only two files may be generated in total:
143 1. Summary Markdown
144 2. CSV
145- No temporary/intermediate/auxiliary files may be left behind (including extracted text dumps, caches, logs, images, backups). If conversion/extraction requires intermediate artifacts, keep them in memory or ensure all non-target files are deleted before final output.
146- Do not use PowerShell to directly write/manipulate CSV/Markdown to avoid encoding/newline issues; always generate and save using UTF-8.
147
148### Reference
149
150- Detailed rules and field descriptions: `references/guide.md`