Literature Review
Overview
Conduct systematic, comprehensive literature reviews following rigorous academic methodology. Search multiple literature databases, synthesize findings thematically, verify all citations for accuracy, and generate professional output documents in markdown and PDF formats.
This skill uses the parallel-web skill (parallel-cli search) as the primary web search tool for broad academic literature discovery, supplemented by specialized database access skills (gget, bioservices, datacommons-client). It provides specialized tools for citation verification, result aggregation, and document generation.
When to Use This Skill
Use this skill when:
- Conducting a systematic literature review for research or publication
- Synthesizing current knowledge on a specific topic across multiple sources
- Performing meta-analysis or scoping reviews
- Writing the literature review section of a research paper or thesis
- Investigating the state of the art in a research domain
- Identifying research gaps and future directions
- Requiring verified citations and professional formatting
Visual Enhancement with Scientific Schematics
⚠️ MANDATORY: Every literature review MUST include at least 1-2 AI-generated figures using the scientific-schematics skill.
This is not optional. Literature reviews without visual elements are incomplete. Before finalizing any document:
- Generate at minimum ONE schematic or diagram (e.g., PRISMA flow diagram for systematic reviews)
- Prefer 2-3 figures for comprehensive reviews (search strategy flowchart, thematic synthesis diagram, conceptual framework)
How to generate figures:
- Use the scientific-schematics skill to generate AI-powered publication-quality diagrams
- Simply describe your desired diagram in natural language
- Nano Banana Pro will automatically generate, review, and refine the schematic
How to generate schematics:
python scripts/generate_schematic.py "your diagram description" -o figures/output.png
The AI will automatically:
- Create publication-quality images with proper formatting
- Review and refine through multiple iterations
- Ensure accessibility (colorblind-friendly, high contrast)
- Save outputs in the figures/ directory
When to add schematics:
- PRISMA flow diagrams for systematic reviews
- Literature search strategy flowcharts
- Thematic synthesis diagrams
- Research gap visualization maps
- Citation network diagrams
- Conceptual framework illustrations
- Any complex concept that benefits from visualization
For detailed guidance on creating schematics, refer to the scientific-schematics skill documentation.
Core Workflow
Literature reviews follow a structured, multi-phase workflow:
Phase 1: Planning and Scoping
Define Research Question: Use PICO framework (Population, Intervention, Comparison, Outcome) for clinical/biomedical reviews
- Example: "What is the efficacy of CRISPR-Cas9 (I) for treating sickle cell disease (P) compared to standard care (C)?"
Establish Scope and Objectives:
- Define clear, specific research questions
- Determine review type (narrative, systematic, scoping, meta-analysis)
- Set boundaries (time period, geographic scope, study types)
Develop Search Strategy:
- Identify 2-4 main concepts from research question
- List synonyms, abbreviations, and related terms for each concept
- Plan Boolean operators (AND, OR, NOT) to combine terms
- Select minimum 3 complementary databases
- Use the parallel-web skill (
parallel-cli search) for initial scoping to quickly gauge the landscape before formal database searches
Set Inclusion/Exclusion Criteria:
- Date range (e.g., last 10 years: 2015-2024)
- Language (typically English, or specify multilingual)
- Publication types (peer-reviewed, preprints, reviews)
- Study designs (RCTs, observational, in vitro, etc.)
- Document all criteria clearly
Phase 2: Systematic Literature Search
Multi-Database Search:
Select databases appropriate for the domain. Always start with parallel-web for broad academic coverage, then supplement with domain-specific databases.
Web-Based Academic Search (parallel-web skill — START HERE):
- Use
parallel-cli search with academic domain filtering for broad scholarly coverage
- Run two searches: academic-focused + general to catch all relevant sources
# Academic-focused search across scholarly sources
parallel-cli search "your research topic" -q "keyword1" -q "keyword2" \
--json --max-results 10 --excerpt-max-chars-total 27000 \
--include-domains "scholar.google.com,arxiv.org,pubmed.ncbi.nlm.nih.gov,semanticscholar.org,biorxiv.org,medrxiv.org,ncbi.nlm.nih.gov,nature.com,science.org,ieee.org,acm.org,springer.com,wiley.com,cell.com,pnas.org,nih.gov" \
-o sources/litreview_<topic>-academic.json
# General search for supplementary sources
parallel-cli search "your research topic" -q "keyword1" -q "keyword2" \
--json --max-results 10 --excerpt-max-chars-total 27000 \
-o sources/litreview_<topic>-general.json
- Use
parallel-cli extract to fetch full content from specific paper URLs or PDFs found in search results
parallel-cli extract "https://arxiv.org/abs/XXXX.XXXXX" --json
Biomedical & Life Sciences:
- Use
gget skill: gget search pubmed "search terms" for PubMed/PMC
- Use
gget skill: gget search biorxiv "search terms" for preprints
- Use
bioservices skill for ChEMBL, KEGG, UniProt, etc.
General Scientific Literature:
- Search arXiv via direct API (preprints in physics, math, CS, q-bio)
- Search Semantic Scholar via API (200M+ papers, cross-disciplinary)
- Use Google Scholar for comprehensive coverage (manual or careful scraping)
Specialized Databases:
- Use
gget alphafold for protein structures
- Use
gget cosmic for cancer genomics
- Use
datacommons-client for demographic/statistical data
- Use specialized databases as appropriate for the domain
Document Search Parameters:
## Search Strategy
### Database: PubMed
- **Date searched**: 2024-10-25
- **Date range**: 2015-01-01 to 2024-10-25
- **Search string**:
("CRISPR"[Title] OR "Cas9"[Title])
AND ("sickle cell"[MeSH] OR "SCD"[Title/Abstract])
AND 2015:2024[Publication Date]
- **Results**: 247 articles
Repeat for each database searched.
Export and Aggregate Results:
Phase 3: Screening and Selection
Deduplication:
python search_databases.py results.json --deduplicate --output unique_results.json
- Removes duplicates by DOI (primary) or title (fallback)
- Document number of duplicates removed
Title Screening:
- Review all titles against inclusion/exclusion criteria
- Exclude obviously irrelevant studies
- Document number excluded at this stage
Abstract Screening:
- Read abstracts of remaining studies
- Apply inclusion/exclusion criteria rigorously
- Document reasons for exclusion
Full-Text Screening:
- Obtain full texts of remaining studies
- Conduct detailed review against all criteria
- Document specific reasons for exclusion
- Record final number of included studies
Create PRISMA Flow Diagram:
Initial search: n = X
├─ After deduplication: n = Y
├─ After title screening: n = Z
├─ After abstract screening: n = A
└─ Included in review: n = B
Phase 4: Data Extraction and Quality Assessment
Extract Key Data from each included study:
- Study metadata (authors, year, journal, DOI)
- Study design and methods
- Sample size and population characteristics
- Key findings and results
- Limitations noted by authors
- Funding sources and conflicts of interest
Assess Study Quality:
- For RCTs: Use Cochrane Risk of Bias tool
- For observational studies: Use Newcastle-Ottawa Scale
- For systematic reviews: Use AMSTAR 2
- Rate each study: High, Moderate, Low, or Very Low quality
- Consider excluding very low-quality studies
Organize by Themes:
- Identify 3-5 major themes across studies
- Group studies by theme (studies may appear in multiple themes)
- Note patterns, consensus, and controversies
Phase 5: Synthesis and Analysis
Create Review Document from template:
cp assets/review_template.md my_literature_review.md
Write Thematic Synthesis (NOT study-by-study summaries):
- Organize Results section by themes or research questions
- Synthesize findings across multiple studies within each theme
- Compare and contrast different approaches and results
- Identify consensus areas and points of controversy
- Highlight the strongest evidence
Example structure:
#### 3.3.1 Theme: CRISPR Delivery Methods
Multiple delivery approaches have been investigated for therapeutic
gene editing. Viral vectors (AAV) were used in 15 studies^1-15^ and
showed high transduction efficiency (65-85%) but raised immunogenicity
concerns^3,7,12^. In contrast, lipid nanoparticles demonstrated lower
efficiency (40-60%) but improved safety profiles^16-23^.
Critical Analysis:
- Evaluate methodological strengths and limitations across studies
- Assess quality and consistency of evidence
- Identify knowledge gaps and methodological gaps
- Note areas requiring future research
Write Discussion:
- Interpret findings in broader context
- Discuss clinical, practical, or research implications
- Acknowledge limitations of the review itself
- Compare with previous reviews if applicable
- Propose specific future research directions
Phase 6: Citation Verification
CRITICAL: All citations must be verified for accuracy before final submission.
Verify All DOIs:
python scripts/verify_citations.py my_literature_review.md
This script:
- Extracts all DOIs from the document
- Verifies each DOI resolves correctly
- Retrieves metadata from CrossRef
- Generates verification report
- Outputs properly formatted citations
Review Verification Report:
- Check for any failed DOIs
- Verify author names, titles, and publication details match
- Correct any errors in the original document
- Re-run verification until all citations pass
Format Citations Consistently:
- Choose one citation style and use throughout (see
references/citation_styles.md)
- Common styles: APA, Nature, Vancouver, Chicago, IEEE
- Use verification script output to format citations correctly
- Ensure in-text citations match reference list format
Phase 7: Document Generation
Generate PDF:
python scripts/generate_pdf.py my_literature_review.md \
--citation-style apa \
--output my_review.pdf
Options:
--citation-style: apa, nature, chicago, vancouver, ieee
--no-toc: Disable table of contents
--no-numbers: Disable section numbering
--check-deps: Check if pandoc/xelatex are installed
Review Final Output:
- Check PDF formatting and layout
- Verify all sections are present
- Ensure citations render correctly
- Check that figures/tables appear properly
- Verify table of contents is accurate
Quality Checklist:
Database-Specific Search Guidance
PubMed / PubMed Central
Access via gget skill:
# Search PubMed
gget search pubmed "CRISPR gene editing" -l 100
# Search with filters
# Use PubMed Advanced Search Builder to construct complex queries
# Then execute via gget or direct Entrez API
Search tips:
- Use MeSH terms:
"sickle cell disease"[MeSH]
- Field tags:
[Title], [Title/Abstract], [Author]
- Date filters:
2020:2024[Publication Date]
- Boolean operators: AND, OR, NOT
- See MeSH browser: https://meshb.nlm.nih.gov/search
bioRxiv / medRxiv
Access via gget skill:
gget search biorxiv "CRISPR sickle cell" -l 50
Important considerations:
- Preprints are not peer-reviewed
- Verify findings with caution
- Check if preprint has been published (CrossRef)
- Note preprint version and date
arXiv
Access via direct API or WebFetch:
# Example search categories:
# q-bio.QM (Quantitative Methods)
# q-bio.GN (Genomics)
# q-bio.MN (Molecular Networks)
# cs.LG (Machine Learning)
# stat.ML (Machine Learning Statistics)
# Search format: category AND terms
search_query = "cat:q-bio.QM AND ti:\"single cell sequencing\""
Semantic Scholar
Access via direct API (requires API key, or use free tier):
- 200M+ papers across all fields
- Excellent for cross-disciplinary searches
- Provides citation graphs and paper recommendations
- Use for finding highly influential papers
Specialized Biomedical Databases
Use appropriate skills:
- ChEMBL:
bioservices skill for chemical bioactivity
- UniProt:
gget or bioservices skill for protein information
- KEGG:
bioservices skill for pathways and genes
- COSMIC:
gget skill for cancer mutations
- AlphaFold:
gget alphafold for protein structures
- PDB:
gget or direct API for experimental structures
Citation Chaining
Expand search via citation networks:
Forward citations (papers citing key papers):
- Use
parallel-cli search to find papers citing a specific work:parallel-cli search "papers citing [Author et al. Year] [paper title]" \
-q "citing" -q "[key author]" \
--json --max-results 10 --excerpt-max-chars-total 27000 \
--include-domains "scholar.google.com,semanticscholar.org,arxiv.org,pubmed.ncbi.nlm.nih.gov" \
-o sources/litreview_forward_citations.json
- Use Google Scholar "Cited by"
- Use Semantic Scholar or OpenAlex APIs
- Identifies newer research building on seminal work
Backward citations (references from key papers):
Citation Style Guide
Detailed formatting guidelines are in references/citation_styles.md. Quick reference:
APA (7th Edition)
- In-text: (Smith et al., 2023)
- Reference: Smith, J. D., Johnson, M. L., & Williams, K. R. (2023). Title. Journal, 22(4), 301-318. https://doi.org/10.xxx/yyy
Nature
- In-text: Superscript numbers^1,2^
- Reference: Smith, J. D., Johnson, M. L. & Williams, K. R. Title. Nat. Rev. Drug Discov. 22, 301-318 (2023).
Vancouver
- In-text: Superscript numbers^1,2^
- Reference: Smith JD, Johnson ML, Williams KR. Title. Nat Rev Drug Discov. 2023;22(4):301-18.
Always verify citations with verify_citations.py before finalizing.
Prioritizing High-Impact Papers (CRITICAL)
Always prioritize influential, highly-cited papers from reputable authors and top venues. Quality matters more than quantity in literature reviews.
Citation Count Thresholds
Use citation counts to identify the most impactful papers:
| Paper Age |
Citation Threshold |
Classification |
| 0-3 years |
20+ citations |
Noteworthy |
| 0-3 years |
100+ citations |
Highly Influential |
| 3-7 years |
100+ citations |
Significant |
| 3-7 years |
500+ citations |
Landmark Paper |
| 7+ years |
500+ citations |
Seminal Work |
| 7+ years |
1000+ citations |
Foundational |
Journal and Venue Tiers
Prioritize papers from higher-tier venues:
- Tier 1 (Always Prefer): Nature, Science, Cell, NEJM, Lancet, JAMA, PNAS, Nature Medicine, Nature Biotechnology
- Tier 2 (Strong Preference): High-impact specialized journals (IF>10), top conferences (NeurIPS, ICML for ML/AI)
- Tier 3 (Include When Relevant): Respected specialized journals (IF 5-10)
- Tier 4 (Use Sparingly): Lower-impact peer-reviewed venues
Author Reputation Assessment
Prefer papers from:
- Senior researchers with high h-index (>40 in established fields)
- Leading research groups at recognized institutions (Harvard, Stanford, MIT, Oxford, etc.)
- Authors with multiple Tier-1 publications in the relevant field
- Researchers with recognized expertise (awards, editorial positions, society fellows)
Identifying Seminal Papers
For any topic, identify foundational work by:
- High citation count (typically 500+ for papers 5+ years old)
- Frequently cited by other included studies (appears in many reference lists)
- Published in Tier-1 venues (Nature, Science, Cell family)
- Written by field pioneers (often cited as establishing concepts)
Best Practices
Search Strategy
- Start with parallel-web: Use
parallel-cli search with academic domains for initial broad coverage before querying specialized databases
- Use multiple databases (minimum 3): Ensures comprehensive coverage — parallel-web counts as one source
- Include preprint servers: Captures latest unpublished findings
- Document everything: Search strings, dates, result counts for reproducibility — save all parallel-cli output to
sources/
- Test and refine: Run pilot searches, review results, adjust search terms
- Sort by citations: When available, sort search results by citation count to surface influential work first
- Use parallel-cli extract: Fetch full content from promising URLs found during search to verify relevance before full-text screening
Screening and Selection
- Use multiple databases (minimum 3): Ensures comprehensive coverage
- Include preprint servers: Captures latest unpublished findings
- Document everything: Search strings, dates, result counts for reproducibility
- Test and refine: Run pilot searches, review results, adjust search terms
Screening and Selection
- Use clear criteria: Document inclusion/exclusion criteria before screening
- Screen systematically: Title → Abstract → Full text
- Document exclusions: Record reasons for excluding studies
- Consider dual screening: For systematic reviews, have two reviewers screen independently
Synthesis
- Organize thematically: Group by themes, NOT by individual studies
- Synthesize across studies: Compare, contrast, identify patterns
- Be critical: Evaluate quality and consistency of evidence
- Identify gaps: Note what's missing or understudied
Quality and Reproducibility
- Assess study quality: Use appropriate quality assessment tools
- Verify all citations: Run verify_citations.py script
- Document methodology: Provide enough detail for others to reproduce
- Follow guidelines: Use PRISMA for systematic reviews
Writing
- Be objective: Present evidence fairly, acknowledge limitations
- Be systematic: Follow structured template
- Be specific: Include numbers, statistics, effect sizes where available
- Be clear: Use clear headings, logical flow, thematic organization
Common Pitfalls to Avoid
- Single database search: Misses relevant papers; always search multiple databases
- No search documentation: Makes review irreproducible; document all searches
- Study-by-study summary: Lacks synthesis; organize thematically instead
- Unverified citations: Leads to errors; always run verify_citations.py
- Too broad search: Yields thousands of irrelevant results; refine with specific terms
- Too narrow search: Misses relevant papers; include synonyms and related terms
- Ignoring preprints: Misses latest findings; include bioRxiv, medRxiv, arXiv
- No quality assessment: Treats all evidence equally; assess and report quality
- Publication bias: Only positive results published; note potential bias
- Outdated search: Field evolves rapidly; clearly state search date
Example Workflow
Complete workflow for a biomedical literature review:
# 1. Create review document from template
cp assets/review_template.md crispr_sickle_cell_review.md
# 2. Start with parallel-web for broad academic search
parallel-cli search "CRISPR Cas9 sickle cell disease gene therapy efficacy" \
-q "CRISPR" -q "sickle cell" -q "gene therapy" \
--json --max-results 10 --excerpt-max-chars-total 27000 \
--include-domains "scholar.google.com,arxiv.org,pubmed.ncbi.nlm.nih.gov,semanticscholar.org,biorxiv.org,nature.com,science.org,cell.com,pnas.org,nih.gov" \
-o sources/litreview_crispr_scd-academic.json
parallel-cli search "CRISPR sickle cell disease clinical trials treatment" \
-q "CRISPR" -q "sickle cell" \
--json --max-results 10 --excerpt-max-chars-total 27000 \
-o sources/litreview_crispr_scd-general.json
# 3. Search specialized databases using appropriate skills
# - Use gget skill for PubMed, bioRxiv
# - Use direct API access for arXiv, Semantic Scholar
# - Export results in JSON format
# 4. Aggregate and process results (combine parallel-cli + database results)
python scripts/search_databases.py combined_results.json \
--deduplicate \
--rank citations \
--year-start 2015 \
--year-end 2024 \
--format markdown \
--output search_results.md \
--summary
# 5. Screen results and extract data
# - Use parallel-cli extract to fetch full content from promising URLs
# - Manually screen titles, abstracts, full texts
# - Extract key data into the review document
# - Organize by themes
# 6. Write the review following template structure
# - Introduction with clear objectives
# - Detailed methodology section
# - Results organized thematically
# - Critical discussion
# - Clear conclusions
# 7. Verify all citations
python scripts/verify_citations.py crispr_sickle_cell_review.md
# Review the citation report
cat crispr_sickle_cell_review_citation_report.json
# Fix any failed citations and re-verify
python scripts/verify_citations.py crispr_sickle_cell_review.md
# 8. Generate professional PDF
python scripts/generate_pdf.py crispr_sickle_cell_review.md \
--citation-style nature \
--output crispr_sickle_cell_review.pdf
# 9. Review final PDF and markdown outputs
Integration with Other Skills
This skill works seamlessly with other scientific skills:
Web Search & Extraction (parallel-web skill — PRIMARY)
- parallel-cli search: Broad academic and general web search with domain filtering — use for initial scoping, finding papers, citation chaining, and supplementary searches
- parallel-cli extract: Fetch full content from paper URLs, journal websites, and preprint servers — use for reading abstracts, extracting reference lists, and verifying paper details
- parallel-cli search --include-domains: Academic-focused search across scholarly domains (arxiv.org, pubmed, nature.com, etc.)
Database Access Skills
- gget: PubMed, bioRxiv, COSMIC, AlphaFold, Ensembl, UniProt
- bioservices: ChEMBL, KEGG, Reactome, UniProt, PubChem
- datacommons-client: Demographics, economics, health statistics
Analysis Skills
- pydeseq2: RNA-seq differential expression (for methods sections)
- scanpy: Single-cell analysis (for methods sections)
- anndata: Single-cell data (for methods sections)
- biopython: Sequence analysis (for background sections)
Visualization Skills
- matplotlib: Generate figures and plots for review
- seaborn: Statistical visualizations
Writing Skills
- brand-guidelines: Apply institutional branding to PDF
- internal-comms: Adapt review for different audiences
Resources
Bundled Resources
Scripts:
scripts/verify_citations.py: Verify DOIs and generate formatted citations
scripts/generate_pdf.py: Convert markdown to professional PDF
scripts/search_databases.py: Process, deduplicate, and format search results
References:
references/citation_styles.md: Detailed citation formatting guide (APA, Nature, Vancouver, Chicago, IEEE)
references/database_strategies.md: Comprehensive database search strategies
Assets:
assets/review_template.md: Complete literature review template with all sections
External Resources
Guidelines:
Tools:
Citation Styles:
Dependencies
Required CLI Tools
# parallel-cli (PRIMARY — for web search and URL extraction)
curl -fsSL https://parallel.ai/install.sh | bash
# Or: uv tool install "parallel-web-tools[cli]"
# Authenticate: parallel-cli auth
Required Python Packages
pip install requests # For citation verification
Required System Tools
# For PDF generation
brew install pandoc # macOS
apt-get install pandoc # Linux
# For LaTeX (PDF generation)
brew install --cask mactex # macOS
apt-get install texlive-xetex # Linux
Check dependencies:
python scripts/generate_pdf.py --check-deps
Summary
This literature-review skill provides:
- Systematic methodology following academic best practices
- Parallel-web powered search using
parallel-cli search for fast, broad academic literature discovery with scholarly domain filtering
- Multi-database integration via existing scientific skills (gget, bioservices, datacommons-client)
- Citation verification ensuring accuracy and credibility
- Professional output in markdown and PDF formats
- Comprehensive guidance covering the entire review process
- Quality assurance with verification and validation tools
- Reproducibility through detailed documentation requirements
Conduct thorough, rigorous literature reviews that meet academic standards and provide comprehensive synthesis of current knowledge in any domain.
1---2name: literature-review-53description: Conduct comprehensive, systematic literature reviews using multiple academic databases (PubMed, arXiv, bioRxiv, Semantic Scholar, etc.). This skill should be used when conducting systematic literature reviews, meta-analyses, research synthesis, or comprehensive literature searches across biomedical, scientific, and technical domains. Creates professionally formatted markdown documents and PDFs with verified citations in multiple citation styles (APA, Nature, Vancouver, etc.).4---56# Literature Review78## Overview910Conduct systematic, comprehensive literature reviews following rigorous academic methodology. Search multiple literature databases, synthesize findings thematically, verify all citations for accuracy, and generate professional output documents in markdown and PDF formats.1112This skill uses the **parallel-web skill** (`parallel-cli search`) as the primary web search tool for broad academic literature discovery, supplemented by specialized database access skills (gget, bioservices, datacommons-client). It provides specialized tools for citation verification, result aggregation, and document generation.1314## When to Use This Skill1516Use this skill when:17- Conducting a systematic literature review for research or publication18- Synthesizing current knowledge on a specific topic across multiple sources19- Performing meta-analysis or scoping reviews20- Writing the literature review section of a research paper or thesis21- Investigating the state of the art in a research domain22- Identifying research gaps and future directions23- Requiring verified citations and professional formatting2425## Visual Enhancement with Scientific Schematics2627**⚠️ MANDATORY: Every literature review MUST include at least 1-2 AI-generated figures using the scientific-schematics skill.**2829This is not optional. Literature reviews without visual elements are incomplete. Before finalizing any document:301. Generate at minimum ONE schematic or diagram (e.g., PRISMA flow diagram for systematic reviews)312. Prefer 2-3 figures for comprehensive reviews (search strategy flowchart, thematic synthesis diagram, conceptual framework)3233**How to generate figures:**34- Use the **scientific-schematics** skill to generate AI-powered publication-quality diagrams35- Simply describe your desired diagram in natural language36- Nano Banana Pro will automatically generate, review, and refine the schematic3738**How to generate schematics:**39```bash40python scripts/generate_schematic.py "your diagram description" -o figures/output.png41```4243The AI will automatically:44- Create publication-quality images with proper formatting45- Review and refine through multiple iterations46- Ensure accessibility (colorblind-friendly, high contrast)47- Save outputs in the figures/ directory4849**When to add schematics:**50- PRISMA flow diagrams for systematic reviews51- Literature search strategy flowcharts52- Thematic synthesis diagrams53- Research gap visualization maps54- Citation network diagrams55- Conceptual framework illustrations56- Any complex concept that benefits from visualization5758For detailed guidance on creating schematics, refer to the scientific-schematics skill documentation.5960---6162## Core Workflow6364Literature reviews follow a structured, multi-phase workflow:6566### Phase 1: Planning and Scoping67681. **Define Research Question**: Use PICO framework (Population, Intervention, Comparison, Outcome) for clinical/biomedical reviews69 - Example: "What is the efficacy of CRISPR-Cas9 (I) for treating sickle cell disease (P) compared to standard care (C)?"70712. **Establish Scope and Objectives**:72 - Define clear, specific research questions73 - Determine review type (narrative, systematic, scoping, meta-analysis)74 - Set boundaries (time period, geographic scope, study types)75763. **Develop Search Strategy**:77 - Identify 2-4 main concepts from research question78 - List synonyms, abbreviations, and related terms for each concept79 - Plan Boolean operators (AND, OR, NOT) to combine terms80 - Select minimum 3 complementary databases81 - **Use the parallel-web skill (`parallel-cli search`) for initial scoping** to quickly gauge the landscape before formal database searches82834. **Set Inclusion/Exclusion Criteria**:84 - Date range (e.g., last 10 years: 2015-2024)85 - Language (typically English, or specify multilingual)86 - Publication types (peer-reviewed, preprints, reviews)87 - Study designs (RCTs, observational, in vitro, etc.)88 - Document all criteria clearly8990### Phase 2: Systematic Literature Search91921. **Multi-Database Search**:9394 Select databases appropriate for the domain. **Always start with parallel-web for broad academic coverage**, then supplement with domain-specific databases.9596 **Web-Based Academic Search (parallel-web skill — START HERE):**97 - Use `parallel-cli search` with academic domain filtering for broad scholarly coverage98 - Run two searches: academic-focused + general to catch all relevant sources99 ```bash100 # Academic-focused search across scholarly sources101 parallel-cli search "your research topic" -q "keyword1" -q "keyword2" \102 --json --max-results 10 --excerpt-max-chars-total 27000 \103 --include-domains "scholar.google.com,arxiv.org,pubmed.ncbi.nlm.nih.gov,semanticscholar.org,biorxiv.org,medrxiv.org,ncbi.nlm.nih.gov,nature.com,science.org,ieee.org,acm.org,springer.com,wiley.com,cell.com,pnas.org,nih.gov" \104 -o sources/litreview_<topic>-academic.json105106 # General search for supplementary sources107 parallel-cli search "your research topic" -q "keyword1" -q "keyword2" \108 --json --max-results 10 --excerpt-max-chars-total 27000 \109 -o sources/litreview_<topic>-general.json110 ```111 - Use `parallel-cli extract` to fetch full content from specific paper URLs or PDFs found in search results112 ```bash113 parallel-cli extract "https://arxiv.org/abs/XXXX.XXXXX" --json114 ```115116 **Biomedical & Life Sciences:**117 - Use `gget` skill: `gget search pubmed "search terms"` for PubMed/PMC118 - Use `gget` skill: `gget search biorxiv "search terms"` for preprints119 - Use `bioservices` skill for ChEMBL, KEGG, UniProt, etc.120121 **General Scientific Literature:**122 - Search arXiv via direct API (preprints in physics, math, CS, q-bio)123 - Search Semantic Scholar via API (200M+ papers, cross-disciplinary)124 - Use Google Scholar for comprehensive coverage (manual or careful scraping)125126 **Specialized Databases:**127 - Use `gget alphafold` for protein structures128 - Use `gget cosmic` for cancer genomics129 - Use `datacommons-client` for demographic/statistical data130 - Use specialized databases as appropriate for the domain1311322. **Document Search Parameters**:133 ```markdown134 ## Search Strategy135136 ### Database: PubMed137 - **Date searched**: 2024-10-25138 - **Date range**: 2015-01-01 to 2024-10-25139 - **Search string**:140 ```141 ("CRISPR"[Title] OR "Cas9"[Title])142 AND ("sickle cell"[MeSH] OR "SCD"[Title/Abstract])143 AND 2015:2024[Publication Date]144 ```145 - **Results**: 247 articles146 ```147148 Repeat for each database searched.1491503. **Export and Aggregate Results**:151 - Export results in JSON format from each database152 - Combine all results into a single file153 - Use `scripts/search_databases.py` for post-processing:154 ```bash155 python search_databases.py combined_results.json \156 --deduplicate \157 --format markdown \158 --output aggregated_results.md159 ```160161### Phase 3: Screening and Selection1621631. **Deduplication**:164 ```bash165 python search_databases.py results.json --deduplicate --output unique_results.json166 ```167 - Removes duplicates by DOI (primary) or title (fallback)168 - Document number of duplicates removed1691702. **Title Screening**:171 - Review all titles against inclusion/exclusion criteria172 - Exclude obviously irrelevant studies173 - Document number excluded at this stage1741753. **Abstract Screening**:176 - Read abstracts of remaining studies177 - Apply inclusion/exclusion criteria rigorously178 - Document reasons for exclusion1791804. **Full-Text Screening**:181 - Obtain full texts of remaining studies182 - Conduct detailed review against all criteria183 - Document specific reasons for exclusion184 - Record final number of included studies1851865. **Create PRISMA Flow Diagram**:187 ```188 Initial search: n = X189 ├─ After deduplication: n = Y190 ├─ After title screening: n = Z191 ├─ After abstract screening: n = A192 └─ Included in review: n = B193 ```194195### Phase 4: Data Extraction and Quality Assessment1961971. **Extract Key Data** from each included study:198 - Study metadata (authors, year, journal, DOI)199 - Study design and methods200 - Sample size and population characteristics201 - Key findings and results202 - Limitations noted by authors203 - Funding sources and conflicts of interest2042052. **Assess Study Quality**:206 - **For RCTs**: Use Cochrane Risk of Bias tool207 - **For observational studies**: Use Newcastle-Ottawa Scale208 - **For systematic reviews**: Use AMSTAR 2209 - Rate each study: High, Moderate, Low, or Very Low quality210 - Consider excluding very low-quality studies2112123. **Organize by Themes**:213 - Identify 3-5 major themes across studies214 - Group studies by theme (studies may appear in multiple themes)215 - Note patterns, consensus, and controversies216217### Phase 5: Synthesis and Analysis2182191. **Create Review Document** from template:220 ```bash221 cp assets/review_template.md my_literature_review.md222 ```2232242. **Write Thematic Synthesis** (NOT study-by-study summaries):225 - Organize Results section by themes or research questions226 - Synthesize findings across multiple studies within each theme227 - Compare and contrast different approaches and results228 - Identify consensus areas and points of controversy229 - Highlight the strongest evidence230231 Example structure:232 ```markdown233 #### 3.3.1 Theme: CRISPR Delivery Methods234235 Multiple delivery approaches have been investigated for therapeutic236 gene editing. Viral vectors (AAV) were used in 15 studies^1-15^ and237 showed high transduction efficiency (65-85%) but raised immunogenicity238 concerns^3,7,12^. In contrast, lipid nanoparticles demonstrated lower239 efficiency (40-60%) but improved safety profiles^16-23^.240 ```2412423. **Critical Analysis**:243 - Evaluate methodological strengths and limitations across studies244 - Assess quality and consistency of evidence245 - Identify knowledge gaps and methodological gaps246 - Note areas requiring future research2472484. **Write Discussion**:249 - Interpret findings in broader context250 - Discuss clinical, practical, or research implications251 - Acknowledge limitations of the review itself252 - Compare with previous reviews if applicable253 - Propose specific future research directions254255### Phase 6: Citation Verification256257**CRITICAL**: All citations must be verified for accuracy before final submission.2582591. **Verify All DOIs**:260 ```bash261 python scripts/verify_citations.py my_literature_review.md262 ```263264 This script:265 - Extracts all DOIs from the document266 - Verifies each DOI resolves correctly267 - Retrieves metadata from CrossRef268 - Generates verification report269 - Outputs properly formatted citations2702712. **Review Verification Report**:272 - Check for any failed DOIs273 - Verify author names, titles, and publication details match274 - Correct any errors in the original document275 - Re-run verification until all citations pass2762773. **Format Citations Consistently**:278 - Choose one citation style and use throughout (see `references/citation_styles.md`)279 - Common styles: APA, Nature, Vancouver, Chicago, IEEE280 - Use verification script output to format citations correctly281 - Ensure in-text citations match reference list format282283### Phase 7: Document Generation2842851. **Generate PDF**:286 ```bash287 python scripts/generate_pdf.py my_literature_review.md \288 --citation-style apa \289 --output my_review.pdf290 ```291292 Options:293 - `--citation-style`: apa, nature, chicago, vancouver, ieee294 - `--no-toc`: Disable table of contents295 - `--no-numbers`: Disable section numbering296 - `--check-deps`: Check if pandoc/xelatex are installed2972982. **Review Final Output**:299 - Check PDF formatting and layout300 - Verify all sections are present301 - Ensure citations render correctly302 - Check that figures/tables appear properly303 - Verify table of contents is accurate3043053. **Quality Checklist**:306 - [ ] All DOIs verified with verify_citations.py307 - [ ] Citations formatted consistently308 - [ ] PRISMA flow diagram included (for systematic reviews)309 - [ ] Search methodology fully documented310 - [ ] Inclusion/exclusion criteria clearly stated311 - [ ] Results organized thematically (not study-by-study)312 - [ ] Quality assessment completed313 - [ ] Limitations acknowledged314 - [ ] References complete and accurate315 - [ ] PDF generates without errors316317## Database-Specific Search Guidance318319### PubMed / PubMed Central320321Access via `gget` skill:322```bash323# Search PubMed324gget search pubmed "CRISPR gene editing" -l 100325326# Search with filters327# Use PubMed Advanced Search Builder to construct complex queries328# Then execute via gget or direct Entrez API329```330331**Search tips**:332- Use MeSH terms: `"sickle cell disease"[MeSH]`333- Field tags: `[Title]`, `[Title/Abstract]`, `[Author]`334- Date filters: `2020:2024[Publication Date]`335- Boolean operators: AND, OR, NOT336- See MeSH browser: https://meshb.nlm.nih.gov/search337338### bioRxiv / medRxiv339340Access via `gget` skill:341```bash342gget search biorxiv "CRISPR sickle cell" -l 50343```344345**Important considerations**:346- Preprints are not peer-reviewed347- Verify findings with caution348- Check if preprint has been published (CrossRef)349- Note preprint version and date350351### arXiv352353Access via direct API or WebFetch:354```python355# Example search categories:356# q-bio.QM (Quantitative Methods)357# q-bio.GN (Genomics)358# q-bio.MN (Molecular Networks)359# cs.LG (Machine Learning)360# stat.ML (Machine Learning Statistics)361362# Search format: category AND terms363search_query = "cat:q-bio.QM AND ti:\"single cell sequencing\""364```365366### Semantic Scholar367368Access via direct API (requires API key, or use free tier):369- 200M+ papers across all fields370- Excellent for cross-disciplinary searches371- Provides citation graphs and paper recommendations372- Use for finding highly influential papers373374### Specialized Biomedical Databases375376Use appropriate skills:377- **ChEMBL**: `bioservices` skill for chemical bioactivity378- **UniProt**: `gget` or `bioservices` skill for protein information379- **KEGG**: `bioservices` skill for pathways and genes380- **COSMIC**: `gget` skill for cancer mutations381- **AlphaFold**: `gget alphafold` for protein structures382- **PDB**: `gget` or direct API for experimental structures383384### Citation Chaining385386Expand search via citation networks:3873881. **Forward citations** (papers citing key papers):389 - Use `parallel-cli search` to find papers citing a specific work:390 ```bash391 parallel-cli search "papers citing [Author et al. Year] [paper title]" \392 -q "citing" -q "[key author]" \393 --json --max-results 10 --excerpt-max-chars-total 27000 \394 --include-domains "scholar.google.com,semanticscholar.org,arxiv.org,pubmed.ncbi.nlm.nih.gov" \395 -o sources/litreview_forward_citations.json396 ```397 - Use Google Scholar "Cited by"398 - Use Semantic Scholar or OpenAlex APIs399 - Identifies newer research building on seminal work4004012. **Backward citations** (references from key papers):402 - Use `parallel-cli extract` to fetch full text of key papers and extract their reference lists:403 ```bash404 parallel-cli extract "https://doi.org/10.xxxx/yyyy" --json405 ```406 - Extract references from included papers407 - Identify highly cited foundational work408 - Find papers cited by multiple included studies409410## Citation Style Guide411412Detailed formatting guidelines are in `references/citation_styles.md`. Quick reference:413414### APA (7th Edition)415- In-text: (Smith et al., 2023)416- Reference: Smith, J. D., Johnson, M. L., & Williams, K. R. (2023). Title. *Journal*, *22*(4), 301-318. https://doi.org/10.xxx/yyy417418### Nature419- In-text: Superscript numbers^1,2^420- Reference: Smith, J. D., Johnson, M. L. & Williams, K. R. Title. *Nat. Rev. Drug Discov.* **22**, 301-318 (2023).421422### Vancouver423- In-text: Superscript numbers^1,2^424- Reference: Smith JD, Johnson ML, Williams KR. Title. Nat Rev Drug Discov. 2023;22(4):301-18.425426**Always verify citations** with verify_citations.py before finalizing.427428### Prioritizing High-Impact Papers (CRITICAL)429430**Always prioritize influential, highly-cited papers from reputable authors and top venues.** Quality matters more than quantity in literature reviews.431432#### Citation Count Thresholds433434Use citation counts to identify the most impactful papers:435436| Paper Age | Citation Threshold | Classification |437|-----------|-------------------|----------------|438| 0-3 years | 20+ citations | Noteworthy |439| 0-3 years | 100+ citations | Highly Influential |440| 3-7 years | 100+ citations | Significant |441| 3-7 years | 500+ citations | Landmark Paper |442| 7+ years | 500+ citations | Seminal Work |443| 7+ years | 1000+ citations | Foundational |444445#### Journal and Venue Tiers446447Prioritize papers from higher-tier venues:448449- **Tier 1 (Always Prefer):** Nature, Science, Cell, NEJM, Lancet, JAMA, PNAS, Nature Medicine, Nature Biotechnology450- **Tier 2 (Strong Preference):** High-impact specialized journals (IF>10), top conferences (NeurIPS, ICML for ML/AI)451- **Tier 3 (Include When Relevant):** Respected specialized journals (IF 5-10)452- **Tier 4 (Use Sparingly):** Lower-impact peer-reviewed venues453454#### Author Reputation Assessment455456Prefer papers from:457- **Senior researchers** with high h-index (>40 in established fields)458- **Leading research groups** at recognized institutions (Harvard, Stanford, MIT, Oxford, etc.)459- **Authors with multiple Tier-1 publications** in the relevant field460- **Researchers with recognized expertise** (awards, editorial positions, society fellows)461462#### Identifying Seminal Papers463464For any topic, identify foundational work by:4651. **High citation count** (typically 500+ for papers 5+ years old)4662. **Frequently cited by other included studies** (appears in many reference lists)4673. **Published in Tier-1 venues** (Nature, Science, Cell family)4684. **Written by field pioneers** (often cited as establishing concepts)469470## Best Practices471472### Search Strategy4731. **Start with parallel-web**: Use `parallel-cli search` with academic domains for initial broad coverage before querying specialized databases4742. **Use multiple databases** (minimum 3): Ensures comprehensive coverage — parallel-web counts as one source4753. **Include preprint servers**: Captures latest unpublished findings4764. **Document everything**: Search strings, dates, result counts for reproducibility — save all parallel-cli output to `sources/`4775. **Test and refine**: Run pilot searches, review results, adjust search terms4786. **Sort by citations**: When available, sort search results by citation count to surface influential work first4797. **Use parallel-cli extract**: Fetch full content from promising URLs found during search to verify relevance before full-text screening480481### Screening and Selection4821. **Use multiple databases** (minimum 3): Ensures comprehensive coverage4832. **Include preprint servers**: Captures latest unpublished findings4843. **Document everything**: Search strings, dates, result counts for reproducibility4854. **Test and refine**: Run pilot searches, review results, adjust search terms486487### Screening and Selection4881. **Use clear criteria**: Document inclusion/exclusion criteria before screening4892. **Screen systematically**: Title → Abstract → Full text4903. **Document exclusions**: Record reasons for excluding studies4914. **Consider dual screening**: For systematic reviews, have two reviewers screen independently492493### Synthesis4941. **Organize thematically**: Group by themes, NOT by individual studies4952. **Synthesize across studies**: Compare, contrast, identify patterns4963. **Be critical**: Evaluate quality and consistency of evidence4974. **Identify gaps**: Note what's missing or understudied498499### Quality and Reproducibility5001. **Assess study quality**: Use appropriate quality assessment tools5012. **Verify all citations**: Run verify_citations.py script5023. **Document methodology**: Provide enough detail for others to reproduce5034. **Follow guidelines**: Use PRISMA for systematic reviews504505### Writing5061. **Be objective**: Present evidence fairly, acknowledge limitations5072. **Be systematic**: Follow structured template5083. **Be specific**: Include numbers, statistics, effect sizes where available5094. **Be clear**: Use clear headings, logical flow, thematic organization510511## Common Pitfalls to Avoid5125131. **Single database search**: Misses relevant papers; always search multiple databases5142. **No search documentation**: Makes review irreproducible; document all searches5153. **Study-by-study summary**: Lacks synthesis; organize thematically instead5164. **Unverified citations**: Leads to errors; always run verify_citations.py5175. **Too broad search**: Yields thousands of irrelevant results; refine with specific terms5186. **Too narrow search**: Misses relevant papers; include synonyms and related terms5197. **Ignoring preprints**: Misses latest findings; include bioRxiv, medRxiv, arXiv5208. **No quality assessment**: Treats all evidence equally; assess and report quality5219. **Publication bias**: Only positive results published; note potential bias52210. **Outdated search**: Field evolves rapidly; clearly state search date523524## Example Workflow525526Complete workflow for a biomedical literature review:527528```bash529# 1. Create review document from template530cp assets/review_template.md crispr_sickle_cell_review.md531532# 2. Start with parallel-web for broad academic search533parallel-cli search "CRISPR Cas9 sickle cell disease gene therapy efficacy" \534 -q "CRISPR" -q "sickle cell" -q "gene therapy" \535 --json --max-results 10 --excerpt-max-chars-total 27000 \536 --include-domains "scholar.google.com,arxiv.org,pubmed.ncbi.nlm.nih.gov,semanticscholar.org,biorxiv.org,nature.com,science.org,cell.com,pnas.org,nih.gov" \537 -o sources/litreview_crispr_scd-academic.json538539parallel-cli search "CRISPR sickle cell disease clinical trials treatment" \540 -q "CRISPR" -q "sickle cell" \541 --json --max-results 10 --excerpt-max-chars-total 27000 \542 -o sources/litreview_crispr_scd-general.json543544# 3. Search specialized databases using appropriate skills545# - Use gget skill for PubMed, bioRxiv546# - Use direct API access for arXiv, Semantic Scholar547# - Export results in JSON format548549# 4. Aggregate and process results (combine parallel-cli + database results)550python scripts/search_databases.py combined_results.json \551 --deduplicate \552 --rank citations \553 --year-start 2015 \554 --year-end 2024 \555 --format markdown \556 --output search_results.md \557 --summary558559# 5. Screen results and extract data560# - Use parallel-cli extract to fetch full content from promising URLs561# - Manually screen titles, abstracts, full texts562# - Extract key data into the review document563# - Organize by themes564565# 6. Write the review following template structure566# - Introduction with clear objectives567# - Detailed methodology section568# - Results organized thematically569# - Critical discussion570# - Clear conclusions571572# 7. Verify all citations573python scripts/verify_citations.py crispr_sickle_cell_review.md574575# Review the citation report576cat crispr_sickle_cell_review_citation_report.json577578# Fix any failed citations and re-verify579python scripts/verify_citations.py crispr_sickle_cell_review.md580581# 8. Generate professional PDF582python scripts/generate_pdf.py crispr_sickle_cell_review.md \583 --citation-style nature \584 --output crispr_sickle_cell_review.pdf585586# 9. Review final PDF and markdown outputs587```588589## Integration with Other Skills590591This skill works seamlessly with other scientific skills:592593### Web Search & Extraction (parallel-web skill — PRIMARY)594- **parallel-cli search**: Broad academic and general web search with domain filtering — use for initial scoping, finding papers, citation chaining, and supplementary searches595- **parallel-cli extract**: Fetch full content from paper URLs, journal websites, and preprint servers — use for reading abstracts, extracting reference lists, and verifying paper details596- **parallel-cli search --include-domains**: Academic-focused search across scholarly domains (arxiv.org, pubmed, nature.com, etc.)597598### Database Access Skills599- **gget**: PubMed, bioRxiv, COSMIC, AlphaFold, Ensembl, UniProt600- **bioservices**: ChEMBL, KEGG, Reactome, UniProt, PubChem601- **datacommons-client**: Demographics, economics, health statistics602603### Analysis Skills604- **pydeseq2**: RNA-seq differential expression (for methods sections)605- **scanpy**: Single-cell analysis (for methods sections)606- **anndata**: Single-cell data (for methods sections)607- **biopython**: Sequence analysis (for background sections)608609### Visualization Skills610- **matplotlib**: Generate figures and plots for review611- **seaborn**: Statistical visualizations612613### Writing Skills614- **brand-guidelines**: Apply institutional branding to PDF615- **internal-comms**: Adapt review for different audiences616617## Resources618619### Bundled Resources620621**Scripts:**622- `scripts/verify_citations.py`: Verify DOIs and generate formatted citations623- `scripts/generate_pdf.py`: Convert markdown to professional PDF624- `scripts/search_databases.py`: Process, deduplicate, and format search results625626**References:**627- `references/citation_styles.md`: Detailed citation formatting guide (APA, Nature, Vancouver, Chicago, IEEE)628- `references/database_strategies.md`: Comprehensive database search strategies629630**Assets:**631- `assets/review_template.md`: Complete literature review template with all sections632633### External Resources634635**Guidelines:**636- PRISMA (Systematic Reviews): http://www.prisma-statement.org/637- Cochrane Handbook: https://training.cochrane.org/handbook638- AMSTAR 2 (Review Quality): https://amstar.ca/639640**Tools:**641- MeSH Browser: https://meshb.nlm.nih.gov/search642- PubMed Advanced Search: https://pubmed.ncbi.nlm.nih.gov/advanced/643- Boolean Search Guide: https://www.ncbi.nlm.nih.gov/books/NBK3827/644645**Citation Styles:**646- APA Style: https://apastyle.apa.org/647- Nature Portfolio: https://www.nature.com/nature-portfolio/editorial-policies/reporting-standards648- NLM/Vancouver: https://www.nlm.nih.gov/bsd/uniform_requirements.html649650## Dependencies651652### Required CLI Tools653```bash654# parallel-cli (PRIMARY — for web search and URL extraction)655curl -fsSL https://parallel.ai/install.sh | bash656# Or: uv tool install "parallel-web-tools[cli]"657# Authenticate: parallel-cli auth658```659660### Required Python Packages661```bash662pip install requests # For citation verification663```664665### Required System Tools666```bash667# For PDF generation668brew install pandoc # macOS669apt-get install pandoc # Linux670671# For LaTeX (PDF generation)672brew install --cask mactex # macOS673apt-get install texlive-xetex # Linux674```675676Check dependencies:677```bash678python scripts/generate_pdf.py --check-deps679```680681## Summary682683This literature-review skill provides:6846851. **Systematic methodology** following academic best practices6862. **Parallel-web powered search** using `parallel-cli search` for fast, broad academic literature discovery with scholarly domain filtering6873. **Multi-database integration** via existing scientific skills (gget, bioservices, datacommons-client)6884. **Citation verification** ensuring accuracy and credibility6895. **Professional output** in markdown and PDF formats6906. **Comprehensive guidance** covering the entire review process6917. **Quality assurance** with verification and validation tools6928. **Reproducibility** through detailed documentation requirements693694Conduct thorough, rigorous literature reviews that meet academic standards and provide comprehensive synthesis of current knowledge in any domain.695