Literature reviews in social science differ from biomedical reviews: gray literature (working papers, policy reports) carries significant weight, quasi-experimental designs are common, and the boundary between "published" and "working" research is fluid. This skill accounts for these disciplinary norms.
For social science reviews, use the SPIDER framework (more flexible than PICO for non-clinical research):
- S (Sample): What population, group, or context?
- PI (Phenomenon of Interest): What intervention, policy, program, or phenomenon?
- D (Design): What study designs to include? (RCT, DID, IV, RDD, qualitative, mixed)
- E (Evaluation): What outcomes or impacts to evaluate?
- R (Research type): Quantitative, qualitative, or mixed methods?
Example: "What is the effect (E: learning outcomes, enrollment) of conditional cash transfer programs (PI) on educational attainment in low- and middle-income countries (S), as measured by experimental and quasi-experimental studies (D, R)?"
For policy evaluation reviews, also consider:
- Implementation context (what makes programs work or fail?)
- Heterogeneity (for whom, where, under what conditions?)
- Cost-effectiveness (what is the cost per unit of impact?)
- Scalability (do effects hold when programs scale?)
Select at minimum 3 complementary databases appropriate for the domain:
Core Social Science Databases:
- Google Scholar: Comprehensive cross-disciplinary coverage, citation tracking
- SSRN: Working papers in economics, finance, law, political science
- NBER Working Papers: Leading economics research (often years before publication)
- JSTOR: Historical and current peer-reviewed articles across social sciences
- Web of Science / Scopus: Citation-indexed peer-reviewed literature
- EconLit: Economics-specific (AEA journals, books, working papers)
Policy and Development:
- J-PAL Evidence: RCTs in development economics
- 3ie Development Evidence Portal: Impact evaluations in international development
- World Bank Open Knowledge Repository: Policy research and working papers
- IMF Working Papers: Macroeconomics and public finance
- OECD iLibrary: Cross-country policy analysis
Specialized by Discipline:
- ERIC: Education research
- PubMed/PMC: Public health, health economics, epidemiology
- PolicyFile: US public policy research
- Campbell Collaboration: Systematic reviews in social sciences
- Cochrane Library: Health intervention reviews (when relevant)
- IZA Discussion Papers: Labor economics
- CEPR Discussion Papers: European economics research
- arXiv (econ, q-fin, stat): Quantitative methods, econometrics
Preprint and Open Access:
- RePEc/IDEAS: Economics working papers and articles
- OSF Preprints: Open science preprints across social sciences
- EdWorkingPapers (Annenberg): Education policy research
Database Selection Strategy:
- Primary database for breadth: Google Scholar or Web of Science
- Working paper repository: NBER, SSRN, or IZA (field-dependent)
- Specialized database: EconLit, ERIC, J-PAL, etc. (topic-dependent)
- Gray literature: World Bank, OECD, or government reports
- Citation chaining: Forward and backward from key papers
Assess each study's methodological quality based on identification strategy and validity:
Study Design Hierarchy (for causal questions):
- RCTs (randomized controlled trials / field experiments)
- Regression discontinuity designs (RDD)
- Instrumental variables (IV)
- Difference-in-differences (DID) with pre-trends evidence
- Propensity score matching / synthetic control
- Observational with controls (OLS with covariates)
- Descriptive / correlational studies
- Qualitative / case studies
Quality Dimensions to Assess:
- Internal validity: How credible is the causal identification?
- Statistical power: Is the sample large enough to detect meaningful effects?
- External validity: How generalizable are the findings?
- Measurement quality: Are key variables well-measured?
- Pre-registration: Was the analysis pre-specified? (reduces p-hacking risk)
- Transparency: Are data and code available for replication?
- Robustness: Do results hold across specifications?
Quality Rating Scale:
- High: Strong identification strategy, adequate power, transparent methods, robust results
- Medium: Reasonable identification with some threats, adequate sample, partial robustness
- Low: Weak identification, small sample, or results sensitive to specification
Red Flags:
- P-values clustered just below 0.05 (possible p-hacking)
- No robustness checks or only one specification reported
- Effect sizes that are implausibly large
- Endogeneity concerns not addressed
- Missing data handled without sensitivity analysis
- Post-hoc subgroup analysis presented as primary finding
Prioritize papers based on methodological rigor, venue quality, and influence:
Citation Count Thresholds (social science norms):
| Paper Age |
Citations |
Classification |
| 0-3 years |
10+ |
Noteworthy |
| 0-3 years |
50+ |
Highly Influential |
| 3-7 years |
50+ |
Significant |
| 3-7 years |
200+ |
Landmark Paper |
| 7+ years |
200+ |
Seminal Work |
| 7+ years |
500+ |
Foundational |
Note: Citation norms vary by subfield. Labor economics papers accumulate citations faster than political theory. Use these as rough guides.
Journal Tiers (Economics):
- Tier 1: American Economic Review, Quarterly Journal of Economics, Econometrica, Journal of Political Economy, Review of Economic Studies
- Tier 2: Review of Economics and Statistics, Journal of the European Economic Association, American Economic Journal (all), Journal of Finance, Journal of Labor Economics, Journal of Public Economics, Journal of Development Economics, Economic Journal
- Tier 3: Respected field journals (Journal of Human Resources, Journal of Health Economics, Journal of Urban Economics, etc.)
Journal Tiers (Political Science):
- Tier 1: American Political Science Review, American Journal of Political Science, Journal of Politics
- Tier 2: Comparative Political Studies, World Politics, International Organization, British Journal of Political Science
- Tier 3: Field journals (Political Analysis, Political Behavior, etc.)
Journal Tiers (Sociology):
- Tier 1: American Sociological Review, American Journal of Sociology
- Tier 2: Social Forces, Demography, Sociology of Education
- Tier 3: Field journals
Working Papers:
- NBER Working Papers carry significant weight (peer network vetting)
- SSRN/IZA papers should be assessed on methodology, not venue
- Check if working papers have been subsequently published
Identifying Seminal Papers:
- Cited by many of the other papers you find (appears across reference lists)
- Introduced a methodology now widely used (e.g., Angrist & Krueger 1991 for IV)
- Published in Tier-1 venue with high citation count
- Referenced in textbooks and survey articles
- Written by researchers recognized as field leaders
Boolean Search Construction:
- AND: narrows (both terms required)
- OR: broadens (either term)
- Quotes: exact phrase ("conditional cash transfer")
- Wildcards: education* matches educational, education, educating
Example for a CCT review:
("conditional cash transfer" OR "CCT" OR "cash transfer program")
AND ("education" OR "school enrollment" OR "attendance" OR "learning outcomes")
AND ("developing countries" OR "low-income" OR "Global South")
Citation Chaining:
- Forward citation search: Find papers citing a key paper (Google Scholar "Cited by")
- Backward citation search: Review references of key papers
- Snowball sampling: Start with 3-5 seminal papers, follow citation networks
- Prioritize papers appearing in multiple reference lists (likely foundational)
Search Refinement:
- Pilot search: Run broad terms, review first 50 results
- Note recurring keywords, author names, and journal names
- Refine search terms based on pilot results
- Run refined search across all selected databases
- Document each iteration for reproducibility
Gray Literature Search:
- NBER: Browse by program (Labor Studies, Public Economics, etc.)
- SSRN: Search by keyword and sort by downloads or citations
- Government reports: Search agency websites directly
- Conference proceedings: ASSA/AEA meetings, APPAM, BREAD
- Dissertations: ProQuest Dissertations for emerging research
Phase 1: Planning and Scoping (Cell 1)
- Define research question using SPIDER framework
- Develop search terms with synonyms and Boolean operators
- Select minimum 3 complementary databases
- Set date range, language, and geographic constraints
- Define inclusion/exclusion criteria (study design, population, outcomes)
- Specify review type: systematic, scoping, narrative, or meta-analysis
Phase 2: Systematic Search and Source Collection (Cell 2)
- Execute search across all selected databases
- Document search strings, dates, and result counts for each database
- Aggregate results and remove duplicates
- Build structured DataFrame of all identified sources
- Record how each source was found (which search, which database)
- Conduct citation chaining from key papers
Phase 3: Screening and Quality Assessment (Cell 3)
- Apply inclusion/exclusion criteria systematically
- Document exclusion reasons with counts
- Assess methodological quality of each included study
- Assign quality rating (high/medium/low)
- Create screening flow diagram (records found -> deduplicated -> screened -> included)
Phase 4: Thematic Synthesis and Gap Analysis (Cell 4)
- Identify 3-6 major themes across included studies
- Synthesize findings within each theme (NOT study-by-study)
- Compare effect sizes and directions across studies
- Weight synthesis by study quality
- Identify consensus findings, contested claims, and gaps
- Note methodological patterns and limitations across studies
Phase 5: Summary, Implications, and References (Cell 5)
- Synthesize key takeaways (what does the weight of evidence suggest?)
- Discuss implications for policy, practice, or future research
- List specific gaps that future research should address
- Acknowledge limitations of the review itself
- Provide properly formatted reference list
Each phase becomes a separate notebook cell. Pure Python code, no decorators or wrappers.
Cell 1 -- Search Strategy:
import pandas as pd
from datetime import date
# LITERATURE REVIEW: [Topic]
# ============================================================
search_strategy = {
"research_question": "What is the effect of [intervention] on [outcome] in [population/context]?",
"framework": "SPIDER",
"sample": "[population or context]",
"phenomenon": "[intervention, policy, or phenomenon]",
"design": "[RCT, DID, IV, RDD, mixed methods, etc.]",
"evaluation": "[outcomes to measure]",
"research_type": "[quantitative, qualitative, mixed]",
"search_terms": [
'("term1" OR "synonym1") AND ("term2" OR "synonym2")',
'"exact phrase" AND (outcome1 OR outcome2)',
],
"databases": [
"Google Scholar",
"NBER Working Papers",
"SSRN",
# Add domain-specific: EconLit, ERIC, J-PAL, 3ie, etc.
],
"date_range": "2010-2025",
"language": "English",
"inclusion_criteria": [
"Peer-reviewed articles or working papers from recognized institutions",
"Empirical studies with quantitative outcome measures",
"Study designs: RCT, DID, IV, RDD, or high-quality observational",
"Population: [specify]",
],
"exclusion_criteria": [
"Purely theoretical or opinion pieces without empirical evidence",
"Studies with sample size < [threshold]",
"Non-peer-reviewed sources without institutional affiliation",
"Studies outside geographic/temporal scope",
],
"review_type": "systematic", # systematic, scoping, narrative, meta-analysis
"date_executed": str(date.today()),
}
# Display search protocol
print("=" * 60)
print("SEARCH PROTOCOL")
print("=" * 60)
for key, value in search_strategy.items():
if isinstance(value, list):
print(f"\n{key}:")
for _item in value:
print(f" - {_item}")
else:
print(f"\n{key}: {value}")
Cell 2 -- Source Collection:
import pandas as pd
# Search Results by Database
# ============================================================
# Document search execution for reproducibility
search_log = [
{"database": "Google Scholar", "date_searched": "2025-01-15",
"search_string": '("conditional cash transfer" OR CCT) AND education',
"results_found": 1240, "after_screening": 45},
{"database": "NBER", "date_searched": "2025-01-15",
"search_string": "conditional cash transfer education",
"results_found": 28, "after_screening": 12},
# ... more databases
]
search_df = pd.DataFrame(search_log)
print("SEARCH EXECUTION LOG")
print("=" * 60)
print(search_df.to_string(index=False))
print(f"\nTotal results: {search_df['results_found'].sum()}")
print(f"After screening: {search_df['after_screening'].sum()}")
# Evidence Table
# ============================================================
sources = [
{
"authors": "Author et al.",
"year": 2020,
"title": "Title of the study",
"journal": "Journal Name or Working Paper Series",
"methodology": "RCT", # RCT/DID/IV/RDD/PSM/observational/qualitative
"sample_size": "N=5,000",
"geographic_scope": "Mexico",
"key_finding": "CCT increased enrollment by 8 percentage points",
"effect_size": "8pp (95% CI: 5-11pp)",
"quality": "high", # high/medium/low
"relevance": "high", # high/medium/low
"source_db": "Google Scholar",
"doi_or_url": "https://doi.org/...",
},
# ... more sources
]
lit_table = pd.DataFrame(sources)
lit_table = lit_table.sort_values(["quality", "relevance", "year"],
ascending=[False, False, False])
# Summary statistics
print(f"\nTotal included studies: {len(lit_table)}")
print(f"\nBy methodology:")
print(lit_table["methodology"].value_counts().to_string())
print(f"\nBy quality rating:")
print(lit_table["quality"].value_counts().to_string())
lit_table
Cell 3 -- Screening Flow and Quality Assessment:
# SCREENING FLOW
# ============================================================
print("SCREENING FLOW DIAGRAM")
print("=" * 60)
flow = {
"Records identified through database searching": 1500,
"Additional records from citation chaining": 45,
"Records after deduplication": 1200,
"Records screened (title/abstract)": 1200,
"Records excluded at screening": 1050,
"Full-text articles assessed": 150,
"Full-text excluded (with reasons)": 108,
"Studies included in synthesis": 42,
}
for _stage, _count in flow.items():
print(f" {_stage}: n = {_count}")
# Exclusion reasons
print("\nExclusion reasons (full-text stage):")
_exclusion_reasons = {
"Wrong population/context": 35,
"Wrong outcome measures": 28,
"Insufficient methodology": 22,
"Duplicate/superseded version": 15,
"Not available in English": 8,
}
for _reason, _n in _exclusion_reasons.items():
print(f" - {_reason}: n = {_n}")
# QUALITY ASSESSMENT
# ============================================================
print("\n" + "=" * 60)
print("QUALITY ASSESSMENT SUMMARY")
print("=" * 60)
_quality_summary = lit_table.groupby(["methodology", "quality"]).size().unstack(fill_value=0)
print(_quality_summary)
# Flag any quality concerns
_low_quality = lit_table[lit_table["quality"] == "low"]
if len(_low_quality) > 0:
print(f"\nWARNING: {len(_low_quality)} low-quality studies included.")
print("These will be noted but down-weighted in synthesis.")
Cell 4 -- Thematic Synthesis:
# THEMATIC SYNTHESIS
# ============================================================
print("=" * 60)
print("EVIDENCE SYNTHESIS")
print("=" * 60)
# Theme 1
print("\n--- Theme 1: [Theme Name] ---")
print("Studies: [Author1 (Year), Author2 (Year), ...]")
print("Finding: [Synthesized finding across studies, not study-by-study]")
print("Strength of evidence: [Strong/Moderate/Weak]")
print("Consistency: [Consistent/Mixed/Contradictory]")
# Theme 2
print("\n--- Theme 2: [Theme Name] ---")
# ... same structure
# Consensus vs. contested
print("\n" + "=" * 60)
print("EVIDENCE MAP")
print("=" * 60)
print("\nConsensus findings (supported by multiple high-quality studies):")
print(" 1. ...")
print(" 2. ...")
print("\nContested or mixed findings:")
print(" 1. ... [Author1 finds X, but Author2 finds Y; difference may be due to ...]")
print("\nKnowledge gaps:")
print(" 1. ... [No studies examine ...]")
print(" 2. ... [Limited evidence on ... subpopulation]")
print(" 3. ... [Methodological gap: no RCTs on ...]")
print("\nMethodological patterns:")
print(f" - Most common design: {lit_table['methodology'].mode().iloc[0]}")
print(f" - Geographic concentration: {lit_table['geographic_scope'].value_counts().head(3).to_string()}")
print(" - Missing designs: [What study types are absent?]")
Cell 5 -- Summary and References:
# SUMMARY AND IMPLICATIONS
# ============================================================
print("=" * 60)
print("SUMMARY")
print("=" * 60)
print("\nKey Takeaways:")
print(" 1. [What does the weight of evidence suggest?]")
print(" 2. [What is the range of effect sizes?]")
print(" 3. [What conditions moderate the effect?]")
print("\nImplications for Policy:")
print(" - ...")
print("\nImplications for Future Research:")
print(" - [Specific gaps to fill]")
print(" - [Methodological improvements needed]")
print(" - [Understudied populations or contexts]")
print("\nLimitations of This Review:")
print(" - [Search limitations: databases not searched, language restriction]")
print(" - [Potential publication bias]")
print(" - [Scope limitations]")
# REFERENCES (APA 7th Edition)
# ============================================================
print("\n" + "=" * 60)
print("REFERENCES")
print("=" * 60)
# Generate formatted references from the evidence table
for _idx, _row in lit_table.sort_values("authors").iterrows():
_ref = f"{_row['authors']} ({_row['year']}). {_row['title']}. {_row['journal']}."
if _row.get('doi_or_url'):
_ref += f" {_row['doi_or_url']}"
print(f"\n{_ref}")
Default: APA 7th Edition (standard for social sciences)
In-text citations:
- Single author: Smith (2023) or (Smith, 2023)
- Two authors: Smith & Jones (2023) or (Smith & Jones, 2023)
- Three or more: Smith et al. (2023) or (Smith et al., 2023)
- Multiple citations: (Angrist, 2001; Duflo et al., 2011; Heckman, 2010)
- With page: (Smith, 2023, p. 42) or (Smith, 2023, pp. 42-45)
Reference list:
- Alphabetical by first author surname
- Hanging indent format
- Include DOI as URL when available
Journal article:
Smith, J. D., Johnson, M. L., & Williams, K. R. (2023). Title of article in sentence case. Title of Periodical in Title Case, 22(4), 301-318. https://doi.org/10.xxx/yyy
Working paper:
Smith, J. D. (2023). Title of paper (NBER Working Paper No. 31234). National Bureau of Economic Research. https://doi.org/10.xxx/yyy
Book:
Author, A. A. (Year). Title of work: Subtitle. Publisher.
Chapter:
Author, A. A. (Year). Title of chapter. In E. E. Editor (Ed.), Title of book (pp. xx-xx). Publisher.
Alternative: Chicago Author-Date (common in political science and sociology) follows similar in-text conventions but differs in reference formatting. Use whichever the user's field prefers.
Store all citation data in the DataFrame for programmatic formatting.
1---2name: lit-review3description: Systematic literature review methodology for social science research. Use when conducting systematic reviews, research synthesis, meta-analyses, scoping reviews, or comprehensive literature searches across economics, political science, sociology, education, public health, and development. Produces structured notebook cells with search strategy, evidence table, thematic synthesis, and gap analysis.4---56<skill_content>78<overview>9A systematic literature review follows a structured, reproducible protocol to identify, evaluate, and synthesize existing research on a topic. This skill enforces methodological rigor by requiring explicit search strategies, transparent inclusion criteria, quality assessment, and structured evidence tables -- all implemented as notebook cells for full reproducibility.1011Literature reviews in social science differ from biomedical reviews: gray literature (working papers, policy reports) carries significant weight, quasi-experimental designs are common, and the boundary between "published" and "working" research is fluid. This skill accounts for these disciplinary norms.12</overview>1314<mandatory_requirements>1516<requirement priority="critical">17 <name>Explicit Search Strategy</name>18 <description>MUST document search terms, databases, date ranges, and inclusion/exclusion criteria BEFORE presenting any findings</description>19 <rationale>Petticrew & Roberts (2006) emphasize that undocumented search strategies make reviews unreproducible and prone to selection bias. In social science, where researcher priors are strong, this discipline is essential</rationale>20 <consequence>Cherry-picked citations that confirm priors rather than representing the field</consequence>21</requirement>2223<requirement priority="critical">24 <name>Structured Evidence Table</name>25 <description>ALL sources MUST be recorded in a structured DataFrame with: authors, year, title, journal/source, methodology, sample_size, geographic_scope, key_finding, effect_size, quality_rating, and relevance</description>26 <rationale>Systematic organization prevents narrative bias and makes gaps in the evidence visible (Tranfield et al., 2003). The structured format enables programmatic analysis of the evidence base</rationale>27 <consequence>Narrative reviews without structure tend to over-weight memorable or recent studies</consequence>28</requirement>2930<requirement priority="critical">31 <name>Source Evaluation</name>32 <description>Each source MUST be assessed for methodological quality: identification strategy, internal validity, external validity, sample size adequacy, and potential threats. Rate as high/medium/low</description>33 <rationale>Not all evidence is equal. Meta-analyses show that effect sizes vary systematically with study quality (Stanley & Doucouliagos, 2012). In social science, identification strategy is the primary quality marker</rationale>34 <consequence>Treating all studies as equally valid produces misleading syntheses</consequence>35</requirement>3637<requirement priority="critical">38 <name>Gap Identification</name>39 <description>The synthesis MUST explicitly identify gaps, contradictions, and unresolved questions in the literature</description>40 <rationale>The primary value of a literature review is mapping what is NOT known, not just summarizing what is. Gap identification motivates new research</rationale>41 <consequence>Review becomes a summary rather than a foundation for new research</consequence>42</requirement>4344<requirement priority="high">45 <name>Gray Literature Inclusion</name>46 <description>MUST search working paper repositories (NBER, SSRN, IZA, CEPR, J-PAL, 3ie) alongside peer-reviewed journals</description>47 <rationale>In economics and policy research, the most current and influential work often circulates as working papers for years before publication. Publication bias means published studies systematically overstate effect sizes (Andrews & Kasy, 2019). Excluding gray literature biases reviews toward significant findings</rationale>48 <consequence>Missing the most current research and introducing publication bias into the review</consequence>49</requirement>5051<requirement priority="high">52 <name>Thematic Synthesis</name>53 <description>Results MUST be organized thematically, NOT as study-by-study summaries. Synthesize across studies within each theme</description>54 <rationale>Study-by-study presentation fails to identify patterns, contradictions, and the weight of evidence. Thematic synthesis is the distinguishing feature of a good review (Braun & Clarke, 2006)</rationale>55 <consequence>Review reads as an annotated bibliography rather than a synthesis</consequence>56</requirement>5758</mandatory_requirements>5960<research_question_framework>6162For social science reviews, use the SPIDER framework (more flexible than PICO for non-clinical research):6364- S (Sample): What population, group, or context?65- PI (Phenomenon of Interest): What intervention, policy, program, or phenomenon?66- D (Design): What study designs to include? (RCT, DID, IV, RDD, qualitative, mixed)67- E (Evaluation): What outcomes or impacts to evaluate?68- R (Research type): Quantitative, qualitative, or mixed methods?6970Example: "What is the effect (E: learning outcomes, enrollment) of conditional cash transfer programs (PI) on educational attainment in low- and middle-income countries (S), as measured by experimental and quasi-experimental studies (D, R)?"7172For policy evaluation reviews, also consider:73- Implementation context (what makes programs work or fail?)74- Heterogeneity (for whom, where, under what conditions?)75- Cost-effectiveness (what is the cost per unit of impact?)76- Scalability (do effects hold when programs scale?)77</research_question_framework>7879<database_guide>8081Select at minimum 3 complementary databases appropriate for the domain:8283Core Social Science Databases:84- Google Scholar: Comprehensive cross-disciplinary coverage, citation tracking85- SSRN: Working papers in economics, finance, law, political science86- NBER Working Papers: Leading economics research (often years before publication)87- JSTOR: Historical and current peer-reviewed articles across social sciences88- Web of Science / Scopus: Citation-indexed peer-reviewed literature89- EconLit: Economics-specific (AEA journals, books, working papers)9091Policy and Development:92- J-PAL Evidence: RCTs in development economics93- 3ie Development Evidence Portal: Impact evaluations in international development94- World Bank Open Knowledge Repository: Policy research and working papers95- IMF Working Papers: Macroeconomics and public finance96- OECD iLibrary: Cross-country policy analysis9798Specialized by Discipline:99- ERIC: Education research100- PubMed/PMC: Public health, health economics, epidemiology101- PolicyFile: US public policy research102- Campbell Collaboration: Systematic reviews in social sciences103- Cochrane Library: Health intervention reviews (when relevant)104- IZA Discussion Papers: Labor economics105- CEPR Discussion Papers: European economics research106- arXiv (econ, q-fin, stat): Quantitative methods, econometrics107108Preprint and Open Access:109- RePEc/IDEAS: Economics working papers and articles110- OSF Preprints: Open science preprints across social sciences111- EdWorkingPapers (Annenberg): Education policy research112113Database Selection Strategy:1141. Primary database for breadth: Google Scholar or Web of Science1152. Working paper repository: NBER, SSRN, or IZA (field-dependent)1163. Specialized database: EconLit, ERIC, J-PAL, etc. (topic-dependent)1174. Gray literature: World Bank, OECD, or government reports1185. Citation chaining: Forward and backward from key papers119</database_guide>120121<quality_assessment>122123Assess each study's methodological quality based on identification strategy and validity:124125Study Design Hierarchy (for causal questions):1261. RCTs (randomized controlled trials / field experiments)1272. Regression discontinuity designs (RDD)1283. Instrumental variables (IV)1294. Difference-in-differences (DID) with pre-trends evidence1305. Propensity score matching / synthetic control1316. Observational with controls (OLS with covariates)1327. Descriptive / correlational studies1338. Qualitative / case studies134135Quality Dimensions to Assess:136- Internal validity: How credible is the causal identification?137- Statistical power: Is the sample large enough to detect meaningful effects?138- External validity: How generalizable are the findings?139- Measurement quality: Are key variables well-measured?140- Pre-registration: Was the analysis pre-specified? (reduces p-hacking risk)141- Transparency: Are data and code available for replication?142- Robustness: Do results hold across specifications?143144Quality Rating Scale:145- High: Strong identification strategy, adequate power, transparent methods, robust results146- Medium: Reasonable identification with some threats, adequate sample, partial robustness147- Low: Weak identification, small sample, or results sensitive to specification148149Red Flags:150- P-values clustered just below 0.05 (possible p-hacking)151- No robustness checks or only one specification reported152- Effect sizes that are implausibly large153- Endogeneity concerns not addressed154- Missing data handled without sensitivity analysis155- Post-hoc subgroup analysis presented as primary finding156</quality_assessment>157158<prioritizing_papers>159160Prioritize papers based on methodological rigor, venue quality, and influence:161162Citation Count Thresholds (social science norms):163| Paper Age | Citations | Classification |164|-----------|-----------|----------------|165| 0-3 years | 10+ | Noteworthy |166| 0-3 years | 50+ | Highly Influential |167| 3-7 years | 50+ | Significant |168| 3-7 years | 200+ | Landmark Paper |169| 7+ years | 200+ | Seminal Work |170| 7+ years | 500+ | Foundational |171172Note: Citation norms vary by subfield. Labor economics papers accumulate citations faster than political theory. Use these as rough guides.173174Journal Tiers (Economics):175- Tier 1: American Economic Review, Quarterly Journal of Economics, Econometrica, Journal of Political Economy, Review of Economic Studies176- Tier 2: Review of Economics and Statistics, Journal of the European Economic Association, American Economic Journal (all), Journal of Finance, Journal of Labor Economics, Journal of Public Economics, Journal of Development Economics, Economic Journal177- Tier 3: Respected field journals (Journal of Human Resources, Journal of Health Economics, Journal of Urban Economics, etc.)178179Journal Tiers (Political Science):180- Tier 1: American Political Science Review, American Journal of Political Science, Journal of Politics181- Tier 2: Comparative Political Studies, World Politics, International Organization, British Journal of Political Science182- Tier 3: Field journals (Political Analysis, Political Behavior, etc.)183184Journal Tiers (Sociology):185- Tier 1: American Sociological Review, American Journal of Sociology186- Tier 2: Social Forces, Demography, Sociology of Education187- Tier 3: Field journals188189Working Papers:190- NBER Working Papers carry significant weight (peer network vetting)191- SSRN/IZA papers should be assessed on methodology, not venue192- Check if working papers have been subsequently published193194Identifying Seminal Papers:1951. Cited by many of the other papers you find (appears across reference lists)1962. Introduced a methodology now widely used (e.g., Angrist & Krueger 1991 for IV)1973. Published in Tier-1 venue with high citation count1984. Referenced in textbooks and survey articles1995. Written by researchers recognized as field leaders200</prioritizing_papers>201202<search_techniques>203204Boolean Search Construction:205- AND: narrows (both terms required)206- OR: broadens (either term)207- Quotes: exact phrase ("conditional cash transfer")208- Wildcards: education* matches educational, education, educating209210Example for a CCT review:211("conditional cash transfer" OR "CCT" OR "cash transfer program")212AND ("education" OR "school enrollment" OR "attendance" OR "learning outcomes")213AND ("developing countries" OR "low-income" OR "Global South")214215Citation Chaining:2161. Forward citation search: Find papers citing a key paper (Google Scholar "Cited by")2172. Backward citation search: Review references of key papers2183. Snowball sampling: Start with 3-5 seminal papers, follow citation networks2194. Prioritize papers appearing in multiple reference lists (likely foundational)220221Search Refinement:2221. Pilot search: Run broad terms, review first 50 results2232. Note recurring keywords, author names, and journal names2243. Refine search terms based on pilot results2254. Run refined search across all selected databases2265. Document each iteration for reproducibility227228Gray Literature Search:229- NBER: Browse by program (Labor Studies, Public Economics, etc.)230- SSRN: Search by keyword and sort by downloads or citations231- Government reports: Search agency websites directly232- Conference proceedings: ASSA/AEA meetings, APPAM, BREAD233- Dissertations: ProQuest Dissertations for emerging research234</search_techniques>235236<workflow>237238Phase 1: Planning and Scoping (Cell 1)239- Define research question using SPIDER framework240- Develop search terms with synonyms and Boolean operators241- Select minimum 3 complementary databases242- Set date range, language, and geographic constraints243- Define inclusion/exclusion criteria (study design, population, outcomes)244- Specify review type: systematic, scoping, narrative, or meta-analysis245246Phase 2: Systematic Search and Source Collection (Cell 2)247- Execute search across all selected databases248- Document search strings, dates, and result counts for each database249- Aggregate results and remove duplicates250- Build structured DataFrame of all identified sources251- Record how each source was found (which search, which database)252- Conduct citation chaining from key papers253254Phase 3: Screening and Quality Assessment (Cell 3)255- Apply inclusion/exclusion criteria systematically256- Document exclusion reasons with counts257- Assess methodological quality of each included study258- Assign quality rating (high/medium/low)259- Create screening flow diagram (records found -> deduplicated -> screened -> included)260261Phase 4: Thematic Synthesis and Gap Analysis (Cell 4)262- Identify 3-6 major themes across included studies263- Synthesize findings within each theme (NOT study-by-study)264- Compare effect sizes and directions across studies265- Weight synthesis by study quality266- Identify consensus findings, contested claims, and gaps267- Note methodological patterns and limitations across studies268269Phase 5: Summary, Implications, and References (Cell 5)270- Synthesize key takeaways (what does the weight of evidence suggest?)271- Discuss implications for policy, practice, or future research272- List specific gaps that future research should address273- Acknowledge limitations of the review itself274- Provide properly formatted reference list275</workflow>276277<output_format>278279Each phase becomes a separate notebook cell. Pure Python code, no decorators or wrappers.280281Cell 1 -- Search Strategy:282```python283import pandas as pd284from datetime import date285286# LITERATURE REVIEW: [Topic]287# ============================================================288289search_strategy = {290 "research_question": "What is the effect of [intervention] on [outcome] in [population/context]?",291 "framework": "SPIDER",292 "sample": "[population or context]",293 "phenomenon": "[intervention, policy, or phenomenon]",294 "design": "[RCT, DID, IV, RDD, mixed methods, etc.]",295 "evaluation": "[outcomes to measure]",296 "research_type": "[quantitative, qualitative, mixed]",297 "search_terms": [298 '("term1" OR "synonym1") AND ("term2" OR "synonym2")',299 '"exact phrase" AND (outcome1 OR outcome2)',300 ],301 "databases": [302 "Google Scholar",303 "NBER Working Papers",304 "SSRN",305 # Add domain-specific: EconLit, ERIC, J-PAL, 3ie, etc.306 ],307 "date_range": "2010-2025",308 "language": "English",309 "inclusion_criteria": [310 "Peer-reviewed articles or working papers from recognized institutions",311 "Empirical studies with quantitative outcome measures",312 "Study designs: RCT, DID, IV, RDD, or high-quality observational",313 "Population: [specify]",314 ],315 "exclusion_criteria": [316 "Purely theoretical or opinion pieces without empirical evidence",317 "Studies with sample size < [threshold]",318 "Non-peer-reviewed sources without institutional affiliation",319 "Studies outside geographic/temporal scope",320 ],321 "review_type": "systematic", # systematic, scoping, narrative, meta-analysis322 "date_executed": str(date.today()),323}324325# Display search protocol326print("=" * 60)327print("SEARCH PROTOCOL")328print("=" * 60)329for key, value in search_strategy.items():330 if isinstance(value, list):331 print(f"\n{key}:")332 for _item in value:333 print(f" - {_item}")334 else:335 print(f"\n{key}: {value}")336```337338Cell 2 -- Source Collection:339```python340import pandas as pd341342# Search Results by Database343# ============================================================344# Document search execution for reproducibility345346search_log = [347 {"database": "Google Scholar", "date_searched": "2025-01-15",348 "search_string": '("conditional cash transfer" OR CCT) AND education',349 "results_found": 1240, "after_screening": 45},350 {"database": "NBER", "date_searched": "2025-01-15",351 "search_string": "conditional cash transfer education",352 "results_found": 28, "after_screening": 12},353 # ... more databases354]355356search_df = pd.DataFrame(search_log)357print("SEARCH EXECUTION LOG")358print("=" * 60)359print(search_df.to_string(index=False))360print(f"\nTotal results: {search_df['results_found'].sum()}")361print(f"After screening: {search_df['after_screening'].sum()}")362363# Evidence Table364# ============================================================365sources = [366 {367 "authors": "Author et al.",368 "year": 2020,369 "title": "Title of the study",370 "journal": "Journal Name or Working Paper Series",371 "methodology": "RCT", # RCT/DID/IV/RDD/PSM/observational/qualitative372 "sample_size": "N=5,000",373 "geographic_scope": "Mexico",374 "key_finding": "CCT increased enrollment by 8 percentage points",375 "effect_size": "8pp (95% CI: 5-11pp)",376 "quality": "high", # high/medium/low377 "relevance": "high", # high/medium/low378 "source_db": "Google Scholar",379 "doi_or_url": "https://doi.org/...",380 },381 # ... more sources382]383384lit_table = pd.DataFrame(sources)385lit_table = lit_table.sort_values(["quality", "relevance", "year"],386 ascending=[False, False, False])387388# Summary statistics389print(f"\nTotal included studies: {len(lit_table)}")390print(f"\nBy methodology:")391print(lit_table["methodology"].value_counts().to_string())392print(f"\nBy quality rating:")393print(lit_table["quality"].value_counts().to_string())394395lit_table396```397398Cell 3 -- Screening Flow and Quality Assessment:399```python400# SCREENING FLOW401# ============================================================402print("SCREENING FLOW DIAGRAM")403print("=" * 60)404405flow = {406 "Records identified through database searching": 1500,407 "Additional records from citation chaining": 45,408 "Records after deduplication": 1200,409 "Records screened (title/abstract)": 1200,410 "Records excluded at screening": 1050,411 "Full-text articles assessed": 150,412 "Full-text excluded (with reasons)": 108,413 "Studies included in synthesis": 42,414}415416for _stage, _count in flow.items():417 print(f" {_stage}: n = {_count}")418419# Exclusion reasons420print("\nExclusion reasons (full-text stage):")421_exclusion_reasons = {422 "Wrong population/context": 35,423 "Wrong outcome measures": 28,424 "Insufficient methodology": 22,425 "Duplicate/superseded version": 15,426 "Not available in English": 8,427}428for _reason, _n in _exclusion_reasons.items():429 print(f" - {_reason}: n = {_n}")430431# QUALITY ASSESSMENT432# ============================================================433print("\n" + "=" * 60)434print("QUALITY ASSESSMENT SUMMARY")435print("=" * 60)436437_quality_summary = lit_table.groupby(["methodology", "quality"]).size().unstack(fill_value=0)438print(_quality_summary)439440# Flag any quality concerns441_low_quality = lit_table[lit_table["quality"] == "low"]442if len(_low_quality) > 0:443 print(f"\nWARNING: {len(_low_quality)} low-quality studies included.")444 print("These will be noted but down-weighted in synthesis.")445```446447Cell 4 -- Thematic Synthesis:448```python449# THEMATIC SYNTHESIS450# ============================================================451print("=" * 60)452print("EVIDENCE SYNTHESIS")453print("=" * 60)454455# Theme 1456print("\n--- Theme 1: [Theme Name] ---")457print("Studies: [Author1 (Year), Author2 (Year), ...]")458print("Finding: [Synthesized finding across studies, not study-by-study]")459print("Strength of evidence: [Strong/Moderate/Weak]")460print("Consistency: [Consistent/Mixed/Contradictory]")461462# Theme 2463print("\n--- Theme 2: [Theme Name] ---")464# ... same structure465466# Consensus vs. contested467print("\n" + "=" * 60)468print("EVIDENCE MAP")469print("=" * 60)470471print("\nConsensus findings (supported by multiple high-quality studies):")472print(" 1. ...")473print(" 2. ...")474475print("\nContested or mixed findings:")476print(" 1. ... [Author1 finds X, but Author2 finds Y; difference may be due to ...]")477478print("\nKnowledge gaps:")479print(" 1. ... [No studies examine ...]")480print(" 2. ... [Limited evidence on ... subpopulation]")481print(" 3. ... [Methodological gap: no RCTs on ...]")482483print("\nMethodological patterns:")484print(f" - Most common design: {lit_table['methodology'].mode().iloc[0]}")485print(f" - Geographic concentration: {lit_table['geographic_scope'].value_counts().head(3).to_string()}")486print(" - Missing designs: [What study types are absent?]")487```488489Cell 5 -- Summary and References:490```python491# SUMMARY AND IMPLICATIONS492# ============================================================493print("=" * 60)494print("SUMMARY")495print("=" * 60)496497print("\nKey Takeaways:")498print(" 1. [What does the weight of evidence suggest?]")499print(" 2. [What is the range of effect sizes?]")500print(" 3. [What conditions moderate the effect?]")501502print("\nImplications for Policy:")503print(" - ...")504505print("\nImplications for Future Research:")506print(" - [Specific gaps to fill]")507print(" - [Methodological improvements needed]")508print(" - [Understudied populations or contexts]")509510print("\nLimitations of This Review:")511print(" - [Search limitations: databases not searched, language restriction]")512print(" - [Potential publication bias]")513print(" - [Scope limitations]")514515# REFERENCES (APA 7th Edition)516# ============================================================517print("\n" + "=" * 60)518print("REFERENCES")519print("=" * 60)520521# Generate formatted references from the evidence table522for _idx, _row in lit_table.sort_values("authors").iterrows():523 _ref = f"{_row['authors']} ({_row['year']}). {_row['title']}. {_row['journal']}."524 if _row.get('doi_or_url'):525 _ref += f" {_row['doi_or_url']}"526 print(f"\n{_ref}")527```528529</output_format>530531<citation_format>532533Default: APA 7th Edition (standard for social sciences)534535In-text citations:536- Single author: Smith (2023) or (Smith, 2023)537- Two authors: Smith & Jones (2023) or (Smith & Jones, 2023)538- Three or more: Smith et al. (2023) or (Smith et al., 2023)539- Multiple citations: (Angrist, 2001; Duflo et al., 2011; Heckman, 2010)540- With page: (Smith, 2023, p. 42) or (Smith, 2023, pp. 42-45)541542Reference list:543- Alphabetical by first author surname544- Hanging indent format545- Include DOI as URL when available546547Journal article:548 Smith, J. D., Johnson, M. L., & Williams, K. R. (2023). Title of article in sentence case. Title of Periodical in Title Case, 22(4), 301-318. https://doi.org/10.xxx/yyy549550Working paper:551 Smith, J. D. (2023). Title of paper (NBER Working Paper No. 31234). National Bureau of Economic Research. https://doi.org/10.xxx/yyy552553Book:554 Author, A. A. (Year). Title of work: Subtitle. Publisher.555556Chapter:557 Author, A. A. (Year). Title of chapter. In E. E. Editor (Ed.), Title of book (pp. xx-xx). Publisher.558559Alternative: Chicago Author-Date (common in political science and sociology) follows similar in-text conventions but differs in reference formatting. Use whichever the user's field prefers.560561Store all citation data in the DataFrame for programmatic formatting.562</citation_format>563564<common_mistakes>565566<mistake severity="critical">567 <what>Presenting findings without documenting search strategy</what>568 <consequence>Review is unreproducible and vulnerable to selection bias</consequence>569 <prevention>ALWAYS create the search strategy cell first, before collecting any sources</prevention>570</mistake>571572<mistake severity="critical">573 <what>Study-by-study summaries instead of thematic synthesis</what>574 <consequence>Review reads as an annotated bibliography, fails to identify patterns or conflicts</consequence>575 <prevention>Organize by theme, synthesize across studies, compare and contrast within themes</prevention>576</mistake>577578<mistake severity="critical">579 <what>Ignoring gray literature (working papers, policy reports)</what>580 <consequence>Publication bias toward significant results. Missing current research that has not yet gone through publication lag</consequence>581 <prevention>Search NBER, SSRN, IZA, J-PAL, and relevant institutional repositories alongside journals</prevention>582</mistake>583584<mistake severity="high">585 <what>Only citing sources that support one position</what>586 <consequence>Confirmation bias undermines the review's credibility and usefulness</consequence>587 <prevention>Document ALL relevant sources found, including contradictory evidence. Present competing findings fairly</prevention>588</mistake>589590<mistake severity="high">591 <what>Treating all studies as equally rigorous</what>592 <consequence>Misleading synthesis that gives observational correlations the same weight as RCT evidence</consequence>593 <prevention>Rate each source for methodological quality and weight synthesis accordingly. An RCT with N=500 is stronger evidence than an OLS study with N=50,000</prevention>594</mistake>595596<mistake severity="high">597 <what>No screening flow documentation</what>598 <consequence>Reader cannot assess how studies were selected, making the review appear arbitrary</consequence>599 <prevention>Document records found, deduplicated, screened, and included with exclusion reasons at each stage</prevention>600</mistake>601602<mistake severity="medium">603 <what>Searching only one database</what>604 <consequence>Incomplete coverage. Different databases index different journals and working paper series</consequence>605 <prevention>Search minimum 3 complementary databases: one broad (Google Scholar), one disciplinary (EconLit/ERIC), one gray literature (NBER/SSRN)</prevention>606</mistake>607608<mistake severity="medium">609 <what>Not documenting search date</what>610 <consequence>Review cannot be updated or reproduced because the evidence base changes over time</consequence>611 <prevention>Record the exact date of each search execution</prevention>612</mistake>613614<mistake severity="medium">615 <what>Conflating statistical significance with importance</what>616 <consequence>Large-sample studies with tiny effects dominate over smaller studies with meaningful effects</consequence>617 <prevention>Report and compare effect sizes, not just p-values. Discuss practical significance alongside statistical significance</prevention>618</mistake>619620</common_mistakes>621622<interpretation_guide>623624<interpreting_results>625- Weight evidence by study quality, not just quantity of studies626- Report effect size ranges across studies, not just average effects627- Note when effects are heterogeneous by context, population, or implementation628- Distinguish between absence of evidence and evidence of absence629- Consider publication bias: significant results are overrepresented in published literature630</interpreting_results>631632<red_flags>633- All included studies find effects in the same direction (possible publication bias)634- Effect sizes shrink as study quality increases (common pattern indicating bias)635- Geographic concentration (findings from one country may not generalize)636- Temporal clustering (field may have evolved since most studies were conducted)637- Funnel plot asymmetry when enough studies for meta-analysis638</red_flags>639640<next_steps>641- Strong consensus with high-quality evidence -> Report with confidence, note remaining gaps642- Mixed evidence -> Investigate sources of heterogeneity (context, methods, populations)643- Limited evidence -> Characterize what is known, emphasize need for more research644- Contradictory high-quality evidence -> Present both sides, analyze why results diverge645</next_steps>646647</interpretation_guide>648649<references>650<paper>Petticrew, M. & Roberts, H. (2006). "Systematic Reviews in the Social Sciences: A Practical Guide." Blackwell Publishing.</paper>651<paper>Tranfield, D., Denyer, D. & Smart, P. (2003). "Towards a Methodology for Developing Evidence-Informed Management Knowledge." British Journal of Management.</paper>652<paper>Stanley, T.D. & Doucouliagos, H. (2012). "Meta-Regression Analysis in Economics and Business." Routledge.</paper>653<paper>Andrews, I. & Kasy, M. (2019). "Identification of and Correction for Publication Bias." American Economic Review.</paper>654<paper>Braun, V. & Clarke, V. (2006). "Using Thematic Analysis in Psychology." Qualitative Research in Psychology.</paper>655<paper>Snyder, H. (2019). "Literature Review as a Research Methodology: An Overview and Guidelines." Journal of Business Research.</paper>656<paper>Waddington, H. et al. (2012). "How to Do a Good Systematic Review of Effects in International Development." Journal of Development Effectiveness.</paper>657</references>658659</skill_content>