Performs meta-analysis with heterogeneity assessment, forest plot generation, and GRADE evidence grading. Use when conducting meta-analyses, assessing heterogeneity, or grading evidence quality.
Meta-analysis quantitatively synthesizes results across multiple studies to produce pooled effect estimates with greater precision than any individual trial. It is the statistical engine behind systematic reviews, health technology assessments, and regulatory benefit-risk evaluations. Poorly conducted meta-analyses — using inappropriate pooling models, ignoring heterogeneity, or failing to assess publication bias — produce misleading conclusions that can harm patients and distort policy. This skill implements Cochrane Handbook methodology, PRISMA reporting standards, and GRADE evidence assessment for defensible quantitative synthesis.
Checkpoint A — Intake and Scoping
Required Intake Questions
Is this meta-analysis part of a registered systematic review (PROSPERO ID)?
What is the clinical question in PICOS format?
What is the primary effect measure (risk ratio, odds ratio, hazard ratio, mean difference, standardized mean difference)?
Are individual participant data (IPD) available, or is this an aggregate-data meta-analysis?
How many studies are expected to be included (impacts model choice and publication-bias assessment)?
Are subgroup analyses or meta-regression pre-specified?
Is a network meta-analysis (NMA) required for indirect comparisons?
What software will be used (R metafor/meta, Stata metan, RevMan, CMA)?
Is a GRADE Summary of Findings table required?
What is the intended audience (regulatory, HTA, journal publication, clinical guideline)?
Required Source Documents
Completed data extraction forms from the systematic review
Risk-of-bias assessments for all included studies
Pre-specified analysis plan (from the systematic review protocol)
Existing meta-analyses on the same topic (for comparison and gap analysis)
Step 1 — Select the Effect Measure
Choose the appropriate summary statistic based on outcome type:
Binary Outcomes
Risk Ratio (RR): Preferred when baseline risk is stable across studies; interpretable as relative risk
Odds Ratio (OR): Required for case-control studies; approximates RR when event rates are low (<10%)
Risk Difference (RD): Absolute measure; useful for NNT calculation; but often heterogeneous across baseline risks
Continuous Outcomes
Mean Difference (MD): When all studies use the same measurement scale
Standardized Mean Difference (SMD): When studies use different scales measuring the same construct (Cohen's d or Hedges' g with small-sample correction)
Time-to-Event Outcomes
Hazard Ratio (HR): Preferred; requires either reported HRs or IPD reconstruction from KM curves using validated methods (Tierney et al., BMJ 2007)
Count/Rate Data
Rate Ratio: Events per person-time; use when follow-up durations differ substantially
Document the rationale for effect-measure selection in the methods section.
Step 2 — Choose the Statistical Model
Fixed-Effect Model
Assumes all studies estimate the same true effect (common-effect assumption)
Appropriate only when studies are clinically and methodologically homogeneous
Weights studies by inverse variance (larger studies contribute more)
Methods: Mantel-Haenszel (binary), inverse variance (continuous), Peto (rare events with balanced arms)
Random-Effects Model
Assumes study-level true effects vary around a mean (distributional assumption)
Appropriate when heterogeneity is expected or observed
Incorporates between-study variance (tau²) into weights
Methods: DerSimonian-Laird (simple but biased tau² estimate), REML (restricted maximum likelihood — preferred), Paule-Mandel, Knapp-Hartung adjustment for CI calculation (recommended when <20 studies)
Model Selection Decision
Default to random-effects in most clinical contexts (heterogeneity is nearly always present)
Present fixed-effect as sensitivity analysis
When tau² estimate is zero, random-effects reduces to fixed-effect
Step 3 — Assess and Quantify Heterogeneity
Heterogeneity assessment is mandatory, not optional:
Cochran's Q test: Chi-squared test for heterogeneity; low power with few studies (p < 0.10 threshold)
I² statistic: Proportion of variability due to between-study differences; 0-40% = low, 30-60% = moderate, 50-90% = substantial, 75-100% = considerable (Cochrane thresholds are overlapping deliberately)
Tau² (τ²): Absolute between-study variance; more informative than I² for clinical interpretation
Prediction interval: Range within which the true effect in a new study would likely fall; far more informative than the confidence interval of the pooled estimate for clinical decision-making
Visual inspection: Examine forest plot for non-overlapping CIs, outliers, and directional inconsistency
When substantial heterogeneity exists, do not simply report the pooled estimate — explore and explain it.
Step 4 — Explore Sources of Heterogeneity
When I² > 40% or clinical heterogeneity is suspected:
Subgroup Analysis
Pre-specified clinical or methodological moderators (dose, population, risk of bias, study design)
Formal test for subgroup differences (interaction test, not separate subgroup p-values)
Minimum 2-3 studies per subgroup for meaningful analysis
Meta-Regression
Requires ≥10 studies (rule of thumb: 10 studies per covariate)
Random-effects meta-regression with Knapp-Hartung modification
Report R² (proportion of between-study variance explained)
Caution: ecological fallacy — study-level associations do not imply patient-level relationships
Sensitivity Analyses
Leave-one-out analysis (recompute removing each study sequentially)
Restrict to low risk-of-bias studies
Restrict to studies with adequate allocation concealment
Compare fixed vs. random effects
Compare effect measures (RR vs. OR)
Step 5 — Assess Publication Bias
Required when ≥10 studies are included in the meta-analysis:
Funnel plot: Scatter plot of effect estimates vs. standard error; asymmetry suggests bias
Egger's test: Regression test for funnel-plot asymmetry (continuous outcomes)
Peters' test: Alternative for binary outcomes (Egger's test is biased for ORs)
Trim-and-fill: Imputes missing studies and recalculates the pooled estimate (provides sensitivity estimate)
Contour-enhanced funnel plot: Adds significance contours to distinguish publication bias from other asymmetry causes
Selection models (Copas, Vevea-Hedges): More sophisticated modeling of publication-selection process
Report the results of bias assessment even if no evidence of bias is found.
Step 6 — Generate Forest Plots and Visualizations
Produce publication-quality figures:
Forest Plot Requirements
Individual study estimates with 95% CIs as horizontal lines/boxes
Box size proportional to study weight
Diamond for pooled estimate (width = 95% CI)
Vertical line at null effect (RR=1 or MD=0)
Numeric columns: study author/year, n per arm, effect estimate, CI, weight
I² and Q-test results displayed below the plot
Prediction interval overlaid on the diamond (when using random effects)
Additional Plots
Funnel plot (with or without contour enhancement)
L'Abbe plot (for binary outcomes — event rates in treatment vs. control)
Galbraith/radial plot (to identify outliers)
Cumulative meta-analysis (chronological accumulation of evidence)
Influence plot (leave-one-out results)
Step 7 — Apply GRADE and Generate Summary of Findings
Rate the certainty of evidence for each outcome using GRADE:
Domain
Downgrade Criteria
Risk of bias
Majority of evidence at high/unclear risk in key domains
Population, intervention, or outcome differs from target question
Imprecision
CI crosses clinical decision threshold; OIS not met
Publication bias
Significant asymmetry; missing studies suspected
Upgrade criteria (for observational studies): large magnitude of effect, dose-response gradient, all plausible confounders would reduce effect.
Final rating: High, Moderate, Low, or Very Low certainty.
Present in a GRADE Summary of Findings (SoF) table with: outcome, number of studies/participants, effect estimate (95% CI), certainty rating, and plain-language interpretation.
Checkpoint B — Analysis Review
Effect measure is appropriate for the outcome type and study designs
Statistical model (fixed/random) is justified and sensitivity analysis includes the alternative
Heterogeneity is quantified (I², tau², prediction interval) and explored
Subgroup analyses and meta-regression are pre-specified in the protocol
Publication-bias assessment is performed (for ≥10 studies)
Forest plots include all required elements (weights, CIs, pooled estimate, heterogeneity statistics)
GRADE assessment is completed for all critical outcomes
Sensitivity analyses are performed and reported
Results are interpreted in context of heterogeneity and certainty
Software, packages, and version numbers are documented
Quality Audit
Data extraction values entering the meta-analysis match the systematic review extraction forms
Effect direction is consistent across all studies (verify coding of events/non-events)
Zero-event studies are handled appropriately (continuity correction or exact methods, not excluded)
Small-study effects are explored beyond simple funnel plots
Prediction intervals are reported alongside confidence intervals for random-effects models
Network meta-analysis (if conducted) assesses transitivity and coherence
All analyses are reproducible from documented code/commands
All [VERIFY] flags have been resolved or escalated
Guidelines
Never pool clinically heterogeneous studies just because statistical heterogeneity is low — clinical judgment precedes statistical combination
Report prediction intervals for all random-effects meta-analyses — the pooled CI alone understates uncertainty
I² is not a measure of the amount of heterogeneity; it is the proportion of variability due to heterogeneity — interpret in context of tau²
Do not use fixed-effect models solely to obtain a smaller p-value when random-effects is more appropriate
Zero-event studies carry information and should not be silently excluded — use Peto method or exact methods
For rare events (<1% event rate), standard methods (Mantel-Haenszel, DerSimonian-Laird) may be unreliable — use exact methods or beta-binomial models
Subgroup analyses are hypothesis-generating unless pre-specified and powered — do not overinterpret
Cite the specific software, package, and version used (e.g., R 4.3, metafor 4.4-0)
Escalate to methodologist when I² >75%, when zero-event handling is complex, or when network meta-analysis shows incoherence
This skill produces statistical analyses — clinical interpretation of pooled effects requires domain-expert collaboration
1---2name: conducting-meta-analyses3description: Performs meta-analysis with heterogeneity assessment, forest plot generation, and GRADE evidence grading. Use when conducting meta-analyses, assessing heterogeneity, or grading evidence quality.4---56# Conducting Meta-Analyses
78## Why This Skill Exists
910Meta-analysis quantitatively synthesizes results across multiple studies to produce pooled effect estimates with greater precision than any individual trial. It is the statistical engine behind systematic reviews, health technology assessments, and regulatory benefit-risk evaluations. Poorly conducted meta-analyses — using inappropriate pooling models, ignoring heterogeneity, or failing to assess publication bias — produce misleading conclusions that can harm patients and distort policy. This skill implements Cochrane Handbook methodology, PRISMA reporting standards, and GRADE evidence assessment for defensible quantitative synthesis.
1112---
1314## Checkpoint A — Intake and Scoping
1516### Required Intake Questions
171. Is this meta-analysis part of a registered systematic review (PROSPERO ID)?
182. What is the clinical question in PICOS format?
193. What is the primary effect measure (risk ratio, odds ratio, hazard ratio, mean difference, standardized mean difference)?
204. Are individual participant data (IPD) available, or is this an aggregate-data meta-analysis?
215. How many studies are expected to be included (impacts model choice and publication-bias assessment)?
226. Are subgroup analyses or meta-regression pre-specified?
237. Is a network meta-analysis (NMA) required for indirect comparisons?
248. What software will be used (R metafor/meta, Stata metan, RevMan, CMA)?
259. Is a GRADE Summary of Findings table required?
2610. What is the intended audience (regulatory, HTA, journal publication, clinical guideline)?
2728### Required Source Documents
29- Completed data extraction forms from the systematic review
30- Risk-of-bias assessments for all included studies
31- Pre-specified analysis plan (from the systematic review protocol)
32- Existing meta-analyses on the same topic (for comparison and gap analysis)
3334---
3536## Step 1 — Select the Effect Measure
3738Choose the appropriate summary statistic based on outcome type:
3940### Binary Outcomes
41- **Risk Ratio (RR)**: Preferred when baseline risk is stable across studies; interpretable as relative risk
42- **Odds Ratio (OR)**: Required for case-control studies; approximates RR when event rates are low (<10%)
43- **Risk Difference (RD)**: Absolute measure; useful for NNT calculation; but often heterogeneous across baseline risks
4445### Continuous Outcomes
46- **Mean Difference (MD)**: When all studies use the same measurement scale
47- **Standardized Mean Difference (SMD)**: When studies use different scales measuring the same construct (Cohen's d or Hedges' g with small-sample correction)
4849### Time-to-Event Outcomes
50- **Hazard Ratio (HR)**: Preferred; requires either reported HRs or IPD reconstruction from KM curves using validated methods (Tierney et al., BMJ 2007)
5152### Count/Rate Data
53- **Rate Ratio**: Events per person-time; use when follow-up durations differ substantially
5455Document the rationale for effect-measure selection in the methods section.
5657---
5859## Step 2 — Choose the Statistical Model
6061### Fixed-Effect Model
62- Assumes all studies estimate the same true effect (common-effect assumption)
63- Appropriate only when studies are clinically and methodologically homogeneous
64- Weights studies by inverse variance (larger studies contribute more)
65- Methods: Mantel-Haenszel (binary), inverse variance (continuous), Peto (rare events with balanced arms)
6667### Random-Effects Model
68- Assumes study-level true effects vary around a mean (distributional assumption)
69- Appropriate when heterogeneity is expected or observed
70- Incorporates between-study variance (tau²) into weights
71- Methods: DerSimonian-Laird (simple but biased tau² estimate), REML (restricted maximum likelihood — preferred), Paule-Mandel, Knapp-Hartung adjustment for CI calculation (recommended when <20 studies)
7273### Model Selection Decision
74- Default to random-effects in most clinical contexts (heterogeneity is nearly always present)
75- Present fixed-effect as sensitivity analysis
76- When tau² estimate is zero, random-effects reduces to fixed-effect
7778---
7980## Step 3 — Assess and Quantify Heterogeneity
8182Heterogeneity assessment is mandatory, not optional:
83841. **Cochran's Q test**: Chi-squared test for heterogeneity; low power with few studies (p < 0.10 threshold)
852. **I² statistic**: Proportion of variability due to between-study differences; 0-40% = low, 30-60% = moderate, 50-90% = substantial, 75-100% = considerable (Cochrane thresholds are overlapping deliberately)
863. **Tau² (τ²)**: Absolute between-study variance; more informative than I² for clinical interpretation
874. **Prediction interval**: Range within which the true effect in a new study would likely fall; far more informative than the confidence interval of the pooled estimate for clinical decision-making
885. **Visual inspection**: Examine forest plot for non-overlapping CIs, outliers, and directional inconsistency
8990When substantial heterogeneity exists, do not simply report the pooled estimate — explore and explain it.
9192---
9394## Step 4 — Explore Sources of Heterogeneity
9596When I² > 40% or clinical heterogeneity is suspected:
9798### Subgroup Analysis
99- Pre-specified clinical or methodological moderators (dose, population, risk of bias, study design)
100- Formal test for subgroup differences (interaction test, not separate subgroup p-values)
101- Minimum 2-3 studies per subgroup for meaningful analysis
102103### Meta-Regression
104- Requires ≥10 studies (rule of thumb: 10 studies per covariate)
105- Random-effects meta-regression with Knapp-Hartung modification
106- Report R² (proportion of between-study variance explained)
107- Caution: ecological fallacy — study-level associations do not imply patient-level relationships
108109### Sensitivity Analyses
110- Leave-one-out analysis (recompute removing each study sequentially)
111- Restrict to low risk-of-bias studies
112- Restrict to studies with adequate allocation concealment
113- Compare fixed vs. random effects
114- Compare effect measures (RR vs. OR)
115116---
117118## Step 5 — Assess Publication Bias
119120Required when ≥10 studies are included in the meta-analysis:
1211221. **Funnel plot**: Scatter plot of effect estimates vs. standard error; asymmetry suggests bias
1232. **Egger's test**: Regression test for funnel-plot asymmetry (continuous outcomes)
1243. **Peters' test**: Alternative for binary outcomes (Egger's test is biased for ORs)
1254. **Trim-and-fill**: Imputes missing studies and recalculates the pooled estimate (provides sensitivity estimate)
1265. **Contour-enhanced funnel plot**: Adds significance contours to distinguish publication bias from other asymmetry causes
1276. **Selection models** (Copas, Vevea-Hedges): More sophisticated modeling of publication-selection process
128129Report the results of bias assessment even if no evidence of bias is found.
130131---
132133## Step 6 — Generate Forest Plots and Visualizations
134135Produce publication-quality figures:
136137### Forest Plot Requirements
138- Individual study estimates with 95% CIs as horizontal lines/boxes
139- Box size proportional to study weight
140- Diamond for pooled estimate (width = 95% CI)
141- Vertical line at null effect (RR=1 or MD=0)
142- Numeric columns: study author/year, n per arm, effect estimate, CI, weight
143- I² and Q-test results displayed below the plot
144- Prediction interval overlaid on the diamond (when using random effects)
145146### Additional Plots
147- Funnel plot (with or without contour enhancement)
148- L'Abbe plot (for binary outcomes — event rates in treatment vs. control)
149- Galbraith/radial plot (to identify outliers)
150- Cumulative meta-analysis (chronological accumulation of evidence)
151- Influence plot (leave-one-out results)
152153---
154155## Step 7 — Apply GRADE and Generate Summary of Findings
156157Rate the certainty of evidence for each outcome using GRADE:
158159| Domain | Downgrade Criteria |
160|--------|--------------------|
161| Risk of bias | Majority of evidence at high/unclear risk in key domains |
162| Inconsistency | I² >60%, unexplained; prediction interval crosses null |
163| Indirectness | Population, intervention, or outcome differs from target question |
164| Imprecision | CI crosses clinical decision threshold; OIS not met |
165| Publication bias | Significant asymmetry; missing studies suspected |
166167Upgrade criteria (for observational studies): large magnitude of effect, dose-response gradient, all plausible confounders would reduce effect.
168169Final rating: High, Moderate, Low, or Very Low certainty.
170171Present in a GRADE Summary of Findings (SoF) table with: outcome, number of studies/participants, effect estimate (95% CI), certainty rating, and plain-language interpretation.
172173---
174175## Checkpoint B — Analysis Review
1761771. [ ] Effect measure is appropriate for the outcome type and study designs
1782. [ ] Statistical model (fixed/random) is justified and sensitivity analysis includes the alternative
1793. [ ] Heterogeneity is quantified (I², tau², prediction interval) and explored
1804. [ ] Subgroup analyses and meta-regression are pre-specified in the protocol
1815. [ ] Publication-bias assessment is performed (for ≥10 studies)
1826. [ ] Forest plots include all required elements (weights, CIs, pooled estimate, heterogeneity statistics)
1837. [ ] GRADE assessment is completed for all critical outcomes
1848. [ ] Sensitivity analyses are performed and reported
1859. [ ] Results are interpreted in context of heterogeneity and certainty
18610. [ ] Software, packages, and version numbers are documented
187188---
189190## Quality Audit
191192- [ ] Data extraction values entering the meta-analysis match the systematic review extraction forms
193- [ ] Effect direction is consistent across all studies (verify coding of events/non-events)
194- [ ] Zero-event studies are handled appropriately (continuity correction or exact methods, not excluded)
195- [ ] Small-study effects are explored beyond simple funnel plots
196- [ ] Prediction intervals are reported alongside confidence intervals for random-effects models
197- [ ] Network meta-analysis (if conducted) assesses transitivity and coherence
198- [ ] All analyses are reproducible from documented code/commands
199- [ ] All [VERIFY] flags have been resolved or escalated
200201---
202203## Guidelines
2042051. Never pool clinically heterogeneous studies just because statistical heterogeneity is low — clinical judgment precedes statistical combination
2062. Report prediction intervals for all random-effects meta-analyses — the pooled CI alone understates uncertainty
2073. I² is not a measure of the amount of heterogeneity; it is the proportion of variability due to heterogeneity — interpret in context of tau²
2084. Do not use fixed-effect models solely to obtain a smaller p-value when random-effects is more appropriate
2095. Zero-event studies carry information and should not be silently excluded — use Peto method or exact methods
2106. For rare events (<1% event rate), standard methods (Mantel-Haenszel, DerSimonian-Laird) may be unreliable — use exact methods or beta-binomial models
2117. Subgroup analyses are hypothesis-generating unless pre-specified and powered — do not overinterpret
2128. Cite the specific software, package, and version used (e.g., R 4.3, metafor 4.4-0)
2139. Escalate to methodologist when I² >75%, when zero-event handling is complex, or when network meta-analysis shows incoherence
21410. This skill produces statistical analyses — clinical interpretation of pooled effects requires domain-expert collaboration
Run npx skillmds add majiayu000/conducting-meta-analyses in your terminal (requires Node.js), paste this page's agent-chat prompt into Claude, Cursor, or any MCP-connected agent, or download the SKILL.md file and copy it into your agent's skills directory.
Performs meta-analysis with heterogeneity assessment, forest plot generation, and GRADE evidence grading. Use when conducting meta-analyses, assessing heterogeneity, or grading evidence quality. It is listed under Coding & Dev Tools on SkillMD.
This skill has not completed SkillMD's automated safety review yet. Capability flags: docs only. SkillMD never runs a skill's scripts for you; review the SKILL.md before installing.
This skill is tagged as working with Claude Code, Claude.ai, OpenAI Codex. SKILL.md is an open format, so most agents that read a skills directory can load it too.
Yes. Installing skills from SkillMD is free, and the skill stays under its author's original license.
majiayu000 (@majiayu000) published this skill. Their other Agent Skills are listed on their SkillMD profile.