# Systematic Review Protocol

> Runs a systematic, scoping or rapid review to a standard somebody else could repeat and get the same set of papers. Produces a registered protocol written before searching, an answerable question in a structured frame, eligibility criteria with a reason attached to each, search strings recorded per database with platform and date and validated against a seed set of known papers, screening by two people with agreement reported, a flow of counts that reconciles, extraction to a piloted form, risk of bias assessed with a named tool that actually changes the synthesis, and a pooling decision made on comparability rather than convenience. Use this skill for a systematic review, scoping review, rapid review, evidence map, meta-analysis, a review chapter that has to survive examination, a review section a funder or ministry has commissioned, or when a referee asks how the papers were selected, how many were screened, or why a particular study is not in the table. Trigger also on vaguer requests such as "review the e

- Skill: `ingridleiria/systematic-review-protocol` (Agent Skill)
- Install (CLI): `npx skillmds@latest add ingridleiria/systematic-review-protocol`
- Raw SKILL.md: https://api.skillmd.com/api/skills/ingridleiria/systematic-review-protocol/raw
- Safety review: pending
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Research & Search
- Author: ingridleiria (https://skillmd.com/u/ingridleiria)
- Updated: 2026-09-17
- Page: https://skillmd.com/skills/ingridleiria/systematic-review-protocol

---


# Systematic Review Protocol

A narrative review reports what its author read. A systematic review reports what exists, and the difference is that a second team, handed the protocol, would assemble the same set of papers. Everything in this skill exists to make that second sentence true, and almost all of the work happens before any paper is read.

The failure this prevents has a specific shape. Somebody searches, finds forty papers, reads them, writes a review, and then, under questioning, cannot say how many records were returned, which ones were excluded and why, or what would have had to be true for a paper that contradicts the conclusion to be found. The review is not necessarily wrong. It is unfalsifiable, which in a field that treats reviews as evidence is worse. The costs land in three places: a desk rejection at any journal that expects a flow diagram, an examiner who asks the one question the chapter cannot answer, and, where the review informs a decision, a recommendation resting on a sample of the literature selected by an unrecorded process that nobody, including its author, can now reconstruct.

The second failure is quieter and more damaging: eligibility criteria decided after the results are visible. A reviewer who narrows the population, or moves the date cutoff, or drops a design, having already seen which papers each choice removes, has run a specification search on a literature. It rarely feels like misconduct while it is happening. The protocol, registered before searching, is what makes it impossible.

## When to use this, and when not to

Use it for any review whose selection process will be scrutinised: a standalone systematic review, a meta-analysis, a scoping review or evidence map, a rapid review commissioned under a deadline, a thesis review chapter that has to be defended, or an evidence section in a report where somebody will ask what was left out.

Use it also when a review already written is challenged. The method runs backwards under protest: the search can be rebuilt and rerun, the counts reconstructed, and the honest answer given, which is sometimes that the original selection cannot be reproduced and the review has to be redone.

Do not use it for the evidence base assembled while writing a paper, where the object is to support an argument accurately rather than to characterise a whole literature. That is `literature-verification`, which enforces the citation standard and builds an evidence matrix without a registered protocol. Do not use it to build a theoretical argument out of a literature, which is `theoretical-framework-review`; a systematic review answers an empirical question about a body of studies, and mistaking one for the other produces a review that counts papers where it should be reconciling constructs.

Do not use it to choose the research question. Where several candidate questions compete, `research-question-ideation` settles that first, because a review scoped around the wrong question is expensive and cannot be salvaged by rescoping halfway. Do not use it as a preregistration for primary data analysis, which is `preregistration-and-analysis-plan`; the two are cousins in logic and different in content.

And do not call something a systematic review because the method was careful. Naming is part of the standard. A review with one screener, two databases and no protocol is a rapid review or a structured narrative review, both of which are legitimate and neither of which should be misdescribed.

## What you need before starting

**A question that can be turned into eligibility criteria.** The test is mechanical: write down two studies that would clearly be included and two that would clearly be excluded, and check that the question alone decides all four. Missing: you have a topic. Write the two or three candidate questions the topic could support, take them to whoever commissioned the work, and get one chosen before anything else happens.

**A decision on review type, made honestly against the resources available.** Systematic, scoping, rapid and evidence map differ in what they promise, not only in effort. Missing: choose by the constraint that binds. Two screeners and eight weeks supports a systematic review; one screener and three weeks supports a rapid review, correctly labelled.

**A second screener, or an explicit decision to proceed without one.** Missing: proceed with one screener, have a second person independently screen a random ten to twenty percent, report agreement on that subset, and state the single-screener limitation in the methods. Do not quietly omit the fact.

**Access to at least two databases appropriate to the field, plus one route to unpublished work if it is eligible.** Missing: use what is available, record exactly what was searched, and state in the limitations which literature the search structurally could not reach. A review that searched one database is not disqualified; a review that does not say so is.

**A seed set of eight to twelve papers you already know belong in the review.** This is the single most useful input and the one most often skipped. It is what the search string is validated against. Missing: build one from citation chasing on two or three papers you do know, before writing any string.

**Somebody who builds search strings, ideally a subject librarian.** Missing: build it yourself in concept blocks, run the recall test against the seed set, and record that the string was not externally reviewed.

**A reference manager and a screening record.** Any tool that deduplicates and holds decisions per record; a spreadsheet is acceptable if it holds one row per record with a decision and a reason. Missing: use a spreadsheet with a fixed column set, defined before screening starts, and never edit a decision without recording who changed it and why.

**The reporting guideline the field expects, and its checklist.** For most fields this is PRISMA 2020, whose checklist and flow diagram are what a journal or an examiner will look for; PRISMA-P covers the protocol itself, PRISMA-ScR the scoping review, and PRISMA-S the reporting of the search. The alternatives are real and field-specific: MOOSE for meta-analyses of observational epidemiology, ENTREQ for qualitative evidence synthesis, ROSES in environmental and conservation work, and the Campbell Collaboration's conduct and reporting standards in social policy, education and criminology. Missing: work to PRISMA 2020, which is the safest default outside those cases, and say which guideline you followed and which you did not.

**A realistic time estimate.** Screening runs at roughly six hundred to a thousand titles and abstracts per person-day, and full-text assessment at fifteen to thirty papers per person-day. Missing: estimate from those rates, and note that first-time reviewers run at roughly half them.

## The method

1. **Name the review type before anything else, and let the resources decide it.** A systematic review promises exhaustiveness within stated limits, dual screening and a risk of bias assessment. A scoping review promises coverage of the shape of a literature, not an effect estimate, and does not require risk of bias. A rapid review promises transparency about the shortcuts taken. An evidence map promises a categorised inventory. The judgement call is whether to promise what you cannot deliver, and the rule is never: relabel the review rather than degrade the method silently, because a rapid review honestly labelled is publishable and a systematic review with hidden shortcuts is retractable.

2. **Frame the question in a structured frame.** For effect questions, population, intervention or exposure, comparator, outcome, and where they bind, study designs, setting, and period. For scoping questions, population, concept, context. Write each element as a sentence, not a keyword. The judgement call is how tightly to specify the outcome, and the rule is that the outcome must be specific enough that two studies measuring different things cannot both satisfy it. "Attainment" fails this test. "Standardised test scores in mathematics or reading, at the end of an academic year" passes.

3. **Write eligibility criteria, each with its reason, and operationalise them.** Population, exposure, comparator, outcomes reported, design, period, language, publication status, and geography where relevant. Every criterion carries one clause of justification, because language and date restrictions are the two a referee always asks about and "resources" is an acceptable answer where it is the true one. Operationalise until a screener does not have to ask: not "adequate sample size" but "twenty or more units per arm". Ambiguity in the criteria is what agreement statistics later measure.

4. **Write the protocol and register it.** Question, criteria, databases and platforms, the draft search string, screening process and number of screeners, the extraction fields, the risk of bias tool, the synthesis approach and the pooling rule, the certainty framework, the team, and the date. Register it where the field has a registry. PROSPERO takes reviews with a health-related outcome, broadly defined, and is the default in health, public health and much of psychology; OSF Registries accepts a protocol in any field; the Campbell Collaboration registers titles and protocols in education, social welfare and criminal justice; INPLASY and Research Registry exist where those do not fit. Write the protocol itself to PRISMA-P, which is the checklist for exactly this document. Where no registry accepts it, deposit it with a timestamp in an institutional or public repository, which achieves the same thing: the date is fixed before the results are visible. The judgement call is how much detail to commit to before you know the literature, and the rule is that anything you would be tempted to change after seeing results must be fixed now: the outcome hierarchy, the pooling rule, and the subgroups.

5. **Build the search in concept blocks, then validate it against the seed set.** One block per question element, controlled vocabulary and free-text synonyms combined within a block, blocks combined across. Include spelling variants, plurals and truncation, and the vocabulary of every discipline that studies the question, because fields name the same construct differently and a string built in one field's language finds one field's papers. Then run the recall test: does the string return every paper in the seed set? If it misses one, the string is wrong, not the seed. Fix the string and rerun. A string that fails the recall test and is used anyway will miss an unknown quantity of the literature, and the failure is undetectable afterwards.

6. **Run each database and record the run.** Platform, exact string as entered including field tags, date run, records returned. Then deduplicate and record the method and the number removed. The judgement call arises when a rerun months later returns different numbers, which it will, and the rule is to record both runs and report the later one as an update with its date, never to overwrite the first.

7. **Pilot the screening before screening.** Both screeners independently assess the same fifty to a hundred records against the criteria, then compare. The purpose is not to measure the screeners; it is to find ambiguity in the criteria. The rule: if agreement is poor at this stage, the criteria are underspecified, so revise them, record the revision as a protocol amendment, and pilot again. Fixing criteria here costs an hour; fixing them after three thousand records costs the screening.

8. **Screen titles and abstracts liberally, then full texts strictly.** At the abstract stage, include anything that might qualify, because the cost of a false include is one full text and the cost of a false exclude is invisible. Two screeners independently; resolve disagreements by discussion, and by a third person where discussion does not settle it. At full text, apply the criteria exactly and record one reason per exclusion, chosen from the criteria list rather than written freehand, because those reasons are reported individually and freehand reasons cannot be tabulated.

9. **Chase citations after screening, not instead of it.** Hand search the reference lists of every included study and the papers citing them, plus the tables of contents of the two or three journals that dominate the included set. Records found this way enter the flow as a separate identification source with their own count. If citation chasing produces many includes that the database search missed, the search string was inadequate: say so, and consider rerunning it.

10. **Extract to a form piloted on three to five studies.** At minimum: citation, country and setting, period, design, unit of analysis, sample size and characteristics, exposure or intervention with its definition, comparator, outcome and its measurement, effect estimate with uncertainty and the model and covariates behind it, funding and conflicts. Two extractors where feasible, or one extracting and one checking every study. The judgement call is what to do when a paper reports several estimates, and the rule must be written in the protocol: name in advance which specification is taken as the paper's headline result, typically the authors' preferred one, and record the others as a sensitivity set rather than choosing per paper as you go.

11. **Contact authors for missing data, and record the outcome.** One email, one reminder, a stated deadline. Record who was contacted, when, and whether they replied, because reviewers ask and because unreplied requests are part of the honest account of what is unknown.

12. **Assess risk of bias per domain with a named tool, and then use it.** Report judgements domain by domain and study by study, never as a single collapsed score, because a score hides which threat is present. Using it means, at minimum, rerunning the synthesis with high-risk studies excluded and reporting whether the conclusion changes. An assessment that appears in an appendix table and never touches the result is decoration, and referees recognise it immediately.

13. **Make the pooling decision explicitly, before computing anything.** The rule: pool only where the studies estimate the same quantity in populations you are willing to treat as exchangeable, using outcome measures that map onto a common scale, with comparators that mean the same thing. Failing any of those, do not pool. Pooling incomparable studies produces a number with no referent, and it is the most common serious error in this literature precisely because the software will always return an answer. Where the decision is close, state it and show both the pooled and the structured version.

14. **Synthesise.** Where pooling is justified: state the effect measure, the model and why that model, how heterogeneity is quantified and what the quantity means, the subgroups and meta-regressions specified in the protocol and only those, and an assessment of small-study effects where there are enough studies to support one, conventionally around ten. Where pooling is not justified, synthesise structurally rather than narratively: group by design or population, tabulate every effect with its uncertainty and direction, describe consistency, and name the studies on each side of any disagreement. "Findings are mixed" is what gets written instead of looking at why they are mixed, and the why is usually visible in the extraction table.

15. **Assess certainty per outcome with a recognised framework and report the reasoning.** Downgrading decisions are judgements and must be shown as such, with the domain and the reason.

16. **Write the deviations table.** Every difference between the registered protocol and what was done, with the reason and the date, and whether the change was made before or after results were visible. That last column is what a careful reader looks for first.

## The pooling decision, in more detail

The four tests, applied in order, and the review stops at the first failure.

**Do the studies estimate the same quantity?** A risk difference and a hazard ratio are not the same quantity, and converting between them requires assumptions that must be stated. Studies estimating an intention-to-treat effect and studies estimating a treatment-on-the-treated effect answer different questions.

**Are the populations exchangeable enough that one average means something?** Not identical: exchangeable. Secondary school students in two countries may be; students and mid-career adults are not.

**Do the outcome measures map onto a common scale?** Standardising two different instruments into a standardised mean difference is legitimate when both measure the same construct and questionable when they do not. The check is whether you would accept the two instruments as substitutes in a single primary study.

**Do the comparators mean the same thing?** A control group receiving nothing and a control group receiving the existing programme produce effects that are not comparable, and this is the test most often skipped because comparators are described briefly in abstracts.

Where two of four fail, the answer is a structured synthesis with a clear statement of why pooling was rejected. That statement is a finding about the literature and belongs in the results, not the limitations.

## Worked example

**Situation.** A doctoral researcher and a colleague were commissioned by a university teaching centre to review whether performance-related pay for schoolteachers improves student attainment. The literature sits across education research, economics of education, and public administration, and the three fields use different vocabulary, different databases and different designs. Twelve weeks were available, two people at roughly half time, with a written report due to the centre's board.

**Task.** Produce a review that could be defended to an audience containing at least one economist and at least one education researcher, both of whom would notice a missing literature. Good meant a reproducible search, a reconciling flow, and a defensible position on whether the estimates could be combined.

**Action.** The question was framed as: among schoolteachers in publicly funded primary and secondary schools, does an individual or group financial incentive tied to measured student performance, compared with fixed salary schedules, change student attainment measured by standardised test scores at the end of an academic year? Designs were restricted to randomised trials and quasi-experimental designs with a stated identification strategy, from 1995 onward, in English or Spanish, published or in a recognised working paper series.

The seed set held eleven papers, six from education journals and five from economics working paper series, assembled by citation chasing from two reviews the team already knew.

The wrong turn came here. The first search string was built with education database vocabulary, with a strong controlled-vocabulary block on teacher incentives and a free-text block on attainment. It returned 2,140 records and looked healthy. The recall test failed: it found nine of the eleven seed papers and missed two, both economics working papers that never used the phrase performance-related pay and instead described a bonus scheme by its programme name and framed the outcome as test score gains. The instinct was to add the two papers by hand and move on, which would have concealed the defect. Instead the string was rebuilt with a third synonym block drawn from economics vocabulary, including bonus, merit pay, incentive pay, pay for performance and the named programmes, and an outcome block that included test score, achievement and learning outcome as free text. The rebuilt string returned 4,318 records across five databases and found all eleven seed papers.

That rebuild cost four days and changed the review. Of the 41 studies eventually included, 14 came from records that the first string would not have returned.

Screening: 4,318 records identified, 1,216 removed as duplicates, 3,102 screened on title and abstract by both reviewers independently. The pilot on the first hundred showed disagreement concentrated entirely on one criterion, whether a school-level bonus counted as an incentive tied to measured performance when the measure was an inspection rating rather than a test score. The criterion was rewritten to exclude non-test-based measures, recorded as a protocol amendment dated before full screening began, and agreement on the second pilot was acceptable. 214 full texts were assessed, 173 excluded with recorded reasons, of which the largest groups were no comparison group at 61 and no student attainment outcome at 44. Citation chasing added 9 records, of which 3 were included. Final set: 41 studies.

Risk of bias used a tool appropriate to non-randomised designs, judged per domain. Eleven studies were high risk on confounding, all of them cross-sectional comparisons of schools that had adopted incentive schemes with those that had not.

The pooling decision took a full day of argument and went against pooling the whole set. The programmes differed on a dimension that mattered: some paid individual teachers on their own students' scores, some paid whole schools on aggregate scores, and theory and the data both suggested these do different things. The comparators also differed, since in four studies the control schools received a non-financial professional development programme. The team pooled a subgroup of 12 randomised individual-incentive studies with test score outcomes, which passed all four tests, and synthesised the rest structurally in a table grouped by incentive design.

**Result.** The pooled subgroup gave a small positive standardised effect with a confidence interval that excluded zero but comfortably included effects too small to matter for policy, and heterogeneity remained substantial after subgrouping. Excluding the three studies at high risk of bias in that subgroup moved the point estimate down by about a fifth without changing the sign. The structured synthesis of the remaining 29 studies showed the group-incentive studies clustering near zero.

The report to the board said, in one sentence, that the evidence supports a small average effect from individual test-linked incentives and does not support the group-incentive designs the centre had been considering. The board changed its proposal. The review took fourteen weeks against an estimate of twelve, and the overrun was entirely the rebuilt search.

### A second scenario, where it goes differently

The same team was later asked by a regional education authority for an evidence summary on classroom observation protocols, in three weeks, one person, for an internal decision that would be taken with or without a review.

The systematic method does not fit that. What was done instead: two databases rather than five, a search string still validated against a seed set of six papers because that test is cheap and load-bearing, one screener with a second person independently checking a random fifteen percent of title and abstract decisions, no grey literature, no risk of bias tool, and no pooling at all. Every one of those shortcuts appeared in a table headed what this review did not do, with the likely direction of the resulting bias stated where it could be reasoned about, which for the missing grey literature meant a probable overstatement of average effects.

The document was labelled a rapid review in its title, its abstract and its first line. It found 38 relevant studies where a full systematic review would probably have found somewhere between fifty and seventy, and it said so.

What changed is only the promise. The recall test, the recorded strings, the recorded reasons for exclusion and the reconciling counts survived the compression, because those four are what make a review checkable and they cost hours rather than weeks. Dual screening, exhaustive sourcing and formal bias assessment are what a rapid review gives up, and naming them is what keeps it honest.

## Output

**The protocol**, written before searching:

```
REVIEW PROTOCOL
Title, review type, registration ID and date
Question, in the structured frame, one sentence per element
Eligibility criteria: table of criterion, specification, reason
Information sources: databases, platforms, grey literature routes, hand searching
Search strategy: full string for the lead database, adaptation rule for the others
Seed set: the papers the string must return
Screening: number of screeners, pilot plan, disagreement rule
Data extraction: the field list, piloting plan, multiple-estimate rule
Risk of bias: tool, domains, who assesses, disagreement rule
Synthesis: pooling rule, effect measure, model, heterogeneity, prespecified subgroups
Certainty: framework and outcomes to be rated
Team, roles, funding, conflicts
```

**The search record**, one row per database run:

| Database | Platform | Date run | Full string | Records returned | Notes |

**The flow counts**, which populate the PRISMA 2020 flow diagram and must reconcile arithmetically:

| Stage | Count |
| Records identified, database search | |
| Records identified, other sources | |
| Duplicates removed | |
| Records screened on title and abstract | |
| Records excluded at screening | |
| Full texts assessed | |
| Full texts excluded, by reason | |
| Studies included | |
| Reports included, where a study has several | |

**The extraction table**, one row per study:

| Study (year) | Country, setting, period | Design and identification | N and unit | Exposure definition | Comparator | Outcome and measure | Effect (CI) | Model and covariates | Funding | Risk of bias by domain |

**The risk of bias summary**, studies as rows and domains as columns, judgements as words rather than symbols or colours.

**The synthesis**, either a pooled estimate with its heterogeneity statistics and prespecified subgroups, or a structured table grouped on the dimension that blocked pooling, plus the certainty rating per outcome with the reason for each downgrade.

**The deviations table:**

| Protocol element | As registered | As done | Reason | Date | Before or after results were visible |

## Failure modes

**Criteria written after the search.** Recognise it by criteria that fit the found literature suspiciously well, such as a date cutoff one year before an inconvenient study. Fix by registering the protocol first; where it is already too late, say plainly that criteria were set after searching and treat the review as exploratory.

**A search string that was never recall-tested.** The most consequential and least visible failure, because there is no signal from inside: a string that misses a third of a literature returns thousands of records and looks fine. Fix by building the seed set and running the test, always, including for a rapid review.

**A single-field vocabulary in a multi-field literature.** Recognise it when the included studies come overwhelmingly from one type of journal. Fix by adding a synonym block in the other field's language and rerunning.

**Screening alone and not saying so.** Fix by disclosing it and adding a checked subsample. The disclosure costs nothing; the concealment is what damages the review when a reader notices there is no agreement statistic.

**Counts that do not reconcile.** Recognise it by adding them up, which referees do. Usually caused by records handled outside the record system, such as papers a coauthor added by email. Fix by having exactly one screening record and entering everything into it.

**Freehand exclusion reasons.** Recognise it when the reasons cannot be tabulated because there are ninety distinct wordings. Fix by fixing a reason list drawn from the criteria before full-text screening starts.

**Risk of bias as decoration.** Recognise it when the assessment table exists and the synthesis section never mentions it. Fix by running the high-risk exclusion sensitivity analysis and reporting the result even when nothing changes.

**Pooling because the software will pool.** Recognise it by a forest plot whose studies differ in population, outcome instrument and comparator. Fix by applying the four tests and reporting the failure as a finding.

**Subgroups invented after seeing heterogeneity.** Recognise it when a subgroup analysis appears that is not in the protocol. Fix by reporting it explicitly as post hoc and exploratory, in the same sentence as the result.

**Silent updates.** Rerunning a search a year later and reporting one set of numbers. Fix by reporting both runs with dates.

## Edge cases

**A literature too small to review systematically.** Where the search returns five eligible studies, the systematic method still applies and produces an honest map of a thin literature, which is a useful result. Do not widen the criteria to reach a target number of papers; report the thinness, since a gap demonstrated by an exhaustive search is stronger evidence than a gap asserted.

**A literature too large to screen.** Where deduplicated records exceed roughly ten thousand and the team is two people, either narrow the question, which is preferable, or use a documented screening prioritisation approach with a stated stopping rule and a validation sample. Never simply screen the first two thousand.

**Machine-assisted screening.** Where a tool ranks or classifies records, treat it as a prioritisation aid and not a screener: record the tool, its version and its settings, and validate it by having a human screen a random sample of what it excluded. Report the error rate found. Where the tool is unavailable, nothing in the method changes; it is an efficiency, not a requirement.

**Non-English literature.** Decide in the protocol which languages are eligible and why. Where a language is eligible but nobody on the team reads it, say how records in it were assessed, and treat machine translation as a screening aid with a stated limitation rather than as a basis for extraction.

**Grey literature and evaluations commissioned by programme funders.** Include them where eligible, and record funding and independence as extraction fields, because the association between who paid and what was found is often the most interesting pattern in the extraction table.

**Preprints and working papers that were later published.** Check every one before finalising. A working paper included in the review may now exist as a published article with a different sample and a different headline number. Cite the published version and note the change where the estimate moved materially.

**Overlapping samples across studies.** Two papers using the same cohort are not two pieces of evidence. Identify overlaps during extraction, group reports by study rather than by paper, and pool at the study level.

**Retracted or corrected studies.** Check retraction status for every included study before finalising, and rerun the synthesis without any retracted study while reporting both.

**A commissioner who wants a conclusion.** Where the review is commissioned and the sponsor has a preferred answer, the protection is the registered protocol and the recorded criteria, agreed with the sponsor before searching. Get that agreement in writing at the start, because it is much harder to obtain once the results are visible.

## Quality bar

- The protocol was registered or timestamped before the first search was run, and every deviation appears in the deviations table with its date and whether results were visible.
- The search string was validated against a named seed set and returned every paper in it.
- Every database run is recorded with platform, exact string, date and records returned, so any of them can be rerun.
- The flow counts reconcile arithmetically, populate the PRISMA 2020 flow diagram or the named alternative, and every full-text exclusion carries a reason from a fixed list.
- Screening used two people, or discloses single screening with a checked subsample and reports agreement.
- Risk of bias is judged per domain per study and visibly changes something in the synthesis.
- The pooling decision is stated with its reasoning, and studies are combined only where the four comparability tests pass.
- The review is labelled as the type it actually is, and its shortcuts are named rather than implied.

## Adapting this to your context

The worked example is an education and economics review with two screeners and eight weeks. The protocol discipline is the method; the databases, appraisal tools and synthesis vocabulary are field choices.

- **The databases.** Two plus a grey literature route, but which two matters. MEDLINE and Embase in health, PsycINFO and Web of Science in psychology, ERIC in education, Scopus in sociology, EconLit and RePEc in economics. Search each platform's controlled vocabulary alongside free text: MeSH, the PsycINFO thesaurus, ERIC descriptors.
- **The appraisal tool.** Name it in the protocol, not later. RoB 2 for randomised trials, ROBINS-I for non-randomised intervention studies, the Newcastle-Ottawa Scale for cohort and case-control designs, CASP or the Mixed Methods Appraisal Tool for qualitative evidence. For certainty, GRADE or GRADE-CERQual.
- **Synthesis when pooling is not justified.** Report it to SWiM, the guideline for synthesis without meta-analysis. Qualitative syntheses have their own methods, thematic synthesis, framework synthesis, meta-ethnography, named in the protocol like any other.
- **The screening rates.** Six hundred to a thousand abstracts a person-day suits structured medical abstracts. Social science abstracts are longer and less standardised, so budget nearer the bottom, lower again for a multilingual set.
- **What not to change.** Register or timestamp the protocol before the first search runs, validate the search string against a seed set, and date every deviation with whether results were visible.

## Related skills

`research-question-ideation` settles which question the review answers before the protocol is written. `literature-verification` is the lighter standard for an evidence base assembled while drafting a paper, and its verification rule applies to every citation this review produces. `theoretical-framework-review` handles the conceptual reconciliation that a review of constructs needs and that a review of effects does not. `preregistration-and-analysis-plan` is the same discipline applied to primary data. `descriptive-statistics-tables` and `academic-tables-booktabs` govern how the extraction and synthesis tables are presented, and `academic-figures-monochrome` governs the forest plot and any evidence map figure. `research-ethics-and-data-protection` applies where the review handles individual participant data. `research-assistant` is the mode to work in when the review is being executed to somebody else's protocol.

