How to Build a Reproducible Systematic Review Search Strategy

MetaSyn Academy diagram showing five stages of reproducible search development: scope, concepts, platform translation, PRESS peer review, and reporting.
Systematic review search strategy and study identification

How to Build a Reproducible Systematic Review Search Strategy

A reproducible systematic-review search is a documented chain of decisions. It starts from clearly defined eligibility criteria, moves through deliberate source selection, combines controlled vocabulary with free-text terms, tests and translates strategies for each platform, records peer review and restrictions, preserves the exact searches that were run, and reports the process transparently.1,2,3 No single query, database count, checklist, or validation score can prove completeness.

The search stage creates the candidate evidence set for the entire review. If a relevant study is missed because the wrong source was chosen, a concept was over-constrained, or syntax was translated incorrectly, later screening cannot recover it.1 The opposite problem also matters: an unnecessarily broad or poorly structured strategy can create a screening burden large enough to delay or destabilize the review. The practical goal is not indiscriminate retrieval; it is high sensitivity with reasonable precision, justified for the question and documented well enough to assess, rerun, and update the work.7,11

Question-to-search development pipeline Eligibility flows through concepts, source choice, terms, testing, translation, peer review, execution, reporting, and updating. A reproducible search is a controlled pipeline Eligibilitycriteria Concepts &sources Terms &logic Test &translate PRESS &execute Report &update Versioned decision record rationale • exact execution • change history • limitations
The pipeline is iterative: test results, peer review, or interface changes can send the team back to an earlier decision without erasing the prior version.6,7

The most reliable workflow treats each arrow in the diagram as a handoff with defined evidence. Eligibility hands the search team a scope statement; the source plan explains coverage; term development preserves provenance; testing records both successes and misses; translation records database and platform separately; peer review creates a recommendation–response trail; execution freezes the exact strategy and yield; and reporting maps the retained record to PRISMA-S.1,9 When one handoff is missing, later stages are forced to rely on inference rather than documented decisions.

Start from eligibility, not from a favourite database

Search planning should be anchored in the review’s eligibility criteria. Cochrane expresses this directly in MECIR C19 for intervention reviews, while JBI and Campbell provide corresponding frameworks for their own review families.1,2,3 The transferable principle is that source selection and search structure follow the evidence sought. The Cochrane requirement to search CENTRAL, MEDLINE, Embase when available, and an applicable Specialized Register belongs to Cochrane intervention reviews; it is not a universal minimum for education, social policy, environmental, qualitative, or other evidence syntheses.

Write down the eligible evidence types before choosing sources. A review that includes trials, economic evaluations, qualitative research, regulatory reports, and unpublished results may need different source families and different searches. A narrow database list copied from another review is not a rationale.1,2

Decision record: for every database, register, website, organization, or supplementary method, state what eligible evidence it can contribute and why another source is not an adequate substitute.3

Translate the question into searchable concepts

Question frameworks are analytical aids, not commands to search every box. A typical intervention search may combine population, intervention, and a study-design filter. Comparator and outcome concepts are often omitted because they are inconsistently described in titles, abstracts, and indexing. Empirical work found lower retrieval potential for some comparator and outcome elements in the studied Cochrane contexts, but that result should guide testing rather than become a universal prohibition.4

For each possible concept, ask:

  • Is the concept required by the eligibility criteria?
  • Is it consistently expressed in searchable metadata?
  • Would omitting it retrieve an unmanageable but still screenable set?
  • Does adding it exclude known eligible studies?
  • Does the review type need a different framework or an iterative search design?

Record concepts you intentionally omit. That negative decision is methodologically useful because it prevents later team members from “tightening” the strategy without understanding the recall cost.

Develop controlled vocabulary and free text together

Controlled vocabulary and free-text searching solve different problems. Subject headings retrieve records indexed under a shared concept even when authors use different wording. Free-text terms retrieve recent unindexed records, wording not represented by the thesaurus, and details expressed in titles or abstracts. Cochrane MECIR C33 therefore requires an appropriate combination for Cochrane intervention reviews, with search strategies customized by database.1

Term development is iterative. Start with protocol language and subject expertise, then inspect controlled-vocabulary records, known relevant studies, prior high-quality strategies, spelling variants, acronyms, hyphenation, older terminology, brand and generic names, and expressions used by different disciplines. Bramer and colleagues describe a structured development method, but it remains one method rather than a universal standard.5

Term decisionWhat to preserveWhy it matters
Controlled headingVocabulary, exact heading, explosion/focus choice, subheadings, date checkedHeadings and indexing practices differ across databases and change over time.
Free-text termTerm, variant, field, phrase/proximity/truncation treatment, provenanceThe same word can behave differently by field and platform.
Rejected termTerm tested, result, and reason for exclusionPrevents reintroducing a term that damaged precision or recall.

Build a logic model, then translate for each platform

A master strategy should preserve the concepts and their relationships, not pretend that one syntax can be pasted everywhere. Database and platform are separate facts: MEDLINE can be searched through PubMed, Ovid, and other interfaces, each with different field tags, phrase behavior, proximity operators, mapping, limits, and export options. Translate concept by concept, then inspect the executed query and test the result.1,6

Translation tools can reduce time and some errors, but they still require human verification. In one comparative study, assisted translation produced fewer errors on average than manual translation, yet errors remained and retrieval varied.6 A translated string is a draft until it has been checked against the current official platform documentation and rerun.

Test the search without claiming validation

Testing is diagnostic, not ceremonial. Known relevant studies can reveal missing terminology, indexing gaps, overly restrictive concepts, or field choices that suppress retrieval. Relative recall compares retrieval against a defined benchmark set. The result is conditional on that benchmark: recovering every benchmark item does not show that all unknown eligible studies were found.7

Useful test records include the benchmark citations, why they are eligible, whether each source retrieved them, which strategy line retrieved them, missed-item analysis, and resulting revisions. Run tests after meaningful changes rather than only at the end.

Use PRESS as structured peer review

The PRESS 2015 guideline organizes peer review around six elements: translation of the question, Boolean and proximity operators, subject headings, text words, spelling and syntax, and limits or filters.8 The reviewer should see the review question, eligibility criteria, target database and platform, full strategy, and relevant context. Record the reviewer, date, recommendations, author response, and final revision so that the peer-review trail is auditable.

PRESS is not certification. A well-conducted peer review can detect errors and improve the strategy, but it does not prove that retrieval is complete or remove the need for review-team judgment.8

Justify restrictions and filters

Date, language, document-type, age, human, and study-design restrictions can all remove eligible records. Use a restriction only when it follows from eligibility or a tested, fit-for-purpose method, and record the source and version of any filter.1 Cochrane warns that filters should be assessed for their development, performance, current accuracy, and platform context, and should not be applied redundantly to pre-filtered resources.

A useful practical rule is simple: if the team cannot explain a restriction in the methods section, it probably does not belong in the search.

Separate documentation from reporting

Documentation is the internal record maintained during development and execution. Reporting is the public account included in the protocol, review, supplement, or repository. PRISMA-S provides 16 reporting items covering databases and platforms, multi-database searching, registers, online resources, citation searching, complete strategies, limits, filters, peer review, dates, deduplication, and updates.9 PRISMA 2020 also requires full strategies for all databases, registers, and websites.10

A recent metaresearch study found that only a small fraction of assessed biomedical database searches reported all six examined PRISMA-S elements, and only one included review was fully reproducible across all its searches.11 These are sample-specific findings, not a universal failure rate, but they illustrate why exact copy-and-paste strategies, platforms, dates, and limits matter.

Update the search as a controlled version

An update is not merely a new date at the end of the strategy. Preserve the prior final strategy, state which sources were rerun, document syntax or filter changes, record the new coverage interval and yield, and explain deviations.1,9 The publication team should be able to distinguish the original search, any rerun before submission, and later update searches.

Diagnose the search at four levels

When a test record is missed or retrieval is unexpectedly large, avoid editing terms at random. Diagnose the problem at the level where it arose:

  1. Scope: the eligibility criteria or searchable interpretation may be unclear.
  2. Concept: an unnecessary concept may be constraining recall, or a necessary concept may be missing.
  3. Representation: headings, text words, spelling variants, proximity, fields, or filters may not express the concept adequately.
  4. Execution: translation, platform interpretation, dates, limits, or copied syntax may differ from the intended strategy.

This hierarchy makes revisions explainable. For example, a known item missed because its intervention is described only by a brand name is a representation problem; a known item missed because the outcome block excludes it is probably a concept problem. Both may be repaired by adding text, but they imply different lessons for the strategy.

Search diagnosis hierarchy Four nested levels move from scope to concept, representation, and execution, with testing evidence feeding the diagnosis. 1 · Scope 2 · Concepts 3 · Representation 4 · Execution ? Use the observed miss or noise to identify the level—then revise and retest.
Testing is most useful when it produces a diagnosis and a traceable revision, not merely a pass label.7,8

Prepare the report from the audit trail

At reporting time, retrieve facts from the controlled record rather than rebuilding them from memory. For each source, provide the database and platform or website, the complete strategy, the last search date, limits and filters with justification, and any peer-review process.9,10 Report deduplication and updates as performed. PRISMA-S is deliberately about transparent reporting; it does not judge whether the chosen source set or search design was optimal.

The practical test is straightforward: if the team cannot paste the exact final search from a dated record, identify the person who approved a change, or explain why a restriction was used, the documentation process is incomplete even if the manuscript contains fluent methods prose.

A practical completion check

Before declaring the search record ready, confirm that another qualified person can answer all of these questions without contacting the original searcher:

  • What evidence was eligible, and which concepts were searched or omitted?
  • Why was each information source selected?
  • Which database–platform combination was used?
  • What exact strategy ran, on what date, with what restrictions and yield?
  • How was performance tested and peer reviewed?
  • What changed after testing or review?
  • How were records deduplicated and how will the search be updated?
  • Where is each PRISMA-S item reported?

If any answer depends on memory, browser history, or an undocumented conversation, the record is not finished.

References

  1. Lefebvre C, Glanville J, Briscoe S, et al. Chapter 4: Searching for and selecting studies. In: Cochrane Handbook for Systematic Reviews of Interventions. Current online version.
  2. Aromataris E, Lockwood C, Porritt K, Pilla B, Jordan Z, eds. JBI Manual for Evidence Synthesis. JBI; 2024. https://doi.org/10.46658/JBIMES-24-01.
  3. MacDonald H, Comer C, Foster M, et al. Searching for studies: a guide to information retrieval for Campbell systematic reviews. Campbell Systematic Reviews. 2024;20(3):e1433. https://doi.org/10.1002/cl2.1433.
  4. Frandsen TF, Nielsen MFB, Lindhardt CL, Eriksen MB. Using the full PICO model as a search tool for systematic reviews resulted in lower recall for some PICO elements. Journal of Clinical Epidemiology. 2020;127:69-75. https://doi.org/10.1016/j.jclinepi.2020.07.005.
  5. Bramer WM, de Jonge GB, Rethlefsen ML, Mast F, Kleijnen J. A systematic approach to searching. Journal of the Medical Library Association. 2018;106(4):531-541. https://doi.org/10.5195/jmla.2018.283.
  6. Clark JM, Sanders S, Carter M, et al. Improving the translation of search strategies using the Polyglot Search Translator. Journal of the Medical Library Association. 2020;108(2):195-207. https://doi.org/10.5195/jmla.2020.834.
  7. Lagisz M, et al. A practical guide to evaluating sensitivity of literature search strings for systematic reviews using relative recall. Research Synthesis Methods. 2025. https://doi.org/10.1017/rsm.2024.6.
  8. McGowan J, Sampson M, Salzwedel DM, Cogo E, Foerster V, Lefebvre C. PRESS Peer Review of Electronic Search Strategies: 2015 Guideline Statement. Journal of Clinical Epidemiology. 2016;75:40-46. https://doi.org/10.1016/j.jclinepi.2016.01.021.
  9. Rethlefsen ML, Kirtley S, Waffenschmidt S, et al. PRISMA-S: an extension to the PRISMA Statement for Reporting Literature Searches in Systematic Reviews. Systematic Reviews. 2021;10:39. https://doi.org/10.1186/s13643-020-01542-z.
  10. Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement. BMJ. 2021;372:n71. https://doi.org/10.1136/bmj.n71.
  11. Rethlefsen ML, et al. Systematic review search strategies are poorly reported and not reproducible. Journal of Clinical Epidemiology. 2024;166:111229. https://doi.org/10.1016/j.jclinepi.2023.111229.

Similar Posts

Leave a Reply

Your email address will not be published. Required fields are marked *