AI for Systematic Reviews: What Researchers Can Automate, What They Must Verify, and What They Must Report

Human-led AI workflow for systematic reviews showing bounded automation, researcher verification and transparent reporting.

AI can help with parts of a systematic review, but no current evidence supports handing it the review end to end. Reliability shifts by task, dataset, tool, prompt, and review context. Define each use in the protocol, test it against an appropriate human standard, verify any output that could change the evidence base, keep an audit trail, and leave the final methodological decisions with qualified people.12

91% Median share of studies missed in search evaluations

A 2025 systematic review found that generative AI missed 68% to 96% of relevant studies across the included search evaluations. Search was the weakest task assessed.2

41 min Median extraction time saved per study

In a prospective study across six reviews, AI-first extraction followed by human verification saved time and matched human-only accuracy.3

0.51 Agreement for overall risk-of-bias judgements

ChatGPT-4o reached moderate weighted agreement with Cochrane judgements across 84 trials, yet sensitivity for high-risk studies was only 53%.5

These results come mainly from health-related review settings. They describe specific evaluations, not fixed performance rates for every model or discipline.

The governing question

Ask what happens when the output is wrong

Speed is easy to measure. Trust is not. A fluent answer might save ten minutes and still damage a review if it drops an eligible study, copies the wrong denominator, or invents a reason for a risk-of-bias call. What the error costs depends on where it enters the workflow.

A weak suggestion during early brainstorming costs nothing to reject. A false exclusion during screening can remove a study from every analysis that follows. A wrong effect estimate can shift a pooled result. And a polished paragraph can hide either mistake, because confident prose tempts authors to stop checking.

That points to a practical rule: the closer an AI output sits to the evidence base or a judgement about that evidence, the stronger the human control needs to be. The rule holds across disciplines, even though the specific safeguard will still depend on the review’s design, its decision context, and what a wrong answer would actually cost.

The author remains responsible

Cochrane, Campbell, JBI, and the Collaboration for Environmental Evidence permit AI use only when review teams can show that it does not compromise methodological rigor or integrity. Their joint position requires human oversight and transparent reporting whenever AI makes or suggests a judgement.1

Name the technology

Automation, machine learning, and generative AI are different methods

Review teams often write that they “used AI” without saying what the system actually did. That phrase is too vague for a methods section. A deterministic deduplication rule, an active-learning screening classifier, and a large language model behave differently and fail differently.

Rule-based automation

Software follows explicit rules, such as matching identifiers or formatting records. Researchers can inspect the rules, exceptions, and logs. Errors usually trace back to incomplete rules or inconsistent source data.

Predictive machine learning

A model classifies or prioritizes records based on training examples. Screening performance depends on the review’s corpus, labels, stopping rule, and threshold. Good results on someone else’s review don’t guarantee good results on yours.

Generative AI and large language models

The system generates text or structured output from prompts and supplied material. It can draft, extract, summarize, or reason in language, but fluent output is no proof of retrieval completeness or factual accuracy.

Retrieval-augmented and agentic systems

These systems pull in external material or run a sequence of tasks on their own. Their performance hinges on source access, retrieval logic, orchestration, and the underlying models. More steps mean more places an error can slip through unnoticed.

A report should name the class of system, the product and version, its input, the task it performed, and the role given to its output. Product names alone make poor method descriptions, because a vendor can swap the model or workflow behind an unchanged interface.

Lifecycle evidence map

AI has no single reliability level across a systematic review

The table below gives a conservative reading of current evidence and guidance. Select a control level to focus on the relevant rows; the full table stays visible and printable if the filter doesn’t run.

Human-control matrix

Treat this as a planning aid, not automatic permission. Your protocol, discipline, institutional policy, journal rules, and local validation still govern the review.

Review stageDefensible assistive roleMain failureMinimum controlEvidence confidence
Question and scopeGenerate candidate concepts or frameworks for discussionIrrelevant, biased, or duplicate questionsResearchers define and approve the final questionLow
ProtocolDraft an outline or check whether planned fields are missingBoilerplate methods and hidden assumptionsAuthors approve every method and pre-specify AI useLow
Search developmentSuggest terms or produce a first query for expert revisionMissed concepts, poor translation, unstable retrievalInformation specialist or expert validates the full strategyModerate
Search executionExplore a topic or supplement formal retrievalUnknown coverage and inaccessible sourcesDo not replace documented database and source searchingLow
Title and abstract screeningPrioritize records or support a reviewer under a tested ruleFalse exclusions, especially when abstracts omit decisive detailsValidate recall and stopping rules in the review; audit exclusionsModerate
Full-text eligibilityExtract passages that may help a reviewer apply criteriaContext loss and inconsistent criteria applicationHumans make and document final eligibility decisionsLow
Data extractionPrepare a structured first extraction for human checkingWrong values, missed fields, and confusion across reportsVerify every field used in description or synthesisModerate
Risk of biasLocate supporting text or prompt reviewers to inspect a domainInvalid judgements when context is absent or ambiguousQualified reviewers make and reconcile judgementsLow
Statistical synthesisDraft code or explain formulas for independent checkingCoding errors and unjustified analytical choicesMethodologist verifies code, data, model, and interpretationLow
Certainty assessmentOrganize evidence already judged by the teamSpurious ratings and invented rationalesReviewers retain all GRADE or equivalent judgementsVery low
Interpretation and reportingEdit language, restructure text, or draft plain-language versionsOverclaim, lost uncertainty, fabricated citationsAuthors verify every claim and own the conclusionsModerate
Living-review surveillanceMonitor new records and prioritize likely updatesModel drift and silent changes in coverageValidate update rules and inspect new decisionsLow

A low confidence rating is not a forecast

It describes the current evidence for the task, not a permanent ceiling on the technology. A newer model may do better. Researchers still need transparent evaluation before moving any task into routine use.

Searching

Search is where enthusiasm most clearly outruns the evidence. In Clark and colleagues’ systematic review of 19 studies, generative AI missed 68% to 96% of relevant studies in the search evaluations, with a median of 91%. Most included studies also carried a high or unclear risk of bias, or applicability concerns.2

An AI system can still help a researcher discover vocabulary, spot a candidate citation, or produce a query worth critiquing. Those uses belong at the exploratory edge of the search. A formal systematic search still needs transparent sources, reproducible strategies, tested retrieval, and documentation another researcher can inspect. PRISMA-S sets the reporting structure for that record.10

Access is its own constraint. A conversational system cannot reach into databases, subscription platforms, grey-literature sources, or local collections it has no access to. A polished list of citations says nothing about what stayed outside its view.

Do not use a chatbot answer as the systematic search

Use generative search to explore language or supplement established methods. Keep the formal evidence-identification process under a documented search plan with source-level records, deduplication procedures, update rules, and expert review where the question warrants it.

Screening

Workload reduction is credible only when missed studies are measured

Machine-learning screening has a longer track record than generative AI. Active-learning systems can reorder records so reviewers see likely inclusions earlier, which can cut effort, but the benefit depends on training decisions and the rule used to stop screening.

The dangerous error is a false exclusion. It removes a study before extraction, appraisal, or synthesis ever happens. Abstracts often leave out the detail that actually decides eligibility, so strong performance on one corpus may not carry over to another. A team using automated prioritization or exclusion should pre-specify its role, test it in the review, set an acceptable recall threshold, and examine what got excluded.

PRISMA 2020 already asks authors to explain how automation was integrated into study selection, name the classifier and version, report training and validation, and show records marked ineligible by automation in the flow diagram where applicable.9 Reporting a software name without the stopping rule or validation leaves the most consequential choice invisible to readers.

Data extraction

Verified extraction has the strongest practical case so far

A prospective study within six ongoing intervention reviews compared an AI-first, human-verified process with human-only extraction. The study covered 9,341 data elements from 63 studies. AI-assisted extraction reached 91.0% accuracy, against 89.0% for human-only extraction, and saved a median of 41 minutes per study. Incorrect values still turned up in 9.0% of AI-assisted cases; human verification was built into the method, not an optional cleanup step.3

A separate feasibility study tested a retrieval-augmented pipeline across experimental, observational, qualitative, and modelling studies. Researchers judged 68% of outputs acceptable overall. Objective fields such as setting and study design performed better than fields requiring interpretation across outcomes or time points.4

Together, these studies support a narrow conclusion: AI can prepare a first extraction for a human to check against the source. They do not support moving values into a meta-analysis unverified. Verification should cover every field that feeds a table, calculation, narrative synthesis, or conclusion, and multiple reports from one study still need study-level linking before either a person or a model can extract coherently.

Methodological judgement

Risk of bias, certainty, and interpretation still need qualified reviewers

Risk-of-bias tools ask reviewers to weigh article text, protocol information, trial conduct, and domain-specific signalling questions together. Missing information matters, and so does the reasoning behind a judgement.

In a 2025 comparison involving 84 randomized trials, ChatGPT-4o reached a weighted kappa of 0.51 for overall RoB 2 judgements against Cochrane consensus assessments. Agreement varied by domain. Sensitivity for identifying high-risk studies was only 53%, while specificity for low-risk studies reached 99%.5 That imbalance could reassure a review team exactly where scrutiny is needed most.

An AI system may locate relevant passages or organize a draft rationale, but the reviewer still has to inspect the source, apply the correct tool, and own the judgement. The same boundary applies to certainty assessment and interpretation. GRADE ratings, synthesis choices, and conclusions depend on the review question and the full evidence record, and a language model does not accept authorship, defend a judgement during peer review, or correct the published record.

MetaSyn working method

Generate, Verify, Document

MetaSyn Academy teaches a three-part rule for AI-assisted work. GVD is an Academy framework, not an official reporting standard, but it gives a review team a repeatable way to keep assistance separate from evidence and judgement.

G

Generate

Give the system one bounded task, the relevant material, a required output format, and a stated prohibition against inventing unavailable information.

V

Verify

Check the output against primary sources, protocol criteria, verified data, and an appropriate human standard. Record the error types alongside the overall accuracy score.

D

Document

Preserve the tool, version, date, inputs, prompts, settings, validation results, corrections, overrides, and the name of the person who approved the output.

Verification must be capable of finding failure

Reading an AI answer and deciding it looks reasonable is not validation. Use source comparison, a labelled benchmark, duplicate human review, reproducible calculations, or another test that can actually expose the error relevant to the task.

Set the rule before seeing the result

Pre-specification closes off convenient exceptions. State the task, comparator, metric, acceptance threshold, checking level, and abandonment rule in the protocol. If the system falls short, fall back to the established method and record the deviation. The guide to protocol amendments and deviations explains how to preserve that decision trail.

Audit trail

Record enough detail for another researcher to understand the intervention

A reproducible AI method needs more than the sentence “ChatGPT was used.” Model versions change, interfaces hide settings, and the same prompt can return a different output tomorrow. The record should show what the system saw, what it produced, and what people changed afterward.

RecordWhat to preserveWhy it matters
System identityProduct, provider, model, version, access route, and date usedThe same product name may point to a different system later.
Intended taskReview stage, exact output, and whether the role was exploratory, assistive, or decision-supportingReaders can judge the consequence of a possible error.
InputsRecords, abstracts, full texts, tables, instructions, examples, and preprocessingPerformance depends on what the system could inspect.
Prompts and settingsFull prompts, templates, structured fields, temperature or equivalent controls, and repeated runsThe instructions are part of the method.
ValidationComparator, sample, metrics, thresholds, error categories, and resultsA reader can assess whether the test matched the claimed use.
Human reviewWho checked the output, what proportion was checked, and how disagreements were resolved“Human in the loop” has no meaning without the loop’s design.
CorrectionsRejected outputs, edits, overrides, failures, and workflow changesCorrections reveal where the system was unreliable.
GovernanceData protection, confidentiality, licences, copyright controls, conflicts, and approvalsMethodological efficiency does not cancel legal or ethical duties.

PRISMA 2020 requires details of automation in selection and data collection. PRISMA-S covers the literature-search record. The joint AI position and RAISE guidance add task-specific expectations for responsible use.1910

Standards and proposals

Use reporting guidance without claiming an endorsement that does not exist

PRISMA is a reporting guideline. It does not certify that an AI method was valid, and a vendor’s claim of “PRISMA compliance” is not an independent performance evaluation.

PRISMA-trAIce, published in 2025, proposes additional items for reporting AI use in systematic literature reviews, covering tool identity, inputs, outputs, human interaction, evaluation, and limitations. Its authors describe it as a foundational proposal and invite a formal consensus process; it has not yet completed that process or become a formally endorsed PRISMA extension.11

A team may use PRISMA-trAIce as an additional reporting aid. Call the checklist what it is, and avoid presenting the review as officially “PRISMA-trAIce compliant.” Official guidance can still change, so check the PRISMA extensions page and the target journal’s policies when preparing the final report.

Data governance

Do not upload material until you know where it goes

Review teams work with licensed articles, unpublished data, peer-review material, author correspondence, and study-level information. A public AI interface may process or retain those materials under terms that differ by provider, account, product, and date.

Check the applicable contract and institutional policy before uploading anything. Record whether data are retained, used for model improvement, transferred across jurisdictions, or accessible to third-party services. Strip out personal or confidential information where you can. If confidentiality cannot be assured, use an approved protected environment, or don’t use the system for that material at all.

ICMJE warns that submitted manuscripts are privileged communications and that uploading them to an AI system can violate confidentiality. It also places responsibility for accuracy, attribution, permissions, and plagiarism on human authors.13 WAME likewise requires disclosure of substantive chatbot use and keeps public responsibility with human authors.14

Before adoption

Five decisions belong in the protocol

1. What exact task will the system perform?

Don’t authorize “AI for the review.” Name the stage, input, output, and downstream decision.

2. What error would damage the review?

Choose metrics that expose that error. Screening needs attention to missed studies; extraction needs field-level errors and missing values.

3. What evidence will justify use?

Use independent evaluations as background, then test performance in the intended review when transfer is uncertain or the stakes are high.

4. Who will check and approve the output?

Name qualified people, the proportion checked, the disagreement process, and the point at which the system gets abandoned.

5. What will readers be able to inspect?

Plan the log, validation report, prompts, versions, corrections, protocol amendments, and final disclosure before production begins.

The systematic review project setup guide covers the wider research infrastructure that should exist before searching begins. The timeline and milestones guide explains where review gates and re-estimation points belong. This editorial adds the AI-specific control layer; it does not replace either method.

A method that can age well

Validate the workflow, not the reputation of a model

Leaderboards age fast. Review methods have to survive the next model release. A defensible protocol therefore avoids permanent claims that one named product is safe, accurate, or best.

Preserve a frozen test set where licensing permits it. Re-run the validation after a major model, prompt, retrieval, interface, or data-source change. Compare error categories alongside the average score, and if performance drops below the pre-specified threshold, stop or redesign the use.

Cochrane’s 2026 platform study makes the same point in practice. Two selected tools are being evaluated within reviews; selection is not endorsement. Cochrane advises authors to apply RAISE guidance to any tool and to demonstrate that its use does not compromise the synthesis.12

The durable advantage is an auditable method

A review team that can test, reject, correct, and report an AI-assisted process is ready for whatever comes next. A team that leans on a product’s reputation has to start over the moment that product changes.

Editorial conclusion

Use AI where its work can be checked

Systematic reviews were never trustworthy simply because every task was manual. They are trustworthy when researchers can show how evidence entered the review, how decisions were made, and where errors were caught. AI does not change that standard.

Current evidence supports assistance in bounded tasks. Verified data extraction is the clearest example. Screening support can cut workload when teams validate performance and guard against false exclusions. Generative search remains a poor substitute for systematic retrieval. Risk-of-bias assessment, certainty judgements, synthesis choices, and conclusions still need qualified human control.

The honest position sits between refusal and surrender. Use the tool. Test it against the work. Keep the evidence chain visible. Then report enough detail that another researcher can understand what the machine did and what the authors decided.

Continue through the Academy
Guided application

A protocol should govern the AI before the AI touches the review

Course 1 teaches protocol development as a documented chain of decisions. Its AI-assisted workflow applies MetaSyn’s Generate, Verify, Document method inside that larger protocol process. Explore the systematic review and meta-analysis protocol course.

Questions researchers ask

Frequently asked questions

Can I use ChatGPT in a systematic review?

Yes, for a defined task that your protocol, institution, data agreements, and target journal permit. Record the model and version, inputs, prompts, date, validation, human checks, and corrections. Don’t treat a general chatbot as a complete search, an autonomous reviewer, a source of unverified data, or an author.

Can AI replace a second reviewer?

No general rule supports replacing a second reviewer across systematic-review tasks. Some screening systems may support prioritization or a tested workflow, but performance must be validated in context and the method must satisfy the review standard being followed. Final eligibility and other judgement-bearing decisions remain human responsibilities under current guidance.

Which systematic review task has the strongest evidence for AI use?

Verified data extraction currently has one of the strongest practical cases. A prospective study within six reviews found that AI-first extraction followed by human verification saved a median of 41 minutes per study and reached accuracy similar to human-only extraction. The result supports an assistive workflow, not unverified extraction.

How should researchers report AI use in a systematic review?

Report the tool, provider, model and version, date, intended task, inputs, prompts or templates, settings, validation method and results, human checking, corrections, failures, protocol deviations, and relevant confidentiality or copyright controls. State who retained responsibility for each methodological decision.

Does PRISMA require reporting of AI and automation?

PRISMA 2020 asks authors to report automation used in study selection and data collection, including how tools were integrated, trained, and validated where applicable. PRISMA-S covers search reporting. AI-specific guidance is still developing, so researchers should also check current RAISE guidance and journal policies.

Is PRISMA-trAIce an officially endorsed PRISMA extension?

No. PRISMA-trAIce is a peer-reviewed proposal for reporting AI use in systematic literature reviews. Its authors invite further consensus work and formal endorsement. Researchers may use it as an additional checklist but should not describe themselves as meeting an official PRISMA-trAIce compliance standard.

Should an AI system be listed as an author?

No. ICMJE and WAME state that AI systems cannot meet authorship requirements. Human authors remain responsible for accuracy, attribution, permissions, disclosure, and the integrity of the published work.

References

Methodological guidance and empirical evidence

  1. Flemyng E, Noel-Storr A, Macura B, et al. Position Statement on Artificial Intelligence Use in Evidence Synthesis Across Cochrane, the Campbell Collaboration, JBI, and the Collaboration for Environmental Evidence. Campbell Syst Rev. 2025;21(4):e70074. doi:10.1002/cl2.70074
  2. Clark J, Barton B, Albarqouni L, et al. Generative artificial intelligence use in evidence synthesis: a systematic review. Res Synth Methods. 2025;16(4):601-619. doi:10.1017/rsm.2025.16
  3. Gartlehner G, Kugley S, Crotty K, et al. Artificial intelligence-assisted data extraction with a large language model: a study within reviews. Ann Intern Med. 2025;178(12):1763-1771. doi:10.7326/ANNALS-25-00739
  4. Simmons Z, Evans B, Harris T, et al. Assessing the feasibility and acceptability of a bespoke large language model pipeline to extract data from different study designs for public health evidence reviews. Cochrane Evid Synth Methods. 2025;3(6):e70061. doi:10.1002/cesm.70061
  5. Taneri PE. Human versus artificial intelligence: comparing Cochrane authors’ and ChatGPT’s risk of bias assessments. Cochrane Evid Synth Methods. 2025;3(5):e70044. doi:10.1002/cesm.70044
  6. Boetje J, van de Schoot R. The SAFE procedure: a practical stopping heuristic for active learning-based screening in systematic reviews and meta-analyses. Syst Rev. 2024;13:81. doi:10.1186/s13643-024-02502-7
  7. Scotti KL, et al. Artificial intelligence and automation in evidence synthesis: an investigation of methods employed in Cochrane, Campbell Collaboration, and Environmental Evidence reviews. Cochrane Evid Synth Methods. 2025;3(5):e70046. doi:10.1002/cesm.70046
  8. Gartlehner G, Nussbaumer-Streit B, Hamel C, et al. Responsible integration of artificial intelligence in rapid reviews: a position statement from the Cochrane Rapid Reviews Methods Group. Cochrane Evid Synth Methods. 2025;3(6):e70063. doi:10.1002/cesm.70063
  9. Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71. doi:10.1136/bmj.n71
  10. Rethlefsen ML, Kirtley S, Waffenschmidt S, et al. PRISMA-S: an extension to the PRISMA Statement for reporting literature searches in systematic reviews. Syst Rev. 2021;10:39. doi:10.1186/s13643-020-01542-z
  11. Holst D, Moenck K, Koch J, Schmedemann O, Schuppstuhl T. Transparent reporting of AI in systematic literature reviews: development of the PRISMA-trAIce checklist. JMIR AI. 2025;4:e80247. doi:10.2196/80247
  12. Cochrane. Cochrane announces selected AI tools for innovative platform study. March 16, 2026. Current online guidance and study notice
  13. International Committee of Medical Journal Editors. Use of artificial intelligence in publishing. Recommendations for the Conduct, Reporting, Editing, and Publication of Scholarly Work in Medical Journals. Accessed July 20, 2026. Current online recommendation
  14. Zielinski C, Winker MA, Aggarwal R, et al. Chatbots, generative AI, and scholarly manuscripts: WAME recommendations. World Association of Medical Editors. Revised May 31, 2023. Accessed July 20, 2026. Current online recommendation
Evidence note: This editorial prioritizes current organizational guidance, reporting standards, systematic reviews, and comparative evaluations. Much of the performance evidence comes from health-related reviews and rapidly changing systems. Numerical results are presented with their study context and should not be transferred to another task, model, field, or review without validation.

Similar Posts

Leave a Reply

Your email address will not be published. Required fields are marked *