AI for Systematic Reviews: What Researchers Can Automate, What They Must Verify, and What They Must Report
AI can help with parts of a systematic review, but no current evidence supports handing it the review end to end. Reliability shifts by task, dataset, tool, prompt, and review context. Define each use in the protocol, test it against an appropriate human standard, verify any output that could change the evidence base, keep an audit trail, and leave the final methodological decisions with qualified people.12
A 2025 systematic review found that generative AI missed 68% to 96% of relevant studies across the included search evaluations. Search was the weakest task assessed.2
In a prospective study across six reviews, AI-first extraction followed by human verification saved time and matched human-only accuracy.3
ChatGPT-4o reached moderate weighted agreement with Cochrane judgements across 84 trials, yet sensitivity for high-risk studies was only 53%.5
These results come mainly from health-related review settings. They describe specific evaluations, not fixed performance rates for every model or discipline.
Ask what happens when the output is wrong
Speed is easy to measure. Trust is not. A fluent answer might save ten minutes and still damage a review if it drops an eligible study, copies the wrong denominator, or invents a reason for a risk-of-bias call. What the error costs depends on where it enters the workflow.
A weak suggestion during early brainstorming costs nothing to reject. A false exclusion during screening can remove a study from every analysis that follows. A wrong effect estimate can shift a pooled result. And a polished paragraph can hide either mistake, because confident prose tempts authors to stop checking.
That points to a practical rule: the closer an AI output sits to the evidence base or a judgement about that evidence, the stronger the human control needs to be. The rule holds across disciplines, even though the specific safeguard will still depend on the review’s design, its decision context, and what a wrong answer would actually cost.
The author remains responsible
Cochrane, Campbell, JBI, and the Collaboration for Environmental Evidence permit AI use only when review teams can show that it does not compromise methodological rigor or integrity. Their joint position requires human oversight and transparent reporting whenever AI makes or suggests a judgement.1
Automation, machine learning, and generative AI are different methods
Review teams often write that they “used AI” without saying what the system actually did. That phrase is too vague for a methods section. A deterministic deduplication rule, an active-learning screening classifier, and a large language model behave differently and fail differently.
Rule-based automation
Software follows explicit rules, such as matching identifiers or formatting records. Researchers can inspect the rules, exceptions, and logs. Errors usually trace back to incomplete rules or inconsistent source data.
Predictive machine learning
A model classifies or prioritizes records based on training examples. Screening performance depends on the review’s corpus, labels, stopping rule, and threshold. Good results on someone else’s review don’t guarantee good results on yours.
Generative AI and large language models
The system generates text or structured output from prompts and supplied material. It can draft, extract, summarize, or reason in language, but fluent output is no proof of retrieval completeness or factual accuracy.
Retrieval-augmented and agentic systems
These systems pull in external material or run a sequence of tasks on their own. Their performance hinges on source access, retrieval logic, orchestration, and the underlying models. More steps mean more places an error can slip through unnoticed.
A report should name the class of system, the product and version, its input, the task it performed, and the role given to its output. Product names alone make poor method descriptions, because a vendor can swap the model or workflow behind an unchanged interface.
AI has no single reliability level across a systematic review
The table below gives a conservative reading of current evidence and guidance. Select a control level to focus on the relevant rows; the full table stays visible and printable if the filter doesn’t run.
Human-control matrix
Treat this as a planning aid, not automatic permission. Your protocol, discipline, institutional policy, journal rules, and local validation still govern the review.
| Review stage | Defensible assistive role | Main failure | Minimum control | Evidence confidence |
|---|---|---|---|---|
| Question and scope | Generate candidate concepts or frameworks for discussion | Irrelevant, biased, or duplicate questions | Researchers define and approve the final question | Low |
| Protocol | Draft an outline or check whether planned fields are missing | Boilerplate methods and hidden assumptions | Authors approve every method and pre-specify AI use | Low |
| Search development | Suggest terms or produce a first query for expert revision | Missed concepts, poor translation, unstable retrieval | Information specialist or expert validates the full strategy | Moderate |
| Search execution | Explore a topic or supplement formal retrieval | Unknown coverage and inaccessible sources | Do not replace documented database and source searching | Low |
| Title and abstract screening | Prioritize records or support a reviewer under a tested rule | False exclusions, especially when abstracts omit decisive details | Validate recall and stopping rules in the review; audit exclusions | Moderate |
| Full-text eligibility | Extract passages that may help a reviewer apply criteria | Context loss and inconsistent criteria application | Humans make and document final eligibility decisions | Low |
| Data extraction | Prepare a structured first extraction for human checking | Wrong values, missed fields, and confusion across reports | Verify every field used in description or synthesis | Moderate |
| Risk of bias | Locate supporting text or prompt reviewers to inspect a domain | Invalid judgements when context is absent or ambiguous | Qualified reviewers make and reconcile judgements | Low |
| Statistical synthesis | Draft code or explain formulas for independent checking | Coding errors and unjustified analytical choices | Methodologist verifies code, data, model, and interpretation | Low |
| Certainty assessment | Organize evidence already judged by the team | Spurious ratings and invented rationales | Reviewers retain all GRADE or equivalent judgements | Very low |
| Interpretation and reporting | Edit language, restructure text, or draft plain-language versions | Overclaim, lost uncertainty, fabricated citations | Authors verify every claim and own the conclusions | Moderate |
| Living-review surveillance | Monitor new records and prioritize likely updates | Model drift and silent changes in coverage | Validate update rules and inspect new decisions | Low |
A low confidence rating is not a forecast
It describes the current evidence for the task, not a permanent ceiling on the technology. A newer model may do better. Researchers still need transparent evaluation before moving any task into routine use.
Generative search is useful for exploration, not proof of completeness
Search is where enthusiasm most clearly outruns the evidence. In Clark and colleagues’ systematic review of 19 studies, generative AI missed 68% to 96% of relevant studies in the search evaluations, with a median of 91%. Most included studies also carried a high or unclear risk of bias, or applicability concerns.2
An AI system can still help a researcher discover vocabulary, spot a candidate citation, or produce a query worth critiquing. Those uses belong at the exploratory edge of the search. A formal systematic search still needs transparent sources, reproducible strategies, tested retrieval, and documentation another researcher can inspect. PRISMA-S sets the reporting structure for that record.10
Access is its own constraint. A conversational system cannot reach into databases, subscription platforms, grey-literature sources, or local collections it has no access to. A polished list of citations says nothing about what stayed outside its view.
Do not use a chatbot answer as the systematic search
Use generative search to explore language or supplement established methods. Keep the formal evidence-identification process under a documented search plan with source-level records, deduplication procedures, update rules, and expert review where the question warrants it.
Workload reduction is credible only when missed studies are measured
Machine-learning screening has a longer track record than generative AI. Active-learning systems can reorder records so reviewers see likely inclusions earlier, which can cut effort, but the benefit depends on training decisions and the rule used to stop screening.
The dangerous error is a false exclusion. It removes a study before extraction, appraisal, or synthesis ever happens. Abstracts often leave out the detail that actually decides eligibility, so strong performance on one corpus may not carry over to another. A team using automated prioritization or exclusion should pre-specify its role, test it in the review, set an acceptable recall threshold, and examine what got excluded.
PRISMA 2020 already asks authors to explain how automation was integrated into study selection, name the classifier and version, report training and validation, and show records marked ineligible by automation in the flow diagram where applicable.9 Reporting a software name without the stopping rule or validation leaves the most consequential choice invisible to readers.
Verified extraction has the strongest practical case so far
A prospective study within six ongoing intervention reviews compared an AI-first, human-verified process with human-only extraction. The study covered 9,341 data elements from 63 studies. AI-assisted extraction reached 91.0% accuracy, against 89.0% for human-only extraction, and saved a median of 41 minutes per study. Incorrect values still turned up in 9.0% of AI-assisted cases; human verification was built into the method, not an optional cleanup step.3
A separate feasibility study tested a retrieval-augmented pipeline across experimental, observational, qualitative, and modelling studies. Researchers judged 68% of outputs acceptable overall. Objective fields such as setting and study design performed better than fields requiring interpretation across outcomes or time points.4
Together, these studies support a narrow conclusion: AI can prepare a first extraction for a human to check against the source. They do not support moving values into a meta-analysis unverified. Verification should cover every field that feeds a table, calculation, narrative synthesis, or conclusion, and multiple reports from one study still need study-level linking before either a person or a model can extract coherently.
Risk of bias, certainty, and interpretation still need qualified reviewers
Risk-of-bias tools ask reviewers to weigh article text, protocol information, trial conduct, and domain-specific signalling questions together. Missing information matters, and so does the reasoning behind a judgement.
In a 2025 comparison involving 84 randomized trials, ChatGPT-4o reached a weighted kappa of 0.51 for overall RoB 2 judgements against Cochrane consensus assessments. Agreement varied by domain. Sensitivity for identifying high-risk studies was only 53%, while specificity for low-risk studies reached 99%.5 That imbalance could reassure a review team exactly where scrutiny is needed most.
An AI system may locate relevant passages or organize a draft rationale, but the reviewer still has to inspect the source, apply the correct tool, and own the judgement. The same boundary applies to certainty assessment and interpretation. GRADE ratings, synthesis choices, and conclusions depend on the review question and the full evidence record, and a language model does not accept authorship, defend a judgement during peer review, or correct the published record.
Generate, Verify, Document
MetaSyn Academy teaches a three-part rule for AI-assisted work. GVD is an Academy framework, not an official reporting standard, but it gives a review team a repeatable way to keep assistance separate from evidence and judgement.
Generate
Give the system one bounded task, the relevant material, a required output format, and a stated prohibition against inventing unavailable information.
Verify
Check the output against primary sources, protocol criteria, verified data, and an appropriate human standard. Record the error types alongside the overall accuracy score.
Document
Preserve the tool, version, date, inputs, prompts, settings, validation results, corrections, overrides, and the name of the person who approved the output.
Verification must be capable of finding failure
Reading an AI answer and deciding it looks reasonable is not validation. Use source comparison, a labelled benchmark, duplicate human review, reproducible calculations, or another test that can actually expose the error relevant to the task.
Set the rule before seeing the result
Pre-specification closes off convenient exceptions. State the task, comparator, metric, acceptance threshold, checking level, and abandonment rule in the protocol. If the system falls short, fall back to the established method and record the deviation. The guide to protocol amendments and deviations explains how to preserve that decision trail.
Record enough detail for another researcher to understand the intervention
A reproducible AI method needs more than the sentence “ChatGPT was used.” Model versions change, interfaces hide settings, and the same prompt can return a different output tomorrow. The record should show what the system saw, what it produced, and what people changed afterward.
| Record | What to preserve | Why it matters |
|---|---|---|
| System identity | Product, provider, model, version, access route, and date used | The same product name may point to a different system later. |
| Intended task | Review stage, exact output, and whether the role was exploratory, assistive, or decision-supporting | Readers can judge the consequence of a possible error. |
| Inputs | Records, abstracts, full texts, tables, instructions, examples, and preprocessing | Performance depends on what the system could inspect. |
| Prompts and settings | Full prompts, templates, structured fields, temperature or equivalent controls, and repeated runs | The instructions are part of the method. |
| Validation | Comparator, sample, metrics, thresholds, error categories, and results | A reader can assess whether the test matched the claimed use. |
| Human review | Who checked the output, what proportion was checked, and how disagreements were resolved | “Human in the loop” has no meaning without the loop’s design. |
| Corrections | Rejected outputs, edits, overrides, failures, and workflow changes | Corrections reveal where the system was unreliable. |
| Governance | Data protection, confidentiality, licences, copyright controls, conflicts, and approvals | Methodological efficiency does not cancel legal or ethical duties. |
PRISMA 2020 requires details of automation in selection and data collection. PRISMA-S covers the literature-search record. The joint AI position and RAISE guidance add task-specific expectations for responsible use.1910
Use reporting guidance without claiming an endorsement that does not exist
PRISMA is a reporting guideline. It does not certify that an AI method was valid, and a vendor’s claim of “PRISMA compliance” is not an independent performance evaluation.
PRISMA-trAIce, published in 2025, proposes additional items for reporting AI use in systematic literature reviews, covering tool identity, inputs, outputs, human interaction, evaluation, and limitations. Its authors describe it as a foundational proposal and invite a formal consensus process; it has not yet completed that process or become a formally endorsed PRISMA extension.11
A team may use PRISMA-trAIce as an additional reporting aid. Call the checklist what it is, and avoid presenting the review as officially “PRISMA-trAIce compliant.” Official guidance can still change, so check the PRISMA extensions page and the target journal’s policies when preparing the final report.
Do not upload material until you know where it goes
Review teams work with licensed articles, unpublished data, peer-review material, author correspondence, and study-level information. A public AI interface may process or retain those materials under terms that differ by provider, account, product, and date.
Check the applicable contract and institutional policy before uploading anything. Record whether data are retained, used for model improvement, transferred across jurisdictions, or accessible to third-party services. Strip out personal or confidential information where you can. If confidentiality cannot be assured, use an approved protected environment, or don’t use the system for that material at all.
ICMJE warns that submitted manuscripts are privileged communications and that uploading them to an AI system can violate confidentiality. It also places responsibility for accuracy, attribution, permissions, and plagiarism on human authors.13 WAME likewise requires disclosure of substantive chatbot use and keeps public responsibility with human authors.14
Five decisions belong in the protocol
1. What exact task will the system perform?
Don’t authorize “AI for the review.” Name the stage, input, output, and downstream decision.
2. What error would damage the review?
Choose metrics that expose that error. Screening needs attention to missed studies; extraction needs field-level errors and missing values.
3. What evidence will justify use?
Use independent evaluations as background, then test performance in the intended review when transfer is uncertain or the stakes are high.
4. Who will check and approve the output?
Name qualified people, the proportion checked, the disagreement process, and the point at which the system gets abandoned.
5. What will readers be able to inspect?
Plan the log, validation report, prompts, versions, corrections, protocol amendments, and final disclosure before production begins.
The systematic review project setup guide covers the wider research infrastructure that should exist before searching begins. The timeline and milestones guide explains where review gates and re-estimation points belong. This editorial adds the AI-specific control layer; it does not replace either method.
Validate the workflow, not the reputation of a model
Leaderboards age fast. Review methods have to survive the next model release. A defensible protocol therefore avoids permanent claims that one named product is safe, accurate, or best.
Preserve a frozen test set where licensing permits it. Re-run the validation after a major model, prompt, retrieval, interface, or data-source change. Compare error categories alongside the average score, and if performance drops below the pre-specified threshold, stop or redesign the use.
Cochrane’s 2026 platform study makes the same point in practice. Two selected tools are being evaluated within reviews; selection is not endorsement. Cochrane advises authors to apply RAISE guidance to any tool and to demonstrate that its use does not compromise the synthesis.12
The durable advantage is an auditable method
A review team that can test, reject, correct, and report an AI-assisted process is ready for whatever comes next. A team that leans on a product’s reputation has to start over the moment that product changes.
Use AI where its work can be checked
Systematic reviews were never trustworthy simply because every task was manual. They are trustworthy when researchers can show how evidence entered the review, how decisions were made, and where errors were caught. AI does not change that standard.
Current evidence supports assistance in bounded tasks. Verified data extraction is the clearest example. Screening support can cut workload when teams validate performance and guard against false exclusions. Generative search remains a poor substitute for systematic retrieval. Risk-of-bias assessment, certainty judgements, synthesis choices, and conclusions still need qualified human control.
The honest position sits between refusal and surrender. Use the tool. Test it against the work. Keep the evidence chain visible. Then report enough detail that another researcher can understand what the machine did and what the authors decided.
Planning, documentation, and practical resources
A protocol should govern the AI before the AI touches the review
Course 1 teaches protocol development as a documented chain of decisions. Its AI-assisted workflow applies MetaSyn’s Generate, Verify, Document method inside that larger protocol process. Explore the systematic review and meta-analysis protocol course.
Frequently asked questions
Can I use ChatGPT in a systematic review?
Yes, for a defined task that your protocol, institution, data agreements, and target journal permit. Record the model and version, inputs, prompts, date, validation, human checks, and corrections. Don’t treat a general chatbot as a complete search, an autonomous reviewer, a source of unverified data, or an author.
Can AI replace a second reviewer?
No general rule supports replacing a second reviewer across systematic-review tasks. Some screening systems may support prioritization or a tested workflow, but performance must be validated in context and the method must satisfy the review standard being followed. Final eligibility and other judgement-bearing decisions remain human responsibilities under current guidance.
Which systematic review task has the strongest evidence for AI use?
Verified data extraction currently has one of the strongest practical cases. A prospective study within six reviews found that AI-first extraction followed by human verification saved a median of 41 minutes per study and reached accuracy similar to human-only extraction. The result supports an assistive workflow, not unverified extraction.
How should researchers report AI use in a systematic review?
Report the tool, provider, model and version, date, intended task, inputs, prompts or templates, settings, validation method and results, human checking, corrections, failures, protocol deviations, and relevant confidentiality or copyright controls. State who retained responsibility for each methodological decision.
Does PRISMA require reporting of AI and automation?
PRISMA 2020 asks authors to report automation used in study selection and data collection, including how tools were integrated, trained, and validated where applicable. PRISMA-S covers search reporting. AI-specific guidance is still developing, so researchers should also check current RAISE guidance and journal policies.
Is PRISMA-trAIce an officially endorsed PRISMA extension?
No. PRISMA-trAIce is a peer-reviewed proposal for reporting AI use in systematic literature reviews. Its authors invite further consensus work and formal endorsement. Researchers may use it as an additional checklist but should not describe themselves as meeting an official PRISMA-trAIce compliance standard.
Should an AI system be listed as an author?
No. ICMJE and WAME state that AI systems cannot meet authorship requirements. Human authors remain responsible for accuracy, attribution, permissions, disclosure, and the integrity of the published work.
Methodological guidance and empirical evidence
- Flemyng E, Noel-Storr A, Macura B, et al. Position Statement on Artificial Intelligence Use in Evidence Synthesis Across Cochrane, the Campbell Collaboration, JBI, and the Collaboration for Environmental Evidence. Campbell Syst Rev. 2025;21(4):e70074. doi:10.1002/cl2.70074
- Clark J, Barton B, Albarqouni L, et al. Generative artificial intelligence use in evidence synthesis: a systematic review. Res Synth Methods. 2025;16(4):601-619. doi:10.1017/rsm.2025.16
- Gartlehner G, Kugley S, Crotty K, et al. Artificial intelligence-assisted data extraction with a large language model: a study within reviews. Ann Intern Med. 2025;178(12):1763-1771. doi:10.7326/ANNALS-25-00739
- Simmons Z, Evans B, Harris T, et al. Assessing the feasibility and acceptability of a bespoke large language model pipeline to extract data from different study designs for public health evidence reviews. Cochrane Evid Synth Methods. 2025;3(6):e70061. doi:10.1002/cesm.70061
- Taneri PE. Human versus artificial intelligence: comparing Cochrane authors’ and ChatGPT’s risk of bias assessments. Cochrane Evid Synth Methods. 2025;3(5):e70044. doi:10.1002/cesm.70044
- Boetje J, van de Schoot R. The SAFE procedure: a practical stopping heuristic for active learning-based screening in systematic reviews and meta-analyses. Syst Rev. 2024;13:81. doi:10.1186/s13643-024-02502-7
- Scotti KL, et al. Artificial intelligence and automation in evidence synthesis: an investigation of methods employed in Cochrane, Campbell Collaboration, and Environmental Evidence reviews. Cochrane Evid Synth Methods. 2025;3(5):e70046. doi:10.1002/cesm.70046
- Gartlehner G, Nussbaumer-Streit B, Hamel C, et al. Responsible integration of artificial intelligence in rapid reviews: a position statement from the Cochrane Rapid Reviews Methods Group. Cochrane Evid Synth Methods. 2025;3(6):e70063. doi:10.1002/cesm.70063
- Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71. doi:10.1136/bmj.n71
- Rethlefsen ML, Kirtley S, Waffenschmidt S, et al. PRISMA-S: an extension to the PRISMA Statement for reporting literature searches in systematic reviews. Syst Rev. 2021;10:39. doi:10.1186/s13643-020-01542-z
- Holst D, Moenck K, Koch J, Schmedemann O, Schuppstuhl T. Transparent reporting of AI in systematic literature reviews: development of the PRISMA-trAIce checklist. JMIR AI. 2025;4:e80247. doi:10.2196/80247
- Cochrane. Cochrane announces selected AI tools for innovative platform study. March 16, 2026. Current online guidance and study notice
- International Committee of Medical Journal Editors. Use of artificial intelligence in publishing. Recommendations for the Conduct, Reporting, Editing, and Publication of Scholarly Work in Medical Journals. Accessed July 20, 2026. Current online recommendation
- Zielinski C, Winker MA, Aggarwal R, et al. Chatbots, generative AI, and scholarly manuscripts: WAME recommendations. World Association of Medical Editors. Revised May 31, 2023. Accessed July 20, 2026. Current online recommendation