Professional Evidence Synthesis in the Age of AI: Skills, Accountability and Human Judgment
The professional question is no longer whether AI can assist
Evidence checked: 12 August 2026
AI can already search, classify, extract, draft, code and coordinate sequences of research tasks. That changes professional evidence synthesis, but not in the simple way suggested by either extreme. The credible position is neither “AI can run the review” nor “a human must redo everything the machine touches.” The professional problem is deciding what can be delegated, what evidence is needed before delegation, where verification belongs, and who remains answerable when the result is wrong.
The current cross-organization position is unusually clear on the last point. Cochrane, the Campbell Collaboration, JBI and the Collaboration for Environmental Evidence state that evidence synthesists remain ultimately responsible for their synthesis, including the decision to use AI. Their 2025 position statement permits AI and automation when methodological rigor and integrity are protected, requires human oversight, and calls for transparent reporting when AI makes or suggests judgments.1
That is a professional standard of responsibility, not a claim that every output must be reread line by line. The associated RAISE initiative separates responsibilities across evidence synthesists, methodologists, tool developers, organizations, publishers and other actors, and provides separate guidance on developing, evaluating, selecting and using AI tools.2 The framework is guidance. It does not create a licence, statutory duty or universal numerical performance threshold.
This article therefore sits inside the wider Professional Practice in Evidence Synthesis collection. Its subject is not which button to press. It is what professional competence looks like when part of the work is performed by systems whose outputs may be useful, unstable, difficult to inspect or deceptively plausible.
Figure 1. The responsibility stack is an evidence-informed conceptual synthesis. The formal source supports human oversight and ultimate responsibility; the division into execution, operational control and professional accountability is used here to make those responsibilities visible.12
Accountability does not mean doing every task yourself
Professional accountability is compatible with delegation. Review teams have always divided work among information specialists, statisticians, subject experts, junior reviewers and senior methodologists. AI adds another source of delegated execution, but unlike a human collaborator it cannot accept authorship responsibility, explain its conduct during peer review, correct the published record or be professionally sanctioned.
ICMJE’s current guidance makes the authorship boundary explicit. AI-assisted technologies should not be listed as authors because they cannot take responsibility for accuracy, integrity and originality. Human authors remain responsible for submitted material that involved AI, must review AI-generated content, and must ensure appropriate attribution and absence of plagiarism.3 ICMJE also requires disclosure of AI-assisted technologies used in producing submitted work. Importantly, it says nondisclosure may require corrective action and may be considered misconduct in some circumstances. That is more precise than claiming that every undisclosed use is automatically research misconduct.
The professional consequence is straightforward. When AI contributes to a review, authors need to know enough about its role to defend the decision to use it, understand the evidence supporting that use, and explain what controls were applied. “The software did it” is not a methodological rationale.
Human oversight should follow the risk, not a ritual
“Human in the loop” sounds reassuring but says almost nothing. A human can be present and still fail to detect the important error. Effective oversight depends on what the AI is doing, how visible its mistakes are, how far an error can propagate, and what evidence already exists for the tool in that task and context.
Current empirical evidence strongly supports task-specific caution. A 2026 benchmark of LLMs extracting data from full-text randomized trials found high precision but consistently incomplete recall. Performance changed substantially by information type and prompting strategy, leading the authors to recommend different automation levels according to task complexity and risk rather than one universal rule.4 A systematic review of generative AI across evidence-synthesis activities likewise found large differences across tasks, tools and study settings.5
This supports a risk-proportionate professional model. It is a synthesis of current guidance and empirical findings, not an official five-level standard. The question is not “Did a human look at it?” The question is “Was the control capable of finding the error that matters?”
Figure 2. The USE–VERIFY–VALIDATE–ESCALATE–STOP sequence is an evidence-informed conceptual translation of current risk-sensitive guidance, not a formally validated evidence-synthesis standard. It avoids the equally weak assumptions that all AI outputs require identical checking or that validation elsewhere automatically transfers to the current review.
Validation and verification solve different problems
The distinction is central to professional competence. Validation asks whether a tool or AI-assisted workflow is fit for an intended purpose. It examines performance across a defined set of cases, ideally using an appropriate reference standard. Verification asks whether a particular output is correct: whether this study was eligible, this value was extracted accurately, this citation exists, or this statistical result matches the data.
A tool can be well validated and still produce a wrong individual output. Conversely, checking a handful of outputs does not establish that a system is valid across the whole corpus. Local calibration may be needed when the review’s terminology, eligible-study prevalence, document structure or decision rules differ materially from the evidence used to evaluate the tool. Monitoring becomes relevant when the process is repeated over time, especially in living or continuously updated evidence systems.
No current evidence supports one universal sensitivity, recall or error threshold for every AI-assisted review task. A screening workflow, extraction system and code-generation assistant create different errors with different consequences. Acceptability therefore needs to be prespecified against the risk of the specific task rather than borrowed from a memorable percentage.
Figure 3. Validation is workflow- or system-level evidence; verification concerns specific outputs. The terminology is used here in an operational sense consistent with current AI evaluation guidance. The appropriate combination depends on the task and risk.24
Professional AI competence is more than prompt writing
The most durable professional skills are unlikely to be tied to the syntax of one chatbot interface. Model names, prompt techniques and product features change too quickly. The more defensible competency architecture has three layers.
Baseline AI literacy
Understand that performance is task- and context-dependent. Recognize hallucination and omission risk. Know that fluent output is not evidence of accuracy. Protect sensitive information, understand the need for disclosure, and know when a tool’s limitations require escalation.
Role-relevant capability
Be able to evaluate whether an AI-supported method fits a particular review task, design or interpret a pilot evaluation, define verification procedures, document exceptions, and supervise AI-supported work that falls within your professional responsibility.
Specialist AI-methods competence
Design or benchmark classifiers, extraction systems, APIs, retrieval pipelines or agents; evaluate error distributions and transferability; build auditable workflows; and advise teams when ordinary review expertise is not enough to evaluate the computational method.
This model does not require every evidence synthesist to become a software engineer. It requires people to understand enough about the systems they use to recognize when the methodological responsibility exceeds their own competence.
Evidence Synthesis Professional Competency and Career Roadmap Turn AI literacy, validation, supervision and professional judgment into explicit development priorities, alongside the rest of your evidence-synthesis capability profile.Human judgment needs a better definition than “humans are better”
Some judgments remain human because they are normative. Recommendation panels may weigh values, acceptability, equity or trade-offs that cannot be reduced to extracting the statistically most likely answer. Other judgments remain human because context matters: whether populations are sufficiently similar, whether a methodological departure changes interpretation, or whether a client’s requested shortcut would make the evidence product misleading.
A third group is different. AI may currently be weak at a task, but that does not prove the task is inherently human. The 2026 data-extraction evidence is a good example. Current models still omit or confuse important information, yet performance is improving and differs substantially by field type and prompt design.4 Those are current technological limits, not philosophical boundaries.
Professional writing should preserve this distinction. “AI cannot exercise methodological judgment” is too absolute. A system can already apply some explicit rules or detect familiar methodological features. The stronger claim is that consequential contextual and normative judgments still require qualified people to own the interpretation, especially where the system’s performance has not been demonstrated.
Agentic evidence synthesis raises a different supervision problem
Agentic systems do more than generate one response. They may plan a sequence, retrieve papers, call search or code tools, pass outputs to other agents, revise intermediate results and decide which step comes next. That changes the failure surface. The professional is no longer checking one answer; they may be supervising a chain of transformations.
The evidence is moving quickly. A 2026 systematic review of automated meta-analysis documented substantial expansion of automation across the review process while emphasizing continuing integration and reliability challenges.8 Peer-reviewed work on retrieval-augmented scientific literature synthesis also shows that agents can retrieve and synthesize scientific material with far stronger grounding than an ordinary unaided language model, but scientific question answering is not equivalent to a protocol-driven systematic review.9
Evidence-synthesis-specific agentic results are more preliminary. The 2026 AgentSLR preprint reports substantial acceleration in specialized epidemiological reviews, but its evaluation includes human-in-the-loop validation and should not be interpreted as evidence for unmonitored autonomy.10 A July 2026 AutoSynthesis preprint goes further toward an end-to-end meta-analysis workflow, yet recovered only 71.4% of studies from its human-conducted benchmark.11 That is evidence of rapid technical progress. It is not evidence that publication-grade autonomous systematic reviews are solved.
Figure 4. The error cascade is a conceptual risk model. Agentic evidence-synthesis research is still emerging, so the figure illustrates a plausible propagation mechanism rather than a quantified probability of failure. Current end-to-end evaluations nevertheless show why final-output inspection alone cannot establish completeness.1011
Supervision must protect independent reasoning
AI can improve a task and still make supervision harder. Automation-bias research predates generative AI and shows that people can over-rely on computer-generated advice, including accepting incorrect recommendations or reducing independent information seeking.12 The direct evidence in systematic reviewing is still limited, so it would be too strong to claim that AI is already deskilling evidence synthesists.
The training risk is nevertheless credible. A junior reviewer who sees an AI-generated eligibility decision before applying the criteria independently may learn to verify a suggestion rather than construct the judgment. The same problem can occur with prewritten risk-of-bias rationales or statistical code. Supervisors therefore need to decide not only whether the final answer is correct, but whether the learning design still gives trainees enough independent practice to understand how the answer is reached.
The opposite possibility also matters. AI can create practice examples, explain code, surface alternative interpretations and give rapid feedback. The professional teaching question is not “AI or no AI.” It is whether the technology is strengthening reasoning or quietly replacing the experience through which that reasoning develops.
Data governance begins before the upload
Evidence-synthesis teams may handle licensed full text, unpublished manuscripts, confidential sponsor documents, individual participant data, peer-review material and proprietary analyses. The relevant question is not whether a product calls itself “enterprise AI.” The team needs to know the actual contractual and technical conditions: retention, provider training use, access control, jurisdiction, logging, subprocessors and institutional approval.
ICMJE explicitly warns that using AI in manuscript handling can violate confidentiality and places responsibility on people to protect privileged material.3 The same professional reasoning applies more broadly. If the user does not have authority to disclose a document to an external processor, access to the document does not itself create permission to upload it.
Copyright, licences, privacy law, ethics approvals and client contracts are different constraints. They should not be collapsed into a universal rule such as “subscription PDFs can never be processed by AI.” The defensible rule is narrower: inspect the applicable licence, data agreement, institutional policy and legal context before transferring protected material to another system.
Governance, Ethics & Stewardship Place individual professional responsibility within the wider institutional framework for research integrity, responsible technology, stewardship and accountable evidence practice.Transparency should show what the AI was allowed to influence
Reporting requirements are evolving, but the direction is consistent. The joint position statement requires full and transparent reporting when AI or automation makes or suggests judgments.1 ICMJE requires authors to disclose whether AI-assisted technologies were used in producing submitted work and to explain how they were used.3 RAISE provides more detailed recommendations for selecting, evaluating and describing AI-supported methods.2
A useful methods record therefore identifies the system, the task, the material it could inspect, the role given to its output, the relevant validation evidence, the human control and the corrections that followed. More detailed reporting may also include prompts, settings, code and local evaluation. Collaboration for Environmental Evidence now provides living AI-reporting guidance that asks for tool identity, purpose, validation, limitations, ethical considerations and, for prompt-based systems, the prompts used.13 That is sector-specific guidance, not proof that every journal currently requires the identical reporting fields.
This page does not replace the workflow-level AI guide
The professional question and the workflow question overlap, but they are not the same. This article asks what capabilities and responsibilities sit around AI-enabled evidence work. The existing AI workflow guide addresses the operational question of which review tasks may be automated, which outputs need verification, how automation should be validated in the review, and how it should be reported.
AI for Systematic Reviews: Automate, Verify, Report Move from professional accountability to the workflow-level question: what researchers can automate, what they must verify, and what they should report inside a systematic review.The professional role is changing, but the labour-market story is not settled
It is reasonable to expect more evidence professionals to spend time evaluating tools, designing quality controls, resolving AI-human disagreements, protecting data and maintaining provenance. RAISE already treats AI use as a multi-role problem rather than a single-user skill, and new agentic systems make supervision and infrastructure more visible parts of evidence work.2
It is not yet defensible to say that AI will eliminate junior systematic-review roles, that “prompt engineer” is becoming a core evidence profession, or that manual reviewing will disappear. Those are labour-market predictions for which current evidence is inadequate. The safer conclusion is narrower: as more execution becomes automatable, the comparative value of validation, methodological judgment, quality assurance, data stewardship and transparent decision ownership is likely to increase. That remains a reasonable inference, not a settled workforce forecast.
Professional evidence synthesis in the AI era
The strongest professional position is neither defensive nor credulous. AI should be used where its role is methodologically justified and its errors can be controlled. It should be rejected or constrained when the team cannot establish fitness for purpose, cannot protect the input data, cannot inspect consequential failures, or lacks the expertise to supervise the method.
The professional evidence synthesist of the AI era is therefore not simply the person who knows the newest tool. It is the person who can distinguish a useful automation from an unsupported shortcut; validation from verification; technical fluency from methodological validity; and an impressive demonstration from evidence strong enough to change a real review workflow.
Agentic systems will make that distinction more important. As tools gain the ability to plan and act across several stages, professional oversight must move from checking isolated answers to governing the evidence chain. The responsibility is to know what entered the system, what transformed it, which decisions were automated, what was checked, what remained uncertain and who was prepared to stop the process when the evidence no longer justified it.
That is the durable capability. Models will change. Interfaces will change. Benchmarks will improve. The obligation to make consequential evidence work inspectable and defensible is much less likely to disappear.
Frequently asked questions
Must a human manually verify every AI output in a systematic review?
Can an AI agent now conduct a publication-grade systematic review autonomously?
Does using AI without disclosure automatically count as research misconduct?
Is prompt engineering a core professional evidence-synthesis skill?
Is enterprise AI automatically safe for confidential review data?
References
- Flemyng E, Noel-Storr A, Macura B, Gartlehner G, Thomas J, Meerpohl JJ, et al. Position statement on artificial intelligence (AI) use in evidence synthesis across Cochrane, the Campbell Collaboration, JBI and the Collaboration for Environmental Evidence 2025. Environ Evid. 2025;14:20. doi:10.1186/s13750-025-00374-5 .
- Thomas J, Flemyng E, Noel-Storr A, et al. Responsible use of AI in Evidence SynthEsis (RAISE): guidance and recommendations. Open Science Framework. 2025-2026 living guidance. doi:10.17605/OSF.IO/FWAUD .
- International Committee of Medical Journal Editors. Use of artificial intelligence in publishing. In: Recommendations for the Conduct, Reporting, Editing, and Publication of Scholarly Work in Medical Journals. Updated January 2026. Available from: ICMJE Recommendations .
- Li L, Mathrani A, Susnjak T. What level of automation is “good enough”? A benchmark of large language models for meta-analysis data extraction. Res Synth Methods. 2026;17(4):671-692. doi:10.1017/rsm.2025.10066 .
- Clark J, Barton B, Albarqouni L, et al. Generative artificial intelligence use in evidence synthesis: a systematic review. Res Synth Methods. 2025;16(4):601-619. doi:10.1017/rsm.2025.16 .
- Gartlehner G, Kugley S, Crotty K, et al. Artificial intelligence-assisted data extraction with a large language model: a study within reviews. Ann Intern Med. 2025;178(12):1763-1771. doi:10.7326/ANNALS-25-00739 .
- Taneri PE. Human versus artificial intelligence: comparing Cochrane authors’ and ChatGPT’s risk of bias assessments. Cochrane Evid Synth Methods. 2025;3(5):e70044. doi:10.1002/cesm.70044 .
- Li L, Mathrani A, Susnjak T. Transforming evidence synthesis: a systematic review of the evolution of automated meta-analysis in the age of AI. Res Synth Methods. 2026;17:403-450. doi:10.1017/rsm.2025.10065 .
- Asai A, et al. Synthesizing scientific literature with retrieval-augmented language models. Nature. 2026. doi:10.1038/s41586-025-10072-4 .
- Padarha S, Kearns RO, Naidoo T, et al. AgentSLR: automating systematic literature reviews in epidemiology with agentic AI. arXiv. 2026. arXiv:2603.22327 . Preprint; interpret as emerging evidence.
- AutoSynthesis: an agentic system for automated meta-analysis. arXiv. 2026. arXiv:2607.15247 . Preprint; interpret as emerging evidence.
- Goddard K, Roudsari A, Wyatt JC. Automation bias: a systematic review of frequency, effect mediators, and mitigators. J Am Med Inform Assoc. 2012;19(1):121-127. doi:10.1136/amiajnl-2011-000089 .
- Collaboration for Environmental Evidence. Artificial Intelligence Reporting Guidance: reporting guidance for the use of artificial intelligence in environmental evidence synthesis. Living guidance. Available from: Collaboration for Environmental Evidence .
Evidence note. AI capability changes faster than most evidence-synthesis standards. Performance findings above describe specific systems, datasets and research contexts and should not be transferred automatically to another review. Agentic-system evidence is particularly immature; preprints are identified as such. The durable claims in this article therefore rely primarily on professional responsibility, fit-for-purpose validation, risk-sensitive oversight, transparent reporting and accountable human decision ownership.