Which Risk of Bias Tool Should You Use? RoB 2, ROBINS-I, ROBIS, AMSTAR 2 and JBI
Critical appraisal methodology
A scholarly guide to preserving the construct, evidence, and decision trail behind critical appraisal.
The name of an instrument is not its construct
The phrase “risk-of-bias tool” suggests a coherent family of interchangeable instruments. In practice, available methods evaluate different objects for different purposes. RoB 2 asks whether a specified randomized-trial result may be materially biased. ROBIS asks whether flaws in the conduct or interpretation of a systematic review place its findings at risk of bias. AMSTAR 2 asks whether a systematic review of healthcare interventions has critical or non-critical methodological weaknesses. A JBI checklist may support critical appraisal of a design without claiming the same domain-based construct or unit.
Selection should begin with a sentence that could appear in your methods section: “We will assess [construct] in [evidence object] at the level of [unit] to inform [decision].” If your team cannot complete that sentence, comparing instrument names is premature. This exercise forces a clear separation of internal validity, applicability, reporting completeness, methodological conduct, and certainty. It also reveals whether a single review contains multiple appraisal questions that require different methods. Accessing free systematic review and meta-analysis templates can help structure these initial protocol definitions.
Construct clarity is essential because downstream reporting inherits it. A review cannot legitimately report “low risk of bias” if the chosen checklist merely measured broad reporting thoroughness or relevance. Nor should a ROBIS rating for a systematic review be applied to the primary trials inside it. Preserving these distinctions ensures readers understand what your conclusions cover and what remains unresolved.1,7,8
Begin intervention and exposure appraisal with the target causal effect
For randomized trials, relevant bias mechanisms depend heavily on your effect of interest. RoB 2 distinguishes between the effect of assignment to intervention (intention-to-treat) and the effect of adhering to intervention (per-protocol). Post-randomization deviations carry different weight under each target. Reviewers must specify the result, outcome measure, time point, and analysis population before answering signaling questions. Stating that “this trial is low risk” is too broad to describe what was actually appraised.2,3
ROBINS-I evaluates a non-randomized intervention study as an attempt to mimic a target randomized trial. Confounding, selection, classification, missing data, outcome measurement, and selective reporting are assessed against that specified causal contrast. The 2016 instrument remains the established standard, while the ROBINS-I v2 material adds algorithms and triage guidance. If you use the draft version, note that status explicitly in your protocol.4,5
ROBINS-E addresses exposure effects in non-randomized follow-up studies and similarly requires a defined causal effect and result. It is not a generic observational study checklist. Because “cohort” studies can evaluate interventions, exposures, prognostic factors, or predictive models, selecting a tool by study design label alone often leads to errors.6
Protocol authors should define target causal contrasts when selecting appraisal tools. If your causal question changes during analysis, re-evaluate your tool selection to prevent mismatches between ratings and pooled estimates.
Diagnostic accuracy and prediction models require estimate-level discipline
Diagnostic accuracy reviews should align with QUADAS-3, which introduced an ideal test-accuracy trial framework, defined synthesis questions, and result-level assessments. If your protocol was designed around QUADAS-2, ensure your extraction structures and summary figures match the updated framework.9
Comparative test accuracy requires additional tools. QUADAS-C evaluates bias in head-to-head test comparisons and is used alongside QUADAS-3 rather than as a standalone replacement. Indirect comparisons between separate studies do not turn a review into a QUADAS-C assessment simply because multiple tests are analyzed.9,10
Prediction models involve a different appraisal logic. PROBAST evaluates risk of bias and applicability in prediction model studies, while PROBAST+AI updates these standards for regression and artificial intelligence models, separating model-development quality from validation risk of bias. Match your tool to the evidence object rather than the clinical topic.11,12
ROBIS and AMSTAR 2 answer different review-level questions
ROBIS evaluates risk of bias in systematic review findings across four domains: eligibility criteria, study identification, data collection/appraisal, and synthesis methods. It suits guideline authors and overview reviewers determining whether review flaws systematically distorted conclusions.7
AMSTAR 2 critically appraises systematic reviews of healthcare interventions by rating confidence based on critical and non-critical weaknesses. Its developers explicitly state it should not produce an overall numeric score. Differences between ROBIS and AMSTAR 2 ratings occur because risk of bias and overall methodological quality are related but distinct constructs.8
Select your tool based on your objective: evaluating potential bias in review findings, appraising methodological quality, or satisfying organizational standards. If using both, report their findings separately rather than attempting to average them into a single score.
Specialist designs require dedicated appraisal frameworks
Not every study design fits standard risk-of-bias tools. Prognostic factors, disease prevalence, qualitative research, mixed methods, interrupted time series, animal models, and measurement properties each require specialized manuals. Finding a checklist online does not establish its validity or suitability for your review.
Specialist selection follows the same core principles: Identify the target inference, key bias mechanisms, assessment unit, and reviewer qualifications required. When standard tools fall short, consulting a methodologist is a proper procedural step rather than a project setback.
COSMIN methods illustrate these boundaries for patient-reported outcomes and measurement properties, requiring dedicated appraisal logic. Similar boundaries apply across diagnostic, prognostic, and environmental research fields.
Numeric quality scores obscure critical methodological flaws
Summing checklist items into a single numeric score assumes that items carry equal weight and that high scores in one area compensate for severe flaws in another. In practice, applying different quality scales to identical studies can yield contradictory conclusions about study quality and effect sizes.14
Domain-based appraisal systems preserve non-compensable flaws. A trial can excel in reporting and blinding yet suffer from fatal selection bias. Reviewers must evaluate specific bias mechanisms rather than relying on summary scores to generate an artificial label.
This principle aligns with structured tools like RoB 2, which use explicit logic to map domain responses to overall risk ratings. This mapping is domain-specific and auditable, ensuring overall judgments reflect actual methodological risk.
Tool versions and licensing shape review reproducibility
Citing a tool without a version number creates ambiguity when domain rules or algorithms evolve. Archive the exact official manual used, record your access date, and document local implementation guidelines. For living reviews facing major tool updates, decide prospectively whether to maintain the original tool version or re-appraise existing studies, documenting your rationale.
Check licensing terms before embedding appraisal tools into digital software or public templates. Many instruments allow free academic use but restrict commercial distribution or derivative modifications. Review templates should teach selection and log decisions while linking to official developer repositories.
Report local protocol rules as implementation guidelines rather than tool modifications. If your team drops domains or converts signaling questions into numeric points, report the result as a custom checklist rather than citing the original instrument’s name.
JBI checklists support critical appraisal within specific review frameworks
JBI provides critical appraisal checklists across diverse study designs, making them useful when conducting reviews under JBI methodology or evaluating designs lacking dedicated domain-based tools. However, a checklist score should not be treated as a direct substitute for a result-level risk-of-bias assessment.
Avoid converting JBI checklist responses into percentage cutoffs unless supported by official guidance. Distinguish item-level checklist responses from formal risk-of-bias algorithms, and prespecify any study exclusion rules in your protocol.
If organizational policies mandate a specific tool, record both the compliance requirement and the methodological fit to maintain transparency for readers.
Piloting validates tool selection before full assessment
A theoretically sound tool can fail in practice if reviewers struggle to identify target results, locate source data, or apply domain rules consistently. Test your chosen tool on a pilot sample of challenging studies before starting full assessment.
Use pilot results to refine local coding instructions and reviewer training without altering the core tool structure. If persistent errors stem from a fundamental mismatch between the tool and your study designs, update your tool choice in your protocol before production appraisal begins.
Document your pilot sample, observed disagreements, and resulting protocol clarifications to demonstrate that your appraisal workflow is fully calibrated.
Focus on target inferences before selecting tools
Selecting an appraisal tool requires a clear sequence of decisions: evidence design, target construct, assessment unit, target result, governing framework, tool version, implementation rules, reviewer expertise, and analytical goals. This structured path ensures your chosen tool supports your research question.
Documenting your tool selection rationale clarifies why an instrument was chosen, what its ratings mean, and how appraisal results informed your synthesis. That audit trail provides greater scientific value than simply naming a familiar tool.15
References and evidence scope
Methodological literature supporting critical appraisal selection, construct preservation, and PRISMA 2020 compliance. Retrieve official appraisal manuals directly from developer repositories.
- Boutron I, Page MJ, Higgins JPT, Altman DG, Lundh A, Hróbjartsson A. Chapter 7: Considering bias and conflicts of interest among the included studies. In: Higgins JPT, Thomas J, Chandler J, Cumpston M, Li T, Page MJ, et al, editors. Cochrane Handbook for Systematic Reviews of Interventions. Version 6.5. Cochrane; 2024. https://www.cochrane.org/authors/handbooks-and-manuals/handbook/current/chapter-07
- Higgins JPT, Savović J, Page MJ, Elbers RG, Sterne JAC. Chapter 8: Assessing risk of bias in a randomized trial. In: Higgins JPT, Thomas J, Chandler J, Cumpston M, Li T, Page MJ, et al, editors. Cochrane Handbook for Systematic Reviews of Interventions. Version 6.5. Cochrane; 2024. https://www.cochrane.org/authors/handbooks-and-manuals/handbook/current/chapter-08
- Sterne JAC, Savović J, Page MJ, Elbers RG, Blencowe NS, Boutron I, et al. RoB 2: a revised tool for assessing risk of bias in randomised trials. BMJ. 2019;366:l4898. https://doi.org/10.1136/bmj.l4898
- Sterne JAC, Hernán MA, Reeves BC, Savović J, Berkman ND, Viswanathan M, et al. ROBINS-I: a tool for assessing risk of bias in non-randomised studies of interventions. BMJ. 2016;355:i4919. https://doi.org/10.1136/bmj.i4919
- Risk of Bias Development Group. ROBINS-I Version 2, November 2025 draft [Internet]. Bristol: Risk of Bias tools; 2025. https://www.riskofbias.info/welcome/robins-i-v2
- Higgins JPT, Morgan RL, Rooney AA, Taylor KW, Thayer KA, Silva RA, et al. A tool to assess risk of bias in non-randomized follow-up studies of exposure effects (ROBINS-E). Environ Int. 2024;186:108602. https://doi.org/10.1016/j.envint.2024.108602
- Whiting P, Savović J, Higgins JPT, Caldwell DM, Reeves BC, Shea B, et al. ROBIS: a new tool to assess risk of bias in systematic reviews was developed. J Clin Epidemiol. 2016;69:225-234. https://doi.org/10.1016/j.jclinepi.2015.06.005
- Shea BJ, Reeves BC, Wells G, Thuku M, Hamel C, Moran J, et al. AMSTAR 2: a critical appraisal tool for systematic reviews that include randomised or non-randomised studies of healthcare interventions, or both. BMJ. 2017;358:j4008. https://doi.org/10.1136/bmj.j4008
- Whiting PF, Tomlinson E, Rutjes AWS, Davenport CF, Yang B, Westwood ME, et al. QUADAS-3: a revised tool for the quality assessment of diagnostic test accuracy studies. Ann Intern Med. 2026. https://doi.org/10.7326/ANNALS-25-02104
- Yang B, Mallett S, Takwoingi Y, Davenport CF, Hyde C, Whiting PF, et al. QUADAS-C: a tool for assessing risk of bias in comparative diagnostic accuracy studies. Ann Intern Med. 2021;174(11):1592-1599. https://doi.org/10.7326/M21-2234
- Wolff RF, Moons KGM, Riley RD, Whiting PF, Westwood M, Collins GS, et al. PROBAST: a tool to assess the risk of bias and applicability of prediction model studies. Ann Intern Med. 2019;170(1):51-58. https://doi.org/10.7326/M18-1376
- Moons KGM, Damen JAA, Kaul T, Hooft L, Andaur Navarro C, Dhiman P, et al. PROBAST+AI: an updated quality, risk of bias, and applicability assessment tool for prediction models using regression or artificial intelligence methods. BMJ. 2025;388:e082505. https://doi.org/10.1136/bmj-2024-082505
- Tomlinson E, Cooper C, Davenport C, Rutjes AWS, Leeflang M, Mallett S, Whiting P. Common challenges and suggestions for risk of bias tool development: a systematic review of methodological studies. J Clin Epidemiol. 2024;171:111370. https://doi.org/10.1016/j.jclinepi.2024.111370
- Jüni P, Witschi A, Bloch R, Egger M. The hazards of scoring the quality of clinical trials for meta-analysis. JAMA. 1999;282(11):1054-1060. https://doi.org/10.1001/jama.282.11.1054
- Page MJ, McKenzie JE, Bossuyt PM, Boutron I, Hoffmann TC, Mulrow CD, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71. https://doi.org/10.1136/bmj.n71