How to Report Risk of Bias and Critical Appraisal Results in Systematic Reviews
Critical appraisal methodology
A scholarly guide to preserving the construct, evidence, and decision trail behind critical appraisal.
Start reporting before the assessment results exist
The reporting architecture should be specified in your protocol because your chosen unit and tool dictate your data structure. A result-level tool requires identifiers for outcome, measurement, time point, and analysis; a study-level checklist requires a different row structure. Adding these identifiers only when drafting the manuscript invites aggregation errors.
Your methods section should name the instrument and exact version, detail review-specific implementation guidance, describe reviewer independence and consensus rules, declare automation use, and state how judgments informed synthesis. Distinguish study-level risk of bias from synthesis-level bias due to missing results.2,3 Accessing free systematic review and meta-analysis templates helps establish these reporting structures early in your protocol.
Local modifications require explicit disclosure. Dropping domains, altering response options, or calculating summary scores changes the underlying construct. Do not cite an official instrument without disclosing material deviations.
Preserve assessment unit, result identity, and denominator
Your report should allow readers to map every displayed judgment back to the specific assessed object. For RoB 2, this is a specific trial result; for ROBINS-I, it is a result estimating a defined intervention effect. A single study can yield multiple assessments with different judgments. Collapsing them to the most severe rating exaggerates one synthesis while obscuring another.6,7
Counts and percentages must state their exact denominator. “Six of 15 results” and “six of ten studies” convey different information, even if drawn from the same review. When domains include not-applicable categories or missing details, report those counts rather than silently shifting denominators.
Your master dataset should maintain stable identifiers linking study, report source, result, synthesis grouping, and summary figures. This relational structure supports corrections and keeps manuscript tables synchronized with underlying judgments.
Use visualizations as maps, not verdicts
Traffic-light plots show individual judgments across domains, while weighted bar charts illustrate overall distributions. Both work well when labels, legends, and denominators are explicitly defined. The robvis tool provides standardized formats for several instruments, though authors remain responsible for checking data mappings and explaining assessment units.5
Color alone is inaccessible and implies moral ranking. Figures should incorporate text labels or symbols, accessible color palettes, clear legends, and accompanying summary tables or repository data files. Red should not substitute for a defined high-risk rating, and custom categories should not map to standard colors without explanation.
Avoid combining instruments with incompatible constructs into a single visual scale. Display mixed study designs in separate panels with tool-specific legends and narrative summaries.
Write a pattern-based narrative, not a list of labels
Your results narrative should highlight common bias mechanisms, cite supporting evidence, identify affected syntheses, and note where reporting was incomplete. Avoid declaring an entire body of literature “low quality” when the tool evaluated a narrower construct.
Effective sentences combine numerator, denominator, assessment unit, and rationale: for instance, noting that four of nine primary outcome results carried concerns regarding selective reporting because prespecified analysis plans were unavailable. Distinguish missing information from clear evidence of bias.
Do not average domain judgments into a composite score. Quality scales yield inconsistent rankings and allow strengths in unrelated areas to mask severe methodological flaws.8
Explain what changed because of the appraisal
Connect critical appraisal findings directly to your analytical and interpretive decisions, such as running sensitivity analyses, stratifying presentations, downgrading certainty, omitting pooling, or qualifying conclusions. Specify these actions for each assessment unit, following your protocol.
Sensitivity analyses should state both the removal rule and its impact on pooled estimates. Stating that results remained robust is insufficient when effect sizes, confidence intervals, or heterogeneity shifted. When exclusions are post hoc, declare them as exploratory rather than confirmatory.
Study-level appraisal and missing-results bias evaluate different issues. A set of low-risk studies does not eliminate publication bias across an entire synthesis. Address both layers in your final certainty statement without merging them into a single label.1,2
Risk of bias informs certainty but does not equal certainty
A primary-study risk-of-bias judgment addresses potential systematic error in a specific result. Certainty frameworks evaluate broader outcome-level domains, including inconsistency, indirectness, imprecision, and publication bias. Show how appraisal ratings inform the risk-of-bias domain without presenting a study-level rating as the certainty score for the entire evidence base.
When multiple results contribute to an outcome, apply a transparent rule to carry result-level ratings into body-of-evidence assessments. Consider relative study weights, sensitivity tests, and the direction of plausible bias rather than relying on a simple majority vote.
Keep risk-of-bias tables and certainty tables connected yet conceptually distinct. Readers can then determine whether low certainty stems from biased studies, sparse data, inconsistency, or reporting bias.
Construct the manuscript as a chain of verifiable claims
Your methods section should state the instrument, version, assessment unit, reviewer workflow, consensus process, and analytical applications. Every element should match your protocol and project record, with any protocol amendments explicitly declared.
Your results section should report the appraised objects, exact denominators, domain patterns, and overall judgments. Provide a manuscript table or supplementary file allowing readers to inspect individual ratings. Visual figures can summarize these data but must not introduce categories absent from the dataset.
Your synthesis section should identify which assessments contributed to each analysis and how ratings altered your analytical strategy. If sensitivity tests excluded high-risk results, report the criteria and exact counts removed. If pooling was omitted, explain how appraisal informed your narrative synthesis.
Your discussion should interpret plausible bias impacts without treating qualitative ratings as exact numerical offsets. Distinguish positive evidence of bias mechanisms from uncertainty caused by incomplete reporting or limitations outside the tool’s scope.
Abstracts should summarize key appraisal findings only when units and denominators remain clear. Replace phrases like “most studies were low quality” with precise statements linked to specific tools and primary domain patterns.
Anticipate the questions a methodological reviewer will ask
Peer reviewers should be able to confirm that your appraisal tool matched your study design and target result, that two authors reviewed independently, how disagreements were resolved, whether local guidance was added, and where complete ratings can be inspected.
Reviewers should also be able to verify every denominator, distinguish study-level from synthesis-level bias, and track how ratings influenced your conclusions. If your text leaves these questions open, formatting adjustments cannot fix the missing evidence trail.
Reconcile all counts across your abstract, main text, tables, figures, certainty assessments, and supplementary files before submission. Verify that category names match official tool versions and that citations point to correct manuals. This consistency check provides essential methodological quality control.
Missing-results bias needs its own synthesis-level account
Selective non-reporting within studies informs study-level tools, but missing results across studies require a dedicated synthesis-level evaluation. Included studies can be well conducted individually while the overall synthesis remains biased because unfavorable trials were withheld.
Report the search methods used to locate protocols, registry records, unpublished reports, or small-study patterns, noting any limitations. Funnel plot asymmetry is not a direct test of publication bias and may reflect clinical heterogeneity, chance, or methodological variation. Interpret asymmetry in light of study counts and characteristics.
Your narrative should explain how missing-results concerns altered confidence in each synthesis, avoiding generic statements disconnected from your data.
Living reviews need versioned appraisal and reporting
Review updates may incorporate new studies, update existing records, or transition to revised appraisal tools. Document which tool version produced each judgment and whether earlier assessments were updated. Mixing tool versions without clear policies creates artificial trends that reflect tool updates rather than changes in study quality.
Your dataset should record assessment dates, tool versions, and review cycle numbers. When ratings change because new documents emerge, log that update separately from re-interpretations. This distinction helps readers compare successive review versions.
Update reports should highlight shifts in appraisal findings and show whether overall conclusions changed. Re-publishing static figures without explaining updates reduces the utility of living evidence reviews.
Conclusion: make every judgment traceable to its consequence
Rigorous reporting maintains a clear chain: target construct, appraisal instrument, assessment unit, domain evidence, exact denominator, synthesis impact, and final conclusion. Tables provide detailed data, figures illustrate patterns, and narrative explains why those patterns matter. None is sufficient on its own.
When this chain remains explicit, readers can distinguish bias in specific results from missing evidence across syntheses or broader study limitations. Your review becomes easier to audit, update, and apply in clinical guidelines and policy decisions.
References and evidence scope
Methodological literature supporting critical appraisal reporting and PRISMA 2020 compliance. Retrieve official appraisal manuals directly from developer repositories.
- Boutron I, Page MJ, Higgins JPT, Altman DG, Lundh A, Hróbjartsson A. Chapter 7: Considering bias and conflicts of interest among the included studies. In: Higgins JPT, Thomas J, Chandler J, Cumpston M, Li T, Page MJ, et al, editors. Cochrane Handbook for Systematic Reviews of Interventions. Version 6.5. Cochrane; 2024. https://www.cochrane.org/authors/handbooks-and-manuals/handbook/current/chapter-07
- Page MJ, Higgins JPT, Sterne JAC. Chapter 13: Assessing risk of bias due to missing evidence in a meta-analysis. In: Higgins JPT, Thomas J, Chandler J, Cumpston M, Li T, Page MJ, et al, editors. Cochrane Handbook for Systematic Reviews of Interventions. Version 6.5. Cochrane; 2024. https://www.cochrane.org/authors/handbooks-and-manuals/handbook/current/chapter-13
- Page MJ, McKenzie JE, Bossuyt PM, Boutron I, Hoffmann TC, Mulrow CD, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71. https://doi.org/10.1136/bmj.n71
- Page MJ, Moher D, Bossuyt PM, Boutron I, Hoffmann TC, Mulrow CD, et al. PRISMA 2020 explanation and elaboration: updated guidance and exemplars for reporting systematic reviews. BMJ. 2021;372:n160. https://doi.org/10.1136/bmj.n160
- McGuinness LA, Higgins JPT. Risk-of-bias VISualization (robvis): an R package and Shiny web app for visualizing risk-of-bias assessments. Res Synth Methods. 2021;12(1):55-61. https://doi.org/10.1002/jrsm.1411
- Sterne JAC, Savović J, Page MJ, Elbers RG, Blencowe NS, Boutron I, et al. RoB 2: a revised tool for assessing risk of bias in randomised trials. BMJ. 2019;366:l4898. https://doi.org/10.1136/bmj.l4898
- Sterne JAC, Hernán MA, Reeves BC, Savović J, Berkman ND, Viswanathan M, et al. ROBINS-I: a tool for assessing risk of bias in non-randomised studies of interventions. BMJ. 2016;355:i4919. https://doi.org/10.1136/bmj.i4919
- Jüni P, Witschi A, Bloch R, Egger M. The hazards of scoring the quality of clinical trials for meta-analysis. JAMA. 1999;282(11):1054-1060. https://doi.org/10.1001/jama.282.11.1054