GRADE Certainty-of-Evidence Assessment Worksheet for Systematic Reviews

GRADE certainty assessment is not a score applied after meta-analysis. It is an outcome-level judgment about how confidently the true effect lies on one side of a chosen threshold, or within a chosen range, after considering limitations in the body of evidence.1,2

That sequence matters. Before rating risk of bias, inconsistency, indirectness, imprecision or publication bias, reviewers must know what claim they are trying to be certain about. Core GRADE 2 therefore places the target of certainty and the relevant threshold before the domain judgments. GRADE Guidance 41 discontinued categories of contextualization and recommends referring directly to the chosen threshold or range used to define the target of certainty.2,3

Working principleThis page is a documentation worksheet, not an automated GRADE engine. Use it to record one outcome-level assessment, the evidence considered for each domain, and the rationale for the final category. It deliberately does not choose the threshold, decide whether a concern is serious, diagnose publication bias, or calculate the final certainty category. Those remain methodological and clinical judgments.1,8
Evidence basis checked August 2026Core GRADE 1–5, GRADE Guidance 38 and 41, and Cochrane MECIR C74–C75. The tool checks documentation and internal consistency only; it does not validate a reviewer’s GRADE judgment.

Begin with the outcome and the threshold, not with the domain labels

Certainty is rated separately for each outcome because the relevant studies, precision, directness and risk-of-bias profile may differ across outcomes. In Core GRADE, randomized controlled trial evidence starts at high certainty and non-randomized studies of interventions start at low certainty. The five rating-down domains then address distinct reasons why that starting confidence may need to be reduced.1,4

Risk of bias

Would limitations in the studies contributing most information plausibly distort the outcome estimate?

Inconsistency

Do study results vary in a way that changes interpretation relative to the chosen threshold, without a credible explanation?

Indirectness

Does the available evidence meaningfully mismatch the target population, intervention, comparator, outcome or setting?

Imprecision

Does the confidence interval cross a threshold that would change the conclusion about the effect?

Publication bias

Are there converging reasons to suspect that the available studies systematically omit unfavorable or null evidence?

Rating up

For methodologically rigorous non-randomized evidence, is there a sufficiently large magnitude of effect or a credible dose-response gradient to consider rating certainty up?

Outcome-level GRADE certainty assessment workflow A compact four-stage workflow moves from the outcome and chosen threshold to the evidence-design starting certainty, domain judgments, and the documented final certainty category. A separate note emphasizes that the worksheet supports documentation but does not replace human judgment. Outcome + threshold What claim is being rated? Starting certainty Evidence design Domain judgments Risk of bias · inconsistency · indirectness Imprecision · publication bias · eligible rating up Final certainty High · Moderate · Low · Very low Human judgment remains central: the worksheet checks documentation, not the correctness of the GRADE rating. Outcome-level GRADE certainty workflow on mobile A compact vertical sequence moves from outcome and threshold to starting certainty, domain judgments, final certainty, and a reminder that the worksheet does not automate methodological judgment. 1. Outcome + threshold Define the claim 2. Starting certainty Evidence design 3. Domain judgments RoB · inconsistency · indirectness · imprecision · publication bias · rating up 4. Final certainty High · Moderate · Low · Very low Documentation aid, not an automated GRADE verdict. Methodological judgment remains with the reviewers.
Figure 1. A defensible GRADE assessment begins with the outcome and the chosen threshold or range, records the design-based starting certainty, evaluates the relevant domains, and ends with a documented reviewer judgment of final outcome-level certainty.

What the worksheet should record without pretending to decide

A reproducible assessment needs more than five dropdown menus. Risk of bias depends on the contribution and results of studies at different risk levels. Inconsistency is judged from unexplained variability in study results, including point estimates, confidence-interval overlap and their relation to the chosen threshold; I² is supportive rather than determinative. Indirectness compares the target PICO with the evidence PICO. Imprecision asks how the confidence interval relates to the chosen threshold or thresholds. Publication bias is a structured judgment based on converging evidence rather than a funnel-plot verdict.2,4,5,6

Current Core GRADE rating-up scopeCore GRADE 4 describes two rating-up situations for non-randomized studies of interventions: a large magnitude of effect and a credible dose-response gradient. This Core worksheet therefore does not include a separate upgrading item based on the predicted direction of residual confounding. Dose-response upgrading requires a credibility assessment and should be by no more than one level.4,7

Interactive working worksheet

GRADE certainty-of-evidence assessment worksheet

Document one outcome at a time. The worksheet checks documentation completeness and internal consistency, but every domain judgment and the final certainty category remain reviewer decisions. Entries are processed in your browser and are not sent by this worksheet to MetaSyn Academy.

1. Outcome and target of certainty

State the outcome, comparison and threshold or range that defines the claim being assessed.

Required foundation
GRADE Guidance 41 retired the minimally/partly/fully contextualized labels.

2. Evidence design and starting certainty

Record the design-based starting point. Core GRADE starts randomized controlled trial evidence at high certainty and non-randomized studies of interventions at low certainty; the worksheet does not assign the starting category for you.

Starting point
Core GRADE: randomized evidence starts high; NRSI starts low. If only one randomized trial contributes, inconsistency across studies cannot be assessed, but the other certainty domains still require judgment.

3. Rating-down domains

Choose the judgment, then document the evidence and reasoning. “Serious” and “very serious” judgments require a rationale.

Five domains
Risk of biasBody-of-evidence judgment
InconsistencyInterpret variability relative to the threshold
IndirectnessCompare target PICO with evidence PICO
ImprecisionConfidence interval relative to threshold
Core GRADE 2 considers rating down one or two levels for imprecision; two levels are commonly considered when the CI includes both important benefit and important harm.
Core GRADE may invoke OIS when the CI does not cross the chosen threshold but sparse information could still undermine precision.
Publication biasStructured suspicion, not mechanical detection

4. Rating-up considerations for non-randomized evidence

Core GRADE 4 describes two rating-up situations for methodologically rigorous NRSI: large magnitude of effect and a credible dose-response gradient. Rating up remains a reviewer judgment and is not calculated by this worksheet.

NRSI only
Core GRADE 4: for appropriate NRSI, RR >2.0 or <0.5 supports considering +1; RR >5.0 or <0.2 supports considering +2. Similar thresholds may be considered for ORs and HRs.
Document the effect measure and interval used to support any large-effect upgrade.
GRADE Guidance 38 recommends rating up by no more than one level after establishing credibility.
Core GRADE boundaryThis Core worksheet includes the two rating-up situations described in Core GRADE 4: large magnitude of effect and dose-response gradient. It does not add a separate upgrading item based on the predicted direction of residual confounding.

5. Final outcome-level certainty and audit trail

Assign the final category yourself after considering the domain judgments. The worksheet will not calculate it.

Human judgment
Complete the worksheet, then select Check documentation. Warnings identify missing documentation or internal inconsistencies; they do not decide the GRADE rating.

Documented assessment summary

This preview is a working record, not an automatically generated GRADE verdict.

Live preview
DomainJudgmentRationale

A complete record should explain the path, not merely display the destination

The value of a certainty worksheet lies in its audit trail. Another reviewer should be able to see which outcome was assessed, which threshold defined the target of certainty, which studies contributed the evidence, what was considered for each domain, and why the final category was chosen. Cochrane MECIR states that certainty should be assessed for each outcome using the five GRADE considerations, that all certainty assessments should be justified and documented, and that ideally two people should assess certainty independently and reach consensus on downgrading decisions.8

The completed worksheet should support, rather than replace, the concise certainty statement that appears in a Summary of Findings table. Keep the working record detailed enough to audit; keep the manuscript-facing certainty statement concise enough to interpret.

Methodological boundary

This worksheet is designed for Core GRADE assessment of intervention-effect evidence. Network meta-analysis, diagnostic test accuracy, prognosis and qualitative evidence require specialist GRADE guidance, extensions or distinct frameworks. GRADE-CERQual, for example, addresses confidence in qualitative evidence-synthesis findings and is not interchangeable with Core GRADE certainty assessment.

Get every MetaSyn template free, including this one.

Leave your email and I’ll send this resource as an editable Word file and a printable PDF, plus access to the smart online version. You’ll also get every new template as it’s finished. No noise, just the resources.

GRADE certainty assessment

Frequently asked questions

What does “certainty of evidence” mean in current GRADE?

Current GRADE frames certainty around confidence in a specified claim about the true effect, commonly defined by a chosen threshold or range. The target of certainty should therefore be stated explicitly for the outcome being assessed.1,2

Should I still classify an assessment as minimally, partly or fully contextualized?

No. GRADE Guidance 41 discontinued categories of contextualization and recommends referring directly to the threshold or range used to define the target of certainty.3

Does a high I² automatically mean serious inconsistency?

No. Core GRADE 3 does not treat a heterogeneity statistic as an automatic rating rule. Reviewers should consider the pattern and magnitude of between-study differences and whether unexplained variability changes the certainty claim relative to the chosen threshold.5

Can this worksheet calculate the final GRADE certainty automatically?

No. The worksheet can check whether required documentation is present and flag internal inconsistencies, but the threshold, domain judgments, reasons for rating down or up, and final certainty category remain reviewer judgments.1,8

Can non-randomized evidence be rated up because residual confounding would predictably reduce the observed effect?

Not as a separate rating-up criterion in current Core GRADE. The Core GRADE framework retains large magnitude of effect and a credible dose-response gradient as the relevant rating-up situations for non-randomized intervention evidence.1,4,7

How many reviewers should assess certainty?

For Cochrane intervention reviews, MECIR states that certainty assessments should be justified and documented and that, ideally, two people should assess certainty independently and reach consensus on downgrading decisions.8

Does low or very low certainty mean that there is no effect?

No. Certainty describes confidence in the evidence supporting a specified claim about the true effect. It is not the effect estimate itself. An estimated benefit or harm can be supported by low-certainty evidence, while a small or null effect can be supported by high-certainty evidence.1,2

Methodological sources

References

These sources support the Core GRADE concepts, domain judgments, rating-up considerations and review process used in the worksheet above.

  1. Guyatt G, Agoritsas T, Brignardello-Petersen R, et al. Core GRADE 1: overview of the Core GRADE approach. BMJ. 2025;389:e081903. doi:10.1136/bmj-2024-081903.
  2. Guyatt G, Zeng L, Brignardello-Petersen R, et al. Core GRADE 2: choosing the target of certainty rating and assessing imprecision. BMJ. 2025;389:e081904. doi:10.1136/bmj-2024-081904.
  3. Hultcrantz M, Schünemann HJ, Mustafa RA, et al. GRADE Certainty Ratings: Thresholds Rather Than Categories of Contextualization (GRADE Guidance 41). Ann Intern Med. 2025;178(8):1183–1186. doi:10.7326/ANNALS-25-00548.
  4. Guyatt G, Wang Y, Eachempati P, et al. Core GRADE 4: rating certainty of evidence—risk of bias, publication bias, and reasons for rating up certainty. BMJ. 2025;389:e083864. doi:10.1136/bmj-2024-083864.
  5. Guyatt G, Schandelmaier S, Brignardello-Petersen R, et al. Core GRADE 3: rating certainty of evidence—assessing inconsistency. BMJ. 2025;389:e081905. doi:10.1136/bmj-2024-081905.
  6. Guyatt G, Iorio A, De Beer H, et al. Core GRADE 5: rating certainty of evidence—assessing indirectness. BMJ. 2025;389:e083865. doi:10.1136/bmj-2024-083865.
  7. Murad MH, Verbeek J, Schwingshackl L, et al. GRADE Guidance 38: updated guidance for rating up certainty of evidence due to a dose-response gradient. J Clin Epidemiol. 2023;164:45–53. doi:10.1016/j.jclinepi.2023.09.011.
  8. Cochrane. Methodological Expectations of Cochrane Intervention Reviews (MECIR), C74–C75: assessing and justifying certainty of the body of evidence. Cochrane MECIR C74–C75.