How to Assess Certainty of Evidence with GRADE in a Systematic Review

MetaSyn Academy guide to assess certainty of evidence with grade in a systematic review, illustrating the decisions documented by GRADE

Assessing certainty with GRADE asks a different question from estimating an effect. Meta-analysis summarizes what the included studies estimate; GRADE asks how confidently the true effect supports a specified claim for each outcome after limitations in the body of evidence are considered.1,2

This distinction is easy to blur. A narrow confidence interval does not guarantee high certainty if the underlying studies are biased or indirect. A heterogeneous meta-analysis does not automatically imply low certainty if the variation does not change interpretation relative to the relevant threshold. A strong-looking association from non-randomized evidence may still be vulnerable to bias and confounding. GRADE makes these reasons explicit rather than compressing them into a numerical quality score.1,3,4

Core GRADE framework. The seven-part Core GRADE series published in 2025 summarizes the essential approach for intervention-effect evidence. It retains five rating-down domains, risk of bias, inconsistency, indirectness, imprecision and publication bias, while clarifying how those judgments should be made. GRADE Guidance 41 also discontinued categories of contextualization and recommends referring directly to the threshold or range that defines the target of certainty.1,2,5
Purpose of this guide. This page explains the methodological reasoning behind outcome-level GRADE certainty judgments. If you need a structured working record for documenting those judgments, use the GRADE Certainty-of-Evidence Assessment Worksheet for Systematic Reviews.

1. Certainty begins with a claim about an outcome

Core GRADE requires reviewers to choose the target of the certainty rating before judging imprecision. Depending on the question, the relevant threshold may be a minimal important difference (MID) or the null. The location of the point estimate relative to that threshold defines the claim in which certainty is being rated,for example, confidence in an important effect, in little or no important effect, or in the existence of a non-null effect.2,5

Certainty is therefore outcome-specific. Mortality, symptoms, serious adverse events and quality of life can have different certainty ratings within the same review because the contributing studies, measurement limitations, directness and precision may differ by outcome. A single certainty label for an entire review conceals that structure.1

Five GRADE rating-down domains for an outcome-level certainty claim An outcome-level certainty claim defined relative to a threshold or range is assessed across risk of bias, inconsistency, indirectness, imprecision and publication bias. Outcome-level certainty claim Define the outcome and the threshold or range before interpreting the certainty domains Risk of bias Study limitations Inconsistency Unexplained variability Indirectness Target PICO mismatch Imprecision CI versus threshold Publication bias Missing evidence Each domain asks a different question; none is a numerical score. Five GRADE rating-down domains A compact mobile display of an outcome-level certainty claim followed by the five GRADE rating-down domains. Outcome-level certainty claim Defined relative to a threshold or range Risk of bias Study limitations Inconsistency Unexplained variability Indirectness Target PICO mismatch Imprecision CI versus threshold Publication bias Missing evidence Five distinct threats to confidence, not five points in a score.
Figure 1. The five rating-down domains address different threats to confidence in an outcome-level certainty claim. They should be judged separately and documented explicitly rather than converted into a numerical quality score.

2. The starting point is a convention, not a verdict

Core GRADE retains design-based starting points: randomized controlled trial evidence starts at high certainty, while non-randomized studies of interventions (NRSI) start at low certainty. The starting point does not determine the final category; certainty can be rated down for limitations, and methodologically rigorous NRSI may be considered for rating up when current criteria are met.1,3

When only one study contributes to an outcome, there is no between-study variability to examine. Inconsistency across studies therefore cannot be assessed from that single study, while the other certainty domains still require judgment.4

3. Risk of bias moves from study judgments to the body of evidence

Risk of bias is often misunderstood because study-level tools and GRADE answer different questions. Core GRADE first considers risk of bias in individual studies and then examines the body of evidence: how much low- and high-risk evidence contributes to the pooled estimate and whether their results differ importantly. Tools such as RoB 2 or ROBINS-I can inform study-level assessment, but the GRADE decision is whether study limitations in the body of evidence justify rating certainty down.3

The likely direction of bias can also matter. Core GRADE explicitly asks whether an expected direction of bias would weaken or, in some situations, reinforce the inference being made. The purpose is not to count or average risk-of-bias flags; it is to judge whether study limitations threaten the particular certainty claim.3

4. Inconsistency is not another name for I²

Core GRADE 3 defines inconsistency as unexplained variability in results across studies. Reviewers first inspect the magnitude of differences in point estimates, overlap of confidence intervals and the relationship of study estimates to the chosen threshold. For binary outcomes, Core GRADE generally focuses on consistency of relative effects; for continuous outcomes, it focuses on absolute effects.4

I² can support the judgment but may be misleading, particularly when confidence intervals are narrow. Reviewers should also consider prespecified hypotheses for effect modification and the credibility of subgroup effects. If a subgroup effect is credible and substantial, presenting separate estimates and certainty ratings may be more informative than downgrading one pooled estimate.4

5. Indirectness asks whether the evidence answers the target question

Core GRADE 5 defines indirectness as a mismatch between the target PICO and the PICO of the best available evidence. Differences in population, intervention, comparator or outcome matter when they are likely to produce an important difference in the magnitude of effect. Surrogate outcomes are a particularly important source of possible indirectness.6

Not every mismatch warrants rating down. The reviewer must judge how likely the mismatch is to change the answer to the target question. Core GRADE also distinguishes PICO-related indirectness from indirect comparisons in network meta-analysis; network comparisons require network-specific certainty methods rather than a simple transfer of pairwise rules.6

6. Imprecision is primarily a threshold-and-confidence-interval problem

Core GRADE 2 makes the chosen threshold and the 95% confidence interval central to imprecision. If the interval crosses the threshold that defines the certainty claim, reviewers rate down for imprecision. They may rate down two levels when the interval is wide enough to include both an important benefit and an important harm. Earlier GRADE Guidance 34 also emphasized threshold-and-CI reasoning; Core GRADE 2 provides the current simplified Core approach.2,7

Confidence intervals interpreted against benefit, null and harm thresholds Three hypothetical confidence intervals illustrate a precise important benefit, uncertainty about whether benefit is important, and a very wide interval compatible with important benefit and important harm. More benefit ← → More harm Benefit MID Null Harm MID A CI stays beyond benefit MID B CI crosses benefit MID C Important benefit and harm remain possible Threshold-relative GRADE imprecision Three hypothetical confidence intervals are compared with benefit MID, null and harm MID thresholds. Benefit ← → Harm Benefit MID Null Harm MID A Beyond MID B Crosses MID C CI includes important benefit and important harm
Figure 2. Hypothetical confidence intervals illustrate Core GRADE imprecision logic. Crossing the threshold that defines the certainty target supports rating down; a very wide interval compatible with both important benefit and important harm may justify rating down two levels.2

Optimal information size (OIS) is not ignored in Core GRADE 2, but it is used in a specific situation. If the confidence interval does not cross the chosen threshold yet the observed effect is unusually large, reviewers should consider whether the total sample size or number of events meets the OIS. If it does not, rating down for imprecision may still be appropriate.2

7. Publication bias is a judgment about missing evidence, not a funnel-plot diagnosis

Core GRADE 4 treats publication bias as a judgment informed by multiple signals. Reviewers consider known unpublished studies, whether the available studies are small, the role of industry sponsorship, and whether funnel-plot or statistical assessments are feasible and informative. Statistical tests of funnel-plot asymmetry have important false-positive and false-negative limitations and generally require at least 10 studies to be useful.3

Funnel-plot asymmetry is therefore not synonymous with publication bias. It can reflect publication bias, but it can also arise from heterogeneity or other small-study effects. Core GRADE recommends integrating the available evidence rather than treating one plot or test as a diagnosis.3

8. Rating up is narrower under Core GRADE than under older guidance

Core GRADE considers rating up for methodologically rigorous NRSI when the evidence has not already been rated down to very low certainty. For large magnitude of effect, a relative risk above 2.0 or below 0.5 supports considering an upgrade of one level; a relative risk above 5.0 or below 0.2 supports considering two levels, with similar thresholds for odds ratios and hazard ratios. These are considerations, not automatic upgrades.3

The other Core GRADE rating-up route is a credible dose-response gradient. GRADE Guidance 38 asks reviewers to evaluate the credibility of the gradient, including the analytical approach, confounding, ecological bias, consistency and supporting indirect evidence. When a gradient is judged credible, upgrading for dose-response is capped at one level.8

The older rating-up criterion based solely on the predictable direction of plausible residual confounding is not retained as a separate Core GRADE rating-up criterion.1,3

9. The final category is reasoned from the domains; it is not the sum of them

GRADE uses four final certainty categories: high, moderate, low and very low. The path from the design-based starting point to the final category is transparent, but the domains are not averaged into a numerical score. Reviewers should document each reason for rating down or up and avoid double-counting the same underlying problem across domains.1

Avoid averaging domains

“Serious” risk of bias and “not serious” imprecision do not average to a middle score. Each domain addresses a different threat to certainty.

Keep the rationale visible

The final category should remain traceable to the evidence considered and to explicit reasons for every rating-down or rating-up decision.

10. Reproducibility requires independent judgment and an audit trail

Cochrane MECIR C74–C75 requires assessment of certainty for each outcome and requires all certainty judgments to be justified and documented. MECIR states that, ideally, two people should assess certainty independently and reach consensus on downgrading decisions.9

That audit trail matters because GRADE contains structured judgments rather than an automatic scoring algorithm. Another trained reviewer should be able to identify the target of certainty, the evidence considered for each domain and the rationale linking those judgments to the final category.

AI-assisted assessment requires the same boundary. A 2025 BMJ editorial on AI-assisted certainty rating argues that gains in efficiency should not compromise trustworthiness. AI can support parts of the evidence workflow, but human responsibility for the methodological judgments and recommendations remains essential.10

11. Difficult cases need scope discipline

The Core GRADE series summarized here focuses on comparative intervention-effect evidence. Other evidence structures may require additional or different guidance. Network meta-analysis introduces indirect comparisons and network-specific certainty considerations; prognostic evidence uses dedicated GRADE guidance; qualitative evidence uses GRADE-CERQual. The appropriate response is to use the relevant specialist framework rather than force every evidence structure into a pairwise intervention template.

12. Worked reasoning: one hypothetical outcome

Consider a hypothetical review of six randomized trials comparing inhaled corticosteroid plus a long-acting beta-agonist with inhaled corticosteroid alone in adults with moderate persistent asthma. The critical outcome is exacerbation requiring oral corticosteroids over 12 months. Assume a comparator risk of 250 per 1,000, a pooled RR of 0.68 (95% CI 0.55–0.78), and a prespecified MID of 50 fewer exacerbations per 1,000. The point estimate corresponds to 80 fewer per 1,000, and the confidence interval corresponds approximately to 113 fewer to 55 fewer per 1,000. All values are fabricated for teaching and are not derived from real trials.

Judgment Hypothetical reasoning
Starting certainty High, because the body consists of randomized trials.
Risk of bias Not serious: most information is assumed to come from low-risk trials and the estimate is stable in a hypothetical sensitivity analysis.
Inconsistency Not serious: point estimates and confidence intervals are assumed to be compatible relative to the threshold; a hypothetical I² of 22% is supportive information rather than the decision rule.
Indirectness Not serious: a minor hypothetical population difference is judged unlikely to change the effect for the target PICO.
Imprecision Not serious in this example: the entire absolute-effect confidence interval remains beyond the prespecified MID of 50 fewer per 1,000.
Publication bias Not strongly suspected in the hypothetical example after considering registration and study patterns; six studies are acknowledged as too few for a reliable asymmetry test.
Final certainty High in this hypothetical example because no rating-down domain is judged serious. Rating up is not applicable because randomized evidence already starts at high certainty.

The educational point is not that these labels follow automatically from the numbers. A different baseline risk, MID, risk-of-bias contribution, pattern of study results or evidence about missing studies could change the judgments while leaving the relative effect estimate unchanged.

13. Certainty should feed a Summary of Findings table, not stand in for a recommendation

Once the outcome-level judgment is complete, the final category and a concise explanation can be transferred into a GRADE Summary of Findings Table Template for Systematic Reviews. Core GRADE 6 emphasizes presenting effects and certainty in a compact, interpretable Summary of Findings format while retaining the reasoning needed to understand the judgment.11

High-certainty evidence does not by itself imply a strong recommendation. Core GRADE 7 treats certainty as one input to a broader Evidence-to-Decision process that also considers benefits and harms, values, resources, equity, acceptability and feasibility.12

Conclusion

A defensible GRADE assessment is a chain of explicit judgments. Define the outcome and the target of certainty. Start from the evidence design. Judge risk of bias, inconsistency, indirectness, imprecision and publication bias for the body of evidence, considering rating up only when current criteria apply. Then assign one outcome-level certainty category and preserve the rationale so another reviewer can understand the path.

GRADE certainty methods

Frequently asked questions

What are the five GRADE domains that can lower certainty?

Risk of bias, inconsistency, indirectness, imprecision and publication bias. They address different threats to confidence in an outcome-level certainty claim and should not be averaged into a numerical score.1

What threshold should I use when rating certainty?

The threshold depends on the target of certainty. Core GRADE 2 commonly uses the minimal important difference or the null, and GRADE Guidance 41 recommends stating the chosen threshold or range directly rather than assigning a contextualization category.2,5

Can I rate down for inconsistency just because I² is high?

No. I² is supportive information and may be misleading. Core GRADE 3 prioritizes visual and clinical interpretation of variability, including differences in point estimates, confidence-interval overlap and the relationship of study estimates to the chosen threshold.4

What is the role of optimal information size in current GRADE imprecision assessment?

The confidence interval relative to the chosen threshold is the primary Core GRADE criterion. OIS becomes especially relevant when the interval does not cross the threshold but the observed effect is unusually large and the total sample size or number of events is limited.2

Can non-randomized evidence be rated up?

Yes, in appropriate circumstances. Core GRADE considers large or very large effects and credible dose-response gradients as rating-up situations for methodologically rigorous NRSI. Dose-response upgrading is capped at one level.3,8

Does high-certainty evidence automatically support a strong recommendation?

No. Certainty is one input to recommendation development. Evidence-to-Decision frameworks also consider benefits and harms, values, resources, equity, acceptability and feasibility.12

Methodological sources

References

Primary GRADE guidance and current Cochrane methodological standards were used wherever available. References are ordered by first citation in the article.

  1. Guyatt G, Agoritsas T, Brignardello-Petersen R, Mustafa RA, Rylance J, Foroutan F, et al. Core GRADE 1: overview of the Core GRADE approach. BMJ. 2025;389:e081903. doi:10.1136/bmj-2024-081903.
  2. Guyatt G, Zeng L, Brignardello-Petersen R, Prasad M, De Beer H, Murad MH, et al. Core GRADE 2: choosing the target of certainty rating and assessing imprecision. BMJ. 2025;389:e081904. doi:10.1136/bmj-2024-081904.
  3. Guyatt G, Wang Y, Eachempati P, Iorio A, Murad MH, Hultcrantz M, et al. Core GRADE 4: rating certainty of evidence, risk of bias, publication bias, and reasons for rating up certainty. BMJ. 2025;389:e083864. doi:10.1136/bmj-2024-083864. Correction: BMJ. 2025;390:r1468. doi:10.1136/bmj.r1468.
  4. Guyatt G, Schandelmaier S, Brignardello-Petersen R, De Beer H, Prasad M, Murad MH, et al. Core GRADE 3: rating certainty of evidence, assessing inconsistency. BMJ. 2025;389:e081905. doi:10.1136/bmj-2024-081905.
  5. Hultcrantz M, Schünemann HJ, Mustafa RA, Rind DM, Murad MH, Mayer M, et al. GRADE Certainty Ratings: Thresholds Rather Than Categories of Contextualization (GRADE Guidance 41). Ann Intern Med. 2025;178(8):1183–1186. doi:10.7326/ANNALS-25-00548.
  6. Guyatt G, Iorio A, De Beer H, Owen A, Agoritsas T, Murad MH, et al. Core GRADE 5: rating certainty of evidence, assessing indirectness. BMJ. 2025;389:e083865. doi:10.1136/bmj-2024-083865.
  7. Zeng L, Brignardello-Petersen R, Hultcrantz M, Mustafa RA, Murad MH, Iorio A, et al. GRADE Guidance 34: update on rating imprecision using a minimally contextualized approach. J Clin Epidemiol. 2022;150:216–224. doi:10.1016/j.jclinepi.2022.07.014.
  8. Murad MH, Verbeek J, Schwingshackl L, Filippini T, Vinceti M, Akl EA, et al. GRADE Guidance 38: updated guidance for rating up certainty of evidence due to a dose-response gradient. J Clin Epidemiol. 2023;164:45–53. doi:10.1016/j.jclinepi.2023.09.011.
  9. Cochrane. Methodological Expectations of Cochrane Intervention Reviews (MECIR), C74–C75: assessing the certainty of the body of evidence and justifying assessments. Current online standard. Cochrane MECIR C74–C75.
  10. Wu J, Guyatt G, Yao L. Efficiency must not compromise trustworthiness in rating certainty and formulating recommendations in AI era. BMJ. 2025;389:r1105. doi:10.1136/bmj.r1105.
  11. Guyatt G, Yao L, Murad MH, Hultcrantz M, Agoritsas T, De Beer H, et al. Core GRADE 6: presenting the evidence in summary of findings tables. BMJ. 2025;389:e083866. doi:10.1136/bmj-2024-083866.
  12. Guyatt G, Vandvik PO, Iorio A, Agarwal A, Yao L, Eachempati P, et al. Core GRADE 7: principles for moving from evidence to recommendations and decisions. BMJ. 2025;389:e083867. doi:10.1136/bmj-2024-083867.

Similar Posts

Leave a Reply

Your email address will not be published. Required fields are marked *