GRADE Evidence-Certainty Assessment and Summary of Findings

A meta-analysis can estimate an average effect with impressive numerical precision and still leave a decision-maker unsure how much confidence to place in that estimate. Statistical synthesis and certainty assessment answer different questions. The Grading of Recommendations Assessment, Development and Evaluation (GRADE) approach was created to make that second question explicit, structured and open to scrutiny.1,2,25

GRADE is not a score added after the analysis is complete. It is a framework for judging the certainty of a body of evidence for each important outcome, communicating the basis for that judgment, and keeping the assessment of evidence separate from the later process of formulating recommendations.14,25,31 That separation matters because evidence may be highly certain yet support only a conditional recommendation, or uncertain yet still justify a strong recommendation in unusual circumstances. Certainty is one input to a decision, not the decision itself.

Operating principle GRADE asks how certain we are about a specified effect for a specified outcome, population, intervention or exposure, comparator and setting. It does not assign one global grade to a study, a meta-analysis or an entire systematic review.

1Why a pooled effect is not a certainty judgment

A pooled estimate describes what the selected statistical model calculates from the available studies. Its interpretation depends on the estimand, effect measure, included evidence, assumptions about heterogeneity and dependence, and the integrity of the underlying data. None of those features guarantees that the result directly answers the target question or that the confidence interval excludes effects that would lead to different decisions.

Consider two reviews reporting the same risk ratio and the same nominal 95% confidence interval. In one, the contributing trials may be well conducted, consistent, directly applicable and comprehensively reported. In the other, allocation may be inadequately concealed, results may vary without a credible explanation, the population may differ from the target population, and several registered outcomes may be missing. The numerical summary can look similar while the justified confidence in it is very different. GRADE makes the reasons for that difference visible.

The framework therefore begins before any certainty label is assigned. Reviewers must define the question and identify outcomes that matter to patients, service users, policy-makers or other affected groups. Certainty is then assessed separately for each outcome because the contributing studies, measurement quality, missingness, precision and applicability may differ across outcomes within the same comparison.2,3,5

Risk of bias

Concerns about systematic error in the studies contributing to a particular result.

Certainty of evidence

Confidence in the effect estimate or in its position relative to a decision-relevant threshold for one outcome.

Summary of Findings

A concise presentation of effects, participant and study counts, certainty ratings and explanations for prioritized outcomes.

Recommendation

A judgment about what should be done after considering benefits, harms and other decision criteria in addition to certainty.

From a review question to a recommendation A five-stage pathway moves from a defined question and prioritized outcomes to synthesized effects, outcome-level certainty judgments, a Summary of Findings table, and a separate Evidence-to-Decision process. Question PICO & target outcomes defined Synthesis Relative & absolute effects Certainty Outcome-level GRADE judgment Summary Effects, certainty & explanations Decision Benefits, harms & context The certainty judgment and the recommendation are connected, but they are not the same judgment. From a review question to a recommendation A vertical five-stage pathway moves from the question through synthesis, certainty assessment and evidence summary to a separate decision process. 1. Question PICO & target outcomes defined 2. Synthesis Relative & absolute effects 3. Certainty Outcome-level GRADE judgment 4. Summary Effects, certainty & explanations 5. Decision Benefits, harms & context Certainty and recommendation are connected, but not identical.
Figure 1. GRADE places a transparent certainty assessment between statistical synthesis and decision-making. A Summary of Findings table communicates the evidence; an Evidence-to-Decision framework considers what action, if any, should follow.

2What GRADE means by certainty

GRADE uses four categories: high, moderate, low and very low certainty. These are ordered summaries of confidence, not numerical probabilities and not points on a psychometric scale. In contemporary GRADE thinking, the target of the judgment should be stated clearly. Reviewers may be assessing confidence that the true effect lies on one side of a threshold, within a specified range, or sufficiently close to the estimated effect for the intended purpose.1,2,10,26

HighVery confident that the true effect is close to the estimate or target conclusion.
ModerateModerately confident; the true effect is probably close, but may be meaningfully different.
LowLimited confidence; the true effect may be substantially different.
Very lowVery little confidence; the true effect is likely to be substantially different.

The language attached to these categories should be adapted to the target of the assessment. If the question is whether an intervention produces an important benefit, a high-certainty judgment concerns confidence about that threshold-based conclusion. If the question is the magnitude of an effect, the judgment concerns how closely the estimate is expected to represent the true effect. Stating the target avoids a common ambiguity in which reviewers assign a category without explaining what, exactly, they are certain about.

For conventional intervention-effect questions, randomized trials usually begin at high certainty and non-randomized studies at low certainty in the original GRADE framework. The starting point is then reconsidered through explicit domains. This design-based starting rule should not be transferred mechanically to every evidence question. Diagnostic accuracy, prognosis, prevalence, qualitative findings, network meta-analysis and modelled evidence require guidance adapted to their evidence structures.2,21,22,24

3The domains are reasons for judgment, not automatic penalties

Five domains can lower certainty in a body of evidence for an outcome. Each asks a different question. Their purpose is to identify why confidence should change, not to generate a total score. A concern may justify no downgrade, one level or, when sufficiently serious, more than one level. The rationale must be tied to the evidence contributing to that outcome and to the target conclusion.

Risk of biasCould limitations in study design or conduct systematically distort the effect estimate?
InconsistencyDo effect estimates differ in direction or magnitude beyond a credible, explained pattern?
IndirectnessDoes the evidence differ from the target population, intervention, comparator, outcome or setting?
ImprecisionDoes the confidence interval include effects that would support materially different conclusions?
Publication biasIs important evidence likely to be missing because results, outcomes or analyses were selectively unavailable?

Risk of bias

The risk-of-bias domain concerns the studies and results that contribute to the outcome, not the reputation of a study design in the abstract. Reviewers examine design-specific judgments, the amount and weight of evidence at different risk levels, the direction of likely bias and whether sensitivity analyses change the interpretation. Averaging domain labels or converting appraisal tools into numerical quality scores discards information that GRADE needs.5,28

Inconsistency

Inconsistency is not synonymous with a large I2. The judgment considers the direction and magnitude of effects, confidence-interval overlap, absolute as well as relative effects, plausible effect modifiers and whether prespecified explanations account for the variation. A high heterogeneity statistic can coexist with effects that support the same practical conclusion; a low statistic can conceal an important difference when evidence is sparse or imprecise.6,7,27

Indirectness

Evidence is indirect when it does not fully match the target question. Differences in population, intervention or exposure, comparator, outcome and setting may matter if they are expected to change the effect or its interpretation. A surrogate outcome is not automatically invalid, but it must be identified as a substitute for a patient-important outcome and its credibility considered. Indirect comparisons add another form of indirectness when the interventions of interest have not been compared directly.8,29

Imprecision

Imprecision is a decision-relevant judgment about the range of effects compatible with the data. The question is not simply whether the confidence interval crosses the no-effect value. Reviewers ask whether it includes effects on both sides of a threshold that would lead to different conclusions, and whether the information size is adequate for the claimed certainty. Updated guidance distinguishes minimally contextualized and contextualized approaches and refines when a two-level downgrade may be justified.9,10,26

Publication bias and missing evidence

Publication bias is one form of bias due to missing evidence. Funnel-plot asymmetry can raise concern, but it neither proves nor excludes publication bias. Registries, protocols, regulatory records, discrepancies between planned and reported outcomes, author correspondence and the pattern of small versus large studies may provide more direct evidence. The judgment should reflect the likelihood that missing results would materially change the conclusion.11,28

Three considerations may raise certainty in appropriate bodies of non-randomized evidence: a large magnitude of effect, a credible dose-response gradient, and plausible residual confounding that would reduce an observed effect or create an effect when none was observed. These are not rewards for an impressive result. They require rigorous evidence, attention to bias and precision, and a reasoned explanation of why the pattern increases confidence.12,13,28

Large effectThe magnitude is sufficiently large, credible and precise that common residual biases are unlikely to explain it.
Dose-response gradientThe gradient is analytically credible, not an ecological or confounded pattern, and supports no more than the justified upgrade.
Opposing residual confoundingAll plausible residual confounding would diminish the observed effect or would create a false effect in the opposite direction.
Do not double-count the same limitation. A wide confidence interval may inform imprecision; unexplained differences among study effects may inform inconsistency. The same underlying concern should not be used repeatedly to lower certainty under several labels unless each domain captures a genuinely distinct problem.
DomainEvidence to examineQuestion the explanation must answer
Risk of biasDesign-specific appraisal, contribution of studies, direction of bias, sensitivity analysesWhy could the contributing evidence systematically overestimate or underestimate this outcome?
InconsistencyEffect direction and magnitude, interval overlap, absolute effects, heterogeneity, modifiersWhy does the unexplained variation weaken the target conclusion?
IndirectnessPICO and setting match, surrogate status, directness of comparisonsWhy might the available evidence behave differently in the target question?
ImprecisionConfidence interval, decision thresholds, optimal information size, event countsWhich materially different conclusions remain compatible with the data?
Publication biasRegistries, protocols, missing outcomes, small-study patterns, sponsor and author recordsWhat evidence suggests that unavailable results could alter the conclusion?

A defensible GRADE assessment leaves a visible reasoning trail. Reviewers need structured ways to record outcome definitions, domain judgments, thresholds, effect estimates and explanations, and readers need guidance that shows how those records should be interpreted. The resources placed between this section and the continuation of the article support those different tasks without replacing methodological judgment.

Evidence certainty and decision frameworks

Assess certainty and communicate evidence with GRADE

Move from synthesized effects to transparent certainty judgments. Assess GRADE domains for each critical outcome, explain every rating decision, build Summary of Findings tables, and connect evidence to recommendations without collapsing judgment into a score.

GRADE Evidence-Certainty Assessment and Summary of Findings

GRADE

Assess certainty at the outcome level, communicate effects transparently in Summary of Findings tables, and connect evidence judgments to structured decision-making.

GRADE Summary of Findings Table template for systematic reviews

Present relative and absolute effects, participant counts, certainty judgments, and concise explanations in a clear manuscript-ready table.

GRADE certainty-of-evidence assessment worksheet for systematic reviews

Work through risk of bias, inconsistency, indirectness, imprecision, publication bias, upgrading factors, and support for every outcome-level judgment.

GRADE evidence-to-decision framework template for evidence synthesis

Translate evidence into recommendations by documenting benefits, harms, values, resources, equity, acceptability, and feasibility without collapsing judgment into a score.

4From domain judgments to one outcome-level rating

The final certainty category is not calculated by adding domain points. Reviewers begin at the starting level appropriate to the evidence structure, examine each domain, and decide whether a concern is serious enough to rate certainty down or whether a justified consideration supports rating it up. The resulting outcome-level category should reflect the combined, documented judgments without treating domains as arithmetic units or allowing one favourable feature to cancel an unrelated limitation.

Two teams can examine the same evidence and reach different ratings without either having made an arithmetic error. GRADE deliberately contains judgment. Reproducibility therefore depends less on pretending that judgment can be removed and more on making the target, evidence, thresholds and rationale explicit enough for another team to understand, challenge and, where appropriate, reproduce the conclusion.3,4,25

Define the target before judging the interval

Certainty can be assessed with different degrees of contextualization. In a non-contextualized assessment, the focus is confidence in the effect estimate itself, with limited reference to a specific decision threshold. A minimally contextualized assessment asks whether the effect is important or unimportant relative to a chosen threshold. A fully contextualized assessment considers the magnitude of all important benefits and harms and the thresholds that would alter a recommendation.10,26

These approaches are not interchangeable labels. They determine how imprecision and sometimes inconsistency are interpreted. A confidence interval crossing zero may be unimportant if every plausible effect would lead to the same practical conclusion. Conversely, an interval entirely on one side of zero may remain seriously imprecise if it spans effects that would lead to different decisions. The threshold must be justified rather than chosen after seeing the result.

Estimate focused

Non-contextualized

How confident are we in the estimated effect, without committing to a specific decision threshold?

Threshold focused

Minimally contextualized

How confident are we that the true effect is important, trivial or absent relative to a specified threshold?

Decision focused

Fully contextualized

How confident are we in the balance of all critical benefits and harms in relation to a decision? This requires explicit outcome importance and thresholds across the full evidence profile.

Explain the judgment at the level where it was made

A useful explanation identifies the domain, the evidence that caused concern, the direction or consequence of the limitation and the number of levels changed. “Downgraded for imprecision” is incomplete. “Downgraded one level because the confidence interval includes both a clinically important benefit and little or no benefit, with information below the prespecified optimal information size” is auditable. The same standard applies to decisions not to downgrade when an apparent concern does not threaten the target conclusion.

Explanations should also reveal whether the judgment was made independently by more than one reviewer, how disagreements were resolved, and whether the assessment changed after correction of data or clarification from study authors. The GRADE guidance group’s requirements emphasize domain-by-domain judgments, the standard four certainty categories and evidence tables that preserve the methods and reasons underlying the rating.4

5Evidence profiles and Summary of Findings tables serve different depths of communication

GRADE produces two closely related forms of evidence presentation. An evidence profile retains the detailed domain judgments and their explanations for each outcome. A Summary of Findings table presents the main decision-relevant results in a compact format: the outcome, anticipated absolute effects, relative effect where appropriate, number of participants and studies, certainty rating and concise explanations.1,3,14,30

The table is not a decorative summary placed at the end of a review. Its structure forces decisions about which outcomes matter, which time points and effect measures are appropriate, what baseline risk should be used, and how relative effects translate into absolute differences. These choices should be made consistently with the protocol and the target question, not selected to make a result appear more favourable.

FeatureEvidence profileSummary of Findings table
Primary purposePreserve the detailed certainty assessment and rationale.Communicate the most important effects and certainty judgments concisely.
Domain detailShows judgments for risk of bias, inconsistency, indirectness, imprecision, publication bias and relevant upgrading factors.Usually presents the final certainty category with footnoted explanations.
Effect presentationMay contain detailed evidence and calculations supporting the assessment.Prioritizes relative and absolute effects, study and participant counts, and the clearest decision-relevant interpretation.
AudienceReview teams, guideline panels and readers auditing the judgment.Clinicians, policy-makers, researchers, patients and other readers who need a concise evidence summary.

Outcome selection is part of the method

Cochrane guidance commonly recommends seven or fewer critical and important outcomes for a principal Summary of Findings table. This is a communicability limit, not a quota and not permission to omit inconvenient outcomes. Outcomes should be prioritized before results are known, with attention to patient importance, benefits, harms, follow-up and the decision the review is intended to inform.3,14,30

Surrogate outcomes require particular care. If a surrogate is used in place of a patient-important outcome, the substitution should be explicit and the resulting indirectness considered. A statistically precise surrogate effect does not automatically provide high-certainty evidence about the health outcome that matters. Likewise, an outcome measured with different instruments or at different time points may require careful definition before evidence is combined or displayed.

Absolute effects need a defensible baseline risk

Relative effects often transfer more consistently across baseline risks than absolute effects, but decisions frequently depend on absolute benefits and harms. For binary outcomes, a relative risk is commonly applied to a specified control-group risk to estimate the absolute difference. The chosen baseline risk may come from the included studies, a representative population or a relevant risk stratum. Its source and applicability should be stated because different baseline risks produce different absolute effects even when the relative effect is unchanged.14

For continuous, time-to-event, diagnostic and other outcomes, the presentation method must match the measure and the decision. A standardized mean difference may need translation into a familiar instrument or threshold. A hazard ratio does not directly state how many events are prevented at a particular time. Software can perform calculations, but it cannot decide which baseline risk, threshold or presentation is scientifically appropriate.

Footnotes are part of the evidence, not an optional annotation. A certainty symbol without an explanation conceals the reasoning. Concise footnotes should state why certainty changed, identify the evidence behind the judgment and allow the reader to trace the conclusion back to the review methods and results.

6Certainty and recommendations must remain separate

A certainty assessment asks what confidence can be placed in the evidence for an outcome. A recommendation asks what should be done. The second judgment requires more than the first. GRADE Evidence-to-Decision (EtD) frameworks bring the relevant considerations into one transparent structure: the priority of the problem, desirable and undesirable effects, certainty of evidence, values and preferences, resource requirements, equity, acceptability and feasibility. The exact criteria vary with clinical, public-health, health-system, coverage and other decision types.1618,31

High-certainty evidence of a small benefit does not automatically justify a strong recommendation if harms, costs, inequity or variable preferences are substantial. Conversely, low-certainty evidence usually limits confidence in a recommendation, although GRADE recognizes exceptional situations in which a strong recommendation may still be justified. The panel must state the exceptional rationale clearly; the recommendation does not erase the uncertainty in the evidence.31

Certainty assessment and Evidence-to-Decision framework The upper panel shows an evidence question leading to effect estimates and outcome-level certainty. A separate lower panel combines certainty with benefits and harms, values, resources, equity, acceptability and feasibility to reach a recommendation. Evidence Assessment Defined Question and prioritized outcomes Effect Estimates relative and absolute Certainty Rating for each outcome Evidence-to-Decision Deliberation Benefits & Harms Values & Preferences Resources & Equity Acceptability & Feasibility Direction and strength of recommendation Certainty assessment and Evidence-to-Decision framework A vertical diagram separates evidence assessment from the later Evidence-to-Decision deliberation that produces a recommendation. Evidence Assessment 1. Defined Question PICO & outcomes 2. Effect Estimates Relative & absolute 3. Certainty Rating For each outcome Evidence-to-Decision Deliberation Benefits & Harms Values & Preferences Resources & Equity Acceptability & Feasibility Direction & Strength of Recommendation Certainty informs the decision; it does not determine it.
Figure 2. The evidence assessment and the recommendation are sequential but distinct. An EtD framework prevents certainty from being treated as a proxy for the balance of consequences or for contextual judgments.

Strong and conditional recommendations

A strong recommendation indicates that a panel is sufficiently confident that the desirable consequences outweigh the undesirable consequences, or vice versa, for almost all people in the specified context. A conditional recommendation indicates that the best choice may vary because the balance is close, preferences differ, evidence is uncertain or implementation factors matter. The wording should state the direction, strength, population and relevant conditions. A conditional recommendation is not weak science, and a strong recommendation is not proof of a large effect.17,31

Structured EtD use can improve transparency because the panel must show which criteria were considered and how each influenced the conclusion. Empirical work has found associations between EtD use and higher-quality, more credible guideline recommendations, although such studies do not prove that the framework alone causes better decisions.18,19

7Reproducibility depends on process, not software alone

GRADEpro GDT supports evidence profiles, Summary of Findings tables and EtD frameworks, and is the standard platform used for Cochrane GRADE tables. It can organize calculations, judgments, team roles and outputs. It cannot decide whether an outcome is critical, whether a surrogate is credible, which threshold matters, whether heterogeneity is important, or how missing evidence affects confidence. Those remain scientific judgments that must be justified.15

A rigorous workflow normally includes advance specification of outcomes and methods, training and calibration of assessors, independent judgment or a clearly described verification process, reconciliation of disagreements, and preservation of the original reasons for each decision. Rapid reviews may use streamlined approaches, such as one assessor with verification, when resources are constrained, but the core domains, outcome-level ratings, standard terminology and explanations should remain intact.20

Report explicitlyWhy it matters
The target question, population, intervention or exposure, comparator, outcomes, time points and settingReaders cannot judge directness or thresholds without knowing the intended target.
Which outcomes were critical or important and how they were selectedOutcome selection should not be driven by favourable or statistically significant results.
The starting certainty and guidance used for the evidence typeDifferent questions and evidence structures may require different GRADE adaptations.
Every domain judgment and reason for rating down, rating up or making no changeThe final category is interpretable only through its reasoning trail.
Decision thresholds, baseline risks and contextualization approachThese choices determine how imprecision and absolute effects are interpreted.
Reviewer process, disagreements, amendments, software and versionTransparent process makes the assessment auditable and updateable.

Common misapplications

One rating for the whole review

Certainty differs by outcome and sometimes by comparison, time point or population. A single review-wide label conceals those differences.

A numerical GRADE score

Adding points across domains implies a measurement precision and compensation rule that the framework does not support.

Automatic rules from statistics

No fixed I2, p value, event count or confidence-interval crossing can replace a contextual domain judgment.

Low certainty means no effect

Low certainty means limited confidence in the conclusion. It does not convert an estimated benefit, harm or null result into proof of no effect.

High certainty means recommend

Certainty does not determine the balance of benefits and harms, resources, values, equity, acceptability or feasibility.

Software-generated certainty

A completed table can still be methodologically weak if the outcomes, thresholds, domain explanations or evidence inputs are poorly chosen.

8Use the framework that matches the evidence structure

The core GRADE logic has been extended or adapted for evidence structures that differ from conventional pairwise intervention reviews. GRADE-CERQual assesses confidence in findings from qualitative evidence syntheses using methodological limitations, coherence, adequacy of data and relevance. It is related to GRADE but is not a substitute label for quantitative certainty.22

Network meta-analysis introduces additional issues, including the contribution of direct and indirect evidence, transitivity, incoherence and network-wide imprecision. CINeMA operationalizes six domains for confidence in network estimates and can support structured assessment, but it should not be treated as automatically interchangeable with a GRADE Working Group assessment. Comparative studies have found meaningful disagreement between approaches, particularly in complex networks.21,23

Complex interventions, diagnostic tests, prognostic factors, prevalence estimates and modelled evidence also require question-specific guidance. The correct response to a non-standard evidence structure is not to force it into an intervention template. Reviewers should identify the applicable methodological guidance, state any adaptations and avoid claiming generic “GRADE use” when essential domains or definitions have been altered.24

Scope statement. This article explains the architecture of GRADE certainty assessment, Summary of Findings tables and Evidence-to-Decision reasoning. Applying GRADE to a specific review still requires the current guidance for that review question, outcome type, study design and decision context.

Conclusion

Meta-analysis produces estimates. GRADE asks what confidence those estimates can carry for each important outcome and requires the reasons for that confidence to be documented. Summary of Findings tables communicate the effects and certainty in a form that readers can use. Evidence-to-Decision frameworks then place that evidence alongside benefits, harms, values, resources, equity, acceptability and feasibility before a recommendation is made.25,30,31

The strength of the approach lies neither in four labels nor in a software-generated table. It lies in disciplined separation: study limitations are separated from certainty; certainty is separated from recommendation; statistical signals are separated from automatic verdicts; and every consequential judgment is linked to evidence and an explanation. When those separations are preserved, GRADE does what a defensible evidence framework should do: it makes uncertainty visible without turning it into vagueness, and it makes judgment explicit without pretending that judgment is calculation.

The frequently asked questions and complete reference list below form the final section of this same article.

Get the Resource Infrastructure Free Forever

Get Free Lifetime Access

GRADE certainty of evidence

Frequently asked questions

What exactly does GRADE rate?

GRADE rates the certainty of a body of evidence for a specified outcome and target question. It does not give one quality grade to an individual study, a meta-analysis or an entire systematic review. Different outcomes in the same review can receive different certainty ratings because they may rely on different studies, measurements, event counts and assumptions.1,2,25

Is risk of bias the same as certainty of evidence?

No. Risk of bias concerns systematic error in the studies contributing to a result. Certainty of evidence is broader and also considers inconsistency, indirectness, imprecision, publication bias and, where applicable, factors that may raise certainty. A low-risk-of-bias body of evidence can still have low certainty because it is indirect or seriously imprecise.2,5,28

Does conducting a meta-analysis automatically increase certainty?

No. Meta-analysis can increase statistical precision when compatible evidence is combined appropriately, but it does not remove bias, indirectness, unexplained inconsistency or missing evidence. Pooling more studies can produce a narrow confidence interval around a biased or poorly targeted estimate. Certainty depends on the complete body of evidence and the target conclusion, not on whether a pooled diamond was produced.

Does low-certainty evidence mean that there is no effect?

No. Low certainty means that confidence in the estimated effect or threshold-based conclusion is limited and that the true effect may be substantially different. It is compatible with benefit, harm or little effect. “Low certainty” should not be rewritten as “no evidence,” “no effect” or “the intervention does not work.”

Does high-certainty evidence require a strong recommendation?

No. A recommendation also depends on the balance of benefits and harms, values and preferences, resources, equity, acceptability and feasibility. High-certainty evidence may support a conditional recommendation when those considerations vary or the net benefit is small. In exceptional situations, a strong recommendation may be made despite low certainty, but the reason should be explicit.1618,31

How many outcomes should a Summary of Findings table contain?

Cochrane guidance commonly recommends seven or fewer critical and important outcomes for a principal comparison. This is a communication guideline, not a quota. The table should contain the outcomes most important for understanding benefits and harms, selected without knowledge of which results are favourable.3,14,30

Must every concern lead to a downgrade?

No. A domain changes certainty only when the concern is serious enough to weaken confidence in the target conclusion. Reviewers should explain both downgrading decisions and important decisions not to downgrade. Mechanical rules based on one statistic or checklist response are inconsistent with the judgment-based structure of GRADE.

Can GRADEpro GDT make the certainty judgment automatically?

No. GRADEpro GDT organizes evidence profiles, Summary of Findings tables, judgments and Evidence-to-Decision frameworks. Reviewers still have to define the outcomes, select thresholds and baseline risks, interpret each domain and justify the final rating. Software improves organization; it does not replace methodological judgment.15

Are GRADE-CERQual and CINeMA interchangeable with standard GRADE?

No. GRADE-CERQual addresses confidence in findings from qualitative evidence syntheses. CINeMA addresses confidence in network meta-analysis estimates through a network-specific domain structure. Both are related to GRADE, but each has its own definitions and procedures. The selected framework should match the evidence structure and be reported explicitly.2123

Source record

References

References follow Vancouver order of first citation and use superscript in-text numbering consistent with Nature-style presentation. Institutional guidance is linked to the official version used for verification.

  1. Guyatt G, Oxman AD, Akl EA, Kunz R, Vist G, Brozek J, et al. GRADE guidelines: 1. Introduction—GRADE evidence profiles and summary of findings tables. J Clin Epidemiol. 2011;64(4):383–394. doi:10.1016/j.jclinepi.2010.04.026.
  2. Balshem H, Helfand M, Schünemann HJ, Oxman AD, Kunz R, Brozek J, et al. GRADE guidelines: 3. Rating the quality of evidence. J Clin Epidemiol. 2011;64(4):401–406. doi:10.1016/j.jclinepi.2010.07.015.
  3. Schünemann HJ, Higgins JPT, Vist GE, Glasziou P, Akl EA, Skoetz N, Guyatt GH. Chapter 14: Completing “Summary of findings” tables and grading the certainty of the evidence. In: Higgins JPT, Thomas J, Chandler J, Cumpston M, Li T, Page MJ, Welch VA, editors. Cochrane Handbook for Systematic Reviews of Interventions. Current online version. Cochrane; updated August 2023. Available from: Cochrane Handbook Chapter 14.
  4. Schünemann HJ, Brennan S, Akl EA, Hultcrantz M, Alonso-Coello P, Xia J, et al. The development methods of official GRADE articles and requirements for claiming the use of GRADE: a statement by the GRADE guidance group. J Clin Epidemiol. 2023;159:79–84. doi:10.1016/j.jclinepi.2023.05.010.
  5. Guyatt GH, Oxman AD, Vist G, Kunz R, Brozek J, Alonso-Coello P, et al. GRADE guidelines: 4. Rating the quality of evidence—study limitations (risk of bias). J Clin Epidemiol. 2011;64(4):407–415. doi:10.1016/j.jclinepi.2010.07.017.
  6. Guyatt GH, Oxman AD, Kunz R, Woodcock J, Brozek J, Helfand M, et al. GRADE guidelines: 7. Rating the quality of evidence—inconsistency. J Clin Epidemiol. 2011;64(12):1294–1302. doi:10.1016/j.jclinepi.2011.03.017.
  7. Guyatt G, Zhao Y, Mayer M, Briel M, Mustafa R, Izcovich A, et al. GRADE guidance 36: updates to GRADE’s approach to addressing inconsistency. J Clin Epidemiol. 2023;158:70–83. doi:10.1016/j.jclinepi.2023.03.003.
  8. Guyatt GH, Oxman AD, Kunz R, Woodcock J, Brozek J, Helfand M, et al. GRADE guidelines: 8. Rating the quality of evidence—indirectness. J Clin Epidemiol. 2011;64(12):1303–1310. doi:10.1016/j.jclinepi.2011.04.014.
  9. Guyatt GH, Oxman AD, Kunz R, Brozek J, Alonso-Coello P, Rind D, et al. GRADE guidelines 6. Rating the quality of evidence—imprecision. J Clin Epidemiol. 2011;64(12):1283–1293. doi:10.1016/j.jclinepi.2011.01.012.
  10. Zeng L, Brignardello-Petersen R, Hultcrantz M, Mustafa RA, Murad MH, Iorio A, et al. GRADE Guidance 34: update on rating imprecision using a minimally contextualized approach. J Clin Epidemiol. 2022;150:216–224. doi:10.1016/j.jclinepi.2022.07.014.
  11. Guyatt GH, Oxman AD, Montori V, Vist G, Kunz R, Brozek J, et al. GRADE guidelines: 5. Rating the quality of evidence—publication bias. J Clin Epidemiol. 2011;64(12):1277–1282. doi:10.1016/j.jclinepi.2011.01.011.
  12. Guyatt GH, Oxman AD, Sultan S, Glasziou P, Akl EA, Alonso-Coello P, et al. GRADE guidelines: 9. Rating up the quality of evidence. J Clin Epidemiol. 2011;64(12):1311–1316. doi:10.1016/j.jclinepi.2011.06.004.
  13. Murad MH, Verbeek J, Schwingshackl L, Filippini T, Vinceti M, Akl EA, et al. GRADE guidance 38: updated guidance for rating up certainty of evidence due to a dose-response gradient. J Clin Epidemiol. 2023;164:45–53. doi:10.1016/j.jclinepi.2023.09.011.
  14. Guyatt GH, Oxman AD, Santesso N, Helfand M, Vist G, Kunz R, et al. GRADE guidelines: 12. Preparing Summary of Findings tables—binary outcomes. J Clin Epidemiol. 2013;66(2):158–172. doi:10.1016/j.jclinepi.2012.01.012.
  15. Cochrane GRADEing Methods Group. GRADEpro GDT. Cochrane Methods. Available from: GRADEpro GDT. User documentation available from: GRADEpro GDT User Guide.
  16. Alonso-Coello P, Schünemann HJ, Moberg J, Brignardello-Petersen R, Akl EA, Davoli M, et al. GRADE Evidence to Decision frameworks: a systematic and transparent approach to making well informed healthcare choices. 1: Introduction. BMJ. 2016;353:i2016. doi:10.1136/bmj.i2016.
  17. Andrews J, Guyatt G, Oxman AD, Alderson P, Dahm P, Falck-Ytter Y, et al. GRADE guidelines: 14. Going from evidence to recommendations: the significance and presentation of recommendations. J Clin Epidemiol. 2013;66(7):719–725. doi:10.1016/j.jclinepi.2012.03.013.
  18. Li SA, Alexander PE, Reljic T, Cuker A, Nieuwlaat R, Wiercioch W, et al. Evidence to Decision framework provides a structured “roadmap” for making GRADE guidelines recommendations. J Clin Epidemiol. 2018;104:103–112. doi:10.1016/j.jclinepi.2018.09.007.
  19. Meneses-Echavez JF, Bidonde J, Montesinos-Guevara C, Amer YS, Loaiza-Betancur AF, Tellez-Tinjaca LA, et al. Using Evidence to Decision frameworks led to guidelines of better quality and more credible and transparent recommendations. J Clin Epidemiol. 2023;162:38–46. doi:10.1016/j.jclinepi.2023.07.013.
  20. Gartlehner G, Nussbaumer-Streit B, Devane D, Kahwati L, Viswanathan M, King VJ, et al. Rapid reviews methods series: guidance on assessing the certainty of evidence. BMJ Evid Based Med. 2024;29(1):50–54. doi:10.1136/bmjebm-2022-112111.
  21. Nikolakopoulou A, Higgins JPT, Papakonstantinou T, Chaimani A, Del Giovane C, Egger M, et al. CINeMA: an approach for assessing confidence in the results of a network meta-analysis. PLoS Med. 2020;17(4):e1003082. doi:10.1371/journal.pmed.1003082.
  22. Lewin S, Booth A, Glenton C, Munthe-Kaas H, Rashidian A, Wainwright M, et al. Applying GRADE-CERQual to qualitative evidence synthesis findings: introduction to the series. Implement Sci. 2018;13(Suppl 1):2. doi:10.1186/s13012-017-0688-3.
  23. Noori A, Sadeghirad B, Thabane L, Bhandari M, Guyatt GH, Busse JW. The GRADE Working Group and CINeMA approaches provided inconsistent certainty of evidence ratings for a network meta-analysis of opioids for chronic noncancer pain. J Clin Epidemiol. 2024;169:111276. doi:10.1016/j.jclinepi.2024.111276.
  24. Montgomery P, Movsisyan A, Grant SP, Macdonald G, Rehfuess EA. Considerations of complexity in rating certainty of evidence in systematic reviews: a primer on using the GRADE approach in global health. BMJ Glob Health. 2019;4(Suppl 1):e000848. doi:10.1136/bmjgh-2018-000848.
  25. Guyatt G, et al. Core GRADE 1: overview of the Core GRADE approach. BMJ. 2025;389:e081903. doi:10.1136/bmj-2024-081903.
  26. Guyatt G, et al. Core GRADE 2: choosing the target of certainty rating and assessing imprecision. BMJ. 2025;389:e081904. doi:10.1136/bmj-2024-081904.
  27. Guyatt G, et al. Core GRADE 3: rating certainty of evidence—assessing inconsistency. BMJ. 2025;389:e081905. doi:10.1136/bmj-2024-081905.
  28. Guyatt G, et al. Core GRADE 4: rating certainty of evidence—risk of bias, publication bias, and reasons for rating up certainty. BMJ. 2025;389:e083864. doi:10.1136/bmj-2024-083864.
  29. Guyatt G, et al. Core GRADE 5: rating certainty of evidence—assessing indirectness. BMJ. 2025;389:e083865. doi:10.1136/bmj-2024-083865.
  30. Guyatt G, et al. Core GRADE 6: presenting the evidence in Summary of Findings tables. BMJ. 2025;389:e083866. doi:10.1136/bmj-2024-083866.
  31. Guyatt G, et al. Core GRADE 7: principles for moving from evidence to recommendations and decisions. BMJ. 2025;389:e083867. doi:10.1136/bmj-2024-083867.