Network Meta-analysis Assumptions Checklist

Practical appraisal resourceBy Dr. Esmaeel Saeedy RobatEvidence checked: 9 August 2026

A connected treatment network is not evidence that the network is credible. Network meta-analysis can combine direct and indirect treatment effects only when its clinical and methodological assumptions are sufficiently plausible. The difficult part is that the most important assumption, transitivity, cannot be established by a single statistical test. It has to be argued from the decision question, treatment definitions, study eligibility, and the distribution of variables that could modify relative treatment effects.1,2,18


This checklist is designed for that appraisal step. It asks whether the evidence network is structurally usable, whether the treatment comparisons are sufficiently comparable to support indirect inference, whether heterogeneity and direct-indirect disagreement were examined appropriately, and whether bias and treatment rankings were interpreted with adequate caution. It deliberately does not calculate an NMA-validity score. Current NMA frameworks rely on structured judgments because several consequential concerns cannot be reduced to defensible numerical thresholds.3,4,8,16


Use it before trusting the league table. The checklist can be applied prospectively when planning an NMA, during methodological peer review, or retrospectively when appraising a published analysis. It is an appraisal aid, not a substitute for an NMA statistician, domain expert, CINeMA assessment, GRADE judgment, RoB NMA assessment, or ROB-MEN assessment.

What the checklist is testing

Three ideas are often blurred together. Heterogeneity concerns variation in effects among studies informing a comparison. Transitivity concerns whether different sets of studies can support valid indirect comparisons, especially whether plausible treatment-effect modifiers are distributed sufficiently similarly across those sets. Consistency, also called coherence in some frameworks, concerns agreement between direct and indirect evidence where both are available. These ideas interact, but one cannot stand in for another.1,5,6


Connectivity and credibility are different questions Two equally connected triangular treatment networks are shown. In the left network, effect modifier distributions are similar across comparisons. In the right network, one comparison has a different severity distribution, creating a transitivity concern despite identical connectivity. Connectivity tells us whether treatments can be linked It does not establish the credibility of the indirect comparison Connected + plausible comparability Connected + transitivity concern A B C Severity and prior treatment distributions are similar across A–B, A–C and B–C trials A B C B–C trials enrol a systematically different severity group: geometry is unchanged, credibility is not Network structure creates indirect pathways. Clinical and methodological comparability determines whether those pathways are defensible.
Figure 1. Identical network geometry can support very different credibility judgments. Connectivity is necessary for conventional network estimation across all nodes, but it does not establish transitivity.

Why treatment-effect modifiers deserve their own attention

A prognostic factor predicts outcome regardless of treatment. A treatment-effect modifier changes the relative effect of one treatment compared with another. The distinction matters because transitivity is threatened by systematic differences in modifiers across the sets of trials that form an indirect comparison. Baseline severity, treatment line, prior therapy, baseline risk, dose, follow-up, co-interventions and trial era can be plausible modifiers in some clinical settings, but they are not automatically modifiers in every network. They should be identified from clinical and biological reasoning, previous evidence, subgroup knowledge and decision context rather than selected only because a post-hoc interaction happens to reach statistical significance.2,5


No universal balance threshold exists. Neither Stage 1 nor the final evidence audit identified a validated cut-off that converts a difference in an effect modifier into “acceptable” or “unacceptable” transitivity. The safe question is whether clinically important imbalance is present, how much information is missing, and whether the resulting uncertainty changes the credibility of the indirect estimate.4,17

Three distinct appraisal layers in network meta-analysis Three horizontal panels show heterogeneity within comparisons, transitivity across comparisons, and consistency between direct and indirect evidence. Arrows show that these layers inform one another but answer different questions. Do not use one assumption as a proxy for another 1 Heterogeneity: within a treatment comparison Do studies estimating the same contrast vary clinically, methodologically or statistically? 2 Transitivity: across different treatment comparisons Are plausible treatment-effect modifiers sufficiently comparable to support indirect inference? 3 Consistency / coherence: direct versus indirect evidence Where both sources exist, do they agree within the limits of statistical power and uncertainty? Low heterogeneity does not prove transitivity. A non-significant inconsistency test does not prove validity.
Figure 2. Heterogeneity, transitivity and consistency operate at different levels of the evidence structure. A defensible appraisal keeps them separate, then considers how concerns interact.

How to use the checklist

Complete the domains in order. Start with the decision question and network structure, because later judgments are meaningless if the treatment nodes or target contrast are unclear. Then examine comparability and effect modifiers before looking at statistical incoherence. A statistical test is a diagnostic aid, not a substitute for the conceptual assessment. Finally, examine bias, missing evidence and ranking interpretation before deciding whether the network warrants specialist review or cautious use.1,6,7,8


No important concern identified
Available information does not reveal a material concern for this checkpoint.
Concern identified
A feature could materially weaken one or more network estimates and needs explanation or sensitivity analysis.
Insufficient information
Reporting is inadequate to make the judgment. Do not convert missing information into reassurance.
Not applicable / not assessable
The checkpoint cannot apply to the network structure, such as direct-indirect incoherence testing in a loop-free network.
Requires specialist review
The issue depends on statistical assumptions or modeling choices that need methodologist/statistician assessment.

Interactive NMA assumptions checklist

Record the appraisal and rationale for each checkpoint. Your entries are stored only in this browser when local storage is available; nothing is transmitted to MetaSyn Academy.

Important: “No important concern identified” is not a pass mark. It means that no material concern was identified from the information available for that checkpoint. The final interpretation remains network- and contrast-specific.

1. Decision scope and network geometry

Establish what the network is trying to estimate and what its geometry permits before appraising assumptions.

1.1 Is the decision question, population, intervention set, comparator structure, outcome and target contrast clearly defined?

An NMA can be statistically estimable yet irrelevant to the actual decision if nodes, population or outcome do not match the intended question.

1.2 Are treatment nodes defined consistently enough that clinically different interventions are not merged without justification?

Dose, formulation, treatment class, combination components and co-interventions can change what a node represents.

1.3 Is the network connected for the contrasts being estimated, and is the presence or absence of closed loops explicitly described?

Connectivity determines which conventional NMA contrasts are estimable. Closed loops determine whether direct-indirect disagreement can be examined empirically.

1.4 If multi-arm trials are present, is their correlated contribution acknowledged and handled by the NMA model?

Treating comparisons from the same multi-arm trial as independent can distort uncertainty.


2. Evidence comparability and transitivity

This is the clinical-methodological core of the checklist. Transitivity is judged, not statistically certified.

2.1 Is it clinically plausible that the interventions could have been jointly randomized within the target population and decision context?

Joint randomizability is a useful conceptual expression of transitivity. Major differences in indication, treatment line or eligibility may make an indirect comparison implausible.

2.2 Were plausible treatment-effect modifiers identified before interpreting the network results?

Clinical knowledge, biological rationale, prior evidence and decision context should drive modifier selection. Statistical interaction significance is not a prerequisite for plausibility.

2.3 Were the distributions of plausible treatment-effect modifiers compared across the different treatment comparisons?

The relevant comparison is across sets of trials contributing to different edges, not merely between treatment arms within a trial.

2.4 Is incomplete reporting of important effect modifiers acknowledged rather than interpreted as evidence of balance?

Aggregate trial reports often omit variables needed for transitivity appraisal. Unreported information creates uncertainty; it does not justify a reassuring judgment.

2.5 Are outcome definitions, measurement windows and follow-up periods sufficiently comparable for the intended indirect contrasts?

Differences can act as clinical or methodological effect modifiers and may explain apparent direct-indirect disagreement.


3. Heterogeneity and incoherence

Keep within-comparison variation separate from direct-indirect disagreement. Both require contextual interpretation.

3.1 Were clinically or methodologically important sources of heterogeneity examined within direct comparisons?

Statistical heterogeneity is only one part of variation. Study design, population, intervention and outcome differences can matter even when an I² value appears small.

3.2 Was between-study heterogeneity estimated and interpreted with appropriate uncertainty rather than by a fixed numerical threshold?

There is no universal I² or τ² cut-off that validates or invalidates an NMA. Sparse comparisons can make heterogeneity estimates very imprecise.

3.3 Is it clear whether the network structure permits empirical assessment of direct-indirect disagreement?

Loop-free star or tree networks can support indirect estimates, but they do not provide the closed-loop structure required for ordinary direct-versus-indirect incoherence checks.

3.4 Where assessable, were suitable local and/or global approaches used to examine incoherence, with their limitations acknowledged?

Node/side splitting, loop-specific methods and design-by-treatment approaches answer different diagnostic questions. No single test proves the network is coherent.

3.5 If important heterogeneity or incoherence was identified, were plausible explanations investigated without selectively deleting inconvenient evidence?

Subgroup analysis, network meta-regression, sensitivity analysis or revised node definitions may be informative, but statistical adjustment cannot automatically repair an implausible clinical network.


4. Evidence credibility and interpretation

These checkpoints route to specialist tools where needed; they do not duplicate them.

4.1 Was risk of bias in the conduct, synthesis and interpretation of the NMA itself considered?

The 2025 RoB NMA tool assesses limitations in how an NMA was assembled, analyzed and interpreted. It is distinct from study-level risk-of-bias tools.

4.2 Was risk of bias from missing evidence considered at the network-estimate level?

ROB-MEN is designed for missing evidence in NMA and feeds the reporting-bias domain of CINeMA. It does not replace RoB NMA.

4.3 Are consequential model choices and sensitivity analyses reported clearly enough for specialist review?

Effect scale, fixed/random-effects choice, heterogeneity assumptions, multi-arm handling and Bayesian priors/convergence can matter. This checklist records whether they are transparent; it does not recalculate them.

4.4 If treatments were ranked, were rankings interpreted alongside effect magnitude, uncertainty, clinical relevance and evidence credibility?

SUCRA, P-scores and rank probabilities answer ranking questions; they do not by themselves establish clinically important superiority.

4.5 Were network estimates taken forward to an appropriate certainty/credibility assessment and reported against current NMA reporting guidance?

CINeMA and GRADE-NMA formalize certainty judgments. PRISMA-NMA 2015 remains the current published reporting extension while its update is still in development.11,12


Overall unresolved concerns

Summarize the network-specific concerns that should travel with the interpretation. Do not convert them into a score.


Appraisal summary

Complete the checklist to generate a narrative list of concerns, insufficient information, non-assessable items and specialist-review points. No aggregate validity score will be calculated.

Local privacy note: when supported, this tool saves form entries only in your current browser using local storage. It does not upload or transmit research data.

How to interpret the output

The summary is designed to preserve uncertainty rather than collapse it. A network can be connected and still have a serious transitivity concern. A star network can be clinically credible even though incoherence is structurally unavailable for assessment. A network can show little statistical heterogeneity while important treatment-effect modifiers remain imbalanced. Conversely, statistical heterogeneity or a local inconsistency signal may be explainable after examining outcome timing, node definitions, study design or other modifiers. None of these situations is well represented by a single total score.3,4,6


Structural barrier

A key desired contrast may be unavailable to conventional NMA when treatments sit in disconnected components. That does not make all synthesis impossible, but it changes the estimand and may require separate networks or specialist population-adjustment methods with additional assumptions.

Serious credibility concern

Major effect-modifier imbalance, incompatible treatment definitions, influential high-risk evidence or unexplained direct-indirect disagreement may materially weaken specific network estimates. The appropriate response depends on the problem; “run meta-regression” is not a universal repair.

Potentially investigable issue

Sparse evidence, heterogeneity, outcome timing differences or influential trials may be explored through sensitivity analyses, alternative node definitions, subgrouping or meta-regression when the data support those analyses. Residual uncertainty should remain visible.

Structurally untestable

In a loop-free network, ordinary incoherence testing is unavailable. Record it as not assessable rather than “no inconsistency.” The transitivity appraisal becomes more important because there is no direct-indirect diagnostic cross-check.


Worked appraisal example

Imagine an NMA comparing three psychological interventions, A, B and C. All three connect through usual care, so the network is structurally connected. The A-versus-usual-care trials mainly enrolled participants with mild symptoms and no previous therapy; the B-versus-usual-care trials enrolled participants with more severe, treatment-resistant symptoms; the C-versus-usual-care trials span both groups. Baseline severity and prior treatment are plausible effect modifiers. The network diagram looks clean, but the distributions are not comparable enough to justify a reassuring transitivity judgment without further investigation.

If the network has no closed loops, incoherence is not assessable. If it has closed loops and the node-splitting test is non-significant, that still does not erase the clinical imbalance. The result should be recorded as a transitivity concern, with any sensitivity or network meta-regression analysis described as an investigation rather than proof that the problem has disappeared. If SUCRA then ranks B first, the ranking needs to be read alongside the effect estimate, its interval, the transitivity concern and the certainty assessment.4,13,14,15


Methodological boundary

This checklist is intentionally narrower than a full NMA appraisal framework. It does not reproduce RoB NMA’s 17 items, ROB-MEN’s missing-evidence workflow, CINeMA’s contribution matrix, GRADE-NMA’s certainty process, or the statistical diagnostics needed to fit and compare NMA models. Those tools answer related but different questions. T31 is the point where a reviewer makes the network’s assumptions and unresolved concerns explicit before interpretation.8,9,10

For the conceptual reasoning behind each checkpoint, use the paired guide, Network Meta-Analysis Assumptions: Transitivity, Consistency and Heterogeneity. For the broader HTA/HEOR decision context, return to Evidence Synthesis for HTA and HEOR: Comparative Effectiveness and Decision Support.


Get every MetaSyn template free, including this one.

Leave your email and I’ll send this resource as an editable Word file and a printable PDF, plus access to the smart online version. You’ll also get every new template as it’s finished. No noise, just the resources.

Methodological scope note. This checklist appraises assumptions and interpretation readiness. It is not a replacement for fitting or auditing an NMA model, study-level risk-of-bias assessment, RoB NMA, ROB-MEN, CINeMA, GRADE-NMA, or statistical consultation. “No important concern identified” means that the available information did not reveal a material concern for that checkpoint; it does not certify validity.

References

  1. Chaimani A, Caldwell DM, Li T, Higgins JPT, Salanti G. Chapter 11: Undertaking network meta-analyses. In: Higgins JPT, Thomas J, Chandler J, et al, editors. Cochrane Handbook for Systematic Reviews of Interventions. Current online version. Cochrane. https://www.cochrane.org/authors/handbooks-and-manuals/handbook/current/chapter-11
  2. Jansen JP, Naci H. Is network meta-analysis as valid as standard pairwise meta-analysis? It all depends on the distribution of effect modifiers. BMC Med. 2013;11:159. https://doi.org/10.1186/1741-7015-11-159
  3. Spineli LM, Kalyvas C, Yepes-Nuñez JJ, et al. Low awareness of the transitivity assumption in complex networks of interventions: a systematic survey from 721 network meta-analyses. BMC Med. 2024;22:112. https://doi.org/10.1186/s12916-024-03322-1
  4. Spineli LM. An empirical study on 209 networks of treatments revealed intransitivity to be common and multiple statistical tests suboptimal to assess transitivity. BMC Med Res Methodol. 2024;24:301. https://doi.org/10.1186/s12874-024-02436-7
  5. Brignardello-Petersen R, Tomlinson G, Florez I, et al. Grading of Recommendations Assessment, Development, and Evaluation concept article 5: addressing intransitivity in a network meta-analysis. J Clin Epidemiol. 2023;160:151-159. https://doi.org/10.1016/j.jclinepi.2023.06.010
  6. Higgins JPT, Jackson D, Barrett JK, Lu G, Ades AE, White IR. Consistency and inconsistency in network meta-analysis: concepts and models for multi-arm studies. Res Synth Methods. 2012;3(2):98-110. https://doi.org/10.1002/jrsm.1044
  7. Jackson D, Boddington P, White IR. The design-by-treatment interaction model: a unifying framework for modelling loop inconsistency in network meta-analysis. Res Synth Methods. 2016;7(3):329-332. https://doi.org/10.1002/jrsm.1188
  8. Nikolakopoulou A, Higgins JPT, Papakonstantinou T, et al. CINeMA: an approach for assessing confidence in the results of a network meta-analysis. PLoS Med. 2020;17(4):e1003082. https://doi.org/10.1371/journal.pmed.1003082
  9. Chiocchia V, Nikolakopoulou A, Higgins JPT, et al. ROB-MEN: a tool to assess risk of bias due to missing evidence in network meta-analysis. BMC Med. 2021;19:304. https://doi.org/10.1186/s12916-021-02166-3
  10. Lunny C, Higgins JPT, White IR, et al. Risk of Bias in Network Meta-Analysis (RoB NMA) tool. BMJ. 2025;388:e079839. https://doi.org/10.1136/bmj-2024-079839
  11. Hutton B, Salanti G, Caldwell DM, et al. The PRISMA Extension Statement for Reporting of Systematic Reviews Incorporating Network Meta-analyses of Health Care Interventions: checklist and explanations. Ann Intern Med. 2015;162(11):777-784. https://doi.org/10.7326/M14-2385
  12. Veroniki AA, Tricco AC, Rangira D, et al. Updating the PRISMA reporting guideline for network meta-analysis: a scoping review. J Clin Epidemiol. 2025;188:111985. https://doi.org/10.1016/j.jclinepi.2025.111985
  13. Chiocchia V, White IR, Salanti G. The complexity underlying treatment rankings: how to use them and what to look at. BMJ Evid Based Med. 2023;28(3):180-182. https://doi.org/10.1136/bmjebm-2021-111904
  14. Salanti G, Nikolakopoulou A, Efthimiou O, Mavridis D, Egger M, White IR. Introducing the Treatment Hierarchy Question in Network Meta-Analysis. Am J Epidemiol. 2022;191(5):930-938. https://doi.org/10.1093/aje/kwab278
  15. Curteis T, Wigle A, Michaels CJ, Nikolakopoulou A. Ranking of treatments in network meta-analysis: incorporating minimally important differences. BMC Med Res Methodol. 2025;25:67. https://doi.org/10.1186/s12874-025-02499-0
  16. Noori A, Sadeghirad B, Thabane L, et al. The GRADE Working Group and CINeMA approaches provided inconsistent certainty of evidence ratings for a network meta-analysis of opioids for chronic noncancer pain. J Clin Epidemiol. 2024;169:111276. https://doi.org/10.1016/j.jclinepi.2024.111276
  17. Wang Y, Xia R, Pericic TP, et al. How do network meta-analyses address intransitivity when assessing certainty of evidence: a systematic survey. BMJ Open. 2023;13:e075212. https://doi.org/10.1136/bmjopen-2023-075212
  18. Brignardello-Petersen R, Florez ID, Yepes-Nuñez JJ, et al. Introduction to network meta-analysis: understanding what it is, how it is done, and how it can be used for decision-making. Am J Epidemiol. 2025;194(3):837-843. https://doi.org/10.1093/aje/kwae260

Current-status note: the official PRISMA site continues to list PRISMA-NMA 2015 as the published NMA extension. An update is under development; the 2025 scoping review identified candidate reporting items, but no finalized replacement checklist was publicly available at the evidence cut-off used for this page.

Frequently asked questions

Does a connected network mean the NMA is valid?

No. Connectivity is a structural requirement for estimating contrasts across a conventional connected NMA, but it does not establish transitivity, low bias, or credible interpretation.

Can transitivity be tested statistically?

Not definitively. Statistical approaches can explore dissimilarity or direct-indirect disagreement, but transitivity remains primarily a clinical and methodological judgment about whether the treatment comparisons are sufficiently comparable, especially with respect to plausible treatment-effect modifiers.

Is a star-shaped network automatically unreliable?

No. A star network can support indirect comparisons if transitivity is plausible. Its limitation is that ordinary direct-versus-indirect incoherence cannot be assessed when there are no closed loops.

Does a non-significant inconsistency test prove consistency?

No. Inconsistency tests often have limited power, especially in sparse networks. Failure to detect disagreement does not establish network validity.

What is the difference between RoB NMA and ROB-MEN?

RoB NMA assesses potential bias in how an NMA was assembled, analyzed and interpreted. ROB-MEN specifically assesses risk of bias due to missing evidence in network meta-analysis. They address different constructs and neither replaces the other.

Should I use an I² threshold to reject an NMA?

No universal I² threshold validates or invalidates an NMA. Heterogeneity should be interpreted in context, with attention to the effect scale, clinical variation, number of studies and uncertainty in the heterogeneity estimate.

Can SUCRA or P-scores tell me which treatment is best?

They summarize a treatment hierarchy under a particular ranking question. They should be interpreted with effect magnitude, confidence or credible intervals, clinical importance and evidence credibility rather than used as stand-alone evidence of superiority.

Is PRISMA-NMA 2015 still current?

Yes, it remains the current published PRISMA extension for NMA at this page’s evidence cut-off. An update is in development, but an unfinished draft should not be treated as the reporting standard.