NETWORK META-ANALYSIS ASSUMPTIONS RESOURCE
Network Meta-analysis Assumptions Checklist
Appraise network geometry, similarity, transitivity, consistency, effect modifiers, heterogeneity, and model reporting before interpreting indirect or mixed comparisons.
A connected treatment network is not evidence that the network is credible. Network meta-analysis can combine direct and indirect treatment effects only when its clinical and methodological assumptions are sufficiently plausible. The difficult part is that the most important assumption, transitivity, cannot be established by a single statistical test. It has to be argued from the decision question, treatment definitions, study eligibility, and the distribution of variables that could modify relative treatment effects.1,2,18
This checklist is designed for that appraisal step. It asks whether the evidence network is structurally usable, whether the treatment comparisons are sufficiently comparable to support indirect inference, whether heterogeneity and direct-indirect disagreement were examined appropriately, and whether bias and treatment rankings were interpreted with adequate caution. It deliberately does not calculate an NMA-validity score. Current NMA frameworks rely on structured judgments because several consequential concerns cannot be reduced to defensible numerical thresholds.3,4,8,16
What the checklist is testing
Three ideas are often blurred together. Heterogeneity concerns variation in effects among studies informing a comparison. Transitivity concerns whether different sets of studies can support valid indirect comparisons, especially whether plausible treatment-effect modifiers are distributed sufficiently similarly across those sets. Consistency, also called coherence in some frameworks, concerns agreement between direct and indirect evidence where both are available. These ideas interact, but one cannot stand in for another.1,5,6
Why treatment-effect modifiers deserve their own attention
A prognostic factor predicts outcome regardless of treatment. A treatment-effect modifier changes the relative effect of one treatment compared with another. The distinction matters because transitivity is threatened by systematic differences in modifiers across the sets of trials that form an indirect comparison. Baseline severity, treatment line, prior therapy, baseline risk, dose, follow-up, co-interventions and trial era can be plausible modifiers in some clinical settings, but they are not automatically modifiers in every network. They should be identified from clinical and biological reasoning, previous evidence, subgroup knowledge and decision context rather than selected only because a post-hoc interaction happens to reach statistical significance.2,5
How to use the checklist
Complete the domains in order. Start with the decision question and network structure, because later judgments are meaningless if the treatment nodes or target contrast are unclear. Then examine comparability and effect modifiers before looking at statistical incoherence. A statistical test is a diagnostic aid, not a substitute for the conceptual assessment. Finally, examine bias, missing evidence and ranking interpretation before deciding whether the network warrants specialist review or cautious use.1,6,7,8
Available information does not reveal a material concern for this checkpoint.
A feature could materially weaken one or more network estimates and needs explanation or sensitivity analysis.
Reporting is inadequate to make the judgment. Do not convert missing information into reassurance.
The checkpoint cannot apply to the network structure, such as direct-indirect incoherence testing in a loop-free network.
The issue depends on statistical assumptions or modeling choices that need methodologist/statistician assessment.
Interactive NMA assumptions checklist
Record the appraisal and rationale for each checkpoint. Your entries are stored only in this browser when local storage is available; nothing is transmitted to MetaSyn Academy.
Appraisal summary
Complete the checklist to generate a narrative list of concerns, insufficient information, non-assessable items and specialist-review points. No aggregate validity score will be calculated.
How to interpret the output
The summary is designed to preserve uncertainty rather than collapse it. A network can be connected and still have a serious transitivity concern. A star network can be clinically credible even though incoherence is structurally unavailable for assessment. A network can show little statistical heterogeneity while important treatment-effect modifiers remain imbalanced. Conversely, statistical heterogeneity or a local inconsistency signal may be explainable after examining outcome timing, node definitions, study design or other modifiers. None of these situations is well represented by a single total score.3,4,6
Structural barrier
A key desired contrast may be unavailable to conventional NMA when treatments sit in disconnected components. That does not make all synthesis impossible, but it changes the estimand and may require separate networks or specialist population-adjustment methods with additional assumptions.
Serious credibility concern
Major effect-modifier imbalance, incompatible treatment definitions, influential high-risk evidence or unexplained direct-indirect disagreement may materially weaken specific network estimates. The appropriate response depends on the problem; “run meta-regression” is not a universal repair.
Potentially investigable issue
Sparse evidence, heterogeneity, outcome timing differences or influential trials may be explored through sensitivity analyses, alternative node definitions, subgrouping or meta-regression when the data support those analyses. Residual uncertainty should remain visible.
Structurally untestable
In a loop-free network, ordinary incoherence testing is unavailable. Record it as not assessable rather than “no inconsistency.” The transitivity appraisal becomes more important because there is no direct-indirect diagnostic cross-check.
Worked appraisal example
Imagine an NMA comparing three psychological interventions, A, B and C. All three connect through usual care, so the network is structurally connected. The A-versus-usual-care trials mainly enrolled participants with mild symptoms and no previous therapy; the B-versus-usual-care trials enrolled participants with more severe, treatment-resistant symptoms; the C-versus-usual-care trials span both groups. Baseline severity and prior treatment are plausible effect modifiers. The network diagram looks clean, but the distributions are not comparable enough to justify a reassuring transitivity judgment without further investigation.
If the network has no closed loops, incoherence is not assessable. If it has closed loops and the node-splitting test is non-significant, that still does not erase the clinical imbalance. The result should be recorded as a transitivity concern, with any sensitivity or network meta-regression analysis described as an investigation rather than proof that the problem has disappeared. If SUCRA then ranks B first, the ranking needs to be read alongside the effect estimate, its interval, the transitivity concern and the certainty assessment.4,13,14,15
Methodological boundary
This checklist is intentionally narrower than a full NMA appraisal framework. It does not reproduce RoB NMA’s 17 items, ROB-MEN’s missing-evidence workflow, CINeMA’s contribution matrix, GRADE-NMA’s certainty process, or the statistical diagnostics needed to fit and compare NMA models. Those tools answer related but different questions. T31 is the point where a reviewer makes the network’s assumptions and unresolved concerns explicit before interpretation.8,9,10
For the conceptual reasoning behind each checkpoint, use the paired guide, Network Meta-Analysis Assumptions: Transitivity, Consistency and Heterogeneity. For the broader HTA/HEOR decision context, return to Evidence Synthesis for HTA and HEOR: Comparative Effectiveness and Decision Support.
Get every MetaSyn template free, including this one.
Leave your email and I’ll send this resource as an editable Word file and a printable PDF, plus access to the smart online version. You’ll also get every new template as it’s finished. No noise, just the resources.
References
- Chaimani A, Caldwell DM, Li T, Higgins JPT, Salanti G. Chapter 11: Undertaking network meta-analyses. In: Higgins JPT, Thomas J, Chandler J, et al, editors. Cochrane Handbook for Systematic Reviews of Interventions. Current online version. Cochrane. https://www.cochrane.org/authors/handbooks-and-manuals/handbook/current/chapter-11
- Jansen JP, Naci H. Is network meta-analysis as valid as standard pairwise meta-analysis? It all depends on the distribution of effect modifiers. BMC Med. 2013;11:159. https://doi.org/10.1186/1741-7015-11-159
- Spineli LM, Kalyvas C, Yepes-Nuñez JJ, et al. Low awareness of the transitivity assumption in complex networks of interventions: a systematic survey from 721 network meta-analyses. BMC Med. 2024;22:112. https://doi.org/10.1186/s12916-024-03322-1
- Spineli LM. An empirical study on 209 networks of treatments revealed intransitivity to be common and multiple statistical tests suboptimal to assess transitivity. BMC Med Res Methodol. 2024;24:301. https://doi.org/10.1186/s12874-024-02436-7
- Brignardello-Petersen R, Tomlinson G, Florez I, et al. Grading of Recommendations Assessment, Development, and Evaluation concept article 5: addressing intransitivity in a network meta-analysis. J Clin Epidemiol. 2023;160:151-159. https://doi.org/10.1016/j.jclinepi.2023.06.010
- Higgins JPT, Jackson D, Barrett JK, Lu G, Ades AE, White IR. Consistency and inconsistency in network meta-analysis: concepts and models for multi-arm studies. Res Synth Methods. 2012;3(2):98-110. https://doi.org/10.1002/jrsm.1044
- Jackson D, Boddington P, White IR. The design-by-treatment interaction model: a unifying framework for modelling loop inconsistency in network meta-analysis. Res Synth Methods. 2016;7(3):329-332. https://doi.org/10.1002/jrsm.1188
- Nikolakopoulou A, Higgins JPT, Papakonstantinou T, et al. CINeMA: an approach for assessing confidence in the results of a network meta-analysis. PLoS Med. 2020;17(4):e1003082. https://doi.org/10.1371/journal.pmed.1003082
- Chiocchia V, Nikolakopoulou A, Higgins JPT, et al. ROB-MEN: a tool to assess risk of bias due to missing evidence in network meta-analysis. BMC Med. 2021;19:304. https://doi.org/10.1186/s12916-021-02166-3
- Lunny C, Higgins JPT, White IR, et al. Risk of Bias in Network Meta-Analysis (RoB NMA) tool. BMJ. 2025;388:e079839. https://doi.org/10.1136/bmj-2024-079839
- Hutton B, Salanti G, Caldwell DM, et al. The PRISMA Extension Statement for Reporting of Systematic Reviews Incorporating Network Meta-analyses of Health Care Interventions: checklist and explanations. Ann Intern Med. 2015;162(11):777-784. https://doi.org/10.7326/M14-2385
- Veroniki AA, Tricco AC, Rangira D, et al. Updating the PRISMA reporting guideline for network meta-analysis: a scoping review. J Clin Epidemiol. 2025;188:111985. https://doi.org/10.1016/j.jclinepi.2025.111985
- Chiocchia V, White IR, Salanti G. The complexity underlying treatment rankings: how to use them and what to look at. BMJ Evid Based Med. 2023;28(3):180-182. https://doi.org/10.1136/bmjebm-2021-111904
- Salanti G, Nikolakopoulou A, Efthimiou O, Mavridis D, Egger M, White IR. Introducing the Treatment Hierarchy Question in Network Meta-Analysis. Am J Epidemiol. 2022;191(5):930-938. https://doi.org/10.1093/aje/kwab278
- Curteis T, Wigle A, Michaels CJ, Nikolakopoulou A. Ranking of treatments in network meta-analysis: incorporating minimally important differences. BMC Med Res Methodol. 2025;25:67. https://doi.org/10.1186/s12874-025-02499-0
- Noori A, Sadeghirad B, Thabane L, et al. The GRADE Working Group and CINeMA approaches provided inconsistent certainty of evidence ratings for a network meta-analysis of opioids for chronic noncancer pain. J Clin Epidemiol. 2024;169:111276. https://doi.org/10.1016/j.jclinepi.2024.111276
- Wang Y, Xia R, Pericic TP, et al. How do network meta-analyses address intransitivity when assessing certainty of evidence: a systematic survey. BMJ Open. 2023;13:e075212. https://doi.org/10.1136/bmjopen-2023-075212
- Brignardello-Petersen R, Florez ID, Yepes-Nuñez JJ, et al. Introduction to network meta-analysis: understanding what it is, how it is done, and how it can be used for decision-making. Am J Epidemiol. 2025;194(3):837-843. https://doi.org/10.1093/aje/kwae260
Current-status note: the official PRISMA site continues to list PRISMA-NMA 2015 as the published NMA extension. An update is under development; the 2025 scoping review identified candidate reporting items, but no finalized replacement checklist was publicly available at the evidence cut-off used for this page.
Frequently asked questions
Does a connected network mean the NMA is valid?
No. Connectivity is a structural requirement for estimating contrasts across a conventional connected NMA, but it does not establish transitivity, low bias, or credible interpretation.
Can transitivity be tested statistically?
Not definitively. Statistical approaches can explore dissimilarity or direct-indirect disagreement, but transitivity remains primarily a clinical and methodological judgment about whether the treatment comparisons are sufficiently comparable, especially with respect to plausible treatment-effect modifiers.
Is a star-shaped network automatically unreliable?
No. A star network can support indirect comparisons if transitivity is plausible. Its limitation is that ordinary direct-versus-indirect incoherence cannot be assessed when there are no closed loops.
Does a non-significant inconsistency test prove consistency?
No. Inconsistency tests often have limited power, especially in sparse networks. Failure to detect disagreement does not establish network validity.
What is the difference between RoB NMA and ROB-MEN?
RoB NMA assesses potential bias in how an NMA was assembled, analyzed and interpreted. ROB-MEN specifically assesses risk of bias due to missing evidence in network meta-analysis. They address different constructs and neither replaces the other.
Should I use an I² threshold to reject an NMA?
No universal I² threshold validates or invalidates an NMA. Heterogeneity should be interpreted in context, with attention to the effect scale, clinical variation, number of studies and uncertainty in the heterogeneity estimate.
Can SUCRA or P-scores tell me which treatment is best?
They summarize a treatment hierarchy under a particular ranking question. They should be interpreted with effect magnitude, confidence or credible intervals, clinical importance and evidence credibility rather than used as stand-alone evidence of superiority.
Is PRISMA-NMA 2015 still current?
Yes, it remains the current published PRISMA extension for NMA at this page’s evidence cut-off. An update is in development, but an unfinished draft should not be treated as the reporting standard.