Network Meta-Analysis Assumptions: Transitivity, Consistency and Heterogeneity
The hardest question in network meta-analysis is not whether every treatment is connected on a graph. It is whether the trials are comparable enough for the indirect links to mean what the model says they mean. A network can be dense, statistically elegant and still answer the wrong clinical question if the treatment comparisons differ systematically in variables that modify relative effects. That is why transitivity, heterogeneity and consistency have to be examined as distinct parts of one credibility argument rather than as three boxes to tick.1,2
Network meta-analysis extends pairwise meta-analysis by combining direct evidence, indirect evidence and, where available, both together. This creates a richer comparative evidence structure, but it also creates assumptions that ordinary pairwise synthesis does not need in the same way. The most consequential is transitivity: the idea that different treatment comparisons are sufficiently comparable, especially with respect to treatment-effect modifiers, to support indirect inference. Consistency or coherence is related, but it is not the same thing. It is the degree of agreement between direct and indirect evidence where the network structure lets us compare them.1,5,18
Connectivity is necessary, but it is not credibility
A treatment network is usually represented as nodes connected by edges. Nodes represent interventions or intervention classes; edges represent direct comparisons from eligible studies. If every treatment is reachable from every other treatment, the network is connected. That structural property is necessary for a conventional NMA to estimate all pairwise contrasts across the network. If the network breaks into separate components, some cross-component contrasts cannot be estimated without introducing a different method and additional assumptions.1
That is the mathematical question. The clinical question comes next. A placebo-controlled trial of treatment A in newly diagnosed patients and a placebo-controlled trial of treatment B in heavily pretreated patients create a connected A–placebo–B path. They do not automatically create a credible indirect A-versus-B comparison. If prior treatment modifies relative treatment effects, the path is structurally available but clinically suspect. Network geometry therefore belongs at the beginning of appraisal, not at the end of it.
Open, star and disconnected networks are different problems
A star network has a central comparator, often placebo or usual care, with active treatments radiating outward and no closed loops. It can still be connected and therefore estimable. What it cannot provide is an ordinary empirical comparison of direct and indirect evidence between the active treatments, because there is no direct A-versus-B edge to compare with the indirect A–C–B route. In such a network, incoherence is better described as structurally unavailable or not assessable, not “absent.” Transitivity appraisal becomes more important because the network offers no direct-indirect diagnostic cross-check.1,8
A disconnected network is different. If two groups of treatments have no path between them, a conventional NMA cannot estimate all cross-component contrasts under its ordinary evidence structure. Analysts can restrict inference to connected components or consider specialist population-adjustment/external-control approaches, but those methods bring new assumptions. Calling the original network “invalid” is too broad; the precise issue is that the desired cross-component contrast is not identified by the conventional connected network alone.
Transitivity is the central clinical-methodological assumption
Transitivity is often described through the hypothetical multi-arm trial idea: could the competing interventions reasonably have been randomized within the same target population and decision context? This is a useful starting point because it forces the analyst to think beyond the graph. Trials can share a comparator and still differ in treatment line, severity, eligibility, co-interventions, dose, follow-up, outcome definition or study era in ways that change relative treatment effects.1,2
The most operational way to evaluate transitivity is to identify plausible treatment-effect modifiers before interpreting the network and compare their distributions across the sets of studies contributing to different treatment comparisons. The word plausible matters. A variable does not have to demonstrate a statistically significant interaction in the available trials before it deserves attention. Interaction tests are often underpowered, especially with aggregated trial-level data. Clinical knowledge, biological rationale, subgroup evidence, prior research, guideline knowledge and the HTA decision context can all justify prespecifying a potential modifier.4,5
Prognostic factor is not another name for effect modifier
A prognostic variable changes baseline outcome risk. A treatment-effect modifier changes the relative effect of treatment. The two can overlap, but they are not interchangeable. This distinction becomes especially important when reviewers compile long baseline-characteristic tables and assume that every imbalance is equally threatening to transitivity. A difference in a strongly prognostic variable may matter for absolute risk while leaving a relative treatment effect stable; another variable may alter the relative effect directly. The appraisal should focus on clinically plausible modifiers while still documenting important uncertainty where the role of a variable is unclear.
Missing information is an uncertainty, not a reassuring result
Aggregate trial reports often omit variables needed for transitivity assessment. A reviewer may know that previous treatment is likely to modify effect but find it reported in only half of the studies. There is no defensible way to convert that absence into “balanced.” The safer conclusion is that transitivity is insufficiently assessable for that modifier. Empirical research supports this cautious stance: published NMAs continue to underreport transitivity assessments, and study-level statistical approaches for detecting intransitivity have limited performance.3,4,17
Heterogeneity asks a different question
Heterogeneity is often discussed as though it were a single statistic. In practice there are at least three relevant layers. Clinical heterogeneity concerns differences in participants, interventions, outcomes or settings. Methodological heterogeneity concerns study design and conduct. Statistical heterogeneity is the residual variation in estimated effects beyond sampling error and is commonly summarized through a between-study variance such as τ², with I² sometimes reported for direct comparisons. These layers should be interpreted together.1
In NMA, models often use a common heterogeneity variance across treatment comparisons because many comparisons are informed by too few studies to estimate separate variances reliably. That is a modeling assumption, not a clinical truth. The appraisal question is therefore not “Is τ² below the acceptable threshold?” No universal threshold exists. The question is whether the heterogeneity model is transparent, whether clinical and methodological sources were explored, how uncertain the estimate is, and whether the resulting spread changes the interpretation of the treatment effects.
Consistency and incoherence: what the data can test, and what they cannot
If transitivity is plausible, direct and indirect evidence are expected to agree apart from random variation and heterogeneity. Statistical disagreement is usually called inconsistency; CINeMA often uses the term incoherence. The terminology differs, but the methodological point is the same: where the network contains both direct and indirect evidence for a contrast, the analyst can examine whether those sources tell materially different stories.6,8
Local methods examine particular contrasts or loops. Node splitting or side splitting separates direct evidence from the indirect evidence available elsewhere in the network for a selected comparison. Loop-specific approaches examine inconsistency within a closed loop. Global methods, such as the design-by-treatment interaction framework, ask whether inconsistency is evident across the network as a whole. These methods are complementary diagnostic perspectives. There is no universal rule that every NMA should start with one and then mechanically move to the other.6,7
When incoherence is found
Detected disagreement should trigger investigation rather than automatic deletion of whichever study or comparison looks inconvenient. Plausible explanations include differences in effect modifiers, outcome timing, treatment definitions, study design or bias. Sensitivity analysis, revised node definitions, subgrouping or network meta-regression can sometimes clarify the source. But statistical adjustment cannot make fundamentally incomparable populations jointly randomizable after the fact. If a major clinical incompatibility remains, the appropriate conclusion may be to narrow the network or stop interpreting the affected indirect comparison as credible.
Bias in the NMA is broader than bias in its trials
Study-level risk-of-bias tools remain essential, but network meta-analysis creates additional ways in which the review itself can become biased. The 2025 RoB NMA tool was developed specifically to assess limitations in how an NMA was assembled, analyzed and interpreted. Its unit of assessment is the NMA review, not the individual randomized trial. ROB-MEN addresses a different construct: risk of bias due to missing evidence, including unavailable results and unpublished studies, and it is integrated with the reporting-bias domain of CINeMA.9,10
These tools are complementary, not synonyms and not replacements for one another. An NMA can be competently modeled yet vulnerable to missing evidence; another may contain a fairly complete evidence base but be assembled or interpreted in a biased way. A practical assumptions appraisal should therefore ask whether these risks were examined, then hand off to the appropriate specialist framework instead of reproducing it.
Treatment rankings come after credibility, not before it
SUCRA values, P-scores, probabilities of being best and mean ranks can be useful summaries, but they answer ranking questions rather than clinical superiority questions. Two treatments can occupy different ranks while their estimated effects are clinically indistinguishable. Rankings can also differ depending on the metric used and how uncertainty is incorporated. The treatment-hierarchy question should therefore be explicit: what does “better” mean for this outcome and decision?13,14
Recent work has proposed ranking methods that incorporate minimally important differences, which directly addresses the problem of treating trivial effect differences as meaningful separation. These methods are promising but should still be described as developing rather than as a universal replacement for established rank metrics.15
When should the NMA be reconsidered?
There is no single threshold that turns an NMA from acceptable to unacceptable. The better approach is to distinguish structural barriers from credibility concerns and potentially investigable issues.
| Problem | What it means | Reasonable response |
|---|---|---|
| Disconnected evidence for the desired contrast | The conventional network does not identify that cross-component comparison. | Restrict inference to connected components or use a specialist alternative with its additional assumptions; do not manufacture a conventional NMA link. |
| Clinically incompatible treatment nodes or populations | Joint randomizability may be implausible. | Reconsider node definitions, eligibility or the target network. Statistical sophistication cannot substitute for clinical comparability. |
| Major plausible effect-modifier imbalance | Transitivity may be compromised for affected indirect paths. | Investigate clinically and, where data allow, through sensitivity analysis or network meta-regression. Keep residual concern visible. |
| Substantial heterogeneity | Effects vary within comparisons and may be poorly summarized by one network model. | Explore sources and model assumptions; consider alternative synthesis or narrower network if variation is not interpretable. |
| Direct-indirect disagreement | Observed incoherence may reflect effect modification, bias, data problems or model misspecification. | Investigate local and network-wide sources. Do not treat selective exclusion as a default repair. |
| Sparse evidence | Heterogeneity and incoherence are estimated imprecisely; ranks may be unstable. | Use cautious modeling, emphasize intervals and certainty, and avoid overinterpreting non-significant diagnostics. |
| Influential high-risk evidence | Bias in one direct comparison can propagate to multiple network estimates. | Use contribution-aware credibility methods and sensitivity analysis; report which estimates depend on the evidence. |
Where CINeMA, GRADE and PRISMA-NMA fit
CINeMA evaluates confidence in network estimates across six domains: within-study bias, reporting bias, indirectness, imprecision, heterogeneity and incoherence. GRADE-NMA addresses closely related certainty questions using GRADE’s conceptual framework. The two approaches are not identical; empirical comparison has shown that they can produce different certainty judgments for the same network estimates. That is another reason not to treat transitivity or heterogeneity as mechanically scoreable.8,16
For reporting, PRISMA-NMA 2015 remains the current published NMA extension at the evidence cut-off used for this guide. An update is underway. The 2025 scoping review supporting that update identified new candidate reporting needs, including clearer reporting of methods used to assess homogeneity and transitivity and developments in complex NMA methods. Until a replacement guideline is finalized, the published 2015 extension remains the operational baseline, used alongside PRISMA 2020 where appropriate.11,12
Apply the reasoning before interpreting the result
The paired Network meta-analysis assumptions checklist turns the reasoning in this guide into a structured appraisal. Use it to record network geometry, effect-modifier comparability, heterogeneity, incoherence, bias and ranking concerns without calculating an overall validity score. The purpose is to make the reasoning auditable, especially where the evidence is incomplete or a specialist judgment is still needed.
For the larger comparative-effectiveness context, including how NMA relates to direct evidence, ITCs, population adjustment and HTA decision questions, use Evidence Synthesis for HTA and HEOR: Comparative Effectiveness and Decision Support. Pair 31 deliberately stops before the detailed appraisal of anchored and unanchored indirect comparisons, which belongs to the separate ITC guide.
Frequently asked questions
Is transitivity the same as consistency?
No. Transitivity is a clinical and methodological assumption about whether the treatment comparisons are sufficiently comparable for indirect inference. Consistency or coherence concerns statistical agreement between direct and indirect evidence where both can be observed.
Can I prove transitivity by showing balanced baseline tables?
No. Baseline comparisons can support the judgment, especially for prespecified treatment-effect modifiers, but aggregate reporting is incomplete and no universal balance threshold proves transitivity.
Does a star network fail the consistency assumption?
Not automatically. In a loop-free star network, ordinary direct-indirect incoherence is structurally unavailable for assessment. The network may still be credible if transitivity is plausible, but there is less empirical opportunity to challenge that assumption.
Should every NMA use both node splitting and a global inconsistency test?
No universal sequence applies to every network. Local and global approaches answer different diagnostic questions, and the available methods depend on network geometry and data. Their power limitations should be acknowledged.
Is high I² enough to reject an NMA?
No. There is no universal I² cut-off for NMA credibility. Statistical heterogeneity should be interpreted with clinical and methodological variation, the effect scale, the number of studies and uncertainty in the heterogeneity estimate.
Does RoB NMA replace ROB-MEN?
No. RoB NMA evaluates potential bias in how an NMA was conducted and interpreted. ROB-MEN specifically evaluates bias due to missing evidence in network meta-analysis. They address different parts of credibility.
Can meta-regression fix a transitivity problem?
Sometimes it can explore or adjust for measured effect modification when the network contains enough information. It cannot guarantee repair of an implausible network, remove unmeasured effect modification, or create information that was never reported.
Should I report SUCRA if evidence certainty is low?
You may report ranking metrics if they answer a prespecified hierarchy question, but they should not be interpreted as stand-alone evidence of superiority. Effect magnitude, intervals, clinical importance and certainty need to accompany them.
References
- Chaimani A, Caldwell DM, Li T, Higgins JPT, Salanti G. Chapter 11: Undertaking network meta-analyses. In: Higgins JPT, Thomas J, Chandler J, et al, editors. Cochrane Handbook for Systematic Reviews of Interventions. Current online version. Cochrane. https://www.cochrane.org/authors/handbooks-and-manuals/handbook/current/chapter-11
- Jansen JP, Naci H. Is network meta-analysis as valid as standard pairwise meta-analysis? It all depends on the distribution of effect modifiers. BMC Med. 2013;11:159. https://doi.org/10.1186/1741-7015-11-159
- Spineli LM, Kalyvas C, Yepes-Nuñez JJ, et al. Low awareness of the transitivity assumption in complex networks of interventions: a systematic survey from 721 network meta-analyses. BMC Med. 2024;22:112. https://doi.org/10.1186/s12916-024-03322-1
- Spineli LM. An empirical study on 209 networks of treatments revealed intransitivity to be common and multiple statistical tests suboptimal to assess transitivity. BMC Med Res Methodol. 2024;24:301. https://doi.org/10.1186/s12874-024-02436-7
- Brignardello-Petersen R, Tomlinson G, Florez I, et al. Grading of Recommendations Assessment, Development, and Evaluation concept article 5: addressing intransitivity in a network meta-analysis. J Clin Epidemiol. 2023;160:151-159. https://doi.org/10.1016/j.jclinepi.2023.06.010
- Higgins JPT, Jackson D, Barrett JK, Lu G, Ades AE, White IR. Consistency and inconsistency in network meta-analysis: concepts and models for multi-arm studies. Res Synth Methods. 2012;3(2):98-110. https://doi.org/10.1002/jrsm.1044
- Jackson D, Boddington P, White IR. The design-by-treatment interaction model: a unifying framework for modelling loop inconsistency in network meta-analysis. Res Synth Methods. 2016;7(3):329-332. https://doi.org/10.1002/jrsm.1188
- Nikolakopoulou A, Higgins JPT, Papakonstantinou T, et al. CINeMA: an approach for assessing confidence in the results of a network meta-analysis. PLoS Med. 2020;17(4):e1003082. https://doi.org/10.1371/journal.pmed.1003082
- Chiocchia V, Nikolakopoulou A, Higgins JPT, et al. ROB-MEN: a tool to assess risk of bias due to missing evidence in network meta-analysis. BMC Med. 2021;19:304. https://doi.org/10.1186/s12916-021-02166-3
- Lunny C, Higgins JPT, White IR, Dias S, Hutton B, Wright JM, et al. Risk of Bias in Network Meta-Analysis (RoB NMA) tool. BMJ. 2025;388:e079839. https://doi.org/10.1136/bmj-2024-079839
- Hutton B, Salanti G, Caldwell DM, et al. The PRISMA Extension Statement for Reporting of Systematic Reviews Incorporating Network Meta-analyses of Health Care Interventions: checklist and explanations. Ann Intern Med. 2015;162(11):777-784. https://doi.org/10.7326/M14-2385
- Veroniki AA, Tricco AC, Rangira D, et al. Updating the PRISMA reporting guideline for network meta-analysis: a scoping review. J Clin Epidemiol. 2025;188:111985. https://doi.org/10.1016/j.jclinepi.2025.111985
- Chiocchia V, White IR, Salanti G. The complexity underlying treatment rankings: how to use them and what to look at. BMJ Evid Based Med. 2023;28(3):180-182. https://doi.org/10.1136/bmjebm-2021-111904
- Salanti G, Nikolakopoulou A, Efthimiou O, Mavridis D, Egger M, White IR. Introducing the Treatment Hierarchy Question in Network Meta-Analysis. Am J Epidemiol. 2022;191(5):930-938. https://doi.org/10.1093/aje/kwab278
- Curteis T, Wigle A, Michaels CJ, Nikolakopoulou A. Ranking of treatments in network meta-analysis: incorporating minimally important differences. BMC Med Res Methodol. 2025;25:67. https://doi.org/10.1186/s12874-025-02499-0
- Noori A, Sadeghirad B, Thabane L, et al. The GRADE Working Group and CINeMA approaches provided inconsistent certainty of evidence ratings for a network meta-analysis of opioids for chronic noncancer pain. J Clin Epidemiol. 2024;169:111276. https://doi.org/10.1016/j.jclinepi.2024.111276
- Wang Y, Xia R, Pericic TP, et al. How do network meta-analyses address intransitivity when assessing certainty of evidence: a systematic survey. BMJ Open. 2023;13:e075212. https://doi.org/10.1136/bmjopen-2023-075212
- Brignardello-Petersen R, Florez ID, Yepes-Nuñez JJ, et al. Introduction to network meta-analysis: understanding what it is, how it is done, and how it can be used for decision-making. Am J Epidemiol. 2025;194(3):837-843. https://doi.org/10.1093/aje/kwae260