Network Meta-Analysis Assumptions: Transitivity, Consistency and Heterogeneity

MetaSyn Academy guide to network meta-analysis assumptions, illustrating the decisions documented by Network meta-analysis assumptions checklist
Academic methodology guideEvidence checked: 9 August 2026

The hardest question in network meta-analysis is not whether every treatment is connected on a graph. It is whether the trials are comparable enough for the indirect links to mean what the model says they mean. A network can be dense, statistically elegant and still answer the wrong clinical question if the treatment comparisons differ systematically in variables that modify relative effects. That is why transitivity, heterogeneity and consistency have to be examined as distinct parts of one credibility argument rather than as three boxes to tick.1,2


Network meta-analysis extends pairwise meta-analysis by combining direct evidence, indirect evidence and, where available, both together. This creates a richer comparative evidence structure, but it also creates assumptions that ordinary pairwise synthesis does not need in the same way. The most consequential is transitivity: the idea that different treatment comparisons are sufficiently comparable, especially with respect to treatment-effect modifiers, to support indirect inference. Consistency or coherence is related, but it is not the same thing. It is the degree of agreement between direct and indirect evidence where the network structure lets us compare them.1,5,18


One sentence to keep: connectivity tells you whether an indirect route exists; transitivity tells you whether that route is clinically and methodologically plausible; incoherence analysis tells you whether direct and indirect evidence disagree where that disagreement is observable.

Connectivity is necessary, but it is not credibility

A treatment network is usually represented as nodes connected by edges. Nodes represent interventions or intervention classes; edges represent direct comparisons from eligible studies. If every treatment is reachable from every other treatment, the network is connected. That structural property is necessary for a conventional NMA to estimate all pairwise contrasts across the network. If the network breaks into separate components, some cross-component contrasts cannot be estimated without introducing a different method and additional assumptions.1


That is the mathematical question. The clinical question comes next. A placebo-controlled trial of treatment A in newly diagnosed patients and a placebo-controlled trial of treatment B in heavily pretreated patients create a connected A–placebo–B path. They do not automatically create a credible indirect A-versus-B comparison. If prior treatment modifies relative treatment effects, the path is structurally available but clinically suspect. Network geometry therefore belongs at the beginning of appraisal, not at the end of it.


A connected network can still have a transitivity problem A three-treatment network is shown on the left. On the right, two sets of studies contributing to different comparisons have different distributions of a treatment-effect modifier, illustrating that geometry alone does not establish credibility. One network, two different questions 1. Is the network connected? C A B Yes: all nodes share a path 2. Is transitivity plausible? A–C trials mostly treatment-naïve B–C trials mostly treatment-resistant If prior treatment modifies relative effects, the indirect A–B contrast may not be transportable. A path through a common comparator is a structural fact. Whether that path supports a credible indirect effect is a separate judgment.
Figure 1. Network connectivity creates a route for estimation. Transitivity asks whether the studies on that route are comparable enough for indirect inference.

Open, star and disconnected networks are different problems

A star network has a central comparator, often placebo or usual care, with active treatments radiating outward and no closed loops. It can still be connected and therefore estimable. What it cannot provide is an ordinary empirical comparison of direct and indirect evidence between the active treatments, because there is no direct A-versus-B edge to compare with the indirect A–C–B route. In such a network, incoherence is better described as structurally unavailable or not assessable, not “absent.” Transitivity appraisal becomes more important because the network offers no direct-indirect diagnostic cross-check.1,8


A disconnected network is different. If two groups of treatments have no path between them, a conventional NMA cannot estimate all cross-component contrasts under its ordinary evidence structure. Analysts can restrict inference to connected components or consider specialist population-adjustment/external-control approaches, but those methods bring new assumptions. Calling the original network “invalid” is too broad; the precise issue is that the desired cross-component contrast is not identified by the conventional connected network alone.


Transitivity is the central clinical-methodological assumption

Transitivity is often described through the hypothetical multi-arm trial idea: could the competing interventions reasonably have been randomized within the same target population and decision context? This is a useful starting point because it forces the analyst to think beyond the graph. Trials can share a comparator and still differ in treatment line, severity, eligibility, co-interventions, dose, follow-up, outcome definition or study era in ways that change relative treatment effects.1,2


The most operational way to evaluate transitivity is to identify plausible treatment-effect modifiers before interpreting the network and compare their distributions across the sets of studies contributing to different treatment comparisons. The word plausible matters. A variable does not have to demonstrate a statistically significant interaction in the available trials before it deserves attention. Interaction tests are often underpowered, especially with aggregated trial-level data. Clinical knowledge, biological rationale, subgroup evidence, prior research, guideline knowledge and the HTA decision context can all justify prespecifying a potential modifier.4,5


Treatment-effect modifier balance across comparisons Three treatment comparisons are represented as cards with severity distributions. Two cards have similar severity distributions and one differs, creating a concern about transitivity without imposing a numerical threshold. Transitivity asks about plausible treatment-effect modifiers and how their distributions compare across treatment comparisons Example: baseline disease severity. This illustrates reasoning, not a validated balance threshold. A vs common comparator B vs common comparator D vs common comparator Mild Moderate Severe Mild Moderate Severe Mild Moderate Severe Similar pattern Similar pattern Clinically different pattern The judgment is clinical: could the imbalance plausibly change the relative treatment effect? There is no universal numerical cut-off that turns this question into a pass/fail test.
Figure 2. Effect-modifier distributions should be examined across treatment comparisons. The relevant concern is whether differences are clinically capable of changing relative treatment effects, not whether an arbitrary balance threshold is crossed.

Prognostic factor is not another name for effect modifier

A prognostic variable changes baseline outcome risk. A treatment-effect modifier changes the relative effect of treatment. The two can overlap, but they are not interchangeable. This distinction becomes especially important when reviewers compile long baseline-characteristic tables and assume that every imbalance is equally threatening to transitivity. A difference in a strongly prognostic variable may matter for absolute risk while leaving a relative treatment effect stable; another variable may alter the relative effect directly. The appraisal should focus on clinically plausible modifiers while still documenting important uncertainty where the role of a variable is unclear.


Missing information is an uncertainty, not a reassuring result

Aggregate trial reports often omit variables needed for transitivity assessment. A reviewer may know that previous treatment is likely to modify effect but find it reported in only half of the studies. There is no defensible way to convert that absence into “balanced.” The safer conclusion is that transitivity is insufficiently assessable for that modifier. Empirical research supports this cautious stance: published NMAs continue to underreport transitivity assessments, and study-level statistical approaches for detecting intransitivity have limited performance.3,4,17


Heterogeneity asks a different question

Heterogeneity is often discussed as though it were a single statistic. In practice there are at least three relevant layers. Clinical heterogeneity concerns differences in participants, interventions, outcomes or settings. Methodological heterogeneity concerns study design and conduct. Statistical heterogeneity is the residual variation in estimated effects beyond sampling error and is commonly summarized through a between-study variance such as τ², with I² sometimes reported for direct comparisons. These layers should be interpreted together.1


In NMA, models often use a common heterogeneity variance across treatment comparisons because many comparisons are informed by too few studies to estimate separate variances reliably. That is a modeling assumption, not a clinical truth. The appraisal question is therefore not “Is τ² below the acceptable threshold?” No universal threshold exists. The question is whether the heterogeneity model is transparent, whether clinical and methodological sources were explored, how uncertain the estimate is, and whether the resulting spread changes the interpretation of the treatment effects.


Heterogeneity and incoherence are different types of disagreement The upper panel shows several studies estimating the same A versus B comparison with different effects, illustrating heterogeneity. The lower panel shows direct A versus B evidence disagreeing with indirect A versus C versus B evidence, illustrating incoherence. Two forms of disagreement, two different questions Heterogeneity: studies within the same direct comparison disagree Favours A Favours B Incoherence: direct and indirect evidence for the same contrast disagree Direct A vs B effect estimate points left Indirect A–C–B effect estimate points right disagreement Low heterogeneity does not guarantee coherence; observed incoherence can have several clinical, methodological or statistical causes.
Figure 3. Heterogeneity describes variation among studies estimating the same direct contrast. Incoherence describes disagreement between direct and indirect routes for a contrast. They can coexist, but they are not the same problem.

Consistency and incoherence: what the data can test, and what they cannot

If transitivity is plausible, direct and indirect evidence are expected to agree apart from random variation and heterogeneity. Statistical disagreement is usually called inconsistency; CINeMA often uses the term incoherence. The terminology differs, but the methodological point is the same: where the network contains both direct and indirect evidence for a contrast, the analyst can examine whether those sources tell materially different stories.6,8


Local methods examine particular contrasts or loops. Node splitting or side splitting separates direct evidence from the indirect evidence available elsewhere in the network for a selected comparison. Loop-specific approaches examine inconsistency within a closed loop. Global methods, such as the design-by-treatment interaction framework, ask whether inconsistency is evident across the network as a whole. These methods are complementary diagnostic perspectives. There is no universal rule that every NMA should start with one and then mechanically move to the other.6,7


The non-significance trap: a large P value from an inconsistency test does not establish consistency. Sparse networks, small studies and few closed loops can leave these tests with little power. The correct conclusion is that statistical evidence of disagreement was not detected, not that transitivity has been proven.4,6

Local and global incoherence diagnostics A network graph is shown with one triangle highlighted for local assessment and the entire network enclosed for global assessment. Labels explain that the approaches diagnose different scales and neither proves network validity. Local and global diagnostics look at different scales Local: inspect a contrast or loop C A B Node/side splitting asks whether direct A–B agrees with indirect A–C–B evidence. Global: inspect the whole network Design-by-treatment evaluates evidence of inconsistency across the network. Neither scale replaces the clinical transitivity appraisal, and neither a local nor a global non-significant test certifies validity.
Figure 4. Local diagnostics isolate particular direct-indirect disagreements; global models evaluate network-wide evidence of inconsistency. They answer different questions and share power limitations.

When incoherence is found

Detected disagreement should trigger investigation rather than automatic deletion of whichever study or comparison looks inconvenient. Plausible explanations include differences in effect modifiers, outcome timing, treatment definitions, study design or bias. Sensitivity analysis, revised node definitions, subgrouping or network meta-regression can sometimes clarify the source. But statistical adjustment cannot make fundamentally incomparable populations jointly randomizable after the fact. If a major clinical incompatibility remains, the appropriate conclusion may be to narrow the network or stop interpreting the affected indirect comparison as credible.


Bias in the NMA is broader than bias in its trials

Study-level risk-of-bias tools remain essential, but network meta-analysis creates additional ways in which the review itself can become biased. The 2025 RoB NMA tool was developed specifically to assess limitations in how an NMA was assembled, analyzed and interpreted. Its unit of assessment is the NMA review, not the individual randomized trial. ROB-MEN addresses a different construct: risk of bias due to missing evidence, including unavailable results and unpublished studies, and it is integrated with the reporting-bias domain of CINeMA.9,10


These tools are complementary, not synonyms and not replacements for one another. An NMA can be competently modeled yet vulnerable to missing evidence; another may contain a fairly complete evidence base but be assembled or interpreted in a biased way. A practical assumptions appraisal should therefore ask whether these risks were examined, then hand off to the appropriate specialist framework instead of reproducing it.


Treatment rankings come after credibility, not before it

SUCRA values, P-scores, probabilities of being best and mean ranks can be useful summaries, but they answer ranking questions rather than clinical superiority questions. Two treatments can occupy different ranks while their estimated effects are clinically indistinguishable. Rankings can also differ depending on the metric used and how uncertainty is incorporated. The treatment-hierarchy question should therefore be explicit: what does “better” mean for this outcome and decision?13,14


Recent work has proposed ranking methods that incorporate minimally important differences, which directly addresses the problem of treating trivial effect differences as meaningful separation. These methods are promising but should still be described as developing rather than as a universal replacement for established rank metrics.15


Safe interpretation rule: report the effect estimate first, its uncertainty next, the certainty or credibility judgment with it, and the ranking only as a conditional summary. A high SUCRA value from a network with serious transitivity concerns is not strong evidence that the treatment should be preferred.

When should the NMA be reconsidered?

There is no single threshold that turns an NMA from acceptable to unacceptable. The better approach is to distinguish structural barriers from credibility concerns and potentially investigable issues.

ProblemWhat it meansReasonable response
Disconnected evidence for the desired contrastThe conventional network does not identify that cross-component comparison.Restrict inference to connected components or use a specialist alternative with its additional assumptions; do not manufacture a conventional NMA link.
Clinically incompatible treatment nodes or populationsJoint randomizability may be implausible.Reconsider node definitions, eligibility or the target network. Statistical sophistication cannot substitute for clinical comparability.
Major plausible effect-modifier imbalanceTransitivity may be compromised for affected indirect paths.Investigate clinically and, where data allow, through sensitivity analysis or network meta-regression. Keep residual concern visible.
Substantial heterogeneityEffects vary within comparisons and may be poorly summarized by one network model.Explore sources and model assumptions; consider alternative synthesis or narrower network if variation is not interpretable.
Direct-indirect disagreementObserved incoherence may reflect effect modification, bias, data problems or model misspecification.Investigate local and network-wide sources. Do not treat selective exclusion as a default repair.
Sparse evidenceHeterogeneity and incoherence are estimated imprecisely; ranks may be unstable.Use cautious modeling, emphasize intervals and certainty, and avoid overinterpreting non-significant diagnostics.
Influential high-risk evidenceBias in one direct comparison can propagate to multiple network estimates.Use contribution-aware credibility methods and sensitivity analysis; report which estimates depend on the evidence.

Where CINeMA, GRADE and PRISMA-NMA fit

CINeMA evaluates confidence in network estimates across six domains: within-study bias, reporting bias, indirectness, imprecision, heterogeneity and incoherence. GRADE-NMA addresses closely related certainty questions using GRADE’s conceptual framework. The two approaches are not identical; empirical comparison has shown that they can produce different certainty judgments for the same network estimates. That is another reason not to treat transitivity or heterogeneity as mechanically scoreable.8,16


For reporting, PRISMA-NMA 2015 remains the current published NMA extension at the evidence cut-off used for this guide. An update is underway. The 2025 scoping review supporting that update identified new candidate reporting needs, including clearer reporting of methods used to assess homogeneity and transitivity and developments in complex NMA methods. Until a replacement guideline is finalized, the published 2015 extension remains the operational baseline, used alongside PRISMA 2020 where appropriate.11,12


Apply the reasoning before interpreting the result

The paired Network meta-analysis assumptions checklist turns the reasoning in this guide into a structured appraisal. Use it to record network geometry, effect-modifier comparability, heterogeneity, incoherence, bias and ranking concerns without calculating an overall validity score. The purpose is to make the reasoning auditable, especially where the evidence is incomplete or a specialist judgment is still needed.


For the larger comparative-effectiveness context, including how NMA relates to direct evidence, ITCs, population adjustment and HTA decision questions, use Evidence Synthesis for HTA and HEOR: Comparative Effectiveness and Decision Support. Pair 31 deliberately stops before the detailed appraisal of anchored and unanchored indirect comparisons, which belongs to the separate ITC guide.


Frequently asked questions

Is transitivity the same as consistency?

No. Transitivity is a clinical and methodological assumption about whether the treatment comparisons are sufficiently comparable for indirect inference. Consistency or coherence concerns statistical agreement between direct and indirect evidence where both can be observed.

Can I prove transitivity by showing balanced baseline tables?

No. Baseline comparisons can support the judgment, especially for prespecified treatment-effect modifiers, but aggregate reporting is incomplete and no universal balance threshold proves transitivity.

Does a star network fail the consistency assumption?

Not automatically. In a loop-free star network, ordinary direct-indirect incoherence is structurally unavailable for assessment. The network may still be credible if transitivity is plausible, but there is less empirical opportunity to challenge that assumption.

Should every NMA use both node splitting and a global inconsistency test?

No universal sequence applies to every network. Local and global approaches answer different diagnostic questions, and the available methods depend on network geometry and data. Their power limitations should be acknowledged.

Is high I² enough to reject an NMA?

No. There is no universal I² cut-off for NMA credibility. Statistical heterogeneity should be interpreted with clinical and methodological variation, the effect scale, the number of studies and uncertainty in the heterogeneity estimate.

Does RoB NMA replace ROB-MEN?

No. RoB NMA evaluates potential bias in how an NMA was conducted and interpreted. ROB-MEN specifically evaluates bias due to missing evidence in network meta-analysis. They address different parts of credibility.

Can meta-regression fix a transitivity problem?

Sometimes it can explore or adjust for measured effect modification when the network contains enough information. It cannot guarantee repair of an implausible network, remove unmeasured effect modification, or create information that was never reported.

Should I report SUCRA if evidence certainty is low?

You may report ranking metrics if they answer a prespecified hierarchy question, but they should not be interpreted as stand-alone evidence of superiority. Effect magnitude, intervals, clinical importance and certainty need to accompany them.


References

  1. Chaimani A, Caldwell DM, Li T, Higgins JPT, Salanti G. Chapter 11: Undertaking network meta-analyses. In: Higgins JPT, Thomas J, Chandler J, et al, editors. Cochrane Handbook for Systematic Reviews of Interventions. Current online version. Cochrane. https://www.cochrane.org/authors/handbooks-and-manuals/handbook/current/chapter-11
  2. Jansen JP, Naci H. Is network meta-analysis as valid as standard pairwise meta-analysis? It all depends on the distribution of effect modifiers. BMC Med. 2013;11:159. https://doi.org/10.1186/1741-7015-11-159
  3. Spineli LM, Kalyvas C, Yepes-Nuñez JJ, et al. Low awareness of the transitivity assumption in complex networks of interventions: a systematic survey from 721 network meta-analyses. BMC Med. 2024;22:112. https://doi.org/10.1186/s12916-024-03322-1
  4. Spineli LM. An empirical study on 209 networks of treatments revealed intransitivity to be common and multiple statistical tests suboptimal to assess transitivity. BMC Med Res Methodol. 2024;24:301. https://doi.org/10.1186/s12874-024-02436-7
  5. Brignardello-Petersen R, Tomlinson G, Florez I, et al. Grading of Recommendations Assessment, Development, and Evaluation concept article 5: addressing intransitivity in a network meta-analysis. J Clin Epidemiol. 2023;160:151-159. https://doi.org/10.1016/j.jclinepi.2023.06.010
  6. Higgins JPT, Jackson D, Barrett JK, Lu G, Ades AE, White IR. Consistency and inconsistency in network meta-analysis: concepts and models for multi-arm studies. Res Synth Methods. 2012;3(2):98-110. https://doi.org/10.1002/jrsm.1044
  7. Jackson D, Boddington P, White IR. The design-by-treatment interaction model: a unifying framework for modelling loop inconsistency in network meta-analysis. Res Synth Methods. 2016;7(3):329-332. https://doi.org/10.1002/jrsm.1188
  8. Nikolakopoulou A, Higgins JPT, Papakonstantinou T, et al. CINeMA: an approach for assessing confidence in the results of a network meta-analysis. PLoS Med. 2020;17(4):e1003082. https://doi.org/10.1371/journal.pmed.1003082
  9. Chiocchia V, Nikolakopoulou A, Higgins JPT, et al. ROB-MEN: a tool to assess risk of bias due to missing evidence in network meta-analysis. BMC Med. 2021;19:304. https://doi.org/10.1186/s12916-021-02166-3
  10. Lunny C, Higgins JPT, White IR, Dias S, Hutton B, Wright JM, et al. Risk of Bias in Network Meta-Analysis (RoB NMA) tool. BMJ. 2025;388:e079839. https://doi.org/10.1136/bmj-2024-079839
  11. Hutton B, Salanti G, Caldwell DM, et al. The PRISMA Extension Statement for Reporting of Systematic Reviews Incorporating Network Meta-analyses of Health Care Interventions: checklist and explanations. Ann Intern Med. 2015;162(11):777-784. https://doi.org/10.7326/M14-2385
  12. Veroniki AA, Tricco AC, Rangira D, et al. Updating the PRISMA reporting guideline for network meta-analysis: a scoping review. J Clin Epidemiol. 2025;188:111985. https://doi.org/10.1016/j.jclinepi.2025.111985
  13. Chiocchia V, White IR, Salanti G. The complexity underlying treatment rankings: how to use them and what to look at. BMJ Evid Based Med. 2023;28(3):180-182. https://doi.org/10.1136/bmjebm-2021-111904
  14. Salanti G, Nikolakopoulou A, Efthimiou O, Mavridis D, Egger M, White IR. Introducing the Treatment Hierarchy Question in Network Meta-Analysis. Am J Epidemiol. 2022;191(5):930-938. https://doi.org/10.1093/aje/kwab278
  15. Curteis T, Wigle A, Michaels CJ, Nikolakopoulou A. Ranking of treatments in network meta-analysis: incorporating minimally important differences. BMC Med Res Methodol. 2025;25:67. https://doi.org/10.1186/s12874-025-02499-0
  16. Noori A, Sadeghirad B, Thabane L, et al. The GRADE Working Group and CINeMA approaches provided inconsistent certainty of evidence ratings for a network meta-analysis of opioids for chronic noncancer pain. J Clin Epidemiol. 2024;169:111276. https://doi.org/10.1016/j.jclinepi.2024.111276
  17. Wang Y, Xia R, Pericic TP, et al. How do network meta-analyses address intransitivity when assessing certainty of evidence: a systematic survey. BMJ Open. 2023;13:e075212. https://doi.org/10.1136/bmjopen-2023-075212
  18. Brignardello-Petersen R, Florez ID, Yepes-Nuñez JJ, et al. Introduction to network meta-analysis: understanding what it is, how it is done, and how it can be used for decision-making. Am J Epidemiol. 2025;194(3):837-843. https://doi.org/10.1093/aje/kwae260

Similar Posts

Leave a Reply

Your email address will not be published. Required fields are marked *