Indirect Treatment Comparisons Explained: Common Comparators, Assumptions and Credibility
An indirect treatment comparison fails for identifiable reasons. The failure may occur because raw outcomes from unrelated trials are compared as if the patients had been randomized together; because the nominally shared comparator is not actually the same clinical strategy; because plausible treatment-effect modifiers differ across trials; because outcomes or estimands do not describe the same treatment effect; or because an unanchored model cannot account for important prognostic differences. The statistical label – Bucher, MAIC, STC, ML-NMR – does not by itself tell you which of those problems has been solved.1,3,7,18
This guide therefore treats indirect comparison as a problem of causal transport across separate trials, not as a catalogue of statistical methods. The practical question is: what information is being carried from one randomized comparison to another, and what assumptions are required for that transport to remain defensible? That framing separates Pair 32 from a general network meta-analysis guide. Network geometry, incoherence diagnostics and certainty frameworks matter in larger networks, but the distinctive problem here is the credibility of a specific indirect treatment contrast used when head-to-head evidence is absent or incomplete.
1. The naive cross-trial comparison answers the wrong question
Suppose Trial 1 reports a 70% response rate for treatment A and Trial 2 reports a 50% response rate for treatment B. Subtracting 50% from 70% is tempting. It is also not an adjusted indirect comparison. Patients in the two trials were not randomized between A and B. Differences in baseline prognosis, eligibility criteria, supportive care, calendar time or outcome assessment can produce the apparent 20-point gap even if A and B have identical treatment effects.1,2
The common comparator is what allows a simple adjusted ITC to avoid that mistake. If Trial 1 randomizes A versus C and Trial 2 randomizes B versus C, each trial estimates a randomized relative effect against C. The indirect A-versus-B estimate contrasts those relative effects. This does not randomize A against B, but it preserves more of the randomized structure than a comparison of raw A and B outcomes across studies. The Bucher approach remains the clearest conceptual example of that logic.1,4
2. A common comparator is a conditional bridge, not a magic label
It is easy to turn “common comparator” into a diagrammatic criterion: if C appears in both trials, the network is connected. For ITC credibility, that is not enough. C must also be sufficiently comparable as a treatment strategy. A placebo arm may include different background therapies, rescue medication or switching rules. An active comparator may be given at a different dose or in another treatment line. Standard of care may evolve between trial eras. If those differences affect outcomes or relative treatment effects, the anchor is no longer functioning as the same randomized reference across studies.7,8
This distinction matters because population adjustment works on patient characteristics. Reweighting can change the distribution of age, disease severity or prior therapy among patients represented by an IPD dataset. It cannot retroactively change a drug dose, replace an old background regimen with a new one, or make treatment switching rules identical. NICE DSU TSD 18 specifically separates patient covariate adjustment from these treatment-related differences. That is why a “good MAIC balance table” is not sufficient evidence that the common comparator is a credible anchor.7
3. Anchored and unanchored comparisons do not carry the same burden of proof
In an anchored comparison, the randomized legs provide relative treatment effects. Purely prognostic variables influence baseline outcome risk, but when they affect both arms of a randomized trial similarly they do not necessarily bias the relative treatment effect. The cross-trial variables that matter most are therefore plausible treatment-effect modifiers: characteristics that change the relative effect of treatment. If those modifiers differ across the A-versus-C and B-versus-C trial populations, the indirect A-versus-B contrast may no longer transport cleanly.7,8,12
An unanchored comparison is different because there is no common randomized reference to cancel baseline prognostic differences. The model must explain differences in absolute outcomes across populations. Prognostic factors therefore join treatment-effect modifiers in the adjustment set, and the assumption burden rises sharply. Any important unmeasured or unreported prognostic variable can create residual systematic error that the analysis cannot diagnose from the observed data alone. This is why unanchored estimates are not simply anchored estimates with one extra modeling step.7,12,14
4. Four kinds of mismatch should be diagnosed separately
Many appraisals use one broad word “similarity” for every cross-trial difference. That hides the mechanism. A treatment-regimen mismatch is not the same as an effect-modifier imbalance; an endpoint mismatch is not the same as poor overlap in MAIC. Separating the failure mechanisms helps determine whether adjustment is relevant at all.
Treatment-level mismatch
Dose, route, formulation, background therapy, rescue treatment, switching, titration or trial-era standard of care differ. These differences can undermine the common-comparator anchor and generally cannot be repaired by patient reweighting.
Population-level mismatch
The trial populations differ on plausible effect modifiers, or – when unanchored – on prognostic factors. Population-adjusted methods are designed for this class of problem when the relevant variables are measured and available.
Outcome/estimand mismatch
The trials report nominally similar endpoints but use different thresholds, follow-up schedules, censoring rules or intercurrent-event strategies. The estimates may target different clinical effects even if the variable name is identical.
Analysis-level mismatch
Effect measures or estimands do not align; one estimate is conditional and another marginal; overlap is limited; or the model relies on extrapolation that is weakly supported by observed data. This calls for statistical scrutiny rather than a generic similarity judgment.
5. Population adjustment is a targeted repair, not a universal repair
MAIC, STC and ML-NMR are often grouped because each addresses population differences when individual patient data are available for at least part of the evidence base. Their mechanics differ, and historical NICE appraisal practice shows substantial variation in when and how population adjustment has been used.9 MAIC reweights IPD to resemble a target aggregate-data population. STC models outcomes in the IPD and predicts in the target population. ML-NMR extends meta-regression to combine individual- and aggregate-level information across larger connected networks. None of those descriptions implies that one method is universally best.7,10,12
Method choice follows the evidence structure, available data, outcome model and target estimand. MAIC can be reasonable in a pairwise setting with adequate overlap and a defensible set of effect modifiers. Regression-based approaches allow extrapolation, but extrapolation is an assumption, not free information. ML-NMR offers coherent population adjustment across networks and can target decision populations flexibly, but it does not discover unmeasured confounding and does not make incompatible treatment regimens or estimands comparable. The 2024 EU HTA methodological and practical guidance recognises multiple indirect-comparison approaches while requiring the assumptions and limitations of the selected analysis to be made explicit.14,15
6. Effect modifiers and prognostic factors matter for different reasons
A prognostic factor predicts outcome regardless of treatment. A treatment-effect modifier changes the relative effect of one treatment versus another. The distinction is not merely terminological. In an anchored analysis of relative effects, a purely prognostic factor may change baseline risk without changing the randomized treatment contrast, whereas an imbalanced effect modifier can change that contrast across trial populations. In an unanchored comparison, prognostic factors also matter because absolute outcomes from different populations are being connected.7,12
Real data rarely tell you with certainty which variables modify treatment effect. Interaction tests are often underpowered. A sensible appraisal therefore asks how candidate variables were chosen: clinical knowledge, biological rationale, prior subgroup evidence, disease-area literature and regulatory precedent are legitimate inputs. Statistical significance alone is not. Equally, a variable should not be declared irrelevant merely because its observed mean happens to be similar between two trials. The role of the variable and the observed imbalance are different questions.
7. ESS and overlap are diagnostics, not verdicts
MAIC often attracts a single-number interpretation: if the effective sample size falls “too much,” the analysis is declared bad; if it remains high, the analysis is treated as reassuring. Neither conclusion is defensible as a universal rule. ESS summarises the concentration of weights and the precision lost when the IPD are reweighted toward a target population. A large reduction tells you that a smaller subset of the source trial is carrying more of the target-population representation. It can signal limited overlap and statistical fragility, but it does not prove bias.7,11,12
The reverse is equally important. A high retained ESS can coexist with a badly specified adjustment model if a clinically important effect modifier was never included because it was unavailable or overlooked. Weight diagnostics must therefore be read with covariate selection, overlap, balance, sensitivity analyses and the target estimand. Current formal HTA guidance does not provide a universal minimum ESS or percentage-loss threshold that turns this judgment into a pass/fail calculation.
8. Estimands expose hidden incompatibility
Indirect comparisons are often described as if matching the endpoint name and follow-up time were enough. ICH E9(R1) provides a useful reminder that a treatment effect is defined by more than the variable measured. The treatment condition of interest, target population, endpoint, handling of intercurrent events and population-level summary all contribute to the estimand. Two trials can both report progression-free survival and still target different effects if one follows patients regardless of treatment switching while another censors or models switching under a hypothetical strategy.13,16
This issue becomes sharper in population-adjusted ITCs. Marginal and conditional effects are not interchangeable for non-collapsible measures such as odds ratios and hazard ratios. The appraisal does not need to teach collapsibility algebra. It does need to ask whether the estimates being combined have the same population-level interpretation. If one analysis targets a conditional regression coefficient and another a marginal population-average effect, the problem is not fixed by using the same effect-measure abbreviation.12,13
9. Three failure analyses: same method, different credibility
The examples below deliberately focus on reasoning rather than software. They also show why “MAIC was used” or “the comparison was anchored” is not enough to characterize credibility.
Failure analysis A – the statistical method is fine; the comparator is not
Failure analysis B – the anchor is strong; transport is not
Failure analysis C – observed balance looks good; hidden bias remains
9. Why precision cannot rescue a weak identification strategy
An indirect comparison produces an estimate with a standard error, confidence interval or credible interval. Those quantities describe statistical uncertainty under the fitted analysis. They do not tell you whether the cross-trial assumptions were correct. This distinction is easy to lose because a narrow interval looks authoritative. Yet a precisely estimated contrast can still target the wrong effect or be biased by a broken comparator, unmeasured prognosis or incompatible estimands.
The reverse also occurs. A structurally credible anchored comparison can be imprecise because each randomized leg contains only one or two trials. Simulation studies of indirect comparisons have shown the substantial power penalty created by combining uncertainty from separate evidence legs.17 That problem should be described as imprecision, not as evidence that the anchor or transitivity assumption failed. Good appraisal separates four questions: Is the treatment contrast identified by the evidence structure? Are the cross-trial assumptions plausible? Is the estimator appropriate for the target effect? How precise is the resulting estimate?
This separation also explains why a low MAIC ESS is not a general “bias score.” ESS belongs mostly to the precision and overlap story. It can reveal that the target population is represented by a small, heavily weighted subset of the IPD trial. It cannot reveal an unmeasured biomarker, prove that the common comparator is equivalent, or show that the target estimand is correct. A credible appraisal therefore reads statistical diagnostics alongside the causal structure rather than allowing one number to dominate the conclusion.
10. Existing appraisal frameworks point toward structured judgment, not a universal score
Indirect-comparison appraisal did not begin with modern population-adjustment methods. ISPOR’s relevance-and-credibility questionnaire and the published critical-appraisal checklist by Ortega and colleagues were attempts to make reviewers examine the clinical and methodological assumptions rather than accept the statistical output at face value.5,6 Their enduring lesson is useful: credibility depends on several qualitatively different judgments, and those judgments do not combine naturally into a single numerical index.
A missing common comparator, an unreported effect modifier and a wide confidence interval are not interchangeable defects. Giving each “one point” would imply that three minor reporting gaps can equal one fundamental treatment-anchor problem. The arithmetic has no methodological meaning. This is why the paired MetaSyn resource uses narrative statuses and requires a rationale. The reviewer records the mechanism of concern and then states what it does to interpretation.
Newer appraisal tools for network meta-analysis add valuable structure for larger evidence networks, but they should not be copied wholesale into a pairwise ITC guide. Pair 32 deliberately stops before network-level incoherence diagnostics and formal certainty grading. Its contribution is to make the cross-trial bridge inspectable: comparator integrity, treatment-effect transport, target estimand, population-adjustment fit and residual bias.
11. Method selection should follow the failure mechanism
A useful way to choose among indirect-comparison methods is to ask what problem the method is being asked to solve. If two randomized trials share a credible comparator and populations are sufficiently comparable, a simple adjusted indirect comparison may be adequate. If the anchor is credible but plausible effect modifiers are imbalanced and appropriate IPD are available, anchored population adjustment may be warranted. If several treatments and multiple studies form a connected network, ML-NMR may offer a coherent way to model population adjustment across the network. If no randomized anchor exists, the question changes again: the analyst must address absolute-outcome differences, prognostic imbalance and unmeasured confounding risk.
| Observed problem | Methodological response that may be relevant | What still needs checking |
|---|---|---|
| Credible common comparator; no important modifier imbalance | Standard adjusted ITC may be sufficient | Comparator integrity, outcome/estimand alignment, study risk of bias, precision |
| Credible anchor; measured effect modifiers differ | Anchored MAIC, STC or other population adjustment may be considered | Variable selection, overlap, target population, estimand, sensitivity analysis |
| Several linked treatments with mixed IPD/AgD | Network meta-regression / ML-NMR may add coherence and flexibility | Network assumptions, model specification, target estimand, data support |
| No randomized common comparator | Unanchored adjustment or non-randomised comparative methods may be considered | Prognostic factors, effect modifiers, residual confounding, absolute-outcome transport, decision role |
| Comparator treatment strategy differs materially | No population-adjustment method directly fixes the treatment-level mismatch | Whether the anchor can still be defended or the analysis requires a different evidence strategy |
This is not a prescriptive algorithm. The 2025 practical review of ITC method selection similarly emphasizes that method choice depends on the evidence setting rather than a single universal ranking of methods.18 The point is simpler: diagnose the mismatch before choosing the repair. A sophisticated model aimed at the wrong failure mechanism can increase complexity without increasing credibility.
12. Decision-usefulness is a graded conclusion even when the method is binary
Some evidence structures are clearly anchored or unanchored. Some outcomes are clearly incompatible. The final decision use, however, is rarely binary. An ITC may be sufficiently credible for exploratory model scenarios but not for a definitive comparative-effectiveness claim. It may be acceptable as supportive evidence in a rare condition while remaining too assumption-dependent to carry the primary conclusion. It may be structurally strong but too imprecise to distinguish treatments.
A well-written credibility conclusion therefore has two parts. First, it states whether the evidence structure and analysis plausibly identify the requested treatment contrast. Second, it states the role that the estimate can reasonably play given its residual uncertainty. This prevents two common errors: rejecting all indirect evidence merely because it is indirect, and accepting a technically elaborate analysis merely because no obvious implementation error is visible.
13. HTA changes the question from “Is this estimate statistically available?” to “Is it decision-useful?”
HTA bodies often receive ITCs because the evidence programme did not include every relevant head-to-head comparison. Their task is not to reward methodological sophistication; it is to judge whether the estimate can inform a defined reimbursement or clinical-assessment question. The 2024 EU HTA Coordination Group guidance therefore places substantial emphasis on evidence structure, assumptions, choice of synthesis method, heterogeneity, population adjustment and transparent reporting. NICE committees likewise evaluate submitted population-adjusted comparisons rather than applying a blanket rule that the label MAIC is acceptable or unacceptable.14,15
This has two practical consequences. First, jurisdiction-specific guidance should not be converted into universal statistical law. Second, a technically valid analysis can still be poorly aligned with the decision population, comparator or estimand. The question “Can this method be fitted?” is therefore weaker than “Does this analysis estimate the comparison the decision-maker needs, under assumptions that are sufficiently credible for this use?”
14. A useful credibility conclusion names the surviving uncertainty
Weak ITC conclusions often end with a generic caveat: “results should be interpreted with caution.” That phrase is almost content-free. A useful conclusion identifies the mechanism. For example: “The comparison is anchored through a clinically equivalent placebo regimen, but the linked trials differ in treatment line, an a priori effect modifier; population adjustment addresses observed treatment-line imbalance, although biomarker status was unavailable in the comparator trial.” That sentence tells a reviewer what was protected by randomization, what was adjusted, and what remains unresolved.
| Credibility question | Strong answer | Weak answer |
|---|---|---|
| What is the anchor? | Same comparator strategy, dose, background care and switching rules | Same treatment name appears in both publications |
| What threatens transport? | Plausible effect modifiers identified and compared across trials | Baseline tables “look similar” |
| What does population adjustment address? | Measured covariate-distribution differences relevant to the target effect | “MAIC controls for all cross-trial differences” |
| What remains uncertain? | Unavailable modifiers, overlap, extrapolation, unmeasured confounding, estimand mismatch | “No major limitations” after balance is achieved |
15. How the practical resource should be used
The paired Indirect treatment comparison appraisal checklist and evidence table is designed for application, not teaching. It reconstructs the specific comparison, records common-comparator integrity, compares trial populations and covariates, aligns outcomes and estimands, and documents population-adjustment diagnostics. Its output is a narrative residual-uncertainty statement rather than a numerical score. Use the guide when you need to understand why a concern matters; use the resource when you need to document where the concern appears in a specific analysis.
For the broader methodological context – including how direct, indirect and network evidence fit an HTA decision problem – see Evidence Synthesis for HTA and HEOR: Comparative Effectiveness and Decision Support. Detailed assessment of network-level transitivity, heterogeneity and incoherence belongs to the separate NMA assumptions guide rather than being repeated here.
Conclusion
Indirect comparison is defensible when the target treatment effect is clear, the evidence structure supports it, the common comparator is genuinely comparable when an anchor is claimed, cross-trial differences that can modify treatment effects are addressed, outcomes and estimands align, and the remaining uncertainty is carried into interpretation. Population adjustment can strengthen an analysis when the problem is measured population imbalance. It cannot repair every cross-trial difference, and it cannot make unmeasured information observable. The strongest appraisal therefore identifies the failure mechanism before judging the estimate.
Frequently asked questions: edge cases
Can an indirect comparison be preferable to a direct trial for an HTA question?
A direct randomized trial is not automatically decision-relevant if it uses the wrong comparator, population or endpoint. An ITC may be necessary to address the actual decision comparator. That does not make indirect evidence inherently stronger; it means relevance and internal validity are separate dimensions.
Should a covariate be excluded from an anchored adjustment because it is already balanced?
Observed balance and effect-modifier status are different questions. If there is strong prior reason that the variable modifies relative treatment effect, its role should be considered regardless of whether the sample means happen to look similar. The appropriate model choice is a statistical issue, but “already balanced” is not evidence that the variable is biologically irrelevant.
What if two trials use the same outcome but one allows treatment switching?
Check the estimand and analysis strategy. If one trial targets an effect regardless of switching and the other targets a hypothetical no-switching effect, the estimates may not be directly comparable even though the endpoint label is the same.
Can an unanchored ITC support an HTA submission?
It can be considered, particularly when randomized comparative evidence is unavailable, but it carries a stronger assumption burden and greater vulnerability to unmeasured confounding. Its role should reflect that uncertainty rather than being presented as equivalent to a credible anchored randomized comparison.
References
- Bucher HC, Guyatt GH, Griffith LE, Walter SD. The results of direct and indirect treatment comparisons in meta-analysis of randomized controlled trials. J Clin Epidemiol. 1997;50(6):683-691. doi:10.1016/S0895-4356(97)00049-8.
- Song F, Altman DG, Glenny AM, Deeks JJ. Validity of indirect comparison for estimating efficacy of competing interventions: empirical evidence from published meta-analyses. BMJ. 2003;326:472-475. doi:10.1136/bmj.326.7387.472.
- Jansen JP, Fleurence R, Devine B, et al. Interpreting indirect treatment comparisons and network meta-analysis for health-care decision making: report of the ISPOR Task Force on Indirect Treatment Comparisons Good Research Practices: part 1. Value Health. 2011;14(4):417-428. doi:10.1016/j.jval.2011.04.002.
- Hoaglin DC, Hawkins N, Jansen JP, et al. Conducting indirect-treatment-comparison and network-meta-analysis studies: report of the ISPOR Task Force on Indirect Treatment Comparisons Good Research Practices: part 2. Value Health. 2011;14(4):429-437. doi:10.1016/j.jval.2011.01.011.
- Jansen JP, Trikalinos T, Cappelleri JC, et al. Indirect treatment comparison/network meta-analysis study questionnaire to assess relevance and credibility to inform health care decision making. Value Health. 2014;17(2):157-173. doi:10.1016/j.jval.2014.01.004.
- Ortega A, Fraga Fuentes MD, Alegre-del-Rey EJ, et al. A checklist for critical appraisal of indirect comparisons. Int J Clin Pract. 2014;68(10):1181-1189. doi:10.1111/ijcp.12487.
- Phillippo DM, Ades AE, Dias S, Palmer S, Abrams KR, Welton NJ. NICE DSU Technical Support Document 18: Methods for population-adjusted indirect comparisons in submissions to NICE. 2016. NICE Decision Support Unit.
- Phillippo DM, Ades AE, Dias S, Palmer S, Abrams KR, Welton NJ. Methods for population-adjusted indirect comparisons in health technology appraisal. Med Decis Making. 2018;38(2):200-211. doi:10.1177/0272989X17725740.
- Phillippo DM, Ades AE, Dias S, Palmer S, Abrams KR, Welton NJ. Population adjustment methods for indirect comparisons: a review of National Institute for Health and Care Excellence technology appraisals. Int J Technol Assess Health Care. 2019;35(3):221-228. doi:10.1017/S0266462319000333.
- Phillippo DM, Dias S, Ades AE, Welton NJ. Multilevel network meta-regression for population-adjusted treatment comparisons. J R Stat Soc Ser A. 2020;183(3):1189-1210. doi:10.1111/rssa.12579.
- Phillippo DM, Dias S, Ades AE, Welton NJ. Assessing the performance of population adjustment methods for anchored indirect comparisons: a simulation study. Stat Med. 2020;39(30):4885-4911. doi:10.1002/sim.8759.
- Remiro-Azócar A, Heath A, Baio G. Methods for population adjustment with limited access to individual patient data: a review and simulation study. Res Synth Methods. 2021;12(6):750-775. doi:10.1002/jrsm.1511.
- Remiro-Azócar A. Target estimands for population-adjusted indirect comparisons. Stat Med. 2022;41(28):5558-5569. doi:10.1002/sim.9413.
- Health Technology Assessment Coordination Group. Methodological Guideline for Quantitative Evidence Synthesis: Direct and Indirect Comparisons. Adopted 8 March 2024. European Commission.
- Health Technology Assessment Coordination Group. Practical Guideline for Quantitative Evidence Synthesis: Direct and Indirect Comparisons. Adopted 8 March 2024. European Commission.
- International Council for Harmonisation. ICH E9(R1) Addendum on Estimands and Sensitivity Analysis in Clinical Trials. Step 5. European Medicines Agency.
- Mills EJ, Ghement I, O’Regan C, Thorlund K. Estimating the power of indirect comparisons: a simulation study. PLoS One. 2011;6(1):e16237. doi:10.1371/journal.pone.0016237.
- Guo JD, et al. Selection of indirect treatment comparisons for health technology assessments: a practical guide for health economics and outcomes research scientists and clinicians. BMJ Open. 2025;15(3):e091961. Publisher record.