Indirect Treatment Comparison Appraisal Checklist and Evidence Table

Practical appraisal resourceEvidence checked: 9 August 2026

The main appraisal question is not “Was an indirect treatment comparison performed?” It is “Does this particular evidence structure support this particular comparison?” Two trials can share a comparator and still be poor candidates for indirect comparison because the common arm is not clinically equivalent, the populations differ on plausible treatment-effect modifiers, the endpoint or timepoint is not comparable, or the reported estimates target different treatment effects. Conversely, a population-adjusted analysis can look statistically sophisticated while remaining fragile because important variables were unavailable or overlap was weak.1,3,7,14

This resource is built around an evidence dossier rather than a generic quality score. You first reconstruct the comparison: what treatment effect is being requested, which trials supply each leg, whether a randomized anchor exists, and which population the result is meant to inform. You then compare the common comparator, patient characteristics, treatment-effect modifiers, outcomes, estimands, effect measures and any population-adjustment diagnostics side by side. Only after that record is visible do you make a narrative credibility judgment. The tool does not calculate Bucher estimates, MAIC weights, STC predictions or ML-NMR models; it appraises whether a reported analysis has a defensible evidentiary foundation.4,5,8

Why this tool is different from an NMA assumptions checklist. Network meta-analysis appraisal asks whether an evidence network is credible as a network. This resource asks a narrower question: whether a specific cross-trial treatment comparison has a trustworthy anchor, comparable trials, a compatible target effect and an analysis whose assumptions fit the available data. Detailed network incoherence testing belongs elsewhere.

Start with the comparison claim, not the statistical method

An ITC is often introduced by naming its method—Bucher, MAIC, STC, ML-NMR—before stating the treatment effect that decision-makers actually need. That reverses the logic. The target comparison must come first. Record the population, interventions, comparator, endpoint, time horizon and treatment effect of interest. Then ask which evidence can identify that effect. The same dataset can support one estimand reasonably well and another poorly. For example, two oncology trials may both report progression-free survival, yet differ in treatment switching, assessment schedules or censoring strategies. The labels match; the target effects may not.13,16

The first distinction is structural. An anchored comparison links treatment A and treatment B through a common comparator C. Randomization is preserved within A-versus-C and B-versus-C, so the indirect comparison uses relative effects rather than raw arm outcomes. An unanchored comparison has no randomized common comparator linking the evidence. Quantitative analysis may still be possible, but the protection provided by within-trial randomization no longer carries across the comparison, and the burden of modeling cross-trial differences becomes much larger.1,7,12

Four layers that must align before an indirect comparison is credible A vertical four-layer stack shows the target decision, treatment anchor, population comparability, and outcome-estimand alignment. A warning at the side states that a statistical method cannot repair failure in every layer. ITC credibility is a layered problem 1. Decision targetPopulation, treatment contrast, endpoint, time horizon, estimand 2. Treatment anchorIs the shared comparator clinically equivalent across trials? 3. Population bridgeDo plausible effect modifiers support transport of relative effects? 4. Outcome and estimandAre the trials estimating compatible clinical effects? Method comes afterevidence structureMAIC, STC or ML-NMRcannot make a differentdose, endpoint or estimandbecome identical.
Figure 1. Appraise the decision target, anchor, populations and treatment effect before judging the statistical method. Population adjustment is only one layer of the problem.

The common comparator is an anchor only when the common treatment is genuinely comparable

The common comparator does more than connect trial diagrams. In a simple adjusted indirect comparison, it allows relative effects from separate randomized trials to be contrasted without directly comparing raw outcomes across unrelated patient groups. That logic fails if the “same” comparator does not represent the same clinical treatment strategy across trials. A placebo plus modern background therapy can be a different clinical comparator from placebo plus older supportive care. The same active drug at another dose, route, schedule or treatment line may also change the effect observed in the control arm.1,3,7

This is a treatment-level problem, not merely a covariate problem. NICE DSU TSD 18 is explicit that population adjustment addresses differences in patient covariate distributions; it does not adjust away treatment differences such as dosing, formulation, administration, co-treatment, titration or switching. If the anchor itself changes, matching age, disease severity and prior therapy cannot restore the randomized contrast that would have existed under an equivalent comparator.7,8

A common comparator can connect two trials while still failing as a clinical anchor Two trial boxes show treatment A versus comparator C1 and treatment B versus comparator C2. C1 and C2 share a label but differ in background therapy and switching, causing the bridge between trials to fracture. The comparator name is not the comparator strategy Trial 1Trial 2 Treatment AComparator Cmodern background Treatment BComparator Colder background Anchor fracture Rescue therapy allowedSwitching permittedNo rescue therapyDifferent switching rule Patient reweighting cannot turn two different comparator strategies into one.Record regimen and trial-era differences before considering population adjustment.
Figure 2. A nominally shared comparator can fail as an anchor when its dose, background care, rescue therapy or switching strategy differs materially across trials.

Build the evidence dossier before making a credibility judgment

Reading trial reports one at a time encourages memory-based appraisal: “the populations looked similar,” “both used placebo,” or “the same endpoint was reported.” Those impressions are difficult to audit. The dossier below forces the relevant evidence into one record. It is deliberately broader than a baseline-characteristics table because cross-trial credibility can fail at the treatment, population, outcome or analysis level. The aim is not to make every difference disappear. It is to distinguish differences that are likely irrelevant from differences that alter the treatment effect being transported across trials.5,6,14

Indirect treatment comparison evidence dossier

Document the comparison first. Appraise it second. Entries are saved only in this browser when you select Save locally.

No important concern identified
Available evidence supports the comparison on this checkpoint.
Concern identified
A difference or assumption may change credibility and needs explanation.
Insufficient information
The published evidence does not permit a defensible judgment.
Requires specialist review
The issue depends on statistical or clinical expertise beyond this checklist.
A. Comparison identity and evidence structure

Define the treatment effect that the indirect analysis is supposed to inform. Do not start by choosing a method.

B. Common-comparator integrity record

Use this section for anchored comparisons. A shared treatment label is not enough; compare the actual comparator strategy in each trial.

Comparator featureTrial informing ATrial informing BAppraisal note
Name, dose, route
Background / co-treatment
Rescue, switching, titration
Treatment era / standard of care

Does the common comparator represent a sufficiently similar clinical strategy across the linked trials?

If treatment-level differences plausibly change the comparator response or relative effect, patient-level reweighting cannot repair the anchor.

C. Cross-trial population and covariate matrix

Compare variables because of their role in the treatment effect, not simply because they appear in Table 1. In anchored analyses, plausible treatment-effect modifiers are the main concern. In unanchored analyses, prognostic variables also become central because absolute outcomes are compared across populations.7,8,12

VariableTrial informing ATrial informing BRoleWhy it matters

Were plausible treatment-effect modifiers identified using clinical and methodological reasoning rather than statistical significance alone?

Failure to detect a statistical interaction does not show that a variable cannot modify treatment effect. Important variables should be prespecified where possible and justified from the disease and treatment context.

Do the available covariate distributions support the intended transport of treatment effects?

There is no universal numeric threshold for an acceptable imbalance. Record what differs, whether the variable plausibly modifies the treatment effect, and how the analysis addressed the difference.

D. Outcome, timepoint, estimand and effect-measure alignment

The same endpoint name can conceal a different treatment effect. Record how each trial defines and analyses the outcome rather than assuming labels are interchangeable.

FeatureTrial informing ATrial informing BCompatibility note
Outcome definition / threshold
Assessment time / schedule
Intercurrent events
switching, rescue, discontinuation
Effect measure / estimand

Are the trials estimating treatment effects that can be compared on a compatible scale and clinical interpretation?

Mismatched outcome definitions, timepoints, intercurrent-event strategies, or marginal versus conditional effects can make apparently similar results target different quantities.13,16

E. Population-adjustment diagnostic record, if used

Complete this only when MAIC, STC, ML-NMR or another population-adjusted approach is reported. Population-adjusted methods have appeared in heterogeneous HTA settings, including NICE technology appraisals, so the method label alone should not determine credibility.9 Do not use ESS or covariate balance as a substitute for checking whether the chosen variables and target population are appropriate.

Interpret ESS as a diagnostic, not a verdict. A reduced effective sample size reflects the concentration of weighting and loss of precision. It does not, by itself, show that the adjusted estimate is biased or unbiased. A high ESS also does not prove that the right covariates were selected.7,11,12

Does the population-adjustment method fit the evidence structure, target estimand and available data?

No single method is universally preferred. Appraise what the method can identify with the available IPD/aggregate data and what assumptions remain untestable.

F. Residual bias and decision-usefulness

The final judgment should describe what remains uncertain after the analysis, not convert different concerns into one number.

Are trial-level risk-of-bias differences or missing evidence likely to affect one leg of the comparison more than the other?

What uncertainty remains that the reported analysis cannot remove?

Narrative appraisal output

Complete the status fields above, then generate a narrative summary. The output lists concerns and unresolved information; it does not calculate a score.

    Local save uses your browser’s local storage on this device. The page does not send or upload the information you enter.

    How to read the dossier: four failure mechanisms that should not be collapsed into one score

    Broken treatment anchor

    The common comparator differs in regimen, background care, rescue therapy, switching or treatment era. This is not a patient-covariate problem. Population adjustment cannot recreate a comparator that was clinically different across trials.

    Population transport concern

    The linked trials differ on plausible treatment-effect modifiers. In an anchored comparison, this threatens transport of relative effects. In an unanchored comparison, prognostic differences also directly affect the absolute outcomes being compared.

    Outcome or estimand mismatch

    Endpoints share a name but differ in timing, threshold, censoring, intercurrent-event strategy or effect scale. The comparison may then combine estimates of different treatment effects rather than two measurements of the same effect.

    Residual analytical uncertainty

    The method fits the broad evidence structure but relies on limited overlap, unavailable variables, extrapolation or strong modeling assumptions. A method can be correctly implemented and still leave substantial uncertainty.

    Three worked appraisals

    The examples below are illustrative. They show why the evidence dossier should produce a reasoned statement rather than a pass/fail verdict.

    Case 1. Credible randomized anchor with a narrow residual concern

    Evidence structure
    Trial 1 compares A with placebo plus the same background therapy used in Trial 2, which compares B with placebo.
    Population evidence
    Eligibility criteria, disease severity, treatment line and major plausible effect modifiers are similar. One demographic variable differs modestly but there is no strong rationale that it modifies relative treatment effect.
    Outcome/estimand
    Both trials use the same validated endpoint, assessment schedule and treatment-policy strategy.
    Appraisal
    The anchor is credible and the target effects are aligned. The indirect estimate may be decision-useful, while still carrying the lower precision inherent to indirect evidence and the ordinary risk of bias of the underlying trials.

    Case 2. A nominal common comparator that should not be treated as the same anchor

    Evidence structure
    Both trials describe the comparator as “standard care,” but Trial 1 permits a rescue biologic and early switching; Trial 2 does not. The trials were conducted eight years apart after a major change in supportive care.
    Population evidence
    Baseline characteristics are similar enough that a population-adjustment analysis produces excellent covariate balance.
    Outcome/estimand
    The clinical endpoint and follow-up are similar.
    Appraisal
    Good covariate balance does not solve the central problem. The treatment anchor itself has changed. The analysis should be reported as carrying a major comparator-integrity concern; reweighting the patient populations cannot repair differences in the comparator strategy.

    Case 3. Unanchored population adjustment with important unavailable variables

    Evidence structure
    A single-arm study of A is compared with an external study of B. No randomized common comparator links the evidence.
    Population evidence
    IPD are available for A. Age, disease severity and treatment line can be adjusted, but a clinically plausible prognostic factor and biomarker are not reported for the external comparator.
    Diagnostics
    Weighting substantially reduces ESS and several patients receive large weights. Sensitivity analyses vary assumptions about the available covariates but cannot address the unreported variables.
    Appraisal
    The analysis may provide supplementary evidence, but it does not remove the risk of residual systematic error. The dossier should state exactly which variables were unavailable and avoid language implying that balance on observed covariates establishes exchangeability.

    Turn missing information into an explicit evidence finding

    One of the most important outputs of an ITC appraisal is often not a judgment about what the trials show, but a record of what the publications do not allow you to establish. Critical-appraisal research has long shown that indirect comparisons are difficult to judge reliably when clinical and methodological details are poorly reported. A missing characteristic should therefore not be silently treated as balanced, absent or irrelevant. It is an evidence gap in the appraisal itself.5,6

    This matters most when the missing variable has a plausible causal role. Suppose disease stage is a known treatment-effect modifier, but Trial 1 reports its distribution and Trial 2 does not. The correct entry is not “similar populations.” Nor is it automatically “invalid ITC.” The defensible statement is narrower: cross-trial balance on a plausible modifier cannot be assessed from available reporting. That uncertainty may affect the credibility of an anchored analysis; in an unanchored analysis the same missing variable can be even more consequential if it also predicts absolute outcome.

    The same rule applies to treatment and estimand information. If the comparator description does not make clear whether rescue therapy was permitted, record the switching/rescue strategy as insufficiently documented. If the analysis does not state whether an effect is marginal or conditional, do not infer the estimand from the software or model label. If the target trial reports only selected baseline variables, list the clinically important variables that remain unavailable. An appraisal is stronger when it preserves these unknowns instead of converting them into assumptions of equivalence.

    Residual uncertainty remains after observed covariates are adjusted A three-zone diagram separates measured and adjusted variables, measured but unavailable variables, and unmeasured or unknown variables. Only the first zone is directly addressed by standard population adjustment. Observed balance is only part of the credibility problem Measured + availableImportant but unavailableUnmeasured / unknown • reported effect modifiers• reported prognostic factors• covariates used in adjustment• balance / overlap diagnostics • variable known to matter• missing from comparator report• cannot be matched directly• must remain in uncertainty log • unrecorded prognostic factors• unknown effect modifiers• residual confounding risk• not testable from observed balance Adjustment can act hereDocument explicitlyCannot be ruled out
    Figure 3. Population adjustment works with measured, available variables. Missing and unmeasured variables belong in the residual-uncertainty statement rather than being treated as balanced.

    Separate statistical precision from causal credibility

    Indirect estimates are often less precise than direct randomized comparisons because uncertainty from both evidence legs contributes to the final contrast. Simulation work has shown that indirect comparisons can require substantially more evidence to achieve the precision of a direct comparison, particularly when one side of the anchor is supported by few trials.17 That loss of precision is important, but it is not the same problem as bias. A wide interval around a well-anchored estimate and a narrow interval around a confounded unanchored estimate require very different interpretations.

    The checklist therefore records precision-related diagnostics without turning them into a validity grade. A small ESS, extreme MAIC weights, one trial per evidence leg or a wide confidence interval can make an estimate fragile. But a precise estimate can still be causally weak if the comparator changed, a treatment-effect modifier was omitted, or the outcome estimands are incompatible. Conversely, an imprecise estimate can arise from limited sample size even when the structural assumptions are reasonable. This distinction is particularly important in HTA, where a decision-maker needs to know both how uncertain the numerical estimate is and how credible the identification strategy is.2,18

    Completion audit: what must be visible before the ITC leaves appraisal

    Before an indirect estimate is carried into an HTA dossier, economic model, guideline evidence table or manuscript conclusion, the appraisal should leave an auditable trail. The purpose is not bureaucratic completeness. Each item below prevents a different form of ambiguity from being hidden behind the final effect estimate.

    Comparison identity is explicit

    The record names treatment A, treatment B, the target population, endpoint, time horizon and target effect. If an anchor is claimed, the common comparator is named as an actual regimen rather than a generic label.

    Cross-trial differences are classified

    Population differences are separated from treatment-regimen differences and outcome/estimand differences. This prevents a population-adjustment method from being credited with solving a problem it cannot address.

    Unavailable information is preserved

    Important unreported variables, unclear switching rules and unspecified estimands remain visible as unresolved information. They are not converted into “no concern” because no contrary evidence was found.

    The final sentence names the assumption burden

    The conclusion states what makes the estimate usable and what still limits it: for example, credible randomized anchor but uncertain effect-modifier balance; or good observed balance but no randomized anchor and important residual confounding risk.

    The result is intentionally different from a generic “quality checklist.” A user should be able to hand the completed dossier to another reviewer and show exactly where the treatment bridge is sound, where the trials differ, which variables were adjusted, which variables were unavailable, and what part of the uncertainty remains after the statistical analysis. That traceability is the practical value of the resource.

    What this resource can establish—and what it cannot

    A completed dossier can show whether the comparison was defined coherently, whether a randomized anchor exists, whether the common comparator appears clinically equivalent, whether plausible treatment-effect modifiers were considered, whether outcomes and estimands align, and whether a population-adjustment analysis reports the diagnostics needed for interpretation. It can also expose missing information that would otherwise disappear inside a polished statistical result.

    It cannot validate an indirect treatment comparison by itself. It does not replace statistical programming, reproduce the full conduct guidance for MAIC/STC/ML-NMR, or determine whether a model has been correctly estimated. Nor does it convert an unanchored analysis into randomized evidence. Some concerns require specialist statistical review; others require disease-area expertise because the relevance of a difference depends on how treatments work and how outcomes are generated.7,10,14,15

    Do not “repair” an ITC by deleting inconvenient differences from the appraisal. If an important covariate was not reported, record that uncertainty. If the comparator strategy changed, record the anchor problem. If estimands differ, describe the mismatch. The purpose of appraisal is to make the assumption burden visible before the estimate is used in an HTA, model or clinical decision.

    Next step: explain the failure mechanism before judging the estimate

    The paired methodology guide, Indirect Treatment Comparisons Explained: Common Comparators, Assumptions and Credibility, examines why naive cross-trial comparisons fail, what a randomized anchor preserves, how anchored and unanchored assumptions differ, and which kinds of mismatch can or cannot be addressed by population adjustment. For the broader decision context, see Evidence Synthesis for HTA and HEOR: Comparative Effectiveness and Decision Support.

    Get every MetaSyn template free, including this one.

    Leave your email and I’ll send this resource as an editable Word file and a printable PDF, plus access to the smart online version. You’ll also get every new template as it’s finished. No noise, just the resources.

    Frequently asked questions: difficult appraisal cases

    What if the common comparator dose differs only slightly between trials?

    Do not apply an automatic tolerance. Record the difference and ask whether it is clinically plausible that the dose, schedule or background regimen changes the comparator response or the relative treatment effect. If the difference is unlikely to matter, document why. If it could matter, retain it as a credibility concern rather than assuming population adjustment will solve it.

    What if a plausible effect modifier is reported in only one of the two trials?

    Record the variable as unassessable across trials. Its absence from the second report is not evidence of balance. For a population-adjusted analysis, the inability to adjust an important variable should be carried into the residual-uncertainty statement.

    Can an ITC be useful when a potential structural problem remains?

    Sometimes, but the decision use must match the evidence strength. A comparison may still be exploratory or supplementary when no better evidence exists. The appraisal should state the unresolved problem and avoid presenting the estimate as if it had the protection of a credible randomized anchor.

    Should a high retained MAIC effective sample size reassure me?

    Only about precision and weight concentration, and even then in context. A high ESS does not show that the correct covariates were selected, that an unreported modifier is absent, or that the common comparator and estimand are compatible.

    References

    1. Bucher HC, Guyatt GH, Griffith LE, Walter SD. The results of direct and indirect treatment comparisons in meta-analysis of randomized controlled trials. J Clin Epidemiol. 1997;50(6):683-691. doi:10.1016/S0895-4356(97)00049-8.
    2. Song F, Altman DG, Glenny AM, Deeks JJ. Validity of indirect comparison for estimating efficacy of competing interventions: empirical evidence from published meta-analyses. BMJ. 2003;326:472-475. doi:10.1136/bmj.326.7387.472.
    3. Jansen JP, Fleurence R, Devine B, et al. Interpreting indirect treatment comparisons and network meta-analysis for health-care decision making: report of the ISPOR Task Force on Indirect Treatment Comparisons Good Research Practices: part 1. Value Health. 2011;14(4):417-428. doi:10.1016/j.jval.2011.04.002.
    4. Hoaglin DC, Hawkins N, Jansen JP, et al. Conducting indirect-treatment-comparison and network-meta-analysis studies: report of the ISPOR Task Force on Indirect Treatment Comparisons Good Research Practices: part 2. Value Health. 2011;14(4):429-437. doi:10.1016/j.jval.2011.01.011.
    5. Jansen JP, Trikalinos T, Cappelleri JC, et al. Indirect treatment comparison/network meta-analysis study questionnaire to assess relevance and credibility to inform health care decision making. Value Health. 2014;17(2):157-173. doi:10.1016/j.jval.2014.01.004.
    6. Ortega A, Fraga Fuentes MD, Alegre-del-Rey EJ, et al. A checklist for critical appraisal of indirect comparisons. Int J Clin Pract. 2014;68(10):1181-1189. doi:10.1111/ijcp.12487.
    7. Phillippo DM, Ades AE, Dias S, Palmer S, Abrams KR, Welton NJ. NICE DSU Technical Support Document 18: Methods for population-adjusted indirect comparisons in submissions to NICE. 2016. NICE Decision Support Unit.
    8. Phillippo DM, Ades AE, Dias S, Palmer S, Abrams KR, Welton NJ. Methods for population-adjusted indirect comparisons in health technology appraisal. Med Decis Making. 2018;38(2):200-211. doi:10.1177/0272989X17725740.
    9. Phillippo DM, Ades AE, Dias S, Palmer S, Abrams KR, Welton NJ. Population adjustment methods for indirect comparisons: a review of National Institute for Health and Care Excellence technology appraisals. Int J Technol Assess Health Care. 2019;35(3):221-228. doi:10.1017/S0266462319000333.
    10. Phillippo DM, Dias S, Ades AE, Welton NJ. Multilevel network meta-regression for population-adjusted treatment comparisons. J R Stat Soc Ser A. 2020;183(3):1189-1210. doi:10.1111/rssa.12579.
    11. Phillippo DM, Dias S, Ades AE, Welton NJ. Assessing the performance of population adjustment methods for anchored indirect comparisons: a simulation study. Stat Med. 2020;39(30):4885-4911. doi:10.1002/sim.8759.
    12. Remiro-Azócar A, Heath A, Baio G. Methods for population adjustment with limited access to individual patient data: a review and simulation study. Res Synth Methods. 2021;12(6):750-775. doi:10.1002/jrsm.1511.
    13. Remiro-Azócar A. Target estimands for population-adjusted indirect comparisons. Stat Med. 2022;41(28):5558-5569. doi:10.1002/sim.9413.
    14. Health Technology Assessment Coordination Group. Methodological Guideline for Quantitative Evidence Synthesis: Direct and Indirect Comparisons. Adopted 8 March 2024. European Commission.
    15. Health Technology Assessment Coordination Group. Practical Guideline for Quantitative Evidence Synthesis: Direct and Indirect Comparisons. Adopted 8 March 2024. European Commission.
    16. International Council for Harmonisation. ICH E9(R1) Addendum on Estimands and Sensitivity Analysis in Clinical Trials. Step 5. European Medicines Agency.
    17. Mills EJ, Ghement I, O’Regan C, Thorlund K. Estimating the power of indirect comparisons: a simulation study. PLoS One. 2011;6(1):e16237. doi:10.1371/journal.pone.0016237.
    18. Guo JD, et al. Selection of indirect treatment comparisons for health technology assessments: a practical guide for health economics and outcomes research scientists and clinicians. BMJ Open. 2025;15(3):e091961. Publisher record.
    Methodological scope note. This resource appraises the evidentiary foundation and interpretability of an indirect treatment comparison. It does not calculate treatment effects, replace disease-area expertise, reproduce NICE or EU JCA submission requirements, or certify that an analysis is valid. Guidance status and consequential methodological claims were checked through 9 August 2026.