Meta-analysis Heterogeneity, Sensitivity Analysis and Publication Bias

A completed forest plot can look like an ending. The effect estimates have been extracted, a synthesis model has been fitted, and a diamond has been drawn. In practice, this is where a demanding interpretive stage begins. Three distinct questions remain. How much do the underlying effects vary across studies? Would the main conclusion survive defensible changes to important analytical assumptions? Could missing evidence or other small-study mechanisms materially change what is observed?

These questions are related but not interchangeable, and no single statistic answers all three.1,23 A homogeneity test does not establish robustness. A stable pooled estimate does not establish that the evidence base is complete. Funnel-plot asymmetry does not identify its own cause. Publication bias is one form of bias due to missing evidence and one possible cause of small-study effects; the terms should not be used as synonyms.15,23

This guide addresses that diagnostic layer. It assumes that the review question, estimand, eligibility criteria, effect measures, and primary analytical model were defined before interpretation. The broad principles apply across many forms of quantitative evidence synthesis, but the exact models, effect measures, heterogeneity structures, and small-study-effect methods must be chosen for the design and outcome at hand. Methods appropriate for pairwise intervention meta-analysis do not automatically transfer to diagnostic-accuracy, prognostic, prevalence, multivariate, network, or dependent-effect syntheses.

Errors at this stage are consequential. A review that treats an I2 threshold as permission to pool may combine studies that estimate scientifically different quantities. A review that removes an influential study because it changes statistical significance may discard valid information. A review that interprets a non-significant asymmetry test from a small evidence base as proof that publication bias is absent makes an inference the test cannot support. Statistical interpretation therefore requires the same prespecification, checking, and documentation expected of data collection.

Operating principle. A statistical diagnostic is a conditional signal, not a verdict. A defensible conclusion integrates the estimand, study design, clinical and methodological context, risk of bias, direct evidence about missing results, and the limitations of the fitted model.
Figure 1. A defensible interpretation begins with the estimand and scientific compatibility, not with a heterogeneity cutoff. Random-effects modelling does not make incompatible studies compatible.

1. Heterogeneity: variation that no single number settles

Heterogeneity takes several forms. Clinical heterogeneity concerns differences in populations, interventions or exposures, comparators, outcomes, and settings. Methodological heterogeneity concerns differences in design, measurement, conduct, and risk of bias. Statistical heterogeneity is variability in observed effect estimates beyond that expected from sampling error under a specified model.1 Statistical measures can describe evidence of variation, but they do not by themselves identify its clinical or methodological cause.

Cochran’s Q is conventionally calculated as a weighted sum of squared deviations of study estimates from the common-effect estimate. Under the null hypothesis of one common underlying effect and the usual assumptions, Q is compared with a chi-square distribution with k − 1 degrees of freedom.2 With few or imprecise studies, the test has low power, so a non-significant result does not demonstrate homogeneity. With many precise studies, it can detect heterogeneity too small to be decision-relevant. Its p-value therefore should not be used as a rule that permits or prohibits pooling.

A conventional estimator of I2 rescales Q as the estimated percentage of variability in observed effect estimates attributable to heterogeneity rather than sampling error under the model:

I2 = max{0, [Q − (k − 1)] / Q} × 100% This is a relative, precision-dependent statistic. It is not the percentage of studies that are heterogeneous, the probability that heterogeneity exists, or an absolute measure of effect dispersion.

Because I2 is a ratio, the same absolute between-study variance can yield different I2 values when study precision differs. When within-study variances become small and non-zero between-study variance remains, I2 can become high even if the absolute variation is modest on the effect scale.3 It is also imprecisely estimated in small meta-analyses and may be biased, so an interval estimate should be considered where available.4 Fixed labels such as “low”, “moderate”, or “high” can aid description, but they cannot replace judgment about effect direction, magnitude, decision thresholds, and study compatibility.

The absolute heterogeneity parameter is τ2, the between-study variance on the squared analysis scale. Its square root, τ, is the between-study standard deviation on the same analysis scale as the effect estimate.3,5 For a log odds ratio, for example, τ is on the log-odds-ratio scale and should be interpreted or back-transformed accordingly. A point estimate of τ2 equal to zero does not prove that meaningful heterogeneity is absent; boundary estimates and sparse information can conceal substantial uncertainty.

Estimator choice matters because τ2 affects random-effects weights, interval estimates, and prediction. The DerSimonian–Laird method-of-moments estimator remains historically important and is still available in software.6 Comparative studies show that its bias and the coverage of associated normal-approximation intervals can be unsatisfactory in important settings, particularly with few studies, appreciable heterogeneity, or unequal precision.7 Restricted maximum likelihood (REML) and Paule–Mandel often perform better across a wider range of conventional random-effects settings, but no estimator is uniformly optimal; the outcome type, effect measure, sparsity, study count, and model structure remain relevant.5,7

One common plug-in form: μ̂ ± tk−2, 0.975 √[SE(μ̂)2 + τ̂2] This is one approximate construction under a conventional normal random-effects model. Software may use other degrees of freedom or methods. Ratio measures should be analysed on the appropriate logarithmic scale and then back-transformed.
Figure 2. The confidence interval concerns uncertainty about the mean underlying effect. A prediction interval concerns the underlying effect in a new study or setting judged exchangeable with those synthesized. It does not directly predict the noisy observed estimate from a future study, and its nominal coverage is not guaranteed in small or misspecified meta-analyses.8,9

Confidence and prediction intervals answer different questions. A confidence interval quantifies uncertainty about the estimated mean underlying effect. A prediction interval combines uncertainty in that mean with estimated between-study heterogeneity to describe the range in which the underlying effect of a new exchangeable study may lie under the fitted model.8,9 A narrow confidence interval favouring an intervention can coexist with a prediction interval that spans no effect or effects in the opposite direction. This means that the model permits materially different effects across comparable settings; it does not mean that the pooled mean is mathematically contradictory or that every future setting has an equal probability of benefit and harm.

Prediction intervals are model-dependent. Their interpretation assumes that the new setting is sufficiently exchangeable with the included studies and that the random-effects distribution and dependence structure are reasonably specified. Conventional plug-in intervals can have poor coverage with few studies, small heterogeneity, strongly unequal study precision, or departures from the assumed distribution.8 When these conditions are doubtful, the method, assumptions, and uncertainty should be stated rather than presenting the interval as a guaranteed population range.

Confidence intervals for the mean also depend on inferential choices. A conventional Wald-type interval typically uses a normal quantile and a plug-in standard error, treating the estimated heterogeneity as sufficiently well determined. Hartung–Knapp–Sidik–Jonkman-type procedures use a t-distribution and an adjusted variance estimate to improve small-sample inference in many settings, but they are not uniformly superior and can be unstable or unexpectedly narrow unless an appropriate modification is used.10 Authors should report the heterogeneity estimator, interval method, degrees-of-freedom or modification used, software, version, and non-default settings.

2. From describing variation to investigating its possible sources

Once variation is identified, subgroup analysis or meta-regression may investigate prespecified, clinically plausible sources. A statistically significant result in one subgroup and a non-significant result in another does not demonstrate a subgroup difference. The relevant evidence is a formal interaction or between-subgroup comparison, interpreted alongside the size and uncertainty of the difference, the number and distribution of studies, and the credibility of the hypothesis.1

Meta-regression relates study-level characteristics to effect estimates, but the comparison remains observational even when the primary studies are randomized. An association between mean participant age and treatment effect across trials does not establish that older individuals respond differently; study-level age can be confounded with intervention intensity, setting, risk of bias, follow-up, or other trial characteristics. This ecological limitation, low power, measurement error in moderators, multiple testing, and influential studies constrain causal interpretation.11 Prespecified analyses with strong prior rationale may be more credible than post-hoc searches, but they still require cautious language such as “associated with variation” rather than “caused the heterogeneity.”

The familiar suggestion of roughly ten studies per examined covariate is only a rough minimum-information rule, not a validity threshold. More studies may be required when moderators are unbalanced, correlated, measured imprecisely, or tested in multivariable models. Continuous moderators generally preserve more information than arbitrary categorization, but linear specifications can miss non-linear relationships and one or two high-leverage studies can dominate the fitted slope.

A scientific decision precedes all of these models: whether the studies estimate sufficiently compatible quantities for a pooled average to be meaningful. A random-effects model allows underlying effects to vary; it does not make incompatible estimands compatible and does not correct bias in the primary studies. When a common numerical summary would obscure important differences, synthesis without meta-analysis should use a structured, transparent approach rather than informal vote counting by statistical significance.1

META-ANALYSIS ROBUSTNESS RESOURCES

Interrogate heterogeneity, robustness and small-study effects

Use these resources to interpret heterogeneity without automatic thresholds, prespecify sensitivity analyses, test the stability of conclusions and assess possible small-study effects cautiously.

Working resources for heterogeneity, sensitivity analysis and reporting-bias assessment

Protocol

Apply structured worksheets, logs and checklists to integrate statistical output with clinical and methodological judgment, compare alternative analyses and preserve an auditable interpretation.

Meta-analysis heterogeneity interpretation worksheet

Record Q, I², tau-squared, confidence or prediction intervals, clinical diversity, methodological diversity and an evidence-based interpretation from validated software output.

Meta-analysis sensitivity analysis planning and results log

Prespecify alternative assumptions and exclusions, record sensitivity and leave-one-out outputs from validated software and document whether conclusions remain robust.

Funnel plot and small-study effects interpretation checklist

Record applicability conditions, study count, effect measure, visual features, statistical-test output, alternative explanations and cautious conclusions about asymmetry.

3. Sensitivity analysis: testing dependence on defensible assumptions

A sensitivity analysis evaluates whether the interpretation depends materially on plausible analytical choices or uncertain decisions. Relevant dimensions may include the effect measure, synthesis model, τ2 estimator, interval method, missing-data assumptions, treatment of dependent effects, eligibility boundaries, risk-of-bias restrictions, and handling of verified data problems.12 The alternatives must be scientifically defensible; trying many specifications and selectively reporting those that preserve or reverse a preferred conclusion is not a robustness assessment.

Primary specifications and foreseeable alternatives should be prespecified when possible. Unanticipated analyses can still be informative, but they should be labelled post hoc, justified, and reported alongside the prespecified analyses rather than silently replacing them. A complete record of analyses and amendments reduces selective analytical reporting and enables readers to distinguish planned robustness checks from data-responsive exploration.13,20

Leave-one-out analysis and influence diagnostics examine whether results depend disproportionately on particular studies. Studentized deleted residuals assess disagreement between a study and the fitted model; leverage measures identify observations positioned to affect the fit; DFBETAS quantify changes in individual model coefficients after deletion; Cook-type distances summarize broader changes in the fitted model; and Baujat plots display contributions to heterogeneity and influence on the summary estimate.14 These diagnostics answer different questions and need not agree.

Influence is not equivalent to error, bias, or ineligibility. A large, precise study may be influential because it contributes much of the available information. No study should be excluded solely because its removal changes a p-value, narrows heterogeneity, or moves the estimate in a preferred direction. Exclusion requires an independent rationale grounded in eligibility, estimand alignment, design, risk of bias, or verified data integrity. Leave-one-out analysis examines one-study deletion only and cannot identify every influential combination of studies.

Illustrative contour-enhanced funnel plot with observed studies and an asymmetric small-study pattern The vertical dashed line marks the no-effect value. The solid line marks the pooled estimate. Outer blue lines are approximate pseudo-confidence limits around the pooled estimate. Pale orange regions indicate approximate statistical significance relative to the null. An orange observed point is labelled as a small imprecise study, not a missing study. Observed small, imprecise study Potential influence or small-study pattern; the cause is not identified by the plot. No-effect value Pooled estimate Effect estimate Standard error (smaller at top) Approximate p < 0.05 contour Approximate pseudo-limits
Figure 3. Schematic contour-enhanced funnel plot, not empirical data and not to scale. The shaded regions represent approximate two-sided p<0.05 contours relative to the no-effect value under a normal approximation. The blue funnel lines are approximate pseudo-confidence limits around the pooled estimate and do not account for heterogeneity. An asymmetric pattern may indicate small-study effects, but publication bias is only one possible explanation.15,17

4. Small-study effects and the temptation of a single test

Small-study effects describe an association in which smaller or less precise studies tend to report different—often larger—effect estimates than larger, more precise studies. Selective publication can create this pattern, but so can selective non-reporting of outcomes or analyses, differences in risk of bias, clinical or methodological differences associated with study size, sparse data, outliers, chance, and mathematical relationships between an effect measure and its standard error.15,23 A funnel plot displays a pattern; it does not diagnose the mechanism.

Statistical tests of funnel-plot asymmetry are model- and effect-measure-specific. The original Egger regression can behave poorly for some binary effect measures because the effect estimate and its standard error are mathematically related. Harbord- and Peters-type tests were developed as alternatives for binary outcomes under particular conditions, but neither is universally valid. For diagnostic-accuracy meta-analysis, effective-sample-size methods such as the Deeks test are commonly used; standard intervention-meta-analysis tests should not be transferred without justification.15,16

The familiar “ten studies” rule is a rough safeguard, not a requirement that makes a test reliable once reached. With fewer than about ten studies, asymmetry tests usually have very low power; with ten or more, power and calibration may still be poor when study precision varies little, heterogeneity is substantial, sparse data are present, or one study dominates. A non-significant test cannot rule out publication bias, and a significant test cannot prove it.15 Reports should name the exact test, regression specification, predictor, effect scale, sidedness, software, and version.

Contour enhancement overlays regions of statistical significance and can help evaluate whether sparseness is concentrated in non-significant regions, which may be more compatible with selection related to statistical significance. This remains suggestive rather than conclusive: the visual pattern can still arise from heterogeneity, study-level bias, or other mechanisms.17

Trim-and-fill, selection or weight-function models, and p-value-based methods impose different assumptions about which results are unobserved and why. Trim-and-fill estimates what the pooled result would look like under a symmetry-restoring imputation mechanism; it does not recover a uniquely “corrected” or unbiased effect. Selection models are more explicit about selection mechanisms but can be weakly identified and sensitive to model specification. These methods are best treated as scenario-based sensitivity analyses presented alongside the observed synthesis, not as replacements for it.18

Direct evidence about missing results is often more informative than an asymmetry test alone. Relevant evidence includes registry entries without results, discrepancies between protocols and reports, outcomes or analyses known to have been measured but not reported, regulatory or sponsor records, and author correspondence. ROB-ME integrates such evidence with statistical signals to assess risk of bias due to missing evidence.19 An inception cohort or comprehensive audit may support a low-risk judgment despite asymmetry, while documented missing results may support concern even when a funnel plot appears symmetric.

Methodological sequence principle. Move from estimand alignment and primary-data checking to heterogeneity, robustness, small-study patterns, and direct missing-evidence assessment. No single diagnostic should override the design context or the documented evidence trail.

5. Working in order, so no stage erases another

  1. Confirm the estimand and compatibility. Establish that the studies address a sufficiently coherent question and that a pooled summary is interpretable.
  2. Specify the dependence structure. Account for multiple effects, outcomes, time points, or treatment arms from the same study before using conventional inverse-variance formulas.
  3. Describe variation. Report study estimates and intervals, Q, I2 with uncertainty where available, τ2/τ, and a justified prediction interval.
  4. Investigate plausible sources. Use prespecified subgroup analyses or meta-regression with formal interaction tests and non-causal interpretation.
  5. Test robustness. Compare defensible alternative assumptions while preserving the primary analysis and identifying post-hoc work.
  6. Assess influence. Combine leave-one-out results with residual, leverage, Cook-type, DFBETAS, and clinical or methodological review.
  7. Assess small-study patterns. Use plots and tests appropriate to the effect measure, recognizing low power and alternative explanations.
  8. Assess missing evidence directly. Integrate registries, protocols, reports, correspondence, and ROB-ME judgments before drawing conclusions.

Random-effects modelling does not correct bias in primary studies. Stability of a point estimate does not remove interval or model uncertainty. Agreement among several tests applied to the same data is not independent validation when the tests share assumptions. Sparse evidence cannot be repaired by increasingly complex models: with few studies, Q has low power, heterogeneity is poorly estimated, prediction intervals may have inadequate coverage, moderator analyses are weak, influence is concentrated, and asymmetry tests are largely uninformative.

6. Language the evidence can carry

PRISMA 2020 separates the reporting obligations clearly. Methods should describe procedures for investigating heterogeneity (Item 13e), sensitivity analyses (Item 13f), and risk of bias due to missing results (Item 14). Results should report heterogeneity investigations (Item 20b), all sensitivity analyses conducted (Item 20d), and assessments of reporting biases (Item 21). Amendments to registration or protocol information should be explained (Item 24c).20,25

Preferred reporting language. State what was observed and what the method can support: “The estimate varied under the prespecified missing-data scenario”; “the prediction interval crossed the no-effect value”; “a precision–effect association was detected under the specified regression”; or “direct evidence indicated that results were unavailable for registered outcomes.” Avoid claims such as “heterogeneity was acceptable”, “publication bias was absent”, “the outlier was removed because it was influential”, or “trim-and-fill corrected the effect.”

Methodological reporting should specify the effect measure and scale, synthesis model, τ2 estimator, confidence-interval and prediction-interval methods, handling of zero cells or sparse data, covariance or dependence assumptions, sensitivity scenarios, software functions, versions, and non-default arguments. Conventional Q, I2, inverse-variance models, and their intervals generally assume independent effect estimates or a correctly modelled covariance structure.

Multiple outcomes, time points, subscales, comparisons, or treatment arms from the same study create statistical dependence. Treating these estimates as independent can underestimate standard errors and inflate Type I error. Prespecified effect selection, multivariate or multilevel models, known or estimated covariance matrices, or robust variance estimation with appropriate small-sample corrections may be used depending on the estimand and data structure.21,22 Robust variance estimation does not eliminate the need for enough independent study clusters; with few clusters, inference remains fragile.

Effective triangulation combines sources that fail in different ways: quantitative estimates, visual diagnostics, risk-of-bias assessments, clinical and methodological characteristics, protocols, registries, primary reports, and direct missing-result evidence.23 Disagreement is informative. Statistical asymmetry without corresponding evidence of selective non-reporting may reflect heterogeneity or model artefact, while documented missing outcomes remain important even when a funnel plot looks symmetric.

7. Carrying the evidence into certainty

Certainty-of-evidence frameworks should not convert diagnostic outputs into automatic downgrading rules. A high I2, a prediction interval crossing the null, or a significant asymmetry test does not by itself determine the certainty rating. Inconsistency judgments require consideration of the direction and magnitude of study effects, overlap of intervals, clinically or decision-relevant thresholds, plausible explanations, and the population to which the conclusion applies.24

Imprecision and inconsistency should not be double-counted merely because the same wide or threshold-crossing intervals appear in both discussions. Risk of bias due to missing evidence should be informed by the structured missing-evidence assessment rather than by funnel asymmetry alone. The final conclusion should distinguish the estimated mean effect, variation across settings, sensitivity to assumptions, risk of missing-evidence bias, and the residual uncertainty that matters for decisions.

At a publication-ready standard, the objective is not to make every diagnostic agree. It is to show exactly which claims are supported, which depend on assumptions, which are contradicted by other evidence, and which remain unresolved.

Get the Resource Infrastructure Free Forever

Get Free Lifetime Access

Frequently asked questions

Is an I2 below 50% a reliable sign that pooling is appropriate?

No. I2 is a relative, precision-dependent estimate of the proportion of observed variability attributed to heterogeneity under a model; it is not an absolute measure of how far the underlying effects differ. Pooling depends first on estimand compatibility and scientific coherence, then on effect direction and magnitude, τ/τ2, uncertainty in heterogeneity, prediction, and the purpose of the synthesis. No fixed I2 cutoff can make that decision.1,3,4

What is the difference between a confidence interval and a prediction interval?

A confidence interval quantifies uncertainty about the estimated mean underlying effect. A prediction interval describes, under the fitted random-effects model, the range in which the underlying effect in a new exchangeable study or setting may lie. It does not directly predict the observed estimate of a future study, and its nominal coverage is not guaranteed when studies are few, precision is highly unequal, or model assumptions are poor.8,9

Does a random-effects model solve heterogeneity?

No. A random-effects model represents unexplained variation in underlying effects and changes the estimand, weights, and uncertainty. It does not explain heterogeneity, make incompatible studies compatible, or correct bias in the primary studies. The decision to pool still requires a coherent estimand and clinically and methodologically defensible synthesis.1

Can a non-significant test for funnel-plot asymmetry rule out publication bias?

No. Asymmetry tests usually have low power with few studies and can remain weak when study precision varies little or heterogeneity is substantial. A non-significant result is compatible with important selective non-reporting. Direct evidence from registries, protocols, reports, regulatory records, and author correspondence should be integrated through a structured missing-evidence assessment such as ROB-ME.15,19

Does a significant asymmetry test prove publication bias?

No. It indicates an association between effect estimates and a measure of study size or precision under a specified model. Publication bias is one possible cause, but clinical or methodological differences, risk-of-bias differences, sparse data, outliers, effect-measure coupling, and chance may produce the same pattern. The selected test must also be appropriate for the effect measure and data structure.15,16

When is it justified to remove an outlying or influential study?

Statistical influence alone is not a justification. An influential study may be the most precise and informative study in the synthesis. Exclusion requires an independent reason based on eligibility, estimand mismatch, design, risk of bias, or verified data integrity. The primary analysis and transparent sensitivity analysis should show how conclusions change, and no study should be removed merely because it changes statistical significance.14

Does trim-and-fill provide the true effect corrected for publication bias?

No. Trim-and-fill estimates what the pooled result would be under a specific symmetry-restoring imputation mechanism. Heterogeneity, outliers, an incorrect assumed selection direction, and other causes of asymmetry can make the result misleading. It should be reported as a sensitivity analysis under explicit assumptions, not as a corrected or unbiased estimate.18

Which between-study variance estimator should a random-effects meta-analysis use?

No estimator is universally best. REML and Paule–Mandel often perform better than DerSimonian–Laird across many conventional settings, but performance depends on study count, precision, heterogeneity, sparsity, effect measure, and model structure. The estimator should be justified and reported together with the confidence-interval method, prediction-interval method, software, version, and non-default options.58

Do different subgroup p-values prove that subgroup effects differ?

No. A significant effect in one subgroup and a non-significant effect in another does not establish an interaction. The subgroup effects must be compared directly using a formal interaction or between-subgroup test, and the result should be interpreted with its magnitude, uncertainty, prespecification, biological or clinical plausibility, and the number and distribution of studies.1,11

References

  1. Deeks JJ, Higgins JPT, Altman DG, McKenzie JE, Veroniki AA, et al. Chapter 10: Analysing data and undertaking meta-analyses. In: Cochrane Handbook for Systematic Reviews of Interventions. Current online version. Cochrane. Accessed 4 August 2026. Available from: Cochrane Handbook Chapter 10.
  2. Hardy RJ, Thompson SG. Detecting and describing heterogeneity in meta-analysis. Stat Med. 1998;17(8):841–856. doi:10.1002/(SICI)1097-0258(19980430)17:8<841::AID-SIM781>3.0.CO;2-D.
  3. Higgins JPT, Thompson SG. Quantifying heterogeneity in a meta-analysis. Stat Med. 2002;21(11):1539–1558. doi:10.1002/sim.1186.
  4. von Hippel PT. The heterogeneity statistic I2 can be biased in small meta-analyses. BMC Med Res Methodol. 2015;15:35. doi:10.1186/s12874-015-0024-z.
  5. Veroniki AA, Jackson D, Viechtbauer W, et al. Methods to estimate the between-study variance and its uncertainty in meta-analysis. Res Synth Methods. 2016;7(1):55–79. doi:10.1002/jrsm.1164.
  6. DerSimonian R, Laird N. Meta-analysis in clinical trials. Control Clin Trials. 1986;7(3):177–188. doi:10.1016/0197-2456(86)90046-2.
  7. Langan D, Higgins JPT, Jackson D, et al. A comparison of heterogeneity variance estimators in simulated random-effects meta-analyses. Res Synth Methods. 2019;10(1):83–98. doi:10.1002/jrsm.1316.
  8. Partlett C, Riley RD. Random effects meta-analysis: coverage performance of 95% confidence and prediction intervals following REML estimation. Stat Med. 2017;36(2):301–317. doi:10.1002/sim.7140.
  9. Riley RD, Higgins JPT, Deeks JJ. Interpretation of random effects meta-analyses. BMJ. 2011;342:d549. doi:10.1136/bmj.d549.
  10. Jackson D, Law M, Rücker G, Schwarzer G. The Hartung–Knapp modification for random-effects meta-analysis: a useful refinement but are there any residual concerns? Stat Med. 2017;36(25):3923–3934. doi:10.1002/sim.7411.
  11. Thompson SG, Higgins JPT. How should meta-regression analyses be undertaken and interpreted? Stat Med. 2002;21(11):1559–1573. doi:10.1002/sim.1187.
  12. Aung NM, Jurak I, Mehmood S, Axon E. Sensitivity analysis in meta-analysis: a tutorial. Cochrane Evid Synth Methods. 2026;4(1):e70067. doi:10.1002/cesm.70067.
  13. Page MJ, McKenzie JE, Kirkham J, Dwan K, Kramer S, Green S, Forbes A. Bias due to selective inclusion and reporting of outcomes and analyses in systematic reviews of randomised trials of healthcare interventions. Cochrane Database Syst Rev. 2014;(10):MR000035. doi:10.1002/14651858.MR000035.pub2.
  14. Viechtbauer W, Cheung MW-L. Outlier and influence diagnostics for meta-analysis. Res Synth Methods. 2010;1(2):112–125. doi:10.1002/jrsm.11.
  15. Sterne JAC, Sutton AJ, Ioannidis JPA, et al. Recommendations for examining and interpreting funnel plot asymmetry in meta-analyses of randomised controlled trials. BMJ. 2011;343:d4002. doi:10.1136/bmj.d4002.
  16. Deeks JJ, Macaskill P, Irwig L. The performance of tests of publication bias and other sample size effects in systematic reviews of diagnostic test accuracy was assessed. J Clin Epidemiol. 2005;58(9):882–893.
  17. Peters JL, Sutton AJ, Jones DR, Abrams KR, Rushton L. Contour-enhanced meta-analysis funnel plots help distinguish publication bias from other causes of asymmetry. J Clin Epidemiol. 2008;61(10):991–996. doi:10.1016/j.jclinepi.2007.11.010.
  18. Shi L, Lin L. The trim-and-fill method for publication bias: practical guidelines and recommendations based on a large database of meta-analyses. Medicine (Baltimore). 2019;98(23):e15987. doi:10.1097/MD.0000000000015987.
  19. Page MJ, Sterne JAC, Boutron I, et al. ROB-ME: a tool for assessing risk of bias due to missing evidence in systematic reviews with meta-analysis. BMJ. 2023;383:e076754. doi:10.1136/bmj-2023-076754.
  20. Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71. doi:10.1136/bmj.n71.
  21. Tanner-Smith EE, Tipton E, Polanin JR. Handling complex meta-analytic data structures using robust variance estimates: a tutorial in R. J Dev Life Course Criminol. 2016;2(1):85–112. doi:10.1007/s40865-016-0026-5.
  22. Pustejovsky JE, Tipton E. Meta-analysis with robust variance estimation: expanding the range of working models. Prev Sci. 2022;23(3):425–438. doi:10.1007/s11121-021-01246-3.
  23. Page MJ, Higgins JPT, Sterne JAC. Chapter 13: Assessing risk of bias due to missing evidence in a meta-analysis. In: Cochrane Handbook for Systematic Reviews of Interventions. Current online version. Cochrane. Accessed 4 August 2026. Available from: Cochrane Handbook Chapter 13.
  24. Schünemann HJ, Brożek J, Guyatt G, Oxman A, editors. GRADE Handbook: inconsistency. GRADE Working Group. Current online version. Accessed 4 August 2026. Available from: GRADE guidance on inconsistency.
  25. Page MJ, Moher D, Bossuyt PM, et al. PRISMA 2020 explanation and elaboration: updated guidance and exemplars for reporting systematic reviews. BMJ. 2021;372:n160. doi:10.1136/bmj.n160.