META-ANALYSIS SENSITIVITY ANALYSIS PLANNING AND RESULTS RESOURCE
Meta-analysis Sensitivity Analysis Planning and Results Log
Pre-specify alternative assumptions and exclusions, record sensitivity and leave-one-out outputs from validated software, and document whether conclusions remain robust.
The problem: robustness claims without an auditable record
A sensitivity analysis repeats the primary synthesis under one or more defensible alternative decisions, assumptions, or values to determine whether the findings depend on choices that were uncertain or partly arbitrary. Cochrane describes sensitivity analysis as highly desirable and gives examples involving notable assumptions, imputed data, borderline eligibility decisions, studies at high risk of bias, and alternative analytical methods.1,2
The analysis itself is only part of the evidence. A defensible record should also show what was planned, what was actually run, which analyses were added after the results became known, and how each alternative changed the estimate, uncertainty, heterogeneity, or substantive interpretation. PRISMA 2020 separates these responsibilities across methods, results, and protocol-amendment reporting: Item 13f covers sensitivity-analysis methods, Item 20d requires the results of all sensitivity analyses conducted, and Item 24c requires amendments to registration or protocol information to be described and explained.4,5
Selective reporting is not a hypothetical concern. A 2014 methodological review pooling four studies and 485 Cochrane reviews estimated that 38% (95% confidence interval 23% to 54%) added, omitted, upgraded, or downgraded at least one outcome between protocol and publication. The authors cautioned that the underlying reviews predated 2009, so the estimate should not be treated as a current prevalence figure; it remains evidence for why deviations and analytical choices should be traceable.6
Plan before examining results
Define the primary analysis, plausible alternatives, the reason for each alternative, and any substantive decision rule.
Execute and preserve every run
Record model settings, outputs, warnings, dates, files, and whether each run was prespecified or post hoc.
Report results and deviations
Present all sensitivity results, explain protocol amendments, and distinguish numerical stability from decision stability.
The plan states what should happen; the execution log records what did happen; the report makes both visible.
The planning and results log
Meta-analysis sensitivity analysis planning and results log
Use one copy for each outcome and primary synthesis. Complete the planning sections before examining sensitivity results, then record every executed run and its interpretation.
How to interpret the completed log
A complete log supports a robustness claim only when the alternative analyses are scientifically plausible and the interpretation rule is explicit. Numerical stability, uncertainty stability, heterogeneity stability, and decision stability are related but not interchangeable. A result can move only slightly while its confidence interval changes its relation to the no-effect value.
Hypothetical example on an SMD scale: a small estimate shift can change the interval’s relation to the null
Recent empirical context should be interpreted cautiously. A July 2026 preprint examining 358 behavioural-science meta-analyses reported that the median absolute change in Cohen’s d was no greater than 0.047 across the evaluated outlier-handling procedures, while at least one procedure-estimator combination changed statistical-significance status in 11.5% of meta-analyses and smallest-effect-of-interest classification in 15.9%. Because this evidence is a preprint from a particular corpus, it should not be used as a universal robustness threshold.12
Frequently asked questions
These answers clarify the methodological limits of the log. They are designed to support transparent analysis and reporting, not to replace review-specific statistical judgement.
What is a sensitivity analysis in meta-analysis?
A sensitivity analysis repeats the primary synthesis under one or more scientifically defensible alternative decisions, assumptions, or values to assess whether the findings depend on choices that were uncertain or partly arbitrary. It estimates the same target under alternative specifications; it is therefore different from a subgroup analysis, which compares effects across defined groups.
Must every sensitivity analysis be prespecified?
Foreseeable sensitivity analyses should be prespecified whenever possible. Cochrane also recognizes that some relevant issues become apparent only during the review. Such analyses may still be informative, but they should be identified as post hoc, justified, dated, preserved in the analysis log, and reported without selecting results according to their direction or statistical significance.
Is leave-one-out analysis the same as sensitivity analysis?
Leave-one-out analysis is one form of deletion diagnostic and can be used within a sensitivity assessment, but it is not a complete sensitivity-analysis strategy. A full strategy may also examine eligibility decisions, risk-of-bias restrictions, missing-data assumptions, effect-size construction, dependence handling, model choice, heterogeneity estimation, and interval methods.
Does an influential or outlying study justify exclusion?
No. A diagnostic flag is a reason to investigate the study, its data, and the fitted model; it is not an automatic deletion rule. Exclusion is defensible when supported by a documented error, an eligibility violation, or a prespecified methodological restriction. Otherwise, the study should generally remain in the primary synthesis and any omission should be presented as a clearly labelled sensitivity analysis.
Is there a universal cutoff for Cook’s distance or DFBETAS in meta-analysis?
No cutoff is universally valid across meta-analytic models and datasets. Proposed thresholds are screening heuristics whose meaning depends on the model, the number and structure of effect estimates, leverage, and the pattern of the other diagnostics. Flagged cases should therefore be investigated substantively and evaluated through transparent re-analysis rather than deleted mechanically.
How should robustness be judged?
Robustness should be judged across the pooled estimate, confidence interval, prediction interval when appropriate, heterogeneity estimates, convergence or model warnings, and any prespecified substantive decision threshold. An unchanged p-value alone is insufficient, and a changed p-value does not by itself establish an important change in the estimated effect.
What does PRISMA 2020 require for sensitivity analyses?
PRISMA 2020 Item 13f asks authors to describe methods used to assess robustness through sensitivity analyses. Item 20d asks for the results of all sensitivity analyses conducted, and Item 24c asks authors to describe and explain amendments to registration or protocol information. These are reporting requirements; they do not prescribe one universal statistical sensitivity procedure.
Get every MetaSyn template free, including this one.
Leave your email and I’ll send this resource as an editable Word file and a printable PDF, plus access to the smart online version. You’ll also get every new template as it’s finished. No noise, just the resources.
References
- Deeks JJ, Higgins JPT, Altman DG, McKenzie JE, Veroniki AA, eds. Chapter 10: Analysing data and undertaking meta-analyses. In: Cochrane Handbook for Systematic Reviews of Interventions, current online version. Section 10.14, Sensitivity analyses. Chapter last updated November 2024. Cochrane Handbook Chapter 10.
- Aung NM, Jurak I, Mehmood S, Axon E. Sensitivity Analysis in Meta-Analysis: A Tutorial. Cochrane Evidence Synthesis and Methods. 2026;4(1):e70067. doi:10.1002/cesm.70067. The publisher later corrected the article category from “Methods Article” to “Tutorial”; the substantive article and DOI are unchanged.
- Andrade C. Types of Analysis: Planned (prespecified) vs Post Hoc, Primary vs Secondary, Hypothesis-driven vs Exploratory, Subgroup and Sensitivity, and Others. Indian Journal of Psychological Medicine. 2023;45(6):640–641. doi:10.1177/02537176231216842.
- Page MJ, McKenzie JE, Bossuyt PM, Boutron I, Hoffmann TC, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71. doi:10.1136/bmj.n71.
- Page MJ, Moher D, Bossuyt PM, Boutron I, Hoffmann TC, et al. PRISMA 2020 explanation and elaboration: updated guidance and exemplars for reporting systematic reviews. BMJ. 2021;372:n160. doi:10.1136/bmj.n160.
- Page MJ, McKenzie JE, Kirkham J, Dwan K, Kramer S, Green S, Forbes A. Bias due to selective inclusion and reporting of outcomes and analyses in systematic reviews of randomised trials of healthcare interventions. Cochrane Database of Systematic Reviews. 2014;(10):MR000035. doi:10.1002/14651858.MR000035.pub2.
- Viechtbauer W, Cheung MW-L. Outlier and influence diagnostics for meta-analysis. Research Synthesis Methods. 2010;1(2):112–125. doi:10.1002/jrsm.11.
- Baujat B, Mahé C, Pignon J-P, Hill C. A graphical method for exploring heterogeneity in meta-analyses: application to a meta-analysis of 65 trials. Statistics in Medicine. 2002;21(18):2641–2652. doi:10.1002/sim.1221.
- Meng Z, Wang J, Lin L, Wu C. Sensitivity analysis with iterative outlier detection for systematic reviews and meta-analyses. Statistics in Medicine. 2024;43(8):1549–1563. doi:10.1002/sim.10008.
- Aguinis H, Gottfredson RK, Joo H. Best-Practice Recommendations for Defining, Identifying, and Handling Outliers. Organizational Research Methods. 2013;16(2):270–301. doi:10.1177/1094428112470848. General outlier-handling guidance; not specific to meta-analysis.
- Viechtbauer W. Threshold values for Cook’s Distances and DFBETAS. R-sig-meta-analysis mailing list. September 30, 2022. Archived communication. Expert communication from the author of the metafor package; not peer reviewed.
- Havranek T, Irsova Z, Luskova M, Stanley TD. Do decisions about outliers and influential effects matter? Evidence from 358 behavioral science meta-analyses. arXiv preprint. 2026. arXiv:2607.23174. Pre-registered preprint; not peer reviewed at the August 4, 2026 evidence check. Findings are specific to the analysed behavioural-science corpus and procedures.