Meta-analysis Sensitivity Analysis Planning and Results Log

The problem: robustness claims without an auditable record

A sensitivity analysis repeats the primary synthesis under one or more defensible alternative decisions, assumptions, or values to determine whether the findings depend on choices that were uncertain or partly arbitrary. Cochrane describes sensitivity analysis as highly desirable and gives examples involving notable assumptions, imputed data, borderline eligibility decisions, studies at high risk of bias, and alternative analytical methods.1,2

The analysis itself is only part of the evidence. A defensible record should also show what was planned, what was actually run, which analyses were added after the results became known, and how each alternative changed the estimate, uncertainty, heterogeneity, or substantive interpretation. PRISMA 2020 separates these responsibilities across methods, results, and protocol-amendment reporting: Item 13f covers sensitivity-analysis methods, Item 20d requires the results of all sensitivity analyses conducted, and Item 24c requires amendments to registration or protocol information to be described and explained.4,5

Selective reporting is not a hypothetical concern. A 2014 methodological review pooling four studies and 485 Cochrane reviews estimated that 38% (95% confidence interval 23% to 54%) added, omitted, upgraded, or downgraded at least one outcome between protocol and publication. The authors cautioned that the underlying reviews predated 2009, so the estimate should not be treated as a current prevalence figure; it remains evidence for why deviations and analytical choices should be traceable.6

Plan before examining results

Define the primary analysis, plausible alternatives, the reason for each alternative, and any substantive decision rule.

Execute and preserve every run

Record model settings, outputs, warnings, dates, files, and whether each run was prespecified or post hoc.

Report results and deviations

Present all sensitivity results, explain protocol amendments, and distinguish numerical stability from decision stability.

The plan states what should happen; the execution log records what did happen; the report makes both visible.

Figure 1. A three-stage documentation workflow aligned with Cochrane sensitivity-analysis guidance and PRISMA 2020 reporting items 13f, 20d, and 24c.1,4,5
Core principle. Sensitivity analysis is not a search for the specification that produces the preferred result. It is a structured assessment of whether conclusions remain credible under plausible alternative decisions.

The planning and results log

Meta-analysis sensitivity analysis planning and results log

Use one copy for each outcome and primary synthesis. Complete the planning sections before examining sensitivity results, then record every executed run and its interpretation.

Local form: this code does not submit or transmit entries. Values remain only in the current browser page and may be lost if the page is refreshed or closed; print or save a PDF for the project record.
Phase 1 · Before sensitivity results are examined Review, outcome, estimand, and primary specification

Anchor the log to one outcome and one primary synthesis. Record the effect scale and analytical unit so that changes across runs remain interpretable.

Phase 1 · Before sensitivity results are examined Prespecification register: six practical decision domains

These six domains are a practical MetaSyn grouping of the decision examples discussed by Cochrane; they are not an official six-part Cochrane taxonomy. Mark a domain “not applicable” when it is genuinely irrelevant.1

Safeguard: For every planned alternative, specify the rationale, exact change, analysis unit, and expected output. Avoid vague statements such as “alternative methods will be explored if necessary.”
Phase 1 · Before sensitivity results are examined Interpretation rules and decision boundaries

There is no universal percentage change that defines robustness. Any tolerance must be justified for the selected effect scale and decision context.

Phase 2 · As analyses are executed Results log: every run, tagged and retained

Record all executed sensitivity runs, including non-confirming results and analyses not selected for the main manuscript. Preserve the corresponding script and output.

Analysis run 1
Phase 2 · As diagnostics are examined Leave-one-out and influence-diagnostics record

Leave-one-out analyses and influence diagnostics are diagnostic tools, not automatic exclusion rules. Deleted residuals, hat values, Cook’s distance, DFBETAS, and related measures capture different aspects of unusualness or influence; Baujat plots display contribution to heterogeneity against influence on the pooled result. Single-deletion diagnostics may also be affected when several outlying studies are present, which is one reason iterative procedures have been investigated.7,8,9

Phase 2 · After diagnostics Corrections, restrictions, and exclusions

Distinguish a correctable data or coding error from a genuine but unusual study result. If an eligible study is omitted from a sensitivity run, preserve the primary analysis and document the independent methodological rationale.7,10

Phase 2 · Synthesis of the log Robustness interpretation matrix

Separate numerical stability from changes in uncertainty and decision status. Do not define robustness only by whether a P value remains below 0.05.

Decision-boundary caution: A small change in the pooled estimate can still matter when an interval or a substantively defined threshold lies near the decision boundary. Report the numerical results rather than only a binary robustness label.
Phase 2 · Always visible Protocol and analysis-plan deviations

PRISMA 2020 Item 24c requires amendments to registration or protocol information to be described and explained; Item 20d requires results of all sensitivity analyses conducted.4,5

Verification and sign-off
Minimum audit standard: every reported sensitivity result traces to a retained run; every run is identified as prespecified or post hoc; model settings and software are reproducible; corrections and restrictions have documented rationales; and the manuscript reports all sensitivity analyses relevant to the conclusion.

How to interpret the completed log

A complete log supports a robustness claim only when the alternative analyses are scientifically plausible and the interpretation rule is explicit. Numerical stability, uncertainty stability, heterogeneity stability, and decision stability are related but not interchangeable. A result can move only slightly while its confidence interval changes its relation to the no-effect value.

Figure 2. A mathematically explicit hypothetical example. The pooled SMD changes by 0.06, but the alternative 95% confidence interval includes the null. The figure illustrates why the log separates estimate movement from interval and decision status.

Recent empirical context should be interpreted cautiously. A July 2026 preprint examining 358 behavioural-science meta-analyses reported that the median absolute change in Cohen’s d was no greater than 0.047 across the evaluated outlier-handling procedures, while at least one procedure-estimator combination changed statistical-significance status in 11.5% of meta-analyses and smallest-effect-of-interest classification in 15.9%. Because this evidence is a preprint from a particular corpus, it should not be used as a universal robustness threshold.12

Publication wording: replace “the results were robust” with a traceable statement naming the sensitivity analyses, reporting their estimates and intervals, and stating whether any prespecified substantive or statistical decision changed.

Frequently asked questions

These answers clarify the methodological limits of the log. They are designed to support transparent analysis and reporting, not to replace review-specific statistical judgement.

What is a sensitivity analysis in meta-analysis?

A sensitivity analysis repeats the primary synthesis under one or more scientifically defensible alternative decisions, assumptions, or values to assess whether the findings depend on choices that were uncertain or partly arbitrary. It estimates the same target under alternative specifications; it is therefore different from a subgroup analysis, which compares effects across defined groups.

See references 1 and 2.

Must every sensitivity analysis be prespecified?

Foreseeable sensitivity analyses should be prespecified whenever possible. Cochrane also recognizes that some relevant issues become apparent only during the review. Such analyses may still be informative, but they should be identified as post hoc, justified, dated, preserved in the analysis log, and reported without selecting results according to their direction or statistical significance.

See references 1, 3, 4, and 5.

Is leave-one-out analysis the same as sensitivity analysis?

Leave-one-out analysis is one form of deletion diagnostic and can be used within a sensitivity assessment, but it is not a complete sensitivity-analysis strategy. A full strategy may also examine eligibility decisions, risk-of-bias restrictions, missing-data assumptions, effect-size construction, dependence handling, model choice, heterogeneity estimation, and interval methods.

See references 1, 2, and 7.

Does an influential or outlying study justify exclusion?

No. A diagnostic flag is a reason to investigate the study, its data, and the fitted model; it is not an automatic deletion rule. Exclusion is defensible when supported by a documented error, an eligibility violation, or a prespecified methodological restriction. Otherwise, the study should generally remain in the primary synthesis and any omission should be presented as a clearly labelled sensitivity analysis.

See references 7, 9, and 10.

Is there a universal cutoff for Cook’s distance or DFBETAS in meta-analysis?

No cutoff is universally valid across meta-analytic models and datasets. Proposed thresholds are screening heuristics whose meaning depends on the model, the number and structure of effect estimates, leverage, and the pattern of the other diagnostics. Flagged cases should therefore be investigated substantively and evaluated through transparent re-analysis rather than deleted mechanically.

See references 7 and 11.

How should robustness be judged?

Robustness should be judged across the pooled estimate, confidence interval, prediction interval when appropriate, heterogeneity estimates, convergence or model warnings, and any prespecified substantive decision threshold. An unchanged p-value alone is insufficient, and a changed p-value does not by itself establish an important change in the estimated effect.

See references 1, 4, 5, and 12.

What does PRISMA 2020 require for sensitivity analyses?

PRISMA 2020 Item 13f asks authors to describe methods used to assess robustness through sensitivity analyses. Item 20d asks for the results of all sensitivity analyses conducted, and Item 24c asks authors to describe and explain amendments to registration or protocol information. These are reporting requirements; they do not prescribe one universal statistical sensitivity procedure.

See references 4 and 5.

Get every MetaSyn template free, including this one.

Leave your email and I’ll send this resource as an editable Word file and a printable PDF, plus access to the smart online version. You’ll also get every new template as it’s finished. No noise, just the resources.

References

  1. Deeks JJ, Higgins JPT, Altman DG, McKenzie JE, Veroniki AA, eds. Chapter 10: Analysing data and undertaking meta-analyses. In: Cochrane Handbook for Systematic Reviews of Interventions, current online version. Section 10.14, Sensitivity analyses. Chapter last updated November 2024. Cochrane Handbook Chapter 10.
  2. Aung NM, Jurak I, Mehmood S, Axon E. Sensitivity Analysis in Meta-Analysis: A Tutorial. Cochrane Evidence Synthesis and Methods. 2026;4(1):e70067. doi:10.1002/cesm.70067. The publisher later corrected the article category from “Methods Article” to “Tutorial”; the substantive article and DOI are unchanged.
  3. Andrade C. Types of Analysis: Planned (prespecified) vs Post Hoc, Primary vs Secondary, Hypothesis-driven vs Exploratory, Subgroup and Sensitivity, and Others. Indian Journal of Psychological Medicine. 2023;45(6):640–641. doi:10.1177/02537176231216842.
  4. Page MJ, McKenzie JE, Bossuyt PM, Boutron I, Hoffmann TC, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71. doi:10.1136/bmj.n71.
  5. Page MJ, Moher D, Bossuyt PM, Boutron I, Hoffmann TC, et al. PRISMA 2020 explanation and elaboration: updated guidance and exemplars for reporting systematic reviews. BMJ. 2021;372:n160. doi:10.1136/bmj.n160.
  6. Page MJ, McKenzie JE, Kirkham J, Dwan K, Kramer S, Green S, Forbes A. Bias due to selective inclusion and reporting of outcomes and analyses in systematic reviews of randomised trials of healthcare interventions. Cochrane Database of Systematic Reviews. 2014;(10):MR000035. doi:10.1002/14651858.MR000035.pub2.
  7. Viechtbauer W, Cheung MW-L. Outlier and influence diagnostics for meta-analysis. Research Synthesis Methods. 2010;1(2):112–125. doi:10.1002/jrsm.11.
  8. Baujat B, Mahé C, Pignon J-P, Hill C. A graphical method for exploring heterogeneity in meta-analyses: application to a meta-analysis of 65 trials. Statistics in Medicine. 2002;21(18):2641–2652. doi:10.1002/sim.1221.
  9. Meng Z, Wang J, Lin L, Wu C. Sensitivity analysis with iterative outlier detection for systematic reviews and meta-analyses. Statistics in Medicine. 2024;43(8):1549–1563. doi:10.1002/sim.10008.
  10. Aguinis H, Gottfredson RK, Joo H. Best-Practice Recommendations for Defining, Identifying, and Handling Outliers. Organizational Research Methods. 2013;16(2):270–301. doi:10.1177/1094428112470848. General outlier-handling guidance; not specific to meta-analysis.
  11. Viechtbauer W. Threshold values for Cook’s Distances and DFBETAS. R-sig-meta-analysis mailing list. September 30, 2022. Archived communication. Expert communication from the author of the metafor package; not peer reviewed.
  12. Havranek T, Irsova Z, Luskova M, Stanley TD. Do decisions about outliers and influential effects matter? Evidence from 358 behavioral science meta-analyses. arXiv preprint. 2026. arXiv:2607.23174. Pre-registered preprint; not peer reviewed at the August 4, 2026 evidence check. Findings are specific to the analysed behavioural-science corpus and procedures.