Funnel Plot and Small-study Effects Interpretation Checklist

The problem: funnel-plot asymmetry is a signal, not a diagnosis

Funnel plots are often reported as if they directly reveal publication bias. They do not. A funnel plot displays the relationship between study effect estimates and a measure of study size or precision. An asymmetric pattern is evidence of possible small-study effects: smaller or less precise studies tend to show different effects from larger or more precise studies. Selective non-publication is one possible explanation, but so are selective non-reporting of results, greater risk of bias in smaller studies, genuine clinical or methodological differences associated with study size, sparse-data behaviour, effect-measure artefacts, influential observations and chance.1,5,6

The opposite error is equally important. A visually symmetric plot or a non-significant asymmetry test does not establish that the evidence base is complete. Formal tests usually have low power when few studies are available and can remain uninformative when study standard errors vary little. The familiar minimum of ten studies is a rule of thumb below which tests are generally discouraged; it is not a threshold at which a test suddenly becomes reliable.1,5,6

Method choice also matters. The original Egger regression is not a universal test. Some effect measures, notably odds ratios and standardized mean differences, can be correlated with their standard errors, creating artefactual asymmetry. Alternative tests have been proposed for particular binary-outcome settings, while diagnostic-accuracy meta-analysis requires a design-specific approach such as the Deeks effective-sample-size test. None of these methods removes the need to inspect the data, consider heterogeneity, verify assumptions and seek direct evidence about missing results.14,8,9

Figure 1. The attribution problem. A plot can reveal a pattern, but it cannot identify the mechanism that generated it. Interpretation requires competing explanations and direct evidence to be assessed together.1,5

Statistical adjustment does not solve this attribution problem. Trim-and-fill, regression extrapolation, selection models and p-value-based methods estimate what might happen under particular assumptions about selection or small-study effects. Their performance varies across data-generating mechanisms, heterogeneity levels and study counts. They should therefore be presented as sensitivity analyses under explicit assumptions, not as procedures that recover a uniquely corrected or unbiased effect.10,11

Interpretation principle. Establish applicability before testing, describe the observed pattern before attributing it, choose a method appropriate to the effect measure and data structure, and integrate statistical signals with direct evidence about missing studies or results.
Scope. This checklist supports interpretation of funnel plots and small-study-effect analyses in pairwise meta-analysis. Test selection is outcome- and effect-measure-specific. ROB-ME is designed for pairwise syntheses of intervention effects; network meta-analysis requires ROB-MEN, and other synthesis designs may require additional specialist methods.12,13

The interpretation checklist

Funnel plot and small-study effects interpretation checklist

Document applicability, plot construction, visual features, test selection, direct missing-evidence records, sensitivity analyses and a conclusion that does not exceed the evidence.

Use one checklist for one outcome, time point, effect measure and synthesis. This HTML block does not submit or store entries.

Applicability first

1 Synthesis identification and data structure

Identify the exact synthesis being assessed. Do not combine different outcomes, time points, estimands or models in one funnel-plot interpretation.

Applicability first

2 Applicability gate

Formal asymmetry tests are generally discouraged when fewer than about ten independent studies are available or when study precisions are too similar to reveal a size-related trend. Ten studies is a pragmatic minimum, not a guarantee of adequate power or calibration.1,5

Interpretation boundary: a failed gate is a methodological finding, not a reason to run the test anyway. Report why asymmetry could not be evaluated reliably.
Construct and record

3 Funnel-plot construction

The axes, scale, reference line and pseudo-limits determine what the display means. Preserve the exact construction used.

Pseudo-limit caution: conventional triangular limits often represent expected sampling variation around a reference effect under simplified assumptions. They are not ordinary confidence intervals for individual studies and can be misleading when heterogeneity is substantial.1,5
Record before testing

4 Visual features and contour-enhanced reading

Describe the plot before viewing or interpreting a formal-test result. Visual assessment is subjective; independent review or consensus is preferable.

Contour limit: sparsity in non-significant regions can increase the plausibility of significance-related selection; sparsity in highly significant regions can reduce that plausibility. Neither pattern confirms that unobserved studies or results exist.7
Choose the method

5 Formal test selection and output

No single asymmetry test is appropriate for every effect measure. Record the exact test, regression specification and assumptions rather than writing only “Egger’s test”.

Effect-measure-specific considerations
Permitted inference: a test evaluates evidence of a specified association between effect estimates and study size or precision under a model. A significant result does not prove publication bias; a non-significant result does not exclude small-study effects or missing evidence.1,5,6
Weigh explanations

6 Competing-explanations matrix

For each mechanism, distinguish evidence that supports it, evidence that argues against it and information that remains unavailable.

Examine direct evidence

7 Missing-evidence record

Plots and tests examine observed estimates. Registries, protocols, reports, regulatory records and correspondence can provide direct evidence about studies or results that are unavailable for the synthesis.

Use the correct framework

8 Risk of bias due to missing evidence

Use the formal tool and its signalling questions when the synthesis is within scope. This checklist records the outcome; it does not replace completion of ROB-ME or ROB-MEN.

Overall judgment from the completed framework
Scope safeguard: ROB-ME is a structured assessment for risk of bias due to missing evidence in pairwise intervention-effect syntheses; ROB-MEN addresses network meta-analysis. Neither should be claimed as completed when only a funnel plot and asymmetry test have been reviewed.12,13
Sensitivity, not correction

9 Adjustment or sensitivity analyses

Record adjustment analyses only as scenarios under explicit assumptions. Preserve the observed synthesis as the primary empirical result unless the protocol specifies otherwise.

Language safeguard: do not call an adjusted estimate “corrected”, “unbiased” or “the true effect”. Describe it as the result of a specified sensitivity model or scenario.10,11

10 Final conclusion and sign-off

Separate the observed pattern, statistical evidence, direct evidence and formal risk-of-bias judgment. Use language proportionate to the evidence.

Completion standard: applicability is documented; the plot construction and test match the effect measure and data structure; alternative explanations and direct records are evaluated; any adjustment is labelled a sensitivity scenario; and the conclusion distinguishes small-study effects from risk of bias due to missing evidence.

From a visual pattern to an auditable judgment

A completed checklist turns a funnel plot from an impression into a traceable argument. The applicability gate records whether the evidence base can support visual or formal assessment. Plot construction is documented before interpretation. The formal method is tied to the effect measure and data structure. Alternative explanations are evaluated explicitly, and direct records of unavailable studies or results are placed alongside the statistical signal rather than treated as an afterthought.

Figure 2. Schematic contour interpretation, not empirical data and not to scale. Contours alter the relative plausibility of explanations; they do not reveal, count or reconstruct unobserved studies.7

Adjustment analyses belong after this assessment, not before it. Their value lies in showing how conclusions behave under explicit assumptions, not in replacing the observed synthesis. Likewise, the preferred final language is not “the funnel plot showed publication bias” or “publication bias was absent”. A defensible conclusion states whether a small-study pattern was observed, how applicable and informative the chosen test was, which mechanisms were supported or remained plausible, what direct evidence about missing results was identified, and what formal risk-of-bias judgment followed from the appropriate framework.

PRISMA 2020 asks review authors to describe methods for assessing risk of bias due to missing results and to report the resulting assessments. This checklist supports that documentation, but completing it does not by itself establish PRISMA compliance or replace the ROB-ME or ROB-MEN tools.14,15

Archival principle. Preserve the plot, exact software output, applicability decision, competing-explanations matrix, direct missing-evidence record, formal assessment and signed conclusion together. The audit trail is part of the evidence.

Related methodological articles and research tools

Methodological article

Funnel Plots and Egger’s Test: Interpretation, Limitations and Small-Study Effects

A detailed review of funnel-plot construction, effect-measure-specific asymmetry tests, statistical limitations, sensitivity analyses and risk-of-bias assessment.

Read the methodological article →

Methodological overview

Meta-analysis heterogeneity, sensitivity analysis and publication bias

A comprehensive overview of heterogeneity, sensitivity analysis, influence diagnostics, small-study effects and risk of bias due to missing evidence.

Read the methodological overview →

Research templates

Systematic review and meta-analysis templates

Evidence-informed worksheets, diagnostic logs, protocol-planning tools and reporting resources for systematic reviews and meta-analyses.

Browse the research templates →

Get every MetaSyn template free, including this one.

Leave your email and I’ll send this resource as an editable Word file and a printable PDF, plus access to the smart online version. You’ll also get every new template as it’s finished. No noise, just the resources.

References

  1. Sterne JAC, Sutton AJ, Ioannidis JPA, et al. Recommendations for examining and interpreting funnel plot asymmetry in meta-analyses of randomised controlled trials. BMJ. 2011;343:d4002. doi:10.1136/bmj.d4002.
  2. Egger M, Davey Smith G, Schneider M, Minder C. Bias in meta-analysis detected by a simple, graphical test. BMJ. 1997;315(7109):629–634. doi:10.1136/bmj.315.7109.629.
  3. Harbord RM, Egger M, Sterne JAC. A modified test for small-study effects in meta-analyses of controlled trials with binary endpoints. Stat Med. 2006;25(20):3443–3457. doi:10.1002/sim.2380.
  4. Peters JL, Sutton AJ, Jones DR, Abrams KR, Rushton L. Comparison of two methods to detect publication bias in meta-analysis. JAMA. 2006;295(6):676–680. doi:10.1001/jama.295.6.676.
  5. Page MJ, Higgins JPT, Sterne JAC. Chapter 13: Assessing risk of bias due to missing evidence in a meta-analysis. In: Cochrane Handbook for Systematic Reviews of Interventions. Current online version. Cochrane. Accessed 4 August 2026. Available from: Cochrane Handbook Chapter 13 .
  6. Sterne JAC, Gavaghan D, Egger M. Publication and related bias in meta-analysis: power of statistical tests and prevalence in the literature. J Clin Epidemiol. 2000;53(11):1119–1129. doi:10.1016/S0895-4356(00)00242-0.
  7. Peters JL, Sutton AJ, Jones DR, Abrams KR, Rushton L. Contour-enhanced meta-analysis funnel plots help distinguish publication bias from other causes of asymmetry. J Clin Epidemiol. 2008;61(10):991–996. doi:10.1016/j.jclinepi.2007.11.010.
  8. Deeks JJ, Macaskill P, Irwig L. The performance of tests of publication bias and other sample size effects in systematic reviews of diagnostic test accuracy was assessed. J Clin Epidemiol. 2005;58(9):882–893. doi:10.1016/j.jclinepi.2005.01.016.
  9. Zwetsloot PP, van der Naald M, Sena ES, et al. Standardized mean differences cause funnel plot distortion in publication bias assessments. eLife. 2017;6:e24260. doi:10.7554/eLife.24260.
  10. Shi L, Lin L. The trim-and-fill method for publication bias: practical guidelines and recommendations based on a large database of meta-analyses. Medicine (Baltimore). 2019;98(23):e15987. doi:10.1097/MD.0000000000015987.
  11. Carter EC, Schönbrodt FD, Gervais WM, Hilgard J. Correcting for bias in psychology: a comparison of meta-analytic methods. Adv Methods Pract Psychol Sci. 2019;2(2):115–144. doi:10.1177/2515245919847196.
  12. Page MJ, Sterne JAC, Boutron I, et al. ROB-ME: a tool for assessing risk of bias due to missing evidence in systematic reviews with meta-analysis. BMJ. 2023;383:e076754. doi:10.1136/bmj-2023-076754.
  13. Chiocchia V, Nikolakopoulou A, Higgins JPT, et al. ROB-MEN: a tool to assess risk of bias due to missing evidence in network meta-analysis. BMC Med. 2021;19:304. doi:10.1186/s12916-021-02166-3.
  14. Page MJ, McKenzie JE, Bossuyt PM, Boutron I, Hoffmann TC, Mulrow CD, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71. doi:10.1136/bmj.n71.
  15. Page MJ, Moher D, Bossuyt PM, Boutron I, Hoffmann TC, Mulrow CD, et al. PRISMA 2020 explanation and elaboration: updated guidance and exemplars for reporting systematic reviews. BMJ. 2021;372:n160. doi:10.1136/bmj.n160.