Meta-analysis Model Selection and Reporting Checklist

META-ANALYSIS MODEL SELECTION AND REPORTING RESOURCE

Document the inferential target, common-effect or random-effects assumptions, estimator, interval method, heterogeneity handling, sensitivity plans, software, version and reporting decisions.

By Dr. Esmaeel Saeedy Robat, Founder of MetaSyn Academy · Published meta-analyst (Nature Human Behaviour, 2026) Evidence checked: August 2026

The familiar question, “fixed effect or random effects?”, compresses several decisions that answer different methodological questions. A defensible synthesis separates the target of inference from the statistical model, the estimator of between-study variance, the confidence-interval method, the decision to present a prediction interval, and the software implementation. This checklist provides a structured record for those decisions before results are interpreted.

Use boundary. This resource documents a human methodological decision. It does not select a model or estimator, calculate an effect, diagnose heterogeneity, or validate software output. Complete it with the protocol, statistical analysis plan, data dictionary and output from validated statistical software.

Why the model label is only one layer

Model labels are often treated as settings that can be chosen after viewing a heterogeneity statistic. That practice reverses the correct order of reasoning. The first question is what quantity the review intends to estimate. If the included studies are regarded as estimating one common underlying effect, subject only to sampling error, a common-effect model expresses that target. If the review instead treats study effects as different but related values sampled from an assumed distribution, a random-effects model targets the mean of that distribution. These are interpretations, not merely alternative formulas.1,2

The target must be stated in relation to the actual studies and populations. A pooled mean from a random-effects model is not automatically an effect that applies to every setting. Equally, a common-effect result does not become invalid merely because observed estimates differ; sampling error produces variation even under a common-effect assumption. The question is whether the assumption and target are credible enough for the intended inference, and whether remaining diversity is addressed transparently.1,4

Begin with the synthesis objective and estimand

An estimand specifies the quantity the analysis is designed to estimate. It includes the outcome, contrast, population, time point, effect measure and handling of intercurrent events or competing outcomes where relevant. The model decision should be connected to that estimand. For example, an analysis seeking the average intervention effect across meaningfully diverse settings has a different inferential target from an analysis estimating a common pharmacological effect under tightly standardized conditions.1

The synthesis objective should also define the scope of generalization. A random-effects mean estimated from a small, selectively assembled set of studies does not license inference to every possible future study. The assumed distribution describes the effects represented by the synthesis under its sampling and exchangeability assumptions. Those assumptions deserve a written statement rather than an unexplained software selection.2

Common-effect assumptions

Under a common-effect interpretation, each included study estimates the same underlying effect. Differences among observed estimates are attributed to within-study sampling variation. Inverse-variance, Mantel-Haenszel and other methods may implement a common-effect calculation for different data structures. The method and the inferential interpretation should both be reported because identical numerical machinery can support different targets in other fixed-effects formulations.4

The assumption is strongest when the intervention, comparator, outcome definition, follow-up, design and target population are sufficiently aligned. It is not proven by a non-significant heterogeneity test. Such tests often have low power when few studies are available, while trivial differences can become detectable with large evidence bases. A common-effect plan therefore needs a substantive justification and a statement of how unexplained diversity will affect interpretation.1

Random-effects assumptions

A conventional random-effects model assumes that underlying effects vary across studies according to a distribution, commonly taken to be normal on the analysis scale. The pooled estimate describes the mean of that distribution, and the between-study variance quantifies its estimated width. This framework can represent unexplained heterogeneity, but it does not make heterogeneity disappear and it cannot correct bias that differs across studies.2

The distributional assumption is difficult to verify, particularly with few studies. If small studies are systematically different from larger studies, random-effects weighting can give them relatively more influence than a common-effect analysis. That is not a reason to prohibit random-effects models; it is a reason to prespecify checks for asymmetry, influential studies and alternative specifications, and to interpret the mean together with heterogeneity and uncertainty.2,3

Analysis-plan layer stack Six horizontal layers show estimand, model, tau-squared estimator, interval method, software implementation and reporting record. 1. Estimand and population of inference 2. Common-effect or distribution-of-effects model 3. Between-study variance estimator (τ²) and interval 4. Summary-effect and prediction intervals 5. Software, version and settings 6. Output, deviations and report
Figure 1. The model name is one layer in a connected analysis plan. Methodological verification runs across all layers.

Prose equivalent: define the estimand first; state the model and its assumptions; specify how between-study variance and its uncertainty will be estimated; specify summary and prediction intervals; record the software implementation; then preserve outputs, deviations and reporting language.

Between-study variance is an estimated parameter

Random-effects implementation requires an estimate of between-study variance, written as τ2. DerSimonian-Laird, Restricted Maximum Likelihood (REML) and Paule-Mandel are among the available estimators. Their performance varies with the number and size of studies, the effect measure, the true amount of heterogeneity and other features. Reviews of the evidence do not support a universal winner across every setting.3

REML and Paule-Mandel often have attractive statistical properties, while DerSimonian-Laird remains common and can be computationally simple. The responsible reporting choice is to name the estimator, explain why it fits the analysis, and prespecify sensitivity analyses when plausible alternatives could materially affect the result. It is misleading either to leave the estimator hidden as a software default or to declare one historical method categorically forbidden.1,3

Uncertainty around τ2

A point estimate of τ2 can look more certain than it is. With a small evidence base, estimates may be unstable and can equal zero even when important uncertainty remains. Profile-likelihood or Q-profile procedures can provide intervals for the heterogeneity variance in compatible settings. Reporting such an interval can help readers see whether the apparent precision of the heterogeneity estimate is justified.3

The choice of τ2 interval method should be documented separately from the point estimator. A review may use REML for the point estimate and Q-profile for an interval, depending on software support and analysis conditions. The checklist records both fields so that a result cannot be described vaguely as “random effects” without its defining implementation choices.1,3

Summary-effect confidence intervals

The conventional Wald-type interval treats the estimated standard error of the pooled mean in a familiar large-sample way. Hartung-Knapp-Sidik-Jonkman (HKSJ) procedures alter the uncertainty calculation and can improve coverage in some random-effects settings, especially when the number of studies is limited. However, performance is not uniformly superior under every configuration, and modified procedures exist to address particular counterintuitive cases.5

The interval method therefore needs an explicit rationale. Record whether HKSJ is used, which implementation or modification is applied, and what will happen in edge cases such as very few studies, a near-zero heterogeneity estimate or an interval unexpectedly narrower than the conventional alternative. “HKSJ is mandatory” is too broad; “HKSJ was prespecified because these conditions and this software implementation were judged appropriate” is auditable.5

Prediction intervals answer a different question

A confidence interval around a random-effects mean concerns uncertainty about the location of the mean effect. A prediction interval attempts to summarize the range in which an underlying effect for a similar setting might lie, accounting for estimated between-study variation and uncertainty. A narrow confidence interval can coexist with a wide prediction interval because the mean may be estimated precisely while effects vary substantially.1,2

Prediction intervals depend strongly on assumptions, including the distribution of effects, and can behave poorly when few studies inform τ2. Their relevance also depends on whether a “new” setting is sufficiently similar to those represented. The plan should state whether a prediction interval will be reported, which method and scale will be used, the minimum information considered adequate, and how limitations will be explained. Absence of a prediction interval should likewise have a reason rather than reflecting an unchecked default.1

Decision audit from protocol to report A left-to-right path shows protocol target, prespecified methods, validated software output, deviation log and final report. Protocol target Method record Validated software Deviation log Transparent report Every change remains traceable to an analysis decision.
Figure 2. A model decision is complete only when implementation, deviations and reporting remain traceable across all synthesis steps.

Prose equivalent: connect the protocol target to a prespecified method record, implement it in validated software, preserve any deviation with its reason and approval, and report the actual model, estimator, interval methods and software rather than only the intended plan.

Heterogeneity handling and sensitivity planning

Heterogeneity statistics describe aspects of observed inconsistency; they do not choose the model. I2 is the percentage of variability in effect estimates attributable to heterogeneity rather than sampling error under its definition, but its uncertainty can be large and the familiar thresholds are rough conventions. Clinical meaning depends on effect magnitude, direction, outcome and context, not on a threshold alone.1

The plan should identify plausible influential decisions before results are known. Examples include alternative eligible effect measures, alternative τ2 estimators, conventional versus HKSJ intervals, exclusion of a study with a documented data problem, and assumptions for clustered or multi-arm designs. Sensitivity analysis tests robustness to decisions; it should not become a search for a preferred answer. Detailed diagnosis of heterogeneity, subgroup analysis, meta-regression and small-study effects belongs to the separate H07 pathway.

Software is part of the method

Software names alone are insufficient because defaults and available methods change. Record the program, version, package or module, analysis command or settings, estimator, confidence level, transformations, prediction-interval method and output location. RevMan added REML, Q-profile intervals, HKSJ options and prediction intervals in January 2025, illustrating why version-level reporting matters.1

Preserve the analysis file, console output or structured result, not only a copied forest plot. The model checklist should be reconciled against the final output. If the software substituted a default, applied a continuity correction, failed to calculate an interval or handled a zero-heterogeneity case differently from the plan, record the discrepancy and its disposition.1,6

Transparent reporting

Readers need enough information to reproduce the analysis and understand its target. Report the effect measure and analysis scale; common-effect or random-effects interpretation; weighting method; τ2 estimator and any interval; summary-effect interval method; prediction interval method and conditions; heterogeneity statistics; software and version; prespecified sensitivity analyses; and material deviations. PRISMA 2020 supplies the broader reporting context but is not a conduct or quality-assessment tool.6

Reporting should distinguish what was planned from what was done. A deviation is not automatically misconduct or invalidity; unforeseen data can require an adaptation. The problem is an invisible change. Record the trigger, alternative considered, decision maker, date, effect on interpretation and where both original and revised outputs are stored.6

Small evidence bases require explicit caution

When only a few studies are available, several quantities are weakly identified. A heterogeneity test may have little power, τ2 can be estimated as zero despite substantial uncertainty, and normal approximations for a pooled mean may have poor coverage. A random-effects model does not create information that the evidence set lacks. The plan should state how small-study configurations will be recognized and which outputs will be interpreted cautiously.1,3

Hartung-Knapp procedures, alternative τ2 estimators and profile-based intervals can improve particular operating characteristics, but no single option removes all small-sample problems. Some implementations can produce counterintuitive results under special configurations. Record the exact method, any modification and the response to edge cases before results are inspected. If the evidence is too sparse for a stable prediction interval or moderator analysis, state that limitation instead of forcing an estimate.5

Small studies also deserve attention because their estimates can be imprecise and their relative weight may increase under random effects. That observation does not prove reporting bias or justify exclusion. It identifies a sensitivity and interpretation issue. Assessment of small-study effects and publication bias belongs to H07 and requires evidence beyond the weighting pattern.

Model choice does not decide whether pooling is appropriate

Both common-effect and random-effects calculations presuppose that the studies address a sufficiently coherent synthesis question. Statistical accommodation of variation is not a licence to combine incompatible outcomes, interventions, populations or designs. Before completing the model fields, verify the synthesis groups against the protocol and explain why their effects form a meaningful target.1

If study effects answer different questions, the appropriate response may be separate syntheses, structured presentation without pooling, or a redesigned estimand. A mean of incompatible quantities can be mathematically computable and scientifically uninterpretable. The checklist therefore records the population of inference, exchangeability conditions and scope of generalization rather than asking only which software option will be used.1,2

The same principle applies to dependence. Multiple outcomes, time points, intervention arms or effect estimates from one study can be correlated. Treating them as independent can overstate precision. The analysis plan must identify the statistical unit and specify an appropriate multilevel, multivariate, robust-variance or selection strategy when necessary. These choices require qualified review and cannot be inferred from the forest plot.1

Predefine what would count as a robust conclusion

A sensitivity analysis is informative when it targets a credible decision whose uncertainty could change the scientific message. Examples include alternative defensible τ2 estimators, interval procedures, data-correlation assumptions, or treatment of a study with a documented extraction problem. State the expected comparison and interpretation rule in advance. Avoid a large menu of analyses with no hierarchy.6

Robustness is not agreement of point estimates alone. Compare interval width, direction relative to meaningful thresholds, prediction, study influence and the wording of the conclusion. If one defensible analysis supports benefit while another includes important harm, report the disagreement and explain which assumptions differ. Do not select the more favourable specification as primary after viewing results.6

A conclusion can also be stable numerically but fragile scientifically. Similar pooled estimates under several models do not resolve high risk of bias, indirectness or selective reporting. The model record should connect only to the inferential claims it governs, leaving evidence certainty to the appropriate assessment process.

Smart solution: planning and audit checklist

Complete the record before running the primary synthesis, then reconcile it against the final software output. Entries remain in this browser only. Nothing is transmitted.

1. Review and synthesis objective

2. Model and assumptions

3. Between-study variance (τ²)

4. Summary and prediction intervals

5. Heterogeneity and sensitivity boundary

6. Software, output and reporting audit

How to use the completed record

First, compare the checklist with the registered protocol or dated analysis plan. Second, run the prespecified analysis in validated software and preserve the complete output. Third, reconcile every method field against the output, including defaults that may not be obvious in a forest plot. Fourth, document deviations before interpreting whether they changed the result. Finally, transfer the verified method description into the review report and use the paired guide, Fixed-Effect vs Random-Effects Meta-Analysis: Assumptions, Models and Reporting, when the reasoning requires explanation.6

Conclusion

A model decision is defensible when its estimand, assumptions, variance estimator, interval methods, software implementation and reporting are mutually coherent. No heterogeneity threshold can supply that chain of reasoning. Run the prespecified analysis in validated software, preserve outputs and deviations, then quality-check the forest plot before publication.1,6

Get every MetaSyn template free, including this one.

Leave your email and I’ll send this resource as an editable Word file and a printable PDF, plus access to the smart online version. You’ll also get every new template as it’s finished. No noise, just the resources.

Evidence note. This resource prioritizes current Cochrane guidance, PRISMA 2020, and verified statistical methodology publications. Model selection remains conditional on the primary review question, study designs, target estimand, and underlying data properties. No single model or estimator applies universally to all quantitative syntheses.

References and evidence scope

  1. Deeks JJ, Higgins JPT, Altman DG, McKenzie JE, Veroniki AA. Chapter 10: Analysing data and undertaking meta-analyses. In: Higgins JPT, Thomas J, Chandler J, Cumpston M, Li T, Page MJ, et al, editors. Cochrane Handbook for Systematic Reviews of Interventions. Version 6.5. Cochrane; 2024. https://www.cochrane.org/authors/handbooks-and-manuals/handbook/current/chapter-10
  2. Borenstein M, Hedges LV, Higgins JPT, Rothstein HR. A basic introduction to fixed-effect and random-effects models for meta-analysis. Research Synthesis Methods. 2010;1(2):97-111. https://doi.org/10.1002/jrsm.12
  3. Veroniki AA, Jackson D, Viechtbauer W, Bender R, Bowden J, Knapp G, et al. Methods to estimate the between-study variance and its uncertainty in meta-analysis. Research Synthesis Methods. 2016;7(1):55-79. https://doi.org/10.1002/jrsm.1164
  4. IntHout J, Ioannidis JPA, Borm GF. The Hartung-Knapp-Sidik-Jonkman method for random effects meta-analysis is straightforward and considerably outperforms the standard DerSimonian-Laird method. BMC Medical Research Methodology. 2014;14:25. https://doi.org/10.1186/1471-2288-14-25
  5. Higgins JPT, Thompson SG, Spiegelhalter DJ. A re-evaluation of random-effects meta-analysis. Journal of the Royal Statistical Society: Series A (Statistics in Society). 2009;172(1):137-159. https://doi.org/10.1111/j.1467-985X.2008.00552.x
  6. Rice K, Higgins JPT, Lumley T. A re-evaluation of fixed effect(s) meta-analysis. Journal of the Royal Statistical Society: Series A (Statistics in Society). 2018;181(1):205-227. https://doi.org/10.1111/rssa.12275
  7. Page MJ, McKenzie JE, Bossuyt PM, Boutron I, Hoffmann TC, Mulrow CD, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71. https://doi.org/10.1136/bmj.n71
  8. Cochrane. RevMan knowledge base: release notes. Cochrane; 2025. https://documentation.cochrane.org/revman-kb/revman-release-notes-342261851.html

Scope and use boundary

The checklist creates an auditable record of a proposed model decision. It does not calculate an estimate, determine that studies are compatible, repair missing or dependent data, select a statistical model, or certify an analysis. Complex sparse-data, clustered, crossover, multi-arm, non-randomized, and emerging-method decisions require qualified methodological and statistical review.

Questions this checklist is designed to prevent

Should the model be chosen from a heterogeneity test or an I2 threshold?

No. The inferential target and credible assumptions should determine the model. A non-significant heterogeneity test does not establish one common underlying effect, and an observed I2 value is not a model-selection rule. Record heterogeneity statistics as descriptive outputs and explain the model independently.

Does random-effects analysis automatically solve heterogeneity?

No. It represents between-study variation through a statistical distribution and changes the weighting structure. It does not explain why effects vary, correct bias, or make clinically incompatible studies exchangeable. Investigating heterogeneity, subgroup analysis, and meta-regression belong to the next analytical stage.

Is DerSimonian-Laird always the default estimator?

No universal estimator is best in every setting. DerSimonian-Laird, Restricted Maximum Likelihood (REML), and Paule-Mandel have different statistical properties. The estimator, its uncertainty method, and the software implementation should be prespecified and reported. Sensitivity analyses can examine whether a defensible alternative changes the conclusion.

Is a prediction interval the same as the confidence interval for the pooled effect?

No. The confidence interval addresses uncertainty around the selected summary parameter. A prediction interval estimates a range for an underlying effect in a comparable future setting under the random-effects model. Its interpretation depends on the model, the amount and precision of evidence, and assumptions about the distribution of effects.

Can the checklist calculate or approve the analysis?

No. It is a planning and reporting record. It does not select a model or estimator, run a meta-analysis, validate data, or replace statistical review. Calculations must be performed in validated software and checked against the protocol and analysis plan.

Methodology reviewed: August 2026. MetaSyn Academy Reference Framework v2.4.