Fixed-Effect vs Random-Effects Meta-Analysis: Assumptions, Models and Reporting
A rigorous model choice starts with the effect you want to estimate, the population to which you want to generalize, and the assumptions that connect the included studies. Heterogeneity statistics describe evidence; they do not choose the model for you.
Choosing between a common-effect model, often called a fixed-effect model, and a random-effects model is one of the most consequential planning decisions in quantitative evidence synthesis. The choice changes the parameter being estimated, the weights assigned to studies, the uncertainty around the summary, and the scope of the conclusion. It should therefore be an explicit part of the protocol and analysis plan, not a reaction to whether a heterogeneity test happens to cross a significance threshold.1,2 Accessing free systematic review and meta-analysis templates helps establish standardized documentation routines across all synthesis stages.
The terminology itself can obscure the issue. A common-effect model assumes that every study targets one underlying effect and that observed differences arise from sampling error. A random-effects model assumes that study-specific underlying effects vary according to a distribution and commonly estimates the mean of that distribution. Neither label tells you whether the included studies are unbiased, clinically compatible, or correctly analysed. Model choice is a statement about the scientific target and the data-generating structure, not a certificate of quality.
Why a heterogeneity test should not choose the model
A common but weak workflow runs Cochran’s Q test, uses a common-effect model when the result is non-significant, and switches to random effects when it is significant. This procedure confuses lack of evidence for heterogeneity with evidence that all underlying effects are identical. Heterogeneity tests can have low power when few studies are available and excessive sensitivity when many studies are included. Their result also depends on study precision. A threshold-driven switch makes the scientific model contingent on an unstable diagnostic.
The same objection applies to rules based on I2. I2 describes the proportion of observed variability attributable to between-study heterogeneity rather than sampling error under its assumptions. It is affected by within-study precision, uncertainty can be substantial, and the same underlying heterogeneity may yield different I2 values in different evidence sets. It is useful context, but it cannot decide whether the target is one effect or a distribution of effects.
Cochrane guidance consequently recommends that the decision to use a common-effect or random-effects model should not be based on the statistical test for heterogeneity. The review question, study diversity, assumptions, and intended interpretation supply the rationale. Heterogeneity statistics remain important after the model has been prespecified because they describe variation, motivate investigation, and qualify confidence in the summary.1
Start with the estimand
An estimand is the quantity the synthesis is intended to estimate. It includes the outcome, intervention contrast, time point, effect measure, population, and conditions under which the effect is interpreted. In meta-analysis, it must also clarify whether the target is a single common effect, an average across a distribution of study-specific effects, or another well-defined summary. If that question is left implicit, the model label can appear precise while the scientific claim remains ambiguous.1,2
Suppose a review combines trials conducted in similar populations with the same intervention implementation, comparator and follow-up. A common-effect estimand may be defensible if the team argues that the studies estimate the same underlying quantity. Suppose instead that interventions, settings and populations differ in ways expected to modify the effect, yet the studies remain sufficiently related to address a broader question. A distribution-of-effects model may better represent the target. This is not automatic permission to combine everything. Clinical and methodological compatibility must still be justified before any statistical pooling.
The population of inference is equally important. A common-effect analysis may summarize only the included set under the common-effect assumptions. A random-effects average is often interpreted as a mean across a conceptual population of studies, but that generalization requires the included studies to represent that population in a meaningful way. Convenience samples of published studies do not become representative merely because a random-effects command was selected.
What the common-effect model assumes
Under the conventional common-effect model, each observed study estimate equals one common underlying effect plus sampling error. Inverse-variance weighting gives more precise studies more influence because their sampling variances are smaller. If the assumptions are credible, the pooled estimate and its confidence interval address uncertainty about that shared effect.6
This model is sometimes described as appropriate when studies are homogeneous. That wording can be misleading. Observed results never need to be numerically identical, and a non-significant heterogeneity test does not prove the assumption. The justification is substantive and design based: are the studies intended to estimate the same parameter, and are important effect-modifying differences sufficiently controlled or absent? Even when the common-effect model is selected, unexplained disagreement deserves inspection.
A common-effect estimate can be dominated by one or two very large studies. That is expected from its weighting scheme, but it means the summary can primarily reflect the settings represented by those studies. Teams should examine study contributions and determine whether the pooled interpretation matches the review question. A mathematically precise answer to a poorly aligned question is not a useful synthesis.6
What the random-effects model assumes
The conventional random-effects model adds a between-study variance component, usually denoted τ2, to the within-study sampling variance. Each study receives weight inversely related to the sum of these components. As τ2 increases, weights tend to become more similar, so smaller studies generally receive more relative influence than under a common-effect model. This can materially alter the pooled estimate when study effects and study size are associated.5
The usual random-effects summary is an estimate of the mean of an assumed distribution of underlying effects. This interpretation requires more than detecting heterogeneity. The effects should form a scientifically interpretable distribution, and an average should answer a meaningful question. If effects differ because incompatible interventions, outcomes or biases have been combined, the statistical mean may conceal rather than resolve the problem.
Random effects also does not mean that studies were randomly sampled from all possible studies. The model is a working representation of between-study variation. Generalizing beyond the included evidence still requires judgment about transportability, selection, publication processes and the characteristics of future settings.5
Estimating between-study variance
τ2 is not directly observed. It is estimated, often from a modest number of heterogeneous and imprecise studies. Common estimators include DerSimonian-Laird, Restricted Maximum Likelihood (REML), and Paule-Mandel. They differ in bias, efficiency, boundary behaviour, computational method and performance across conditions. No estimator is uniformly best for every review.3
DerSimonian-Laird is historically common and computationally simple, but it can underestimate between-study variance in some settings, particularly with few studies or substantial heterogeneity. Restricted maximum likelihood has attractive statistical properties and is widely available. Paule-Mandel is another well-supported option. The important operational lesson is to name the estimator, prespecify it, verify software defaults, and record whether sensitivity analyses with a defensible alternative change the inference.
Reporting only the pooled effect while omitting τ2 leaves a central part of the model invisible. When possible, report the τ2 estimate, its method and an uncertainty interval. Interpret τ2 on the scale of the effect measure. A numerical value that appears small on one scale may represent important variation on another.3
Confidence intervals and Hartung-Knapp methods
The variance estimator and the confidence-interval method are separate choices. A standard random-effects calculation often uses a normal approximation that treats aspects of heterogeneity estimation too confidently. Hartung-Knapp-Sidik-Jonkman (HKSJ) procedures modify uncertainty around the pooled mean and can provide better coverage in many conditions, especially when the number of studies is small. They are not a universal mechanical remedy: performance depends on implementation and data configuration, and some modified variants are used to avoid counterintuitive narrowing.4
Software interfaces may present these choices under different labels. The methods section should therefore report the procedure rather than only the program name. Record the software, version, package or module, estimator, interval option, and any modification. If the selected implementation differs from the protocol, document why and whether conclusions are sensitive to the change.8
When a prediction interval helps
A confidence interval around the random-effects mean describes uncertainty about that mean. A prediction interval addresses the range in which an underlying effect from a comparable new setting might lie under the fitted model. It can therefore reveal that an apparently precise average coexists with practically important variation, including effects on both sides of the null.
Prediction intervals are most interpretable when a random-effects model is scientifically justified, enough studies inform heterogeneity, and the assumed distribution is plausible. With very few studies, τ2 and the prediction interval can be extremely uncertain. An interval should not be presented as a guaranteed range, nor should it be used to imply that every future setting has equal probability. State the method used and interpret it alongside study characteristics and heterogeneity.
Compare the models without treating them as rivals
| Decision layer | Common-effect model | Random-effects model |
|---|---|---|
| Usual target | One common underlying effect | Mean of a distribution of underlying effects |
| Source of variation | Sampling error around the common effect | Sampling error plus between-study variation |
| Weighting | Inverse within-study variance | Inverse within-study variance plus estimated τ2 |
| Role of small studies | Usually less relative weight | Usually more relative weight as τ2 increases |
| Generalization | Restricted to the common-effect estimand and evidence set | Potentially a conceptual distribution, if representativeness is justified |
| Main vulnerability | False precision if one common effect is implausible | Unstable heterogeneity estimation and misleading averages if studies are incompatible |
A prespecified decision process
First, specify the review question and estimand: outcome, time point, contrast, effect measure and population of inference. Second, assess whether clinical and methodological differences permit a meaningful synthesis. Third, decide whether the target is one common effect or a distribution of effects. Fourth, prespecify the weighting and variance estimator. Fifth, choose the confidence-interval method and conditions for a prediction interval. Sixth, plan sensitivity analyses that test influential assumptions rather than searching for a preferred result.7
Useful sensitivity analyses can compare reasonable τ2 estimators, alternative interval methods, influential-study exclusions justified in advance, and assumptions required to convert or impute data. They should be interpreted as robustness checks. Repeatedly trying specifications and retaining the most favourable result is selective analysis. Preserve the primary analysis and report material deviations and their consequences.
The model plan should also define how heterogeneity will be described and investigated. That includes τ2, I2 and Q where appropriate, along with subgroup analyses or meta-regression justified by plausible effect modifiers. Those methods require their own safeguards against sparse data and multiplicity. They belong to a later analytical step and should not be improvised to rescue an uncomfortable pooled result.
Reporting the analysis so it can be reproduced
A reproducible methods statement identifies the model and its scientific rationale, effect measure and scale, weighting approach, τ2 estimator, summary-effect interval method, prediction-interval method if used, software and version, and prespecified sensitivity analyses. Results should report the number of studies and participants, summary estimate with confidence interval, heterogeneity statistics with uncertainty where possible, prediction interval when justified, and outcomes of sensitivity analyses.7
Do not write only that a random-effects model was used because heterogeneity was high. That sentence leaves the estimand, estimator and interval method unclear and implies an automatic threshold rule. A better explanation states that meaningful variation in underlying effects was anticipated across eligible settings, so the analysis targeted the mean of a distribution; it then names the estimator and uncertainty method.
Likewise, avoid presenting agreement between common-effect and random-effects point estimates as proof of robustness. Similar point estimates can coexist with materially different uncertainty, while both analyses can share the same data limitations. Robustness is a reasoned assessment across assumptions, influence, bias and applicability, not a single numerical coincidence.7
Worked interpretation example
Imagine twelve trials estimate a risk ratio for an intervention at six months. Settings and baseline risk vary, and implementation intensity is expected to modify effects. The protocol therefore targets the mean relative effect across comparable implementation settings and prespecifies a random-effects model. Restricted maximum likelihood estimates τ2, a Hartung-Knapp method forms the confidence interval, and a prediction interval is planned because the review seeks to inform future settings.4
The resulting pooled risk ratio is not interpreted as the effect every setting will experience. The report describes it as the estimated mean of the assumed effect distribution. The confidence interval quantifies uncertainty in that mean; the prediction interval conveys plausible variation in a new comparable setting. If the prediction interval includes important benefit, no effect and harm, the average alone is insufficient for a universal recommendation. Investigating prespecified effect modifiers and considering certainty of evidence become essential.
If a sensitivity analysis using Paule-Mandel produces similar substantive conclusions, the report records that stability. If a conventional normal interval would produce a narrower interval and a different threshold conclusion, the discrepancy is reported, not hidden. The analysis remains anchored to the prespecified primary method.7
Common mistakes and corrections
Mistake: using fixed effect when I2 is below a cut-off
Correction: choose the model from the estimand and assumptions. Report I2 with context and uncertainty as descriptive evidence.
Mistake: assuming random effects is always conservative
Correction: inspect weights and study-size patterns. Greater relative weight for small studies can move the summary and does not guarantee a wider or safer conclusion in every configuration.
Mistake: reporting software without its settings
Correction: name the software version, estimator, interval procedure, prediction-interval method and non-default options.7,8
Mistake: treating the pooled mean as a universal effect
Correction: describe the target precisely and use a prediction interval, study context and effect-modifier evidence where justified.
Mistake: changing the model after seeing the preferred result
Correction: preserve the prespecified primary analysis, label deviations, justify them and report sensitivity analyses transparently.7
Turn the reasoning into an auditable record
Use the Meta-analysis model selection and reporting checklist to document the estimand, assumptions, estimator, interval methods, software and deviations before interpreting results.
Where this step sits in the workflow
Model selection follows data preparation and effect-size computation. It precedes formal investigation of heterogeneity, subgroup analysis, meta-regression and publication-bias assessment. Explore complete guidance in the Meta-analysis methods, effect sizes and forest plots topic library.
References and evidence scope
Methodological guidance supporting model specification, common-effect versus random-effects estimands, between-study variance estimation, and PRISMA 2020 reporting compliance.
- Deeks JJ, Higgins JPT, Altman DG, McKenzie JE, Veroniki AA, editors. Chapter 10: Analysing data and undertaking meta-analyses. In: Higgins JPT, Thomas J, Chandler J, Cumpston M, Li T, Page MJ, et al, editors. Cochrane Handbook for Systematic Reviews of Interventions. Version 6.5. Cochrane; 2024. https://www.cochrane.org/authors/handbooks-and-manuals/handbook/current/chapter-10
- Borenstein M, Hedges LV, Higgins JPT, Rothstein HR. A basic introduction to fixed-effect and random-effects models for meta-analysis. Research Synthesis Methods. 2010;1(2):97-111. https://doi.org/10.1002/jrsm.12
- Veroniki AA, Jackson D, Viechtbauer W, Bender R, Bowden J, Knapp G, et al. Methods to estimate the between-study variance and its uncertainty in meta-analysis. Research Synthesis Methods. 2016;7(1):55-79. https://doi.org/10.1002/jrsm.1164
- IntHout J, Ioannidis JPA, Borm GF. The Hartung-Knapp-Sidik-Jonkman method for random effects meta-analysis is straightforward and considerably outperforms the standard DerSimonian-Laird method. BMC Medical Research Methodology. 2014;14:25. https://doi.org/10.1186/1471-2288-14-25
- Higgins JPT, Thompson SG, Spiegelhalter DJ. A re-evaluation of random-effects meta-analysis. Journal of the Royal Statistical Society: Series A (Statistics in Society). 2009;172(1):137-159. https://doi.org/10.1111/j.1467-985X.2008.00552.x
- Rice K, Higgins JPT, Lumley T. A re-evaluation of fixed effect(s) meta-analysis. Journal of the Royal Statistical Society: Series A (Statistics in Society). 2018;181(1):205-227. https://doi.org/10.1111/rssa.12275
- Page MJ, McKenzie JE, Bossuyt PM, Boutron I, Hoffmann TC, Mulrow CD, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71. https://doi.org/10.1136/bmj.n71
- Cochrane. RevMan knowledge base: release notes. Cochrane; 2025. https://documentation.cochrane.org/revman-kb/revman-release-notes-342261851.html
Build the protocol before the analysis
Course 1 is MetaSyn Academy’s currently available course. It teaches the protocol and decision record that should exist before model selection becomes an analysis step.
Explore How to write a Systematic Review and Meta-Analysis Protocol →See where meta-analysis fits
Course 6 is presented as a future pre-registration pathway within the verified twelve-course curriculum.
Explore the full Meta-Journey →