How to Choose an Effect Size for Meta-Analysis
Choose an effect measure by starting with the outcome construct, available data, target estimand, scale and interpretation, then document the assumptions that make synthesis defensible.
An effect size is not a detachable number that can be selected after extraction. It is an estimate on a defined scale, attached to a particular outcome, comparison, population, time point and statistical model. The Cochrane Handbook distinguishes binary, continuous, ordinal, count or rate, and time-to-event data because the available measures and their interpretations differ across those data structures.1 Treating every result as a generic “effect size” hides those differences at the point where they matter most.
The word size can also mislead. The choice of measure is not a contest to find the largest numerical value. A risk ratio, odds ratio, risk difference, mean difference and standardized mean difference place effects on different scales. Their values cannot be compared as though they were alternative readings from the same ruler. The selection must follow the scientific question, not the apparent magnitude or statistical significance of the result.1,2 Accessing free systematic review and meta-analysis templates helps establish standardized documentation routines across all synthesis stages.
Start with the target effect, not the statistic
Three questions establish the analysis target. What outcome construct matters? What comparison is being made? What effect of that comparison is the review trying to estimate? The first question prevents a shared numeric scale from disguising different concepts. The second fixes the order of groups, exposures or conditions. The third identifies the estimand: the target effect in a defined population, time and analysis context. Prespecified synthesis groups help keep those choices aligned with the review question.3 Only after those questions are clear should the team ask which effect measure can express the target with the available data.
For a randomized intervention review, the target may be the effect of assignment to the intervention at a prespecified follow-up. For a non-randomized association, the target may be a covariate-adjusted contrast. For prognosis, it may be a relative rate over time. Two studies can report the same statistical measure and still estimate different targets because their populations, adjustment sets, time windows or treatment strategies differ. Conversely, studies can report different instruments yet estimate a sufficiently similar construct to support a standardized analysis. The measure does not settle comparability by itself.2,3
Available data matter, but convenience does not have final authority. If the prespecified target requires a risk ratio and reports omit the necessary events, totals or compatible adjusted estimates, the response is to retrieve data, justify a different measure, separate syntheses or acknowledge that quantitative synthesis is not feasible. Quietly changing the analysis because a different column is easier to extract turns a data limitation into an undocumented change of question. Cochrane guidance emphasizes collecting the group summaries, effect estimates and uncertainty needed by the planned synthesis.4
Choosing among measures for continuous outcomes
Mean difference
The mean difference preserves the outcome unit and estimates the average difference between groups. It is usually the most interpretable route when studies use the same instrument, version and unit, and when the scale has the same substantive meaning across populations. If blood pressure is reported in the same unit, for example, a pooled mean difference can be read in that unit. This does not guarantee that pooling is sensible. Follow-up time, measurement protocol, score direction, final values versus changes, and design corrections still need agreement.1
A useful rule is not “same outcome means MD.” It is “same construct, compatible measurement unit and defensible common interpretation may support MD.” A scale called pain can have different ranges or anchors. A laboratory value can be expressed in different units. A change score and a final score can sometimes enter the same mean-difference synthesis, but the variance and assumptions need correct handling. Record the actual route.1,2
Standardized mean difference
The standardized mean difference is commonly used when studies measure the same underlying construct using different scales.1 Standardization divides a group difference by a within-study standard deviation, and small-sample corrections are commonly applied. This produces a common numerical unit, but it does not standardize the construct, study population or instrument quality. Combining a depression inventory with a general wellbeing scale is not justified merely because both can be converted into standard-deviation units.
The denominator also affects meaning. Populations with more variable scores can produce smaller standardized estimates for the same raw difference. Measurement reliability and study design can contribute to that variation. Methodological evaluations of SMD estimators and their sampling variances show that small samples and distributional conditions matter.8 The analysis record should therefore identify whether the software reports Cohen’s d, Hedges’ g or another correction, how the variance was calculated and how score direction was harmonized.
Interpretation should return to the outcome context. Universal small, medium and large labels are not substitutes for a meaningful difference on the actual construct. If possible, re-express the result using a familiar instrument, a minimally important difference or another justified reference. If re-expression requires extra assumptions, report them.1,9
Ratio of means and other routes
A ratio of means expresses a multiplicative contrast and may be useful when the scale has a meaningful zero and proportional change is scientifically interpretable. Empirical work has compared it with MD and SMD across continuous-outcome meta-analyses.6 It is not suitable simply because studies use different scales, and it is not defined sensibly for every measurement scale. The question is whether a ratio such as 0.80 has a stable meaning for the outcome, not whether the software can calculate it.
Choosing among measures for binary outcomes
Binary data commonly support a risk ratio, odds ratio or risk difference. A risk ratio compares probabilities. An odds ratio compares odds. A risk difference subtracts probabilities. Their null values are 1, 1 and 0 respectively. Each can be mathematically valid while answering a different numerical question.1
| Measure | What it expresses | Interpretive strength | Important caution |
|---|---|---|---|
| Risk ratio | Risk in one group divided by risk in the other | Often easier to explain as a relative probability | Depends on event definition and can behave poorly with sparse data or varying risks |
| Odds ratio | Odds in one group divided by odds in the other | Natural for logistic models and some study designs | Must not be interpreted as a risk ratio when events are common |
| Risk difference | Risk in one group minus risk in the other | Direct absolute change for the studied baseline risks | May vary across populations with different comparator risks |
Relative and absolute effects should not be treated as rivals when both are needed for decision-making. A relative measure may be chosen for synthesis, while absolute effects are later estimated for one or more credible comparator risks. That re-expression must preserve the uncertainty and show the assumed baseline risk. The number needed to treat is derived from an absolute effect and baseline risk for interpretation; Cochrane guidance does not recommend pooling NNTs directly.1
Event coding is part of the measure. For odds ratios, reversing event and non-event reciprocates the estimate. For risk differences, it changes the sign. For risk ratios, switching from an adverse event to its complement can change the numerical result and its precision in less intuitive ways. Decide whether the event means recovery, failure, occurrence or absence before extraction, and state which side of the null favours which group.1
Sparse data require more than choosing a label. Zero cells, double-zero trials, small studies and group imbalance can affect estimation and standard errors. Some common corrections can introduce bias or instability. Research on risk-ratio meta-analysis illustrates why blanket rules are unsafe.7 Prespecify the sparse-data method with statistical support and report sensitivity to consequential choices. A free guide cannot replace that case-specific analysis.1,7
Counts, rates, time-to-event outcomes, ordinal data and correlations
Counts allow more than one event per participant. Rates add exposure time, such as events per person-year. A rate ratio compares event rates, while a rate difference expresses an absolute difference in those rates. The analysis needs event counts and person-time or a compatible model-based estimate and uncertainty. Reducing recurrent events to a binary indicator can discard information and answer a different question.1
Time-to-event outcomes incorporate when an event occurs and how censoring is handled. A hazard ratio is common, but it is not a risk ratio over the entire follow-up. Its interpretation depends on the study model and assumptions, including whether a proportional-hazards interpretation is defensible. If studies report only risks at a fixed time, medians or survival curves, recovering a compatible estimate may require specialist methods. Do not relabel one of these summaries as a hazard ratio.1
Ordinal outcomes can be analysed as categories, dichotomized at a threshold, treated as approximately continuous or modelled through proportional odds. Every route has consequences. Thresholds can differ across studies, and dichotomization loses ordering information. A proportional-odds model uses more of the ordinal structure but adds an assumption and may be harder to interpret. The effect measure should follow the outcome definition and available reporting, not a preference for a particular software module.1,2
Correlations are bounded and their sampling properties change near the bounds, so conventional workflows commonly transform them before inverse-variance synthesis. The record should name the transformation and variance method. Recent methodological criticism suggests that standard approaches may be biased in some settings, but emerging evidence should be presented as an advanced qualification rather than a universal replacement rule. When the choice could materially change conclusions, plan sensitivity analysis and specialist review.1,2
Understand the analysis scale, display scale and null
Difference measures use a null value of 0. Ratio measures use a null value of 1 on the original display scale and 0 after logarithmic transformation. Ratio measures such as risk ratios, odds ratios, rate ratios and hazard ratios are generally analysed on the natural-log scale because that scale is symmetric around the null and supports appropriate statistical calculations.1 Software commonly back-transforms the pooled result for display.
Direction must be fixed as carefully as the null. Higher scores can mean improvement on one scale and deterioration on another. An experimental-versus-comparator risk ratio below 1 can represent benefit when the event is harmful and harm when the event is desirable. Forest-plot labels should make the direction visible, but the data record must establish it first. A reversed sign can look statistically plausible and still invert the scientific conclusion.1
Check compatibility, multiplicity and statistical dependence
Pooling requires more than a common measure. Studies must address a sufficiently similar construct and synthesis question. Time points, analysis populations, outcome definitions and adjustment sets can change the estimand. Standardization cannot erase those differences. If compatibility is uncertain, separate the estimates, investigate the source of difference or use narrative synthesis rather than forcing a single pooled number.1,3
Study design changes the variance and sometimes the estimate. Cluster-randomized trials require an analysis that respects clusters. Crossover and paired studies involve correlated observations. Multi-arm studies can reuse a comparator. Multiple outcomes or time points from the same participants are dependent. Treating each estimate as independent can give one study too much influence. Cochrane reporting guidance specifically identifies cluster, crossover, paired and multi-arm structures as unit-of-analysis concerns.2
Multiplicity also requires an outcome-selection rule. A study may report several scales for the same construct, several eligible time points, adjusted and unadjusted estimates, or multiple analysis populations. Selecting the estimate with the smallest P-value is not a solution. Prespecify an outcome hierarchy or use a statistical model that handles dependence. Record deviations when the planned estimate is unavailable.2,5
Questions before combination
- Do the estimates represent the same construct?
- Do they target the same comparison and time?
- Are scale and direction harmonized?
- Is the statistical unit correct?
- Is uncertainty estimated compatibly?
Reasons to pause
- Outcome labels hide different definitions
- Adjustment sets target different conditional effects
- A repeated comparator is counted twice
- Essential variances are reconstructed from weak assumptions
- The chosen measure changes after results are seen
Interpret the estimate on its own scale
Statistical compatibility and practical importance are different judgments. A confidence interval describes uncertainty around the estimate under the model; it does not identify the smallest meaningful effect or establish certainty in the evidence. A relative effect also needs a credible baseline risk before it can describe an absolute change. A standardized effect needs context before standard-deviation units become meaningful. Cochrane guidance recommends planning interpretation, including possible re-expression, at the protocol stage.9
Do not interpret an odds ratio as a risk ratio, a hazard ratio as a risk over a fixed follow-up, or an SMD as a change in the original instrument. Do not call a result important because the confidence interval excludes the null. The effect measure tells you how the estimate is expressed. Clinical, educational, policy or practical importance requires outcome-specific thresholds, consequences, baseline conditions and stakeholder values.1,9
Report the decision so another analyst can reproduce it
A complete methods statement names the outcome and time point, effect measure, comparison order, data source, transformations, variance method, dependence correction, software and version. It explains how different scales or definitions were harmonized, how multiple estimates were selected and how missing statistics were recovered. It also identifies prespecified sensitivity analyses and deviations. PRISMA 2020 provides the reporting framework, but a concise sentence in a manuscript cannot substitute for the underlying analysis record.5
For ratio measures, state that analysis occurred on the log scale and that results were back-transformed where relevant. For SMD, state the estimator and small-sample correction. For adjusted estimates, state which estimate was preferred and why. For cluster or paired designs, state the correction or model. For effect directions, state how scores or events were aligned. These details determine what the pooled number means.1,5
Apply the decision with the paired worksheet
The Effect size selection worksheet for meta-analysis creates one auditable record for each outcome and synthesis group. Use it after defining the target effect and before calculating study-level estimates. It records alternatives, assumptions, design complications, software and approval without acting as a calculator or recommendation engine.
Common errors and why they fail
| Error | Why it fails | Better action |
|---|---|---|
| Choosing from a software default | The menu does not know the construct, estimand or interpretation need. | Specify the target and measure before data entry. |
| Using SMD for different constructs | A common SD unit does not create conceptual equivalence. | Verify construct compatibility or separate outcomes. |
| Calling an OR a RR | Odds and risks differ, especially for common events. | Report and interpret the actual measure. |
| Changing event direction during extraction | The sign or reciprocal can reverse the conclusion. | Freeze coding and audit every study. |
| Ignoring dependence | Repeated participants or comparators can overstate precision. | Use a valid correction or dependent-effects model. |
| Selecting the measure with lower heterogeneity | Apparent homogeneity does not prove that the scale answers the right question. | Choose by scientific target and evaluate robustness transparently. |
| Using universal magnitude labels | Context and outcome meaning disappear. | Interpret against outcome-specific anchors and consequences. |
Worked decisions
Continuous symptoms measured by different instruments
A fictional review evaluates a therapy for the same prespecified anxiety construct at eight to twelve weeks. Studies use three validated scales with different ranges. The team considers MD, SMD and ratio of means. MD would require a common unit that the reports do not share. A ratio is not substantively stable because the scales do not have a common meaningful zero. The team selects a bias-corrected SMD, aligns higher scores to worse symptoms, flags variation in reliability and population SDs, records the software variance method and plans interpretation through a familiar scale and outcome-specific thresholds. The decision is defensible only if the instruments measure the same construct closely enough.1,8
Binary recovery with varying baseline risk
A fictional review studies recovery after a brief intervention. It selects a risk ratio for the primary synthesis because recovery is consistently defined and relative probability matches the protocol question. It does not call the result an odds ratio. Because baseline recovery varies across settings, the team also plans absolute re-expression at several justified comparator risks rather than treating one risk difference as universal. A subgroup with very few events is referred for sparse-data analysis rather than processed through an automatic correction.1,7
Time to relapse with incomplete reporting
A fictional review targets time to first relapse. Several studies report adjusted hazard ratios, while others report only the proportion relapsed at one year. The team does not combine those quantities under a single label. It seeks recoverable log hazard ratios and standard errors, records model and adjustment details and keeps fixed-time risks in a separate synthesis if compatible hazard estimates cannot be obtained. The missing data change feasibility, not the underlying target.1,4
Limitations and specialist boundary
This guide explains the architecture of effect-measure choice. It does not provide a universal decision tree, calculate sampling variances, solve sparse-event problems, recover survival estimates, correct clustered or paired data, or validate emerging estimators. Diagnostic-test accuracy, network meta-analysis, individual participant data and complex multilevel dependence require specialist methods outside this H06 guide. The next H06 guide addresses model assumptions; deep heterogeneity, sensitivity and publication-bias methods belong to H07.1,2
Conclusion
Choose an effect measure by preserving the meaning of the outcome and the target effect. Data type narrows the candidate family. Scale, direction, study design, available uncertainty, compatibility and interpretation determine whether a candidate is defensible. Document the choice before calculation, execute it in validated statistical software and keep the assumptions visible when the result is reported. That sequence prevents a convenient statistic from quietly replacing the scientific question. Explore complete guidance in the Meta-analysis methods, effect sizes and forest plots topic library.1,5
Frequently asked questions
Is effect size the same as estimand?
No. The estimand is the target effect under a defined population, comparison, outcome, time and analysis strategy. The effect measure is the numerical scale used to express an estimate of that target.
Should I always use standardized mean difference when scales differ?
No. The scales must measure the same underlying construct, and the interpretation and standardization assumptions must be defensible. Different scales alone are not enough.
Which is better, a risk ratio or an odds ratio?
Neither is universally better. They express different comparisons and have different properties. Choose according to the question, study designs, reporting, baseline risks, analysis method and intended interpretation.
Can I select the measure that gives the lowest heterogeneity?
Not as a mechanical rule. Effect-measure choice can affect heterogeneity, but apparent homogeneity does not establish that a measure answers the correct scientific question.
Can the same measure be pooled across every eligible study?
Only if the study estimates are sufficiently compatible in construct, comparison, time, analysis population, scale, direction and statistical unit.
References and evidence scope
Methodological guidance supporting effect size selection, continuous and binary effect measures, unit-of-analysis issues, and PRISMA 2020 reporting standards.
- Higgins JPT, Li T, Deeks JJ, editors. Chapter 6: Choosing effect measures and computing estimates of effect. Updated August 2023. In: Higgins JPT, Thomas J, Chandler J, Cumpston M, Li T, Page MJ, et al, editors. Cochrane Handbook for Systematic Reviews of Interventions. Version 6.5. Cochrane; 2024. https://www.cochrane.org/authors/handbooks-and-manuals/handbook/current/chapter-06
- Deeks JJ, Higgins JPT, Altman DG, McKenzie JE, Veroniki AA, editors. Chapter 10: Analysing data and undertaking meta-analyses. In: Higgins JPT, Thomas J, Chandler J, Cumpston M, Li T, Page MJ, et al, editors. Cochrane Handbook for Systematic Reviews of Interventions. Version 6.5. Cochrane; 2024. https://www.cochrane.org/authors/handbooks-and-manuals/handbook/current/chapter-10
- McKenzie JE, Brennan SE, Ryan RE, Thomson HJ, Johnston RV, Thomas J. Chapter 3: Defining the criteria for including studies and how they will be grouped for the synthesis. In: Higgins JPT, Thomas J, Chandler J, Cumpston M, Li T, Page MJ, et al, editors. Cochrane Handbook for Systematic Reviews of Interventions. Version 6.5. Cochrane; 2024. https://www.cochrane.org/authors/handbooks-and-manuals/handbook/current/chapter-03
- Li T, Higgins JPT, Deeks JJ, editors. Chapter 5: Collecting data. In: Higgins JPT, Thomas J, Chandler J, Cumpston M, Li T, Page MJ, et al, editors. Cochrane Handbook for Systematic Reviews of Interventions. Version 6.5. Cochrane; 2024. https://www.cochrane.org/authors/handbooks-and-manuals/handbook/current/chapter-05
- Page MJ, McKenzie JE, Bossuyt PM, Boutron I, Hoffmann TC, Mulrow CD, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71. https://doi.org/10.1136/bmj.n71
- Friedrich JO, Adhikari NKJ, Beyene J. Ratio of means for analyzing continuous outcomes in meta-analysis performed as well as mean difference methods. Journal of Clinical Epidemiology. 2011;64(5):556-564. https://doi.org/10.1016/j.jclinepi.2010.09.016
- Bakbergenuly I, Hoaglin DC, Kulinskaya E. Pitfalls of using the risk ratio in meta-analysis. Research Synthesis Methods. 2019;10(3):398-419. https://doi.org/10.1002/jrsm.1347
- Lin L, Aloe AM. Evaluation of various estimators for standardized mean difference in meta-analysis. Statistics in Medicine. 2021;40(2):403-426. https://doi.org/10.1002/sim.8781
- Schünemann HJ, Vist GE, Higgins JPT, Santesso N, Deeks JJ, Glasziou P, et al. Chapter 15: Interpreting results and drawing conclusions. In: Higgins JPT, Thomas J, Chandler J, Cumpston M, Li T, Page MJ, et al, editors. Cochrane Handbook for Systematic Reviews of Interventions. Version 6.5. Cochrane; 2024. https://www.cochrane.org/authors/handbooks-and-manuals/handbook/current/chapter-15
Build the protocol before the analysis
Course 1 is MetaSyn Academy’s currently available course. It teaches the protocol and decision record that should exist before effect-measure selection becomes an analysis step.
Explore How to write a Systematic Review and Meta-Analysis Protocol →See where meta-analysis fits
Course 6 is presented as a future pre-registration pathway within the verified twelve-course curriculum.
Explore the full Meta-Journey →