EFFECT SIZE SELECTION RESOURCE
Effect Size Selection Worksheet for Meta-analysis
Document the outcome type, construct, measurement scale, direction, candidate effect measure, compatibility assumptions, software output, and rationale without calculating an estimate.
Document the outcome construct, data type, measurement scale, direction, candidate effect measure, compatibility assumptions, software output and rationale before quantitative synthesis.
A meta-analysis cannot repair a mismatch between what studies measured and what a reviewer intends to claim. Before any study-level estimate enters software, the review team needs a controlled record of the outcome, the contrast, the unit, the direction and the assumptions that allow estimates to be compared. Cochrane distinguishes binary, continuous, ordinal, count or rate, and time-to-event outcome data because each type supports different effect measures and different routes to an estimate.1 Treating every result as a generic “effect size” hides those differences at the point where they matter most.
The task is not to find the most sophisticated statistic. It is to select a scale that answers the prespecified synthesis question and can be estimated consistently across eligible studies. A mean difference can preserve the original measurement unit. A standardized mean difference can permit synthesis across different instruments, but changes the unit and introduces assumptions about the meaning of a standard deviation. A risk ratio, odds ratio and risk difference can all be computed from binary outcomes, yet they express different comparisons and behave differently when baseline risk varies. Hazard ratios concern event timing and censoring, not simply whether an event occurred. The right decision therefore depends on the scientific target as well as the data table.1,2 Accessing free systematic review and meta-analysis templates helps establish standardized documentation routines across all synthesis stages.
Why effect-measure selection belongs in the protocol and analysis record
Selection made after viewing study results is vulnerable to convenience and outcome-driven switching. The protocol should identify the intended measure for each outcome and explain important contingencies, such as how different scales, adjusted estimates, rare events, clustered allocation or missing variances will be handled. Prespecified grouping also helps keep the analysis aligned with the review question.3 PRISMA 2020 expects review authors to report the effect measures and synthesis methods used.5 Reporting becomes more credible when the manuscript can be traced to a dated decision record rather than reconstructed after the analysis.
A useful record separates three concepts. The outcome construct is what matters, such as depressive symptoms, mortality or time to relapse. The estimand is the target effect under a defined comparison and population, such as the effect of assignment to an intervention at twelve weeks. The effect measure is the numerical scale used to express an estimate of that target, such as a mean difference or risk ratio. These terms are related, but they are not interchangeable. A review can use a familiar effect measure and still target the wrong effect if the comparison, time point or analysis population does not match the protocol. The extraction plan must also retain the group summaries, estimates and uncertainty needed to implement the selected route.4
Match the measure to the outcome and the intended meaning
Continuous outcomes
Use a mean difference when studies assess the same construct on the same scale and a difference in that scale is meaningful. The result keeps the original unit, which can make interpretation direct. This apparent simplicity does not eliminate judgment. A five-point difference has to refer to comparable versions of the instrument, comparable score direction, a defensible time point and a population for which that difference has the intended meaning. Change scores and final values may sometimes be combined on a mean-difference scale, but the extraction and variance methods must be compatible and prespecified.1
A standardized mean difference is considered when studies measure the same underlying construct using different instruments. Standardization places each difference in standard-deviation units, allowing numerical combination across scales. It does not make unlike constructs equivalent. It also means that the denominator depends on variation within studies, which may reflect population heterogeneity, measurement reliability and study design as well as the underlying effect. The team should record the correction used for small samples, the variance method implemented by the software and the direction assigned to higher scores. Advanced evaluations show that SMD estimation and variance estimation can behave differently under small samples and non-ideal distributions.1,8
Do not convert a standardized estimate into a universal label such as small, medium or large and stop there. Interpretation should return to the outcome, population, instrument properties and a meaningful reference where one is available. A ratio of means is another possible continuous-outcome measure when a multiplicative comparison is scientifically sensible and the scale has a meaningful zero. Empirical work has examined it as an alternative to difference measures, not as a universal replacement.6
Binary outcomes
Risk ratio, odds ratio and risk difference are not alternative names for the same result. A risk ratio compares probabilities, an odds ratio compares odds and a risk difference compares absolute probabilities. The first two have a null value of 1; the risk difference has a null value of 0. Odds ratios can diverge materially from risk ratios when events are common, so interpreting an odds ratio as though it were a risk ratio can exaggerate the apparent change.1
Relative and absolute measures also transport differently across baseline risks. A stable relative effect can imply different absolute effects in populations with different comparator risks. A risk difference may be immediately interpretable in one setting but vary across settings because baseline risk changes. Selection should therefore consider whether the review primarily needs a relative comparison for synthesis, an absolute comparison for a defined population, or both synthesis and later re-expression. The number needed to treat is derived from an absolute difference for interpretation; it is not itself an effect measure to pool directly.1
Sparse events create further constraints. Zero cells, double-zero studies, imbalance and small samples can affect whether a study-level measure and standard error can be estimated reliably. The worksheet records the planned handling and the software method, but it deliberately does not prescribe a continuity correction or a sparse-data model. Those choices require the full analysis context. Methodological work on risk ratios, for example, identifies pitfalls without establishing that one binary measure is best in every review.1,7
Counts, rates, time-to-event outcomes, ordinal data and correlations
A count records how many events a participant experiences; a rate relates events to exposure time. Treating recurrent events as a binary yes/no outcome discards frequency and follow-up information. A rate ratio may be suitable when event counts and person-time or a compatible adjusted estimate are available, but its assumptions and the definition of exposure time must be recorded. Time-to-event data add censoring and follow-up timing. Hazard ratios are common, yet their interpretation depends on how the underlying study model represents hazards and time. An event risk at a fixed time and a hazard ratio are not interchangeable merely because both concern the same clinical event.1
Ordinal outcomes may be analysed through a defensible continuous, binary or proportional-odds route depending on the scale, reporting and model. Dichotomizing a rich scale can lose information and make thresholds differ across studies. Correlations are often transformed before conventional inverse-variance synthesis because their raw sampling distribution is problematic near its bounds. Recent methodological debate about correlation meta-analysis should be treated as an advanced qualification. The core decision record should name the transformation, variance method, software and interpretation rather than claim that an emerging method has already replaced established practice.1,2
Audit compatibility before pooling
Two estimates with the same label are not automatically compatible. A mean difference at four weeks and a mean difference at twelve months may answer different questions. Two risk ratios can reverse direction if one study codes recovery as the event and another codes non-recovery. An SMD for anxiety and an SMD for general psychological distress may use the same unit while measuring different constructs. Compatibility is a scientific judgment supported by the protocol, not a property created by standardization.1
Dependence deserves explicit attention. Cluster-randomized trials assign groups rather than independent individuals. Crossover and paired designs produce correlated observations. Multi-arm trials can contribute the same comparator more than once. Repeated measures and multiple time points can create several eligible estimates from the same participants. If those estimates are treated as independent, their precision can be overstated. The worksheet asks for the design and adjustment so the analysis team can verify that the study-level estimate and variance match the statistical unit.2
Adjusted and unadjusted estimates create another compatibility decision. In non-randomized studies, an adjusted estimate may address confounding better than a crude comparison, but different adjustment sets can target different conditional effects. The record should list the covariates and identify which estimate the protocol prioritizes. A common numeric scale does not guarantee a common estimand.2
Record before analysis
- Outcome definition, unit and time point
- Comparison order and direction of benefit
- Target effect and analysis population
- Data type and available statistics
- Candidate measure and null value
Verify before pooling
- Same construct or a justified grouping
- Compatible scale or defensible standardization
- Correct dependence and variance handling
- Consistent event and direction coding
- Software output matches the prespecified method
How to use the worksheet
- Complete one record for each outcome and synthesis group, not one record for the whole review.
- Describe the construct and estimand before selecting a candidate measure.
- Inventory the statistics and study designs actually available. Do not assume every report contains the needed variance.
- Audit scale, direction, timing and statistical dependence across studies.
- Record the candidate measure, alternatives considered and the reason for the decision.
- Run the calculation in validated statistical software, then return to record the software, version and output location.
- Obtain methodological review for unresolved transformations, dependence, sparse data or incompatible estimands.
Effect size selection decision record
Entries remain in this browser when you select Save locally. Nothing is transmitted to MetaSyn Academy. This form records decisions and does not calculate or validate an effect estimate.
Illustrative record: different depression scales
Fictional example, not an analysis. A review compares a structured support programme with usual care for adult depressive symptoms at twelve weeks. Eligible trials use several validated symptom scales with different numerical ranges. The team defines the construct as self-reported depressive symptom severity, prespecifies the effect of assignment to the programme, confirms that higher values indicate worse symptoms after recoding and proposes a bias-corrected standardized mean difference. The record states that pooling assumes the instruments measure the same underlying construct and that differences in within-study standard deviations do not make the standardized effects substantively incomparable. It names the planned software and variance method, flags one cluster trial for design correction and records that interpretation will use outcome-specific anchors where possible rather than universal magnitude labels.
This record does not establish that the SMD is correct. It makes the assumptions visible so a statistician, methodologist or reviewer can challenge them before the pooled result is produced.
Completion check
A record is ready for analysis only when another qualified team member can reconstruct why the measure was selected and how each eligible study will contribute a compatible estimate and variance. Empty fields are not the only sign of incompleteness. A polished rationale can still fail if it never defines the event, ignores a reversed scale, combines unlike time points, treats correlated estimates as independent or relies on a software default that no one has documented.
- The construct, comparison, population and time point are explicit.
- The estimand is distinguished from the numerical measure.
- The available statistics support the proposed study-level estimate.
- The null value, direction and analysis scale are correct.
- Compatibility across instruments, definitions and designs is justified.
- Transformations and unit-of-analysis corrections have an owner.
- The software, version, output location and review decision are recorded.
- Any departure from the protocol is connected to an amendment or deviation record.
Document uncertainty in the decision itself
An effect-measure record should not create false certainty. Some studies may report insufficient variance information, use unclear event definitions, or provide estimates adjusted for incompatible covariate sets. Record these limitations as unresolved rather than filling gaps from convention. If a recovery method, correlation assumption or conversion is planned, name the required inputs, justify the method, identify who will verify it, and retain the calculation with the analysis output.
Alternative measures may remain defensible. The record should explain why the primary measure best matches the estimand and how a reasonable alternative would change interpretation. Sensitivity analysis is useful when it tests a material assumption, but it should not become a search for the most favourable result. Preserve the prespecified decision, label deviations and report whether alternatives alter the substantive conclusion.5
Interpretation plans belong in the record because they reveal whether the chosen scale serves the review question. State the null, direction, meaningful thresholds and any planned absolute re-expression. For standardized effects, identify outcome-specific anchors or external interpretive evidence where available. For relative effects, specify the baseline risk source needed for absolute translation. This planning prevents a mathematically valid estimate from being reported in language it cannot support.1
Methodological boundary
This resource supports documentation and review of an effect-measure decision. It does not calculate an estimate, recommend a measure automatically, test statistical assumptions, correct a unit-of-analysis error, determine whether studies should be pooled, or validate software output. Use qualified statistical and subject-matter review for complex transformations, sparse data, dependence, non-randomized estimates and emerging methods.1,2
Conclusion
Effect-measure selection is a chain of scientific and statistical decisions, not a menu choice made after data entry. A strong record begins with the outcome and target effect, names the data and design constraints, explains the selected scale, audits compatibility and preserves the software and approval trail. Once that chain is visible, calculation can proceed in validated statistical software without outsourcing judgment to the program.5
Get every MetaSyn template free, including this one.
Leave your email and I’ll send this resource as an editable Word file and a printable PDF, plus access to the smart online version. You’ll also get every new template as it’s finished. No noise, just the resources.
References and evidence scope
- Higgins JPT, Li T, Deeks JJ, editors. Chapter 6: Choosing effect measures and computing estimates of effect. Updated August 2023. In: Higgins JPT, Thomas J, Chandler J, Cumpston M, Li T, Page MJ, et al, editors. Cochrane Handbook for Systematic Reviews of Interventions. Version 6.5. Cochrane; 2024. https://www.cochrane.org/authors/handbooks-and-manuals/handbook/current/chapter-06
- Deeks JJ, Higgins JPT, Altman DG, McKenzie JE, Veroniki AA, editors. Chapter 10: Analysing data and undertaking meta-analyses. In: Higgins JPT, Thomas J, Chandler J, Cumpston M, Li T, Page MJ, et al, editors. Cochrane Handbook for Systematic Reviews of Interventions. Version 6.5. Cochrane; 2024. https://www.cochrane.org/authors/handbooks-and-manuals/handbook/current/chapter-10
- Borenstein M, Hedges LV, Higgins JPT, Rothstein HR. A basic introduction to fixed-effect and random-effects models for meta-analysis. Research Synthesis Methods. 2010;1(2):97-111. https://doi.org/10.1002/jrsm.12
- Veroniki AA, Jackson D, Viechtbauer W, Bender R, Bowden J, Knapp G, et al. Methods to estimate the between-study variance and its uncertainty in meta-analysis. Research Synthesis Methods. 2016;7(1):55-79. https://doi.org/10.1002/jrsm.1164
- Page MJ, McKenzie JE, Bossuyt PM, Boutron I, Hoffmann TC, Mulrow CD, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71. https://doi.org/10.1136/bmj.n71
- Friedrich JO, Adhikari NKJ, Beyene J. Ratio of means for analyzing continuous outcomes in meta-analysis performed as well as mean difference methods. J Clin Epidemiol. 2011;64(5):556-564. https://doi.org/10.1016/j.jclinepi.2010.09.016
- Bakbergenuly I, Hoaglin DC, Kulinskaya E. Pitfalls of using the risk ratio in meta-analysis. Res Synth Methods. 2019;10(3):398-419. https://doi.org/10.1002/jrsm.1347
- Lin L, Aloe AM. Evaluation of various estimators for standardized mean difference in meta-analysis. Stat Med. 2021;40(2):403-426. https://doi.org/10.1002/sim.8781
Scope and use boundary
The worksheet creates an auditable record of a proposed effect-measure decision. It does not calculate an estimate, determine that studies are compatible, repair missing or dependent data, select a statistical model, or certify an analysis. Complex sparse-data, clustered, crossover, multi-arm, non-randomized and emerging-method decisions require qualified methodological and statistical review.
Frequently asked questions
Does the worksheet tell me which effect measure to use?
No. It makes the outcome, estimand, available data, assumptions and rationale visible so the review team can make and review the decision. An automatic recommendation would conceal the contextual judgments the resource is designed to document.
Can I use one record for every outcome in a review?
No. Complete a separate record for each outcome and synthesis group when the construct, time point, comparison, data type or planned measure differs. A review-level policy can be referenced, but each applied decision needs its own traceable record.
Is a standardized mean difference appropriate whenever studies use different scales?
Not automatically. The scales must measure the same underlying construct closely enough for standardization to support the synthesis question. Direction, reliability, population variability, time point and interpretation also need review.
Are odds ratios and risk ratios interchangeable?
No. They compare different quantities. Their numerical values can diverge when events are common, and an odds ratio should not be described as though it were a risk ratio.
Where should the numerical effect estimate be calculated?
Use validated statistical software under a prespecified analysis plan. Record the software, version, options and output location in this worksheet, then verify the result against the source data and analysis specification.