How to Design a Data Extraction Form for a Systematic Review

MetaSyn Academy guide to design a data extraction form for a systematic review, illustrating the decisions documented by Systematic review data
Analytical thesis: The quality of an extraction form is not measured by how many fields it contains. It is measured by whether the form represents the review’s evidence correctly and permits every analysis value to be reconstructed and challenged.

Form design is an exercise in measurement

A data extraction form is often introduced as a table that reviewers complete after study selection. That description understates its scientific role. The form determines which features of the primary studies become variables in the systematic review, how those variables are interpreted, and which distinctions remain visible during synthesis. In this sense, form design is a measurement problem: the review team constructs an operational representation of studies that were not designed to fit its question.

The representation can fail in several ways. A field can be absent although the analysis requires it. A label can merge distinct concepts. A flat layout can confuse the report with the underlying study. A controlled list can impose categories that the evidence does not support. A transcription can be accurate but scientifically wrong because it came from the wrong arm, population or time point. These failures are not solved by adding a generic checklist. They require a model of the evidence and an explicit link between every field and its use.

Methodological authorities converge on several foundations while differing in review-specific detail. Cochrane guidance for intervention reviews emphasizes linked reports, piloted forms, source recording, duplicate outcome collection and preservation of reported and derived data.1 JBI collection requirements depend on the type of synthesis and should be applied through the relevant methodology rather than turned into one universal rule.4 PRISMA 2020 specifies what authors should report about collection but does not certify how a form was designed.2,3 A defensible form therefore combines general controls with review-specific scientific decisions.

Design backwards from the claims the review intends to make

Begin with the planned products: study-characteristics tables, risk-of-bias assessments, effect estimates, subgroup analyses, sensitivity analyses and narrative comparisons. For each product, list the required variables and the exact entity at which each variable exists. An effect estimate may need group-specific sample size, mean and dispersion at a defined time point. A subgroup may require a study-level setting or a participant-level characteristic summarized at arm level. A sensitivity analysis may depend on whether a value was imputed or supplied by an author. The schema should make these requirements explicit before extraction starts.

This backward design exposes analytical promises that cannot yet be implemented. A protocol may name “long-term outcome” without defining the eligible window or how to select among multiple measurements. It may propose an adjusted analysis without specifying which adjustment set takes priority. It may plan to combine cluster-randomised and individually randomised trials without fields for cluster design effects. These are not form-layout questions; they are unresolved synthesis decisions. Form development should return them to the protocol team rather than conceal them inside an extractor note.

Not every potentially interesting characteristic belongs in the form. Collecting variables without a descriptive, analytical or interpretive purpose increases workload and opportunities for inconsistency. Yet parsimony should follow traceability, not precede it. The correct test is whether the field supports a prespecified decision, explains an important boundary, or preserves provenance. If it does neither, it may be omitted. If it does, its definition must be sufficiently precise for reuse.

Represent the evidence as entities and relationships

The most consequential design choice is the unit structure. A report is a document; a study is an investigation; an arm is a group; an outcome is a defined construct and measurement; an estimate is a value for a particular comparison or group at a particular time. A single study can have several reports, arms, outcomes and time points. When these entities are collapsed, fields become repeated or ambiguous and the risk of double counting rises.

A relational design does not require sophisticated database software. Separate worksheets with stable keys can represent the same logic. A report table can link each document to a study ID. A study table can store design and setting once. An arm table can store interventions and group denominators. Outcome and estimate tables can store repeated measurements. The relationships must be documented so that analysts know the expected one-to-many and many-to-many connections.

Some review types require different entities. Diagnostic-accuracy reviews may need index tests, reference standards and thresholds; network meta-analysis may require intervention nodes and multi-arm correlations; individual-participant-data reviews may require dataset and participant provenance; qualitative syntheses may require findings and illustrations. The general principle is not to force every review into the same five tables. It is to identify the units that make the intended claims scientifically coherent.

Define field semantics, not only field names

A field name is a navigation label. The data dictionary supplies the scientific meaning. For each field, specify its definition, entity, data type, permissible values, unit, evidence source, missing-information codes and decision rules. If a field is derived, specify the inputs and transformation. If reviewers must select among several eligible values, state the hierarchy. If the source conflicts with another report, state how the conflict will be preserved and escalated.

Consider “sample size”. A report may provide the number randomised, treated, assessed at baseline, analysed for a specific outcome and included in a model. Each may be correct for a different purpose. A single sample-size field invites reviewers to make unrecorded choices. A well-defined schema names the population and time, identifies the arm and connects the denominator to the estimate. The same principle applies to age, follow-up duration, event counts and adjusted effects.

Controlled vocabularies reduce variation but can create false certainty. “Parallel RCT” may be a stable category; “high dose” may depend on context. Allow a source quotation or explanatory note where classification requires judgment. Preserve the observed description and the review’s coded interpretation as separate fields. This makes it possible to revise the coding without losing what the study reported.

Establish a source hierarchy without erasing disagreement

Companion reports can provide complementary or conflicting information. A protocol may define planned outcomes, a registry may record amendments, and the results article may report the analysis performed. The extraction method should specify which sources are sought and how their roles are interpreted. A source hierarchy can guide routine choices, but it should not automatically discard discrepant evidence. Store each relevant value with its source, then record the resolution.

For every extracted observation, capture a precise location and enough context to prevent the number from becoming detached from its meaning. A table cell may require its row label, column label, unit and footnote. A value in a figure may require the digitisation method and file. An estimate reported only in text may require the sentence and analysis population. Provenance is not fulfilled by attaching the article PDF somewhere in the project folder.

Corrections and derivations should create new states. Store the source value, corrected or transformed value, rule, responsible reviewer and verification. This is essential when converting standard errors to standard deviations, combining groups, deriving change-score dispersion or selecting one time point among several. If a method cannot reproduce the usable value from retained inputs, it is not sufficiently documented.

How extraction errors propagate An upstream entity or source error affects interpretation, calculation and synthesis, whereas a controlled verification loop intercepts the error near its origin. The cost of an error grows downstream Entity mistake wrong study, arm or time Source mistake wrong table or population Meaning mistake unit, statistic or direction Biased analytical result effect estimate or interpretation changes Interception loop stable identifiers → source location → independent entry → discrepancy classification reconciliation → reproducible transformation → locked analysis version Verification is most informative when it explains the origin of disagreement
Figure 1. Errors in entity assignment or source interpretation can survive arithmetic checks and propagate into synthesis. MetaSyn Academy analytical model.5,9

Pilot against diversity, ambiguity and analytical consequence

A random handful of simple studies may demonstrate that the form opens and saves; it does not test whether the schema represents the evidence base. Construct a pilot set that includes likely structural challenges: multiple reports, multiple arms, repeated measures, incomplete outcome reporting, alternative statistics, complex interventions and at least one case requiring a planned derivation. The size should follow the heterogeneity and risk of the review, because no empirical universal number guarantees adequacy.

Reviewers should work independently under the proposed instructions and compare the complete record. Differences must be classified. A transcription error suggests a checking control; a source-selection difference suggests a hierarchy; an entity difference suggests a structural problem; an interpretation difference suggests an operational definition; a repeated calculation difference suggests code rather than manual computation. Treating all disagreements as “reviewer error” prevents the form from improving.

Version every change. State whether previously extracted records are affected and how they will be updated. A label change can be cosmetic, whereas a change in outcome window or analysis population can alter the eligible estimate. Both should be recorded, but only the latter may require re-extraction. The pilot record should allow a later reviewer to understand when the form became stable enough for the main extraction.

Design verification around plausible failure modes

Evidence that single extraction can produce more errors than double extraction supports strong checking, especially for outcome data.5,9 Yet “double extraction” is not one method. Reviewers may independently enter data, one may extract and another verify against the source, or selected fields may receive different treatment. Independence, visibility of the first entry, reconciliation and documentation influence what the control can detect.

High-consequence outcome fields often justify independent duplicate extraction. Complex transformations benefit from tested code and independent review of inputs and expected outputs. Low-ambiguity descriptive fields may use structured extraction with targeted or sampled checking if the applicable standard permits it. Judgment-heavy classifications may need expert adjudication even when they do not enter a meta-analysis directly. The verification plan should name the field families, method, reviewer roles and action when an error is found.

Preserve the two original entries before consensus. The discrepancy itself contains information about the form. If conflicts cluster around one outcome, time window or source, revise the rule and reassess affected records. If discrepancies are simple transpositions, interface constraints or range checks may be appropriate. Verification is most valuable when it changes the system, not when it merely produces a final clean value.

A decision matrix for extraction verification Verification intensity increases with analytical consequence, interpretive ambiguity and transformation burden. Allocate checking by consequence and ambiguity HIGH LOW Analytical consequence and transformation burden → Targeted expert verification complex design labels judgment-rich contextual fields Clarify definitions and adjudicate Independent duplicate extraction effect estimates · denominators · time points derived values · high-impact classifications Preserve both entries, then reconcile Structured single extraction low-ambiguity descriptive fields with range and completeness checks Sample or risk-triggered checking Reproducible computation formulae · conversions · group combinations tested code plus source-level review Verify inputs and expected outputs Interpretive ambiguity →
Figure 2. A verification plan should be justified by field risk rather than treating every item as equally consequential. MetaSyn Academy decision framework.1,5,9

Use automation as a governed operation

Software can pre-populate citation fields, identify candidate sentences, extract tables, check allowed values and reproduce transformations. Reviews of automated extraction demonstrate promise but also heterogeneity of tasks, corpora and evaluation methods.6 Performance should therefore be validated on the review’s own documents and fields. A tool that finds population descriptions accurately may not identify the correct outcome denominator or distinguish adjusted from unadjusted effects.

For each automated operation, record the system, version, prompt or configuration where relevant, input documents, output, threshold or stopping rule, and human verification. Retain the source span offered by the system. If a reviewer cannot relocate and confirm the evidence, the value should not enter the authoritative dataset. Automation must not overwrite previous states or make corrections without a decision record.

Validation should include difficult and negative cases, not only fields the system filled. Missing extractions can be more consequential than incorrect candidates because they may be invisible. Compare against independently reviewed records, stratify performance by field type, and predefine the response to failure. Human oversight is not a decorative final glance; it is the accountable step that decides whether the extracted claim is supported.

Prepare the form for analysis freeze and future update

Before analysis, run structural checks on identifiers, relationships, required fields and allowed values; source checks on high-consequence observations; and reproducibility checks on derived fields. Resolve or explicitly classify open queries. Freeze the dataset, dictionary, form version and transformation code together. A later correction should create a new version with an explanation and an assessment of affected results.

An update team should be able to append new reports, link them to existing studies and apply the current definitions without reconstructing the original logic from column names. Preserve schema migrations and mappings between old and new codes. If the review question or outcome hierarchy changes, distinguish a methodological amendment from a technical migration. These records turn extraction from a one-time spreadsheet exercise into a durable evidence asset.

Data protection applies even when the review uses published studies. Author correspondence may contain personal contact information or unpublished files. Store only what is necessary, use institutional repositories and access controls, and keep public exports free of unnecessary personal data. Auditability does not require indiscriminate disclosure.

Conclusion: design the evidence model, then design the interface

A defensible extraction form begins with the planned claims and the entities that make those claims possible. It defines every field, preserves report–study relationships, separates reported from derived data, tests structural diversity, and directs verification toward consequential failure modes. The interface may be a document, spreadsheet, review platform or database; the methodological architecture should remain stable across those implementations.

The paired Systematic review data extraction form template operationalises this architecture as a practical starting point. It must be adapted to the review design and protocol. Neither the template nor a software validation message certifies that the selected outcome, source or transformation is scientifically correct. That responsibility remains with the review team and its declared method.

Translate this method into a controlled extraction record

Use the paired resource to map report, study, arm, outcome and estimate fields, then export a reviewable provenance record.

Systematic review data extraction form template → Systematic review data extraction and data management →

References

  1. Li T, Higgins JPT, Deeks JJ. Chapter 5: Collecting data. In: Higgins JPT, Thomas J, Chandler J, Cumpston M, Li T, Page MJ, Welch VA, editors. Cochrane Handbook for Systematic Reviews of Interventions. Version 6.5. Cochrane; 2024. Access the current chapter.
  2. Page MJ, McKenzie JE, Bossuyt PM, Boutron I, Hoffmann TC, Mulrow CD, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71. doi:10.1136/bmj.n71.
  3. Page MJ, Moher D, Bossuyt PM, Boutron I, Hoffmann TC, Mulrow CD, et al. PRISMA 2020 explanation and elaboration: updated guidance and exemplars for reporting systematic reviews. BMJ. 2021;372:n160. doi:10.1136/bmj.n160.
  4. Aromataris E, Lockwood C, Porritt K, Pilla B, Jordan Z, editors. JBI Manual for Evidence Synthesis. Adelaide: JBI; 2024. doi:10.46658/JBIMES-24-01.
  5. Buscemi N, Hartling L, Vandermeer B, Tjosvold L, Klassen TP. Single data extraction generated more errors than double data extraction in systematic reviews. J Clin Epidemiol. 2006;59(7):697-703. doi:10.1016/j.jclinepi.2005.11.010.
  6. Jonnalagadda SR, Goyal P, Huffman MD. Automating data extraction in systematic reviews: a systematic review. Syst Rev. 2015;4:78. doi:10.1186/s13643-015-0066-7.
  7. Moher D, Shamseer L, Clarke M, Ghersi D, Liberati A, Petticrew M, et al. Preferred reporting items for systematic review and meta-analysis protocols (PRISMA-P) 2015 statement. Syst Rev. 2015;4:1. doi:10.1186/2046-4053-4-1.
  8. Higgins JPT, Thomas J, Chandler J, Cumpston M, Li T, Page MJ, Welch VA, editors. Cochrane Handbook for Systematic Reviews of Interventions. 2nd ed. Chichester: Wiley; 2019.
  9. Mathes T, Klaßen P, Pieper D. Frequency of data extraction errors and methods to increase data extraction quality: a methodological review. BMC Med Res Methodol. 2017;17:152. doi:10.1186/s12874-017-0431-4.
  10. Li T, Vedula SS, Scherer R, Dickersin K. What comparative effectiveness research is needed? A framework for using guidelines and systematic reviews to identify evidence gaps and research priorities. Ann Intern Med. 2012;156(5):367-377. doi:10.7326/0003-4819-156-5-201203060-00009.

Similar Posts

Leave a Reply

Your email address will not be published. Required fields are marked *