Systematic Review Data Extraction and Data Management

Data Management Methodology

Transforming source reports into an analysis-ready dataset

Design a data-extraction system that links reports to studies, preserves source evidence, governs transformations and produces a dataset fit for synthesis.

Central proposition: Data extraction is not a copying exercise. It is a controlled transformation in which every analytical value must remain connected to the evidence object, source location, interpretation and decision history that produced it.

The smallest transcription error can become a study-level conclusion

After study selection, a systematic review changes form. The evidence base is no longer a set of potentially eligible records; it is a network of reports that must be resolved into studies, study arms, outcomes, time points and analysable estimates. That change is methodologically consequential. A number copied from the wrong arm, a standard error treated as a standard deviation, a follow-up report counted as a second study, or a denominator taken from the randomised rather than analysed population can alter an effect estimate while leaving no visible trace in the final manuscript. Data extraction is therefore the point at which published evidence becomes the review’s own research dataset.

The conversion is not neutral. Reports are written for clinical or scientific communication, not for direct import into a synthesis. The required information may be distributed across the abstract, methods, tables, figures, appendices, registries and companion publications. Definitions can change between reports. Outcome labels may appear identical although scales, time windows or analysis populations differ. Conversely, two labels may describe measurements that are sufficiently comparable for a prespecified synthesis. The extractor must identify the evidence, interpret it under the protocol, and preserve enough context for another reviewer to reconstruct that judgment.

Cochrane guidance consequently treats the study, not the report, as the unit of interest and recommends linking multiple reports of the same study before collection is complete. It also treats a piloted collection form as a bridge between the report and the analysis, not merely as clerical convenience.1,8 PRISMA 2020 asks authors to report how data were collected, how many reviewers participated, whether they worked independently and what automation was used. These are reporting requirements, not a substitute for a defensible collection method.2,3 The method must still fit the review design, protocol and governing standard.

Methodological boundary: No universal extraction form, pilot sample, missing-value code or verification pattern is valid for every review. The structure must be derived from the planned synthesis and adapted to the applicable review methodology.

Define the evidence objects before defining the fields

A rigorous system starts with an entity model. “One row per paper” is attractive because it resembles a bibliography, but it usually confounds documents with investigations. One trial can generate a protocol, registry entry, conference abstract, principal results article, harms report and long-term follow-up. Each report is a source; together they describe one study. If report identity and study identity are stored in the same field, the review can double-count participants, lose complementary information or silently select whichever report is easiest to extract.

The next levels are determined by the scientific question. Parallel trials commonly require arm-level fields; cluster trials require cluster and participant denominators; diagnostic-accuracy reviews require index-test and threshold structure; qualitative syntheses may require findings, illustrations and credibility assessments; prevalence reviews may require population, numerator, denominator and measurement frame. Outcomes then require definition, instrument, metric, direction, time point, analysis population and estimate. These are not decorative metadata. They determine whether observations can be compared and how transformations must be performed.

Stable identifiers allow the system to represent these relationships without forcing all information into one flat row. A report ID links the value to a document. A study ID connects companion reports. Arm, outcome and time-point IDs identify the analytical observation. Where the planned synthesis needs repeated estimates, the combination of these identifiers prevents one value from being mistaken for another. The structure should remain understandable outside a proprietary platform. Identifiers, labels and relationships belong in the exported dataset and its data dictionary.

The evidence-object architecture of data extraction A nested architecture separates reports from studies, arms, outcomes, time points and estimates while retaining links between them. One evidence base, several legitimate units Report document-level source Study investigation-level entity Arm allocated or observed group Outcome × time point measurement context Source evidence page · table · figure registry · supplement Value and interpretation reported value · unit · statistic extractor · decision · verification Analysis field raw or derived with transformation history Stable identifiers preserve every relationship and correction
Figure 1. Extraction is reliable only when the data structure preserves the distinction between documents, investigations and analytical observations. MetaSyn Academy synthesis.1,8

The protocol determines the schema, but the schema tests the protocol

Extraction fields should be traced to a planned descriptive table, risk-of-bias assessment, effect measure, subgroup, sensitivity analysis or narrative synthesis. This traceability prevents indiscriminate collection and exposes gaps in the protocol. If an effect measure requires an analysed denominator but the form records only the number randomised, the synthesis plan and form are misaligned. If a subgroup analysis depends on a participant characteristic but the characteristic has no operational definition, the review is not ready for extraction. Designing the schema is therefore both an implementation step and a methodological rehearsal.

Each field needs more than a label. The data dictionary should state the entity level, definition, permitted format, unit, allowed values, source priority, treatment of “not reported” and “not applicable”, and any derivation rule. A field called “age” is ambiguous until it specifies whether the value is a baseline mean, median, eligibility range or another measure, whether it is arm-level or study-level, and how dispersion is represented. Free text remains appropriate for quotations and explanatory notes, but variables used for filtering or analysis require controlled, explicit values.

The schema should also separate absence from inapplicability. “Not reported” means the information was sought in the defined source set but was not found. “Not applicable” means the field does not logically apply. “Unclear” means information exists but cannot be interpreted reliably. “Not sought” means the search for the field has not been completed. Combining these states into an empty cell destroys information and can make missingness appear random when it actually reflects reporting practices, source availability or review decisions.

Piloting is the practical test of the schema. A varied set of studies should be used to reveal fields that are ambiguous, values that cannot be represented, repeated information that belongs at a different level, and derivations that lack their required inputs. The sample should be chosen for methodological diversity rather than by a universal numerical rule. Revisions should be versioned, approved and applied retrospectively when they affect already extracted records. A pilot ends when the structure can represent the expected evidence and the team can explain its decisions, not when an arbitrary sample size is reached.

Preserve what was reported before recording what will be analysed

Reported, corrected, reconciled and derived values are different evidence states. A report may contain an internally inconsistent total, a figure from which data are digitised, or a statistic that must be transformed. The analysis may require a combined group, change score, log scale or imputed dispersion. Overwriting the reported value with the usable value hides the transformation and prevents independent checking. A controlled record stores the source value and location, the interpreted or corrected value, the transformation rule, the person responsible and the verification outcome.

Provenance should be sufficiently precise to relocate the evidence. “Table 2” may be adequate in a short report but insufficient when multiple supplements or versions exist. Record the report identifier, page or section, table/figure and relevant row or column where practical. For data supplied by an author, record the contact event, response, file or message location, and whether the information clarifies, corrects or supplements the publication. Author-provided data do not become self-authenticating; their relationship to the study and their analytical use still require evaluation.

Calculated fields should be reproducible from stored inputs. Rather than entering only a converted standard deviation, retain the reported standard error, sample size, formula and resulting value. Rather than merging intervention arms without explanation, retain the original arm values and the prespecified combination procedure. Derivation scripts can reduce arithmetic error, but they need version control and test cases. A script that executes successfully may still implement the wrong estimand.

A controlled transformation pathway from source to synthesis An immutable source layer feeds extraction, reconciliation and derivation, with verification and a decision history surrounding every transformation. Do not overwrite the evidence trail Immutable source quotation or transcription with precise location Extracted record typed field and unit extractor and timestamp Reconciled value discrepancy explained consensus preserved Analysis dataset derived fields labelled ready for planned use Controls surrounding every transition stable IDs and data dictionary independent verification proportional to risk reason for correction or derivation version, reviewer and approval history An analysis value without provenance is a claim that cannot be audited
Figure 2. Controlled extraction retains source evidence, reconciled decisions and derived values as related but non-interchangeable records. MetaSyn Academy synthesis.1,5,9

Verification should follow the consequence of error

Methodological evidence shows that independent checking prevents errors, particularly for outcome data that enter meta-analysis. In a study of systematic-review extraction, single extraction produced more errors than double extraction, and methodological reviews continue to document non-trivial error frequencies.5,9 Cochrane intervention-review guidance distinguishes between data domains: duplicate collection of outcome data is mandatory, whereas duplicate collection of study characteristics is highly desirable rather than universally mandatory.1 These classifications should not be generalized to every review family, but they demonstrate an important principle: verification intensity can be aligned with the downstream consequence of a mistake.

A rigorous workflow preserves both original entries before reconciliation. If one reviewer’s entry is simply overwritten, the record cannot show whether the difference arose from transcription, interpretation, source selection or entity linkage. Reconciliation should identify the disputed field, compare source evidence, record the resolution and preserve the final accountable decision. Recurring discrepancies should trigger correction of the form, instructions or training rather than being treated as isolated individual failures.

Machine learning tools can assist with locating text, suggesting values, checking ranges, detecting empty fields or reproducing calculations. Evidence on automated extraction remains task- and system-specific; performance on one corpus does not establish safe autonomous use in another.6 The system should record which operation was automated, its version, the human verification performed and how conflicts were handled. AI-generated values require source-level confirmation before they enter the authoritative dataset. Confidence scores do not replace evidence locations.

A dataset becomes trustworthy through governance, not appearance

Data management should define roles, access, file naming, versioning, change control, backups, retention and the transition from working data to the locked analysis dataset. A clean table can still contain untraceable corrections; a messy-looking working file can contain valuable provenance. Quality is determined by whether the system can answer who entered a value, from which source, under which definition, what changed and why.

Quality checks should operate at several levels. Structural checks test identifiers, permissible codes and required relationships. Range and type checks identify impossible or malformed values. Cross-field checks test internal logic, for example, analysed participants should not normally exceed randomised participants without an explanation. Source verification confirms the evidence itself. Analytical checks reproduce transformations and test whether the dataset matches the planned estimand. No single “validated” badge substitutes for these different controls.

The analysis freeze should be an event, not an informal moment. Record the dataset version, unresolved queries, approved exclusions, transformation code, data dictionary and responsible reviewers. Later corrections should create a new version with a documented reason rather than changing the frozen file invisibly. This allows peer reviewers, collaborators and update teams to understand which data supported a published result and whether a later version differs.

Governance must also preserve negative decisions. A field excluded after piloting, an estimate judged ineligible, or an author-supplied result withheld from analysis can be as important as a retained value. Record the governing rule and reason so that absence from the analysis dataset is not mistaken for oversight. When an update changes the rule, the decision history identifies which studies require reassessment and prevents selective recovery of convenient values.

DATA EXTRACTION RESOURCES

Build a traceable systematic review data-extraction system

Convert included reports into structured, analysis-ready evidence while preserving study–report relationships, source provenance, transformations, verification decisions and missing-data contacts.

Working resources for auditable data extraction and evidence management

Data extraction

Design and pilot extraction forms, define study-characteristics fields, document missing information and author contacts, and preserve every transformation and verification decision.

Systematic review data extraction form template

Design and pilot extraction forms, define study-characteristics fields, document missing information and author contacts, and preserve every transformation and verification decision.

Study characteristics table template

Define the row unit and comparison fields needed to describe included studies without mixing study, arm- and outcome-level information.

Missing data contact log and request templates

Plan precise author queries, document contact attempts and responses, and preserve how additional information changed the analytical record.

Extraction ends when the evidence trail and the analysis dataset agree

The purpose of extraction is not to fill every field. It is to create a dataset in which each value has an unambiguous scientific meaning and a recoverable path back to the evidence. That requires linked report and study identities, level-aware variables, explicit missing-information states, preserved source values, reproducible transformations and accountable reconciliation. When these elements are absent, later statistical precision can conceal earlier uncertainty.

A complete workflow therefore moves in a controlled sequence: define the analytical requirements; model the evidence entities; write the data dictionary; pilot diverse studies; revise and version the form; extract and verify according to the governing method; resolve discrepancies without deleting original entries; manage author-supplied information as new evidence; reproduce derived fields; and freeze a documented analysis version. The order matters because downstream checks cannot fully repair an ill-defined upstream schema.

The resources in this hub support that sequence but do not determine the review’s scientific choices. The extraction form cannot decide which outcomes matter. The characteristics table cannot make incomparable studies comparable. The contact log cannot guarantee a response or eliminate the possibility of selective availability. Each tool must remain subordinate to the protocol, applicable methodological standard and expert judgment.

Boundary of use: Adapt every field and workflow to the review question and design. Preserve the original evidence whenever a value is corrected, transformed or supplied after publication. Do not describe a dataset as verified unless the actual verification method, reviewer roles and unresolved exceptions are recorded.

Build the protocol decisions that extraction must implement

Use the existing Academy course to define the question, outcomes, eligibility rules and planned synthesis before freezing the extraction architecture.

Build your protocol foundation → Explore the 13-course Meta-Journey →

Get the Resource Infrastructure Free Forever

Get Free Lifetime Access

References

  1. Li T, Higgins JPT, Deeks JJ. Chapter 5: Collecting data. In: Higgins JPT, Thomas J, Chandler J, Cumpston M, Li T, Page MJ, Welch VA, editors. Cochrane Handbook for Systematic Reviews of Interventions. Version 6.5. Cochrane; 2024. Access the current chapter.
  2. Page MJ, McKenzie JE, Bossuyt PM, Boutron I, Hoffmann TC, Mulrow CD, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71. doi:10.1136/bmj.n71.
  3. Page MJ, Moher D, Bossuyt PM, Boutron I, Hoffmann TC, Mulrow CD, et al. PRISMA 2020 explanation and elaboration: updated guidance and exemplars for reporting systematic reviews. BMJ. 2021;372:n160. doi:10.1136/bmj.n160.
  4. Aromataris E, Lockwood C, Porritt K, Pilla B, Jordan Z, editors. JBI Manual for Evidence Synthesis. Adelaide: JBI; 2024. doi:10.46658/JBIMES-24-01.
  5. Buscemi N, Hartling L, Vandermeer B, Tjosvold L, Klassen TP. Single data extraction generated more errors than double data extraction in systematic reviews. J Clin Epidemiol. 2006;59(7):697-703. doi:10.1016/j.jclinepi.2005.11.010.
  6. Jonnalagadda SR, Goyal P, Huffman MD. Automating data extraction in systematic reviews: a systematic review. Syst Rev. 2015;4:78. doi:10.1186/s13643-015-0066-7.
  7. Moher D, Shamseer L, Clarke M, Ghersi D, Liberati A, Petticrew M, et al. Preferred reporting items for systematic review and meta-analysis protocols (PRISMA-P) 2015 statement. Syst Rev. 2015;4:1. doi:10.1186/2046-4053-4-1.
  8. Higgins JPT, Thomas J, Chandler J, Cumpston M, Li T, Page MJ, Welch VA, editors. Cochrane Handbook for Systematic Reviews of Interventions. 2nd ed. Chichester: Wiley; 2019.
  9. Mathes T, Klaßen P, Pieper D. Frequency of data extraction errors and methods to increase data extraction quality: a methodological review. BMC Med Res Methodol. 2017;17:152. doi:10.1186/s12874-017-0431-4.
  10. Li T, Vedula SS, Scherer R, Dickersin K. What comparative effectiveness research is needed? A framework for using guidelines and systematic reviews to identify evidence gaps and research priorities. Ann Intern Med. 2012;156(5):367-377. doi:10.7326/0003-4819-156-5-201203060-00009.