How to Extract Measurement-Property Data for PROMs in Systematic Reviews
Extract the evidence packet, not just the coefficient
The safest way to extract PROM measurement-property data is to preserve each result together with the score it belongs to, the population that generated it, the analysis that produced it and the source from which it was obtained. A coefficient separated from those details may be numerically correct and still be unusable. Current COSMIN methodology therefore separates PROM characteristics, study populations, measurement-property results, methodological-quality assessment, result ratings and certainty rather than treating extraction as a single spreadsheet operation.12
The practical rule is simple: preserve what the source actually reported before you interpret, transform or select anything. Record the PROM structure, language identity and score target; bind the result to the correct analysis sample; preserve the analysis or model needed to understand the number; keep a precise source locator; and label every reviewer calculation, reconstruction or author-supplied value as something different from the original report.
Generic data extraction becomes fragile in PROM reviews
Most systematic-review data forms assume that an outcome can be described with a fairly stable sequence: study, group, outcome, time point and numerical result. PROM measurement-property evidence adds several layers that can change the meaning of the result. One questionnaire name may refer to a long form and a short form. A multidimensional PROM may have a total score and several subscales. A translated version may share its structure with the original but have language-specific comprehensibility evidence. A reliability article may report two ICC models. A structural-validity article may test several competing factor models. A single cohort may generate evidence for reliability, measurement error and construct validity.
Flatten all of that into one row too early and the data start to drift. An ICC becomes detached from its model. An alpha for one subscale is later read as evidence for the whole PROM. A translated instrument is accidentally merged with a modified version that also changed the item set. A calculated SDC replaces the original SEM, leaving no way to reproduce the transformation. The problem is rarely that somebody typed the wrong number. More often, the number survived while its identity did not.
COSMIN’s current manual addresses these problems by keeping several stages separate. Step 5 establishes which PROMs are available and extracts their characteristics, feasibility and interpretability. Step 6 then works through measurement-property studies, their populations, results, risk of bias, result ratings, synthesis and certainty.2 The extraction file does not need to imitate the COSMIN workbook exactly, but it should respect those conceptual boundaries.
PROM identity has more than one dimension
The first extraction decision is deceptively difficult: what counts as the same PROM? COSMIN begins from a conservative position. Different versions or modifications may have different measurement properties and are initially considered unique. It also considers each subscale of a multidimensional PROM separately when measurement properties are evaluated.2
But “version” is not one-dimensional. Current Step 5 guidance says different language versions with the same structure can be considered one version at that point. This prevents language translation from automatically being confused with structural modification. The language still matters. COSMIN later recommends separate comprehensibility ratings by language because comprehensibility is language dependent, and cross-cultural validity or measurement invariance explicitly concerns comparisons across groups.2
The practical solution is to preserve several identity dimensions rather than force everything into one version label. Keep the structural or scoring version, language or cultural adaptation, administration mode and score target visible. The exact database implementation is a review-management choice, but those identities answer different questions.
What is being scored?
A short form, changed item set, altered subscale structure or different scoring algorithm can define a genuinely different version whose evidence should not be merged without justification.
In which linguistic form?
Language identity can remain separate from structural identity. This is particularly important when comprehensibility or cross-cultural evidence is language specific.
Which output?
A coefficient for the physical-function subscale does not automatically describe the total score or another subscale, even when all appear under one instrument name.
For whom was it observed?
Measurement properties are empirical findings from particular populations and analysis sets. Later synthesis may judge transferability, but extraction should not erase the source population.
Do not turn every factor model into a different PROM
Structural-validity studies create a special problem because authors may test many models. It is tempting to solve the problem with a rule such as “extract every model” or “take the best-fitting model.” Both are too crude.
The current COSMIN manual says reviewers should evaluate structural validity of the structure and scoring algorithm as the PROM is actually used. When several factor models have been tested, reviewers should focus on models representing that practical structure. COSMIN gives the example of a scale for which a total score is commonly used despite multiple factor structures in the literature; a unidimensional, bifactor or higher-order model may therefore be more relevant to the total-score interpretation than an unrelated exploratory structure that nobody uses for scoring. Other factor structures can still be evaluated but may be less informative.2
This is why extraction eligibility should come from the review protocol and methodological question, not from the size of the fit statistic. If the review asks whether a published four-subscale scoring structure is supported, then the relevant models are those that test that structure. If a paper introduces a revised scoring algorithm that the review also treats as an eligible PROM version, the revised structure may require a distinct record. The decision is about what score is being evaluated, not which model looks best.
Model cherry-picking can happen in both directions. Extracting only the best-fitting model can exaggerate support for a structure. Extracting every exploratory model without a protocol-defined reason can bury the actual scoring question under irrelevant analyses. Define the evidence target first.
Example: three CFA models in one article
A paper tests the instrument’s published three-factor model, a modified three-factor model with correlated residuals and a bifactor model. The abstract reports only the bifactor model because it fits best.
Content validity does not fit a conventional result column
PROM development and content-validity evidence require a different extraction mindset. COSMIN identifies three sources for evaluating content validity: information from concept elicitation and pilot work during PROM development, additional content-validity studies, and the review team’s own ratings of relevance, comprehensiveness and comprehensibility.2
For the first two sources, current COSMIN guidance explicitly says reviewers do not simply extract “the results” about relevance, comprehensiveness and comprehensibility in the same manner as other properties. Instead, they should highlight relevant qualitative or quantitative evidence in the development and content-validity reports, which is then used during the later content-validity rating.2
That distinction matters when designing a form. A forced numeric field encourages the extractor to search for a content-validity index even when the informative evidence is a cognitive interview finding, a recurring misunderstanding of an item, evidence that patients identified a missing concept, or a qualitative judgment about relevance. The extraction record therefore needs an evidence-note structure capable of preserving the aspect being informed, the source group and a precise locator.
There is also no reason to transcribe every patient quotation from every development paper by default. Item-level extraction may be appropriate when a review specifically investigates instrument refinement, short-form development or problematic items. For a broader PROM review, targeted highlighted evidence may be enough. The level should follow the review question.
A number needs its statistical identity
A generic “result” field is usually too weak for PROM evidence. The statistic type, analysis model, uncertainty and denominator can change interpretation. Reliability illustrates this clearly. An ICC is not one parameter. It belongs to a family of models. Agreement and consistency versions answer different questions, and the current COSMIN Risk of Bias methodology gives model choice a substantive role in appraisal.8
Extraction should therefore preserve the exact model when the study reports it. The same principle applies to weighted kappa, factor-model fit statistics, IRT/Rasch parameters, AUCs, correlations, SEMs and SDCs. Do not turn statistical metadata into an improvised quality judgment during extraction. Simply make sure the details survive long enough for the relevant appraisal and result-rating stage.
The analysis sample deserves the same care. COSMIN repeatedly specifies that study-population information should describe the population to which the measurement-property result refers, and for construct-validity comparisons it explicitly notes that the sample size is the sample included in the analysis, with group-specific numbers when groups are compared.2 Total recruitment and analysis N are therefore not interchangeable fields.
| Result | Metadata that often changes its meaning | Extraction mistake to avoid |
|---|---|---|
| ICC | Reliability type, model, agreement/consistency, CI, analysis N, interval and repeated-measure context | Recording only “ICC = 0.82” |
| Cronbach alpha or omega | Exact score/subscale, N, coefficient type and linked dimensionality context | Applying one total-score coefficient to every subscale |
| RMSEA / CFI / TLI / SRMR | Exact model, score structure, estimator where needed, analysis sample | Combining fit statistics from different models into one row |
| SEM / SDC / LoA | Units, calculation/model, repeated-measure context, N and source/derived status | Treating SEM and SDC as interchangeable |
| Correlation | Comparator construct, comparator instrument, expected hypothesis, N and correlation type | Calling every correlation “criterion validity” |
| AUC | What classification or change criterion was used, analysis N and context | Recording an AUC without the criterion it discriminates |
Construct validity starts with the review’s hypotheses
Construct-validity extraction deserves one correction to a common “extract everything” rule. Current COSMIN methodology instructs the review team to decide which hypotheses are relevant before extracting results. The reason is practical: primary-study authors may provide no hypotheses, use vague hypotheses or formulate expectations that conflict with those in other studies. To compare results consistently across studies, the review needs a coherent set of hypotheses and rationales.2
The extractor should therefore find the results that test those relevant hypotheses, not vacuum every correlation from the article. The record needs the comparator or known group, expected direction and magnitude defined by the review, and the observed result. Whether the observation confirms the hypothesis is a later evaluation step. Keeping the expectation and observation separate prevents the data from being retrofitted to the answer.
This distinction also clarifies the role of primary-study hypotheses. They are useful source information and may help the review team understand the authors’ reasoning, but they do not automatically determine the systematic review’s criteria. A consistent synthesis needs hypotheses that are applied across the evidence base.
Source values and reviewer calculations should coexist
Measurement-error evidence makes the provenance problem concrete. COSMIN says reviewers may calculate SEM, SDC or limits of agreement themselves when the correct information is provided and supplies property-specific formulas for doing so.2 The ability to calculate a value does not change its provenance.
Suppose an article reports SEM = 3.4 but no SDC. If the review protocol allows calculation and the correct SEM model and required assumptions are established, the review team may calculate an SDC. The extraction dataset should still contain the source-reported SEM = 3.4 and a separate reviewer-calculated SDC with the formula and inputs. The new value is downstream of the old one.
The same principle applies to transformed correlations, score-direction changes and values estimated from figures. “Mathematically recoverable” is not the same as “reported by the study.” A reader or future reviewer should be able to reproduce the transformation and, if necessary, choose to ignore it and return to the source value.
Missing information needs a vocabulary
Blank cells are dangerous because they hide different states. “Not reported” differs from “not applicable.” A result that exists only in a figure differs from a property that was never measured. An unavailable supplement differs from an unclear statistical model. A value that cannot be extracted differs from a value that an extractor simply missed.
Publication-level missingness should also remain separate from participant-level PROM missingness. A study may report an ICC clearly while also having substantial missing questionnaire responses among participants. The first is a reporting question for the systematic-review data record; the second is part of the primary study’s design and analysis and may later matter for appraisal.
A third situation deserves a flag rather than an automatic conclusion: methods suggest that an analysis was performed but the corresponding result is absent. That may raise a selective-reporting concern, but extraction should record the observation. The formal bias judgment belongs elsewhere.
Multiple reports should be linked without inventing a master source
PROM development and validation often unfold across several reports. A development paper may describe item generation, a later paper may report structural validity, a supplement may contain factor loadings, and a second validation paper may use some of the same participants. PRISMA-COSMIN explicitly uses the language of study reports and asks reviews to cite each included report and describe the included studies.3
The underlying study or sample and the published report therefore need different identifiers in a rigorous extraction system. This prevents the review from counting one cohort twice merely because it appears in two citations. It also lets a result retain the exact report from which it was obtained.
Conflicting sources require more care. A formal corrigendum is designed to correct the publication record. Beyond that, there is no universal rule that a supplement always beats the main text or that the numerical table always beats the narrative. The apparent conflict may disappear when the reviewer discovers different analysis populations, models, rounding conventions or time points.
Extraction and verification do not have one universal configuration
The safest reviewer workflow depends partly on the type of information being collected and the review’s protocol. COSMIN gives explicit guidance for Step 5 PROM characteristics: one reviewer extracts the information and a second checks it. It gives the same recommendation for feasibility and interpretability.2
PRISMA-COSMIN takes a reporting perspective. It asks authors to state how many reviewers collected data from each report, whether multiple reviewers worked independently, how disagreements were handled, how investigators were contacted, whether automation was used and what was done when information across reports was inconsistent.3 Its examples include both one-extractor-plus-verification and independent duplicate extraction. Those examples are not a command that every field must use the same workflow.
For a practical review, more consequential fields deserve stronger checking. PROM/version identity, score target, analysis N, numerical measurement-property results, hypothesis mapping, qualitative content-validity evidence and reviewer calculations can change the synthesis if extracted incorrectly. Bibliographic metadata can often be checked efficiently rather than independently reconstructed twice. That risk-stratified approach is a defensible review-management policy, but it should be labelled as the team’s protocol rather than “the COSMIN rule.”
AI can draft extraction, but the evidence is not PROM-specific
Large language models are increasingly being studied for systematic-review data extraction. A 2025 study within reviews found that an AI-assisted workflow with human verification achieved extraction accuracy comparable to the human-only comparator while reducing extraction time in the study setting.9 That evidence supports interest in assisted workflows, not autonomous extraction of complex psychometric evidence.
The distinction matters here because PROM measurement papers can contain several instruments, several subscales, multiple factor models, different ICC specifications, item-level analyses and several analysis samples in one report. No primary validation located for this guide establishes that an unverified LLM can autonomously extract current COSMIN PROM measurement-property evidence with acceptable reliability.
AI can therefore be useful for bounded assistance: suggesting candidate source passages, drafting bibliographic fields or highlighting places where a statistic may appear. The human reviewer still needs to verify the source, the score identity, the analysis and the value. It should not invent an unreported ICC model, select a preferred CFA, decide that a comparator is a gold standard or silently fill missing psychometric parameters.
PRISMA-COSMIN already anticipates automation. If automation or AI is used during data collection, the review should report how it was used and what validation was performed to understand the risk of incorrect extraction.3
Interpretability and feasibility belong beside the results, not inside them
COSMIN includes interpretability and feasibility in Step 5, but neither is one of the nine measurement properties. Interpretability information can include score distributions, floor and ceiling information, reference values, cut-offs and important-change information. Feasibility can include mode of administration, completion burden, scoring requirements, cost, copyright, availability and other practical characteristics.2
The separation matters because the next methodological question is different. A review may need an MIC to interpret measurement error, but the MIC is not itself a reliability coefficient. Completion time may later influence instrument selection, but a shorter instrument does not become more valid merely because it is easier to administer.
COSMIN specifically recommends extracting MIC information referring to important within-person change even when interpretability is not a primary review aim because it can be needed to evaluate measurement error.2 The extractor should preserve the source terminology and method rather than normalize MIC, MID, responder threshold and similar terms into a single label without checking what the source meant.
Reporting tables are outputs, not the raw extraction architecture
PRISMA-COSMIN 2024 asks systematic reviewers to define what information they collected, describe missing or unclear information, report the data-collection process and present the results of individual measurement-property studies.3 Its Item 26 is especially relevant: for each study, reviewers should present the reported measurement-property result and the rating against predefined quality criteria, and identify results that were computed or estimated rather than reported directly.
That does not mean raw extraction and result rating should occupy the same field. Reporting may display them side by side for readers while the working dataset keeps them conceptually separate. The extraction record owns the source result. The result-rating process adds the later sufficient, insufficient or indeterminate judgment.
The eight COSMIN table templates published in 2025 reinforce that division. Three address PROM characteristics, interpretability and feasibility; two address study characteristics; two organize evaluation of measurement properties; and the eighth presents summary findings and certainty.4 The templates also advise listing versions or subscales separately, placing multiple study reports on separate rows and giving each data point its own cell where possible.
Those are useful reporting principles, but an extraction system has a different job. It must preserve enough granular information to generate several reporting views later. Copying a publication table as the sole raw-data structure can make that difficult when one study contains several analysis models or several linked reports.
Eight failure modes worth checking before synthesis
| Extraction failure | What becomes distorted | Control |
|---|---|---|
| Using the article as the only identity | Separate samples and properties from the same report are conflated. | Link results to the underlying study/sample and property. |
| Using one PROM-version text field for everything | Language, structure and score-target differences become impossible to disentangle. | Preserve structural/scoring identity, language and score separately. |
| Copying only the best CFA model | Structural-validity evidence may be selected by result rather than the scoring structure under review. | Predefine relevant structures/models in the protocol. |
| Recording an ICC without its model | Later reviewers cannot determine what form of reliability was estimated or appraise the method properly. | Capture the reported ICC specification and context. |
| Replacing SEM with reviewer-calculated SDC | Source provenance disappears and the calculation cannot be audited. | Keep both as separately labelled records. |
| Using a blank cell for all missing information | Not reported, not applicable and extractor omission become indistinguishable. | Use explicit missing-information states and explanatory notes. |
| Taking author email as a silent correction | The published value disappears without a formal correction trail. | Preserve publication and author-supplied information separately. |
| Entering feasibility as a score | Extraction begins ranking instruments before the selection framework is applied. | Keep feasibility descriptive and defer comparison. |
The boundary with appraisal and synthesis should be visible
Extraction is the first layer of a longer chain. The result is then appraised for methodological quality, rated against the appropriate measurement-property criteria, summarized with comparable evidence, graded for certainty and eventually used in a fit-for-purpose conclusion about the PROM. Those stages interact, but merging them into one form makes it difficult to tell whether a value came from the study or from the review team’s judgment.
Current COSMIN version 2.0 does not use the old A/B/C labels as its current Step 7 recommendation terminology. It recommends use when high-quality evidence shows all relevant properties are sufficient, recommends against the PROM in its current form when high-quality evidence shows a relevant property is insufficient, and otherwise concludes that a firm recommendation cannot yet be made.2 Pairing extraction directly with an old category label would therefore be both methodologically premature and outdated.
A useful extraction record should survive another reviewer
The most useful test of an extraction system is not whether the original extractor can remember what a field meant next week. It is whether another reviewer can open the record months later and reconstruct the evidence without guessing.
Can they find the source? Can they tell whether a number was reported or calculated? Can they identify the exact score, language and population? Can they distinguish two models from the same article? Can they tell why one correlation was selected for construct validity while another was outside the protocol-defined hypothesis set? Can they see that a missing value was not reported rather than forgotten?
If those questions have clear answers, later appraisal and synthesis become easier to audit. If they do not, the review team will eventually reopen the article and rebuild the missing context manually. Extraction quality is therefore less about creating the largest possible form than preserving the minimum information that prevents later ambiguity.
Related methodological context
Questions about PROM data extraction
What is the basic unit of extraction in a PROM measurement-property review?
There is no single official COSMIN database unit. In practice, the result should remain linked to its report, underlying study/sample, PROM identity, score/subscale, measurement property and relevant analysis. A repeated result record is useful when those relationships would otherwise be lost.
Can different language versions be pooled as the same PROM?
COSMIN says different language versions with the same structure can be considered one version at the Step 5 characteristics stage. That does not mean language can be discarded. Comprehensibility may be evaluated separately by language, and cross-cultural validity or measurement invariance explicitly depends on the compared groups.
Should I extract every alternative factor model in a validation study?
No universal rule requires this. COSMIN emphasizes factor models that represent the structure and scoring algorithm actually used. Additional models should be extracted when they are relevant to the review question or an eligible alternative score structure, not simply because they were reported.
Should a reviewer-calculated value replace the value in the paper?
No. Preserve the source-reported value. Record any reviewer calculation or transformation separately with its inputs, formula and assumptions so that the derivation can be audited or reversed.
How should data supplied by study authors be recorded?
Label the information as author supplied and preserve the correspondence provenance. Author clarification can supplement a publication, but an informal response should not silently erase a different published value or be represented as a formal corrigendum.
Can AI extract PROM measurement-property data automatically?
AI can assist with locating and drafting extraction, but current empirical validation is largely from general systematic-review data extraction rather than autonomous COSMIN PROM extraction. Complex psychometric results still require source-grounded human verification.
References
- Mokkink LB, Elsman EBM, Terwee CB. COSMIN guideline for systematic reviews of patient-reported outcome measures version 2.0. Qual Life Res. 2024;33(11):2929-2939. doi:10.1007/s11136-024-03761-6.
- COSMIN. COSMIN Manual for Systematic Reviews of Patient-Reported Outcome Measures. Version 2.0. Amsterdam UMC; 2024. Official COSMIN manual.
- Elsman EBM, Mokkink LB, Terwee CB, Beaton D, Gagnier JJ, Tricco AC, et al. Guideline for reporting systematic reviews of outcome measurement instruments (OMIs): PRISMA-COSMIN for OMIs 2024. J Clin Epidemiol. 2024;173:111422. doi:10.1016/j.jclinepi.2024.111422.
- Elsman EBM, Boers M, Terwee CB, Beaton D, Abma I, Aiyegbusi OL, Chiarotto A, Haywood K, Matvienko-Sikar K, Mehdipour A, Oosterveer DM, Mokkink LB, Offringa M. Systematic reviews of patient-reported outcome measures (PROMs): table templates for effective communication. Qual Life Res. 2025;34(12):3485-3495. doi:10.1007/s11136-025-04058-y.
- Gagnier JJ, de Arruda GT, Terwee CB, Mokkink LB; Consensus Group. COSMIN reporting guideline for studies on measurement properties of patient reported outcome measures: version 2.0. Qual Life Res. 2025;34(7):1901-1911. doi:10.1007/s11136-025-03950-x.
- de Arruda GT, Terwee CB, Elsman EBM, Avila MA, Gagnier JJ, Mokkink LB; PROM Reporting Group. Explanation & Elaboration document of the COSMIN Reporting Guideline 2.0 for studies on measurement properties of patient-reported outcome measures. Qual Life Res. 2025;34(7):1891-1899. doi:10.1007/s11136-025-03949-4.
- Terwee CB, Prinsen CAC, Chiarotto A, Westerman MJ, Patrick DL, Alonso J, Bouter LM, de Vet HCW, Mokkink LB. COSMIN methodology for evaluating the content validity of patient-reported outcome measures: a Delphi study. Qual Life Res. 2018;27(5):1159-1170. doi:10.1007/s11136-018-1829-0.
- Mokkink LB, de Vet HCW, Prinsen CAC, Patrick DL, Alonso J, Bouter LM, Terwee CB. COSMIN Risk of Bias checklist for systematic reviews of Patient-Reported Outcome Measures. Qual Life Res. 2018;27(5):1171-1179. doi:10.1007/s11136-017-1765-4.
- Gartlehner G, Kugley S, Crotty K, Viswanathan M, Kahwati L, et al. Artificial Intelligence-Assisted Data Extraction With a Large Language Model: A Study Within Reviews. Ann Intern Med. 2025. doi:10.7326/ANNALS-25-00739.