PROMs Measurement-Property Data Extraction Form for Systematic Reviews

By Dr. Esmaeel Saeedy Robat, Founder of MetaSyn Academy · Published meta-analyst (Nature Human Behaviour, 2026) Evidence checked: 11 August 2026 Based on current COSMIN systematic-review methodology
Preserve the evidence before interpreting it

A measurement-property result needs more than a number

A reliability coefficient copied into a spreadsheet can look complete while being almost unusable. To interpret it later, the review team may need to know which PROM version produced it, whether it belongs to the total score or a subscale, which sample entered the analysis, what statistical model was used, whether the value came directly from the paper or was calculated by the reviewers, and exactly where the source information can be found. COSMIN’s current systematic-review methodology separates PROM characteristics, study populations, measurement-property results, risk-of-bias judgments, result ratings and later evidence synthesis for this reason.12

This form is designed around that separation. It records source evidence in a reusable structure without deciding whether the study was methodologically sound, whether the reported measurement property was sufficient, or whether the PROM should ultimately be chosen. Those are later judgments. The extraction task is narrower: preserve enough identity, context and provenance that those later judgments can be made without returning to a pile of PDFs to reconstruct what a number meant.

The form does not calculate a COSMIN quality judgment. It also does not select a preferred PROM, choose the most favourable analysis, or turn a review-team calculation into a source-reported result. Where the current COSMIN method requires methodological appraisal, result rating or certainty assessment, the form keeps a boundary rather than quietly performing the next step.

The smallest useful record is a chain, not a cell

COSMIN distinguishes the published article from the measurement-property study it contains. One article may report several samples, several properties and several analyses. The same PROM name may also refer to versions with different structures or scoring algorithms. Multi-dimensional PROMs require attention to the particular score or subscale because their measurement properties may differ.2

For data management, that means the review team needs more than an article row and a result column. A useful record links the report to the underlying study or sample, the PROM identity, the score being evaluated, the measurement property, the relevant analysis and the result. This relational chain is a practical data-management structure rather than an official COSMIN database specification, but it follows the distinctions that COSMIN asks reviewers to preserve. The 2025 COSMIN reporting templates likewise recommend separate rows for PROM versions and subscales and separate rows when multiple study reports contribute evidence.4

Evidence chain for a PROM measurement-property result A linked sequence runs from a source report through study sample, PROM identity, score target, measurement property and analysis to the final extracted result and provenance. Report article / supplement Study sample / population PROM structure / language Score total / subscale Property reliability, validity, etc. Analysis model / comparison Result value + locator A number becomes reusable evidence only when its context remains attached The chain can be flattened for reporting, but the underlying identities should not be lost.
Figure 1. A practical extraction record links a reported result to the report, population, PROM identity, score and analysis that produced it. The relational structure is a practical data-management approach built around current COSMIN distinctions, not an official COSMIN database model.

Version, language and score need separate identities

Current COSMIN guidance contains an important nuance. Different PROM versions can have different measurement properties, so each version or modification is initially treated as a unique PROM. Yet at the Step 5 characteristics stage, different language versions that have the same structure can be considered one version. COSMIN still recommends language-specific treatment where language matters, particularly for comprehensibility, and cross-cultural validity or measurement invariance necessarily preserves the groups being compared.2

The form therefore does not place everything inside one overloaded “version” field. It records structural or scoring version, language or cultural adaptation, administration mode and score target separately. This lets a translated instrument retain its language identity without automatically pretending that translation created a different scoring structure. It also lets a short form or genuinely modified scoring structure remain distinct even when the family name of the PROM stays the same.

Structural identity

What score is actually produced?

Record the item set, version label or structural description, response format and scoring approach when they distinguish the instrument being evaluated.

Language identity

Which language or adaptation was used?

Keep language and cultural context visible even where the structure and scoring are shared with another language version.

Score identity

Total score or which subscale?

Do not attach one psychometric result to a multidimensional instrument generally when it actually belongs to a particular subscale or score.

Analysis identity

Which model or comparison generated the result?

Preserve the statistical analysis that matters for interpretation. Do not select a result merely because it looks more favourable.

Source-reported and reviewer-created values must remain distinguishable

COSMIN allows reviewers to calculate some measurement-error parameters when the article provides the correct inputs. For example, the current manual gives several formulas for deriving SEM, SDC and limits of agreement under different statistical models.2 A calculation made by the review team is still not the same thing as a value reported by the authors.

The form therefore records a source status for every result. A directly reported coefficient remains labelled as source reported. A value supplied by an author after contact remains author supplied. A calculation performed by the review team remains reviewer calculated. A value estimated from a figure remains reconstructed or digitized. This separation also accords with PRISMA-COSMIN, which asks authors to indicate results that were computed or estimated rather than directly reported.3

Preservation rule: never replace the original reported value with a converted or calculated value. Keep the reported value and create a clearly labelled derivative with the formula, inputs and reason for the transformation.

PROM measurement-property extraction record

Build one traceable evidence record. Add repeated result records when one eligible analysis yields several statistics or when several eligible analyses must remain distinguishable.

Browser-local tool. Nothing entered here is transmitted by this form. The form does not permanently store research data. Use the copy or JSON export action before leaving the page. This tool does not replace the current COSMIN manual or your review protocol.

Identify the report and underlying evidence source

Use your review’s stable identifier for the article, supplement, thesis or other report.
Use the same study ID when several reports describe the same underlying investigation where that linkage is defensible.

Define the PROM identity without collapsing distinct dimensions

A language version does not automatically become a different structural version when its structure is unchanged.

Bind the result to the population that generated it

Use the sample entering the measurement-property analysis, not automatically the number originally enrolled.

Identify the measurement property before extracting results

Choose a measurement property to display the extraction packet that normally needs to remain attached to the result.

Content-validity evidence note

Current COSMIN methodology treats this differently from ordinary numerical-result extraction. Highlight the qualitative or quantitative evidence concerning relevance, comprehensiveness and comprehensibility rather than forcing the study into one coefficient field.2

Enter one or more traceable result records

Add a separate result record when values arise from different eligible models, different score targets, different comparisons, different time points or different source statuses and would be misleading if combined. Your protocol should determine which analyses are eligible. COSMIN does not require reviewers to collect every exploratory factor model simply because it appears in a paper; for structural validity, the current manual emphasizes models that represent the structure and scoring used in practice.2

Make missing information and reviewer-created values visible

Participant-level PROM missingness

Publication-level missing information is not the same as missing questionnaire responses within the primary study. Use this section when participant-level missing items or scores affect interpretation of the result.

Keep interpretability and feasibility descriptive

COSMIN treats interpretability and feasibility as useful characteristics rather than measurement properties. They belong beside the measurement-property evidence, not inside a psychometric quality score.2 Record them when relevant to the review question or later interpretation. Do not rank instruments here.

Add interpretability information
Add feasibility information

Record extraction, checking and reconciliation

Current COSMIN Step 5 recommends that one reviewer extract PROM characteristics and that a second reviewer check the information; it gives the same recommendation for feasibility and interpretability.2 PRISMA-COSMIN does not impose one universal extraction configuration for every data item. It requires reviews to report how many reviewers collected data, whether they worked independently, how disagreements were resolved and whether automation or author contact was used.3

What should travel with each property result?

The correct packet differs by measurement property. The purpose of the notes below is not to reproduce COSMIN criteria or Risk of Bias standards. It is to prevent extraction from stripping away information that later reviewers need to understand what the study actually reported.

Structural validity

Keep the tested structure attached to the fit results

Record the score structure or factor model being tested, the analysis framework and the relevant reported results. CFA, EFA, IRT and Rasch studies do not produce interchangeable evidence packets.

For CFA, model-fit results such as CFI, TLI, RMSEA or SRMR may be relevant when reported. For other frameworks, different parameters become important. Do not convert every possible statistic into a mandatory field.

COSMIN specifically advises reviewers to focus on models representing the structure and scoring algorithm actually used in practice. Other tested structures may be evaluated where informative, but they should not automatically crowd the extraction file merely because they appear in the article.2

Internal consistency

Preserve the coefficient and the score it belongs to

Record the scale or subscale, analysis sample and the reported internal-consistency parameter. Current COSMIN guidance explicitly includes Cronbach alpha, omega and relevant IRT/Rasch reliability parameters.2

Do not treat a coefficient as evidence for every score in the PROM. Structural-validity evidence needed for later interpretation should be linked rather than copied into the result itself.

Cross-cultural validity

Keep the groups and comparison model visible

Record the group factor, language or other relevant grouping variable, sample sizes and the method used to evaluate invariance or DIF.

For multiple-group factor analyses, COSMIN recommends extracting which models were compared, equality constraints and changes in model fit. DIF analyses may require coefficients, test statistics, effect information or item parameters depending on the method.2

Reliability

Do not extract an ICC without its model when the paper reports it

Record the type of reliability, repeated-measure context, analysis N, time interval where relevant, statistic, estimate and confidence interval when reported.

Preserve the ICC model, agreement/consistency specification, number of raters or occasions and weighted-kappa information when available because those details affect later methodological appraisal. An extraction record should preserve these details without judging them.

Measurement error

SEM, SDC and limits of agreement are not synonyms

Record the reported parameter, units, repeated-measure context and analysis sample. COSMIN explicitly asks reviewers to extract SEM, SDC or limits of agreement for continuous scores and 95% confidence intervals and variance components when available.2

If the review team calculates a measurement-error parameter from reported inputs, preserve the source inputs and label the new value as reviewer calculated.

Criterion validity

The comparator needs a name and a justification

Record the accepted gold standard and the result of the comparison, such as a correlation or AUC where relevant. COSMIN notes that a true gold standard is uncommon for many PROM constructs.2

Keep the authors’ terminology separate from the review team’s classification so that a paper cannot silently turn another questionnaire into a gold standard.

Construct validity

Extract only results relevant to the review’s hypotheses

Current COSMIN methodology asks the review team to decide which hypotheses are relevant before extracting the corresponding results. Authors may have used different or contradictory hypotheses.2

Preserve the comparator or known group, expected relationship defined for the review and the observed result. Do not indiscriminately harvest every correlation in the paper.

Responsiveness

Keep change-score evidence distinct from treatment effects

Record the longitudinal comparison, time points, relevant change hypothesis or accepted criterion, analysis population and observed change-score relationship.

A pre-post intervention effect is not automatically evidence about responsiveness. Keep intervention effects, meaningful-change thresholds and measurement-property evidence in their proper roles.

Worked example: three numbers from one repeated-measures study

A validation article reports a test-retest ICC of 0.81 with a 95% confidence interval, a standard error of measurement of 3.4 points and no SDC. The same participants completed the same PROM twice. The paper therefore contains evidence relevant to both reliability and measurement error.

Reliability record ICC = 0.81, with the reported ICC model, confidence interval, analysis sample, interval between measurements and source locator.
Error record Source-reported SEM = 3.4 points, with the score units, analysis sample, SEM model if reported and exact source locator.
Optional derived record If the protocol permits a valid SDC calculation and the necessary model information is available, record the derived SDC separately with formula and inputs. Do not replace the SEM.

The example has one report, one underlying sample and one PROM score but more than one measurement-property result. The form keeps those relationships without converting them into a risk-of-bias judgment or a claim that reliability or measurement error is sufficient.

Conflicting reports should remain visible until they are resolved

A main article, supplement, later correction and author correspondence can disagree. There is no defensible universal rule that a supplementary table always outranks the article narrative or that an informal email automatically replaces a published value. First determine whether the values actually refer to the same model, sample, subgroup and time point. A formal correction is different from unpublished clarification. If the conflict remains consequential, preserve the competing values and document how the review team resolved, or failed to resolve, the discrepancy.

PRISMA-COSMIN specifically asks reviews to report decision rules used when selecting data from multiple reports and the steps taken to resolve inconsistencies across reports.3 The form therefore provides report, study and conflict fields instead of forcing one article to become the unquestioned master record.

Do not use blank cells as a missing-data language

An empty field can mean too many things: the article did not report the statistic, the property was not evaluated, the supplement is unavailable, the value is visible only in a graph, the result is genuinely not applicable, or the extractor simply forgot the field. These states should not look identical.

Publication-level missing information must also remain separate from participant-level missing PROM responses. The latter belongs to the primary study’s design and analysis. A third problem, possible selective non-reporting of a measured property, may have implications for appraisal or interpretation. The extraction record should preserve what was observed in the source and leave the formal methodological judgment to the appropriate appraisal stage.

Record-completion check

  • The report and underlying study/sample are identifiable.
  • The PROM structural or scoring version is distinguishable from its language identity.
  • The exact score or subscale is named.
  • The measurement property was classified before entering the result.
  • The analysis sample, not merely total recruitment, accompanies the result when relevant.
  • The statistical model or comparison is recorded when it changes interpretation.
  • The source value and any reviewer-created derivative remain distinguishable.
  • A precise source locator is recorded in a form appropriate to the source.
  • Missing information is coded rather than represented by an unexplained blank.
  • Conflicting sources remain visible until the review team documents a disposition.
  • Content-validity evidence is not forced into an ordinary coefficient field.
  • Interpretability and feasibility remain descriptive rather than becoming hidden selection scores.
  • Extractor/checker information reflects the workflow actually used.
  • No COSMIN risk-of-bias, result-sufficiency, certainty or PROM-selection judgment has been generated automatically.

The purpose of extraction is recoverability

Good extraction reduces the amount of reconstruction required later. A reviewer looking at the final dataset should be able to tell which report contained the information, which sample generated it, which PROM structure and score were evaluated, what analysis produced the result, what the authors reported, what the review team later calculated, and where any uncertainty remains.

That is also why the final record should remain more granular than a publication table. Reporting tables summarize evidence for readers; the extraction record supports the review team before those summaries exist. The 2025 COSMIN table templates recommend that each data point be placed in its own cell where possible and distinguish PROM characteristics, interpretability, feasibility, study characteristics, evaluation of measurement properties and summary findings.4 A granular extraction record preserves the source information needed to populate those later outputs without copying the official table layouts.

Get every MetaSyn template free, including this one.

Leave your email and I’ll send this resource as an editable Word file and a printable PDF, plus access to the smart online version. You’ll also get every new template as it’s finished. No noise, just the resources.

References

Scope note: This MetaSyn Academy form is an independent extraction aid. It is not an official COSMIN form and does not reproduce COSMIN table templates or Risk of Bias checklist items. It records source evidence and review-team provenance. Users should consult the current COSMIN manual for conduct decisions and PRISMA-COSMIN for reporting requirements. Where the form uses a relational record structure, source-status categories or browser export design, these are MetaSyn implementation decisions intended to preserve traceability rather than official COSMIN requirements.

  1. Mokkink LB, Elsman EBM, Terwee CB. COSMIN guideline for systematic reviews of patient-reported outcome measures version 2.0. Qual Life Res. 2024;33(11):2929-2939. doi:10.1007/s11136-024-03761-6.
  2. COSMIN. COSMIN Manual for Systematic Reviews of Patient-Reported Outcome Measures. Version 2.0. Amsterdam UMC; 2024. Official manual.
  3. Elsman EBM, Mokkink LB, Terwee CB, Beaton D, Gagnier JJ, Tricco AC, et al. Guideline for reporting systematic reviews of outcome measurement instruments (OMIs): PRISMA-COSMIN for OMIs 2024. J Clin Epidemiol. 2024;173:111422. doi:10.1016/j.jclinepi.2024.111422.
  4. Elsman EBM, Boers M, Terwee CB, Beaton D, Abma I, Aiyegbusi OL, Chiarotto A, Haywood K, Matvienko-Sikar K, Mehdipour A, Oosterveer DM, Mokkink LB, Offringa M. Systematic reviews of patient-reported outcome measures (PROMs): table templates for effective communication. Qual Life Res. 2025;34(12):3485-3495. doi:10.1007/s11136-025-04058-y.
  5. Gagnier JJ, de Arruda GT, Terwee CB, Mokkink LB; Consensus Group. COSMIN reporting guideline for studies on measurement properties of patient reported outcome measures: version 2.0. Qual Life Res. 2025;34(7):1901-1911. doi:10.1007/s11136-025-03950-x.
  6. de Arruda GT, Terwee CB, Elsman EBM, Avila MA, Gagnier JJ, Mokkink LB; PROM Reporting Group. Explanation & Elaboration document of the COSMIN Reporting Guideline 2.0 for studies on measurement properties of patient-reported outcome measures. Qual Life Res. 2025;34(7):1891-1899. doi:10.1007/s11136-025-03949-4.
  7. Terwee CB, Prinsen CAC, Chiarotto A, Westerman MJ, Patrick DL, Alonso J, Bouter LM, de Vet HCW, Mokkink LB. COSMIN methodology for evaluating the content validity of patient-reported outcome measures: a Delphi study. Qual Life Res. 2018;27(5):1159-1170. doi:10.1007/s11136-018-1829-0.

Questions about the extraction form

Should every translation be entered as a completely different PROM version?

No. COSMIN notes that different language versions with the same structure can be considered one version at the Step 5 characteristics stage. Language identity should still be preserved because comprehensibility and cross-cultural evidence can be language specific.

Should I extract every statistical model reported in a structural-validity paper?

Not automatically. COSMIN advises reviewers to focus on models that represent the structure and scoring algorithm used in practice. Other models can be evaluated when they are relevant, but the review protocol should define which analyses belong in the evidence set rather than collecting every exploratory model indiscriminately.

Can I calculate an SDC when a paper reports only an SEM?

Sometimes. The current COSMIN manual describes several ways measurement-error parameters can be calculated from reported information. The correct formula depends on the available data and model. If the review team calculates a value, keep the original SEM and record the derived SDC separately with the inputs, formula and assumptions.

Should content-validity interview findings be entered like a numerical reliability result?

No. Current COSMIN guidance recommends highlighting relevant qualitative or quantitative evidence about relevance, comprehensiveness and comprehensibility rather than treating content-validity studies as ordinary coefficient extraction. Keep that source evidence separate from later reviewer ratings.

Does PRISMA-COSMIN require two independent reviewers to extract every field?

No. PRISMA-COSMIN is a reporting guideline and asks authors to state how many reviewers collected data, whether they worked independently, how disagreements were resolved, and whether author contact or automation was used. COSMIN Step 5 specifically recommends one extractor and a second checker for PROM characteristics, feasibility and interpretability.

Does completing the form mean a measurement property is sufficient?

No. Extraction records what the source reported. Methodological-quality assessment, rating a measurement-property result, synthesizing results, grading certainty and deciding whether a PROM is suitable are separate later tasks.