PATIENT-REPORTED OUTCOME MEASURE COMPARISON RESOURCE
Patient-reported Outcome Measure Comparison Worksheet
Compare candidate instruments by construct, population, language, respondent burden, content validity, measurement quality, interpretability, feasibility, and intended use without issuing an automatic selection.
Choose for a defined purpose, not for an imaginary universal ranking
A PROM can be well supported and still be the wrong choice for a particular evidence-synthesis purpose. Another can be easier to administer yet have an evidence gap that matters for the intended interpretation. The useful question is therefore not “Which questionnaire has the highest psychometric score?” It is “Which candidate is best supported for the construct, population, score, language and application we actually need?”
Current COSMIN methodology makes that distinction explicit. After the measurement properties of each PROM have been evaluated, the review can conclude that a PROM can be recommended for use when there is high-quality evidence that all relevant measurement properties are sufficient, recommend against its use in the current form when high-quality evidence shows a relevant property is insufficient, or conclude that the evidence is not yet strong enough for a firm conclusion.12 When more than one PROM is suitable, feasibility and interpretability can legitimately help determine which is most appropriate for the particular application.2
This worksheet starts at that later decision point. It assumes that study-level evidence has already been extracted, appraised and synthesized. It does not ask you to re-enter individual ICCs, factor-loading tables or COSMIN Risk of Bias judgments. Instead, it places the summarized evidence for each candidate beside the practical conditions of the decision, keeps result direction separate from certainty, and leaves the final choice to the review team.
This is not a COSMIN scoring calculator. It does not reproduce an official COSMIN table, calculate a proprietary quality score, average sufficient and insufficient findings, or issue an automated recommendation. The current COSMIN Step 7 conclusion should be taken from your completed review or entered by a qualified reviewer. The final selection recorded here remains a human, context-specific decision.
Start with the decision you are actually trying to make
Instrument comparison becomes unreliable when the candidates arrive before the question. A PROM designed for one construct or population cannot become suitable merely because it has an impressive validation literature elsewhere. COSMIN asks review teams to define the construct, population and context of use clearly because content validity and the interpretation of measurement evidence depend on that scope.2 The same PROM family may also contain a total score, several subscales, a short form and different modified versions whose evidence is not interchangeable.
The first section of the worksheet therefore records the target construct, target population, intended application and exact score or domain required. These fields are not cosmetic metadata. They tell you what “relevant measurement property” means later in the comparison. Criterion validity, for example, may not be relevant when no gold standard exists. Internal consistency is not relevant to a formative measurement model. COSMIN explicitly notes both exceptions in its current Step 7 guidance.2
Other relevance decisions can be more context dependent. A review interested in change over time will naturally care about responsiveness and the interpretation of change. A single-language application may place different practical demands on cross-cultural evidence than a multinational programme. Those judgments should be stated rather than hidden. The worksheet provides a decision-context note so reviewers can explain why a particular property or practical condition matters for this application.
Construct and score target
State the construct precisely and identify the total score, subscale or domain you intend to use. A strong subscale does not validate a different score from the same questionnaire.
Who should the evidence represent?
Record the clinical population and setting. Evidence from a different population may introduce indirectness rather than automatically proving that the PROM is unusable.
Separate translation from equivalence
A translation can exist without language-specific comprehensibility or cross-cultural validity evidence. The worksheet keeps these practical and measurement questions distinct.
State how scores will be used
Selection for future longitudinal research, a defined evidence-synthesis recommendation or routine monitoring may create different priorities even when the underlying evidence base is the same.
PROM comparison workspace
Enter summarized evidence from your completed measurement-property review. Compare candidates without calculating a winner.
Browser-local resource. Entries are not transmitted by this worksheet and are not permanently stored. Copy or export your record before leaving the page. The worksheet assumes that COSMIN appraisal, extraction and evidence synthesis have already been completed elsewhere.
What belongs in a candidate profile?
The strongest input is the end product of a rigorous PROM systematic review, not a collection of attractive individual coefficients. COSMIN separates study quality, individual study results, summarized results and certainty for good reason.2 A candidate profile should therefore record the summarized direction of evidence and the certainty attached to that conclusion. If the evidence is inconsistent or indeterminate, say so. If a property has not been evaluated, use that state rather than borrowing the symbol for an indeterminate result.
The worksheet also asks for the current Step 7 conclusion when one exists. That conclusion has a stronger methodological basis than an improvised comparison score. A PROM with high-quality evidence that all relevant measurement properties are sufficient can be recommended for use. A PROM with high-quality evidence that one relevant property is insufficient should be recommended against in its current form. Everything else remains a no-firm-conclusion situation in the current COSMIN framework.12
There is an important nuance for reviews in which no candidate yet reaches a firm conclusion. COSMIN allows reviewers to identify one or more PROMs that have potential for use while more evidence is developed; such a PROM should have at least very-low-certainty evidence for sufficient content validity, and the review should propose a research agenda.2 “Promising, but more evidence is needed” is therefore available in this worksheet as a human documentation label. It is not presented as a fourth official COSMIN category.
Content validity is not just another box to average
Content validity asks whether the PROM content is relevant to the construct and population, sufficiently comprehensive, and comprehensible as intended. COSMIN describes it as the most important measurement property and requires its assessment for every PROM in a systematic review.24
The stopping condition is specific. If there is high-quality evidence that content validity is insufficient, COSMIN states that the remaining properties need no longer be evaluated because the PROM should not be recommended.2 That is different from weak or incomplete evidence. Very-low-certainty evidence suggesting insufficiency is not the same state as high-certainty evidence showing insufficiency. “No study found” is different again.
For that reason, the worksheet never merges direction and certainty into one traffic-light score. It also allows content validity to be recorded as a summary while preserving notes about relevance, comprehensiveness and comprehensibility. COSMIN recommends reporting those three aspects separately and permits a total summarized content-validity rating; when reviewers consider one aspect more important than another, COSMIN recommends not creating that total summary.2
A practical rule for the worksheet: enter the content-validity conclusion your review has already produced. Do not calculate a new “content validity score” here, and do not let burden, cost or popularity erase high-certainty evidence that the content itself is inadequate.
Structural validity deserves a warning, not an invented automatic ban
COSMIN’s current sequence evaluates structural validity after content validity for reflective measures. High-quality evidence of insufficient structural validity creates a serious interpretive problem because it can become unclear what the scale score represents. The manual says reviewers may decide not to continue evaluating later properties in that situation.2 That language matters. It is not identical to the content-validity stopping statement.
The worksheet therefore records structural validity in the same direction-plus-certainty format as other properties and surfaces a reviewer note rather than calculating a formal exclusion. The exact score is also visible. Evidence against a total score should not automatically contaminate a separately supported subscale, and COSMIN recommends making recommendations for subconstructs separately where appropriate.2
Unknown, indeterminate and insufficient are different decisions
The current COSMIN criteria use sufficient, insufficient and indeterminate result ratings, with inconsistent ratings appearing when results are summarized across evidence sources or studies.3 But a comparison worksheet needs another operational state: “not evaluated.” That means the review has no summarized result for the property. It is not a synonym for an indeterminate result.
This distinction prevents a common selection error. A new PROM may have excellent content-validity work but no reliability study yet. That is an evidence gap. A different PROM may have high-certainty evidence that reliability is insufficient. Those two candidates should not receive the same red cell, the same number of points lost, or the same narrative conclusion.
Evidence quantity also needs careful handling. More validation papers can increase the total information and sometimes improve precision, but study count is not a quality scale. COSMIN grades certainty using risk of bias, inconsistency, imprecision and indirectness; it does not treat publication count as a substitute for those domains.2
Feasibility can legitimately determine the final choice
Feasibility is not a measurement property, but COSMIN explicitly treats it as important when selecting a PROM for a specific application. The current manual mentions patient burden, time to administer, number of items, access to the instrument, licensing and costs, administration mode, available translations and popularity, including use in studies or recommendation in core outcome sets.2
That creates two boundaries worth preserving. First, practical convenience does not turn an insufficient measurement property into a sufficient one. Second, practical information is not trivial. When two PROMs are both defensible on measurement grounds, one may be the better choice because the required language exists, administration fits the setting, interpretability is better developed, licensing is workable or respondent burden is more acceptable.
Popularity needs equally careful wording. Historical uptake, availability of reference values, recommendation in a core outcome set and use across earlier studies can be useful for implementation or comparability. None of those facts proves validity. The worksheet therefore records established use under feasibility rather than granting a psychometric bonus.
Language availability is not the same as cross-cultural validity
A translation can make implementation possible while leaving measurement equivalence uncertain. Content comprehensibility may also differ by language; COSMIN recommends assessing comprehensibility separately for language versions where appropriate.2 Cross-cultural validity or measurement invariance is a separate measurement-property question.
The worksheet therefore records language availability and cross-cultural evidence in different places. It does not claim that a translated PROM is automatically validated across cultures, but it also does not impose an unsupported blanket prohibition on any synthesis involving language versions without invariance testing. The review team should document what evidence exists, what the intended comparison requires, and what uncertainty remains.
Worked comparison: two defensible PROMs and one promising candidate
Imagine a review team choosing a fatigue PROM for future multinational trials. Three fictional candidates remain after the measurement-property review.
The worksheet should not assign North 87 points, Harbour 84 and Field 62. North and Harbour are both measurement-supported candidates. If multilingual implementation is a genuine requirement, Harbour may be preferred despite its greater burden because language availability is decisive for this application. Field can remain a promising research candidate rather than being labelled “poor.” A different single-language project might reasonably choose North instead.
This is what context dependence means in practice. It does not make the decision arbitrary. The evidence, constraints and reasons are all documented. What changes is the purpose.
A decision record should show why the alternatives lost
A transparent comparison is not complete when it names the preferred PROM. It should show what the candidate set was, what evidence was decisive, which limitations remained, and why plausible alternatives were not preferred. PRISMA-COSMIN asks systematic reviews of outcome measurement instruments to report their recommendation methods and resulting recommendations transparently.7
That is why the final part of this worksheet asks for trade-offs, evidence gaps, the candidate named in the decision and a written rationale. The decision label is intentionally human-entered. “No firm selection possible” is a legitimate result. So is identifying a promising PROM with an explicit research agenda when current evidence cannot support a firm recommendation.
Before freezing the comparison
- The construct, population, intended application and score target are explicit.
- Each candidate is tied to the correct version and score or subscale.
- The current COSMIN Step 7 conclusion is entered when available rather than recalculated here.
- Result direction and certainty are separate fields.
- “Not evaluated” is not treated as “insufficient.”
- Properties that are genuinely not relevant are labelled as such rather than scored negatively.
- High-certainty insufficient content validity is not offset by feasibility.
- High-certainty insufficient structural validity is visibly discussed rather than hidden inside a total score.
- Feasibility, interpretability, burden and language availability are treated as legitimate context-specific criteria.
- Popularity or historical use is not presented as evidence of validity.
- No raw study-level extraction or COSMIN Risk of Bias scoring has been recreated.
- No automatic numerical ranking has been generated.
- Alternatives, uncertainty and research gaps remain visible.
- The final selection is a written human judgment tied to the defined purpose.
Related methodological context
Get every MetaSyn template free, including this one.
Leave your email and I'll send this resource as an editable Word file and a printable PDF, plus access to the smart online version. You'll also get every new template as it's finished. No noise, just the resources.
Evidence basis and use notes
Independence notice: This is an original MetaSyn Academy decision-support worksheet. It is not an official COSMIN product and is not produced, endorsed or certified by COSMIN or Amsterdam UMC. It does not reproduce the COSMIN Risk of Bias checklist, criteria tables, review-management workbook or 2025 reporting-table layouts. Consult the current COSMIN guideline and manual when formulating formal methodological conclusions.
Questions about using the comparison worksheet
Does the worksheet automatically choose the best PROM?
No. It deliberately has no weighted score, ranking algorithm or winner function. It structures the decision context, summarized evidence, feasibility and uncertainty, then asks the review team to enter a reasoned human conclusion.
What should I enter when a measurement property has never been evaluated?
Use “Not evaluated.” Do not use “Insufficient.” An evidence gap and evidence of poor performance are different methodological states. If the available information is insufficient to interpret a studied result, “Indeterminate” may be the appropriate summarized result instead.
Can a very feasible PROM still be rejected?
Yes. Feasibility is a legitimate selection consideration, but it does not reverse high-quality evidence that a relevant measurement property is insufficient. High-quality insufficient content validity is particularly important because current COSMIN guidance states that the remaining properties then need no longer be evaluated and the PROM should not be recommended.
Should each subscale be entered as a separate candidate?
When the decision concerns a particular subscale or total score, keep that score identity separate. COSMIN treats subscales separately when measurement properties are evaluated, and different subscales from the same PROM can reach different conclusions.
Can I compare a promising PROM that does not yet have a firm COSMIN conclusion?
Yes, but label the uncertainty clearly. Current COSMIN guidance allows reviewers to identify PROMs with potential for use when firm conclusions cannot yet be drawn, provided sufficient content validity has at least very-low-certainty support, and recommends accompanying that choice with a research agenda.
Does a widely used PROM receive extra points?
No points are awarded. COSMIN includes popularity and recommendation in core outcome sets among feasibility information that may matter to a final choice. Historical use can therefore be relevant to comparability or implementation, but it is not evidence that a PROM has good measurement properties.
Should I re-enter ICCs, alpha coefficients or CFA fit indices here?
Usually not. This worksheet is designed to consume the summarized direction and certainty from a completed measurement-property review. Study-level numerical extraction belongs upstream. Return to raw values only when a specific unresolved decision genuinely requires them.
References
- Mokkink LB, Elsman EBM, Terwee CB. COSMIN guideline for systematic reviews of patient-reported outcome measures version 2.0. Qual Life Res. 2024;33(11):2929-2939. doi:10.1007/s11136-024-03761-6.
- COSMIN. COSMIN Manual for Systematic Reviews of Patient-Reported Outcome Measures. Version 2.0. Amsterdam UMC; 2024. Official manual.
- COSMIN. COSMIN Criteria for Good Measurement Properties. Version 2.0. Amsterdam UMC; 2024. Official criteria.
- Terwee CB, Prinsen CAC, Chiarotto A, Westerman MJ, Patrick DL, Alonso J, Bouter LM, de Vet HCW, Mokkink LB. COSMIN methodology for evaluating the content validity of patient-reported outcome measures: a Delphi study. Qual Life Res. 2018;27(5):1159-1170. doi:10.1007/s11136-018-1829-0.
- Chambers C, et al. Assessing content validity: challenges of conducting systematic reviews of patient-reported outcome measures and recommendations to improve the application of COSMIN guidance. Qual Life Res. 2026. doi:10.1007/s11136-026-04261-5.
- Prinsen CAC, Vohra S, Rose MR, Boers M, Tugwell P, Clarke M, Williamson PR, Terwee CB. How to select outcome measurement instruments for outcomes included in a Core Outcome Set: a practical guideline. Trials. 2016;17:449. doi:10.1186/s13063-016-1555-2.
- Elsman EBM, Mokkink LB, Terwee CB, Beaton D, Gagnier JJ, Tricco AC, et al. Guideline for reporting systematic reviews of outcome measurement instruments (OMIs): PRISMA-COSMIN for OMIs 2024. J Clin Epidemiol. 2024;173:111422. doi:10.1016/j.jclinepi.2024.111422.
- Elsman EBM, Boers M, Terwee CB, Beaton D, Abma I, Aiyegbusi OL, Chiarotto A, Haywood K, Matvienko-Sikar K, Mehdipour A, Oosterveer DM, Mokkink LB, Offringa M. Systematic reviews of patient-reported outcome measures (PROMs): table templates for effective communication. Qual Life Res. 2025;34(12):3485-3495. doi:10.1007/s11136-025-04058-y.
- U.S. Food and Drug Administration. Patient-Focused Drug Development: Selecting, Developing, or Modifying Fit-for-Purpose Clinical Outcome Assessments. Guidance for Industry, FDA Staff, and Other Stakeholders. Final guidance. October 2025. FDA guidance.