How to Choose a Patient-Reported Outcome Measure for Evidence Synthesis
A “best PROM” does not exist outside a defined decision
Choosing a patient-reported outcome measure after an evidence review is not a contest between reliability coefficients. The decision starts with the construct and population, then asks whether the evidence supports the exact score you intend to use. Only after that can interpretability, burden, language, licensing and other practical considerations tell you which defensible candidate is the better fit for the job.
The shortest defensible answer is this: use the current systematic-review conclusion for each PROM as the starting point, keep measurement-property direction separate from certainty, identify which properties are relevant to the intended application, treat high-certainty insufficiency differently from missing evidence, and make feasibility trade-offs only in the open. If no candidate can support a firm decision, say so.
COSMIN’s current manual makes selection the explicit end point of a systematic review when that is the review’s purpose. Step 7 asks reviewers to move from the evidence overview to a transparent recommendation for the most suitable PROM for the construct and population of interest.12 That is already different from the way PROMs are often chosen in practice, where a familiar questionnaire enters the protocol because a previous trial used it or because one coefficient in one paper looked reassuring.
The evidence needed for a choice is layered. Study methods tell us how much trust to place in individual findings. The property results tell us whether the PROM performed sufficiently or insufficiently. Synthesis tells us what the evidence says across studies. Certainty tells us how confident we should be in that summarized conclusion. Selection comes after those stages. Collapsing them into one “quality score” hides the very distinctions that make the decision defensible.
Current COSMIN recommendation logic is simpler than the old A/B/C shorthand
Earlier COSMIN guidance used Category A, B and C language, and those labels remain common in older reviews. They are not the current Step 7 recommendation framework. Version 2.0 instead describes three situations.12
When high-quality evidence shows that all relevant measurement properties are sufficient, the PROM can be recommended for use. When high-quality evidence shows that a relevant property is insufficient, COSMIN recommends against using the PROM in its current form. In the remaining situations, there is not enough evidence for a firm conclusion and more high-quality research is needed.2
The word “relevant” prevents a mechanical all-nine-properties checklist. COSMIN explicitly notes that internal consistency is not relevant when the PROM is based on a formative model, and criterion validity is not relevant when no gold standard exists.2 The selection question is therefore not “Did this PROM get nine green boxes?” It is “For the score and use we care about, what evidence exists for the properties that matter, and how certain is that evidence?”
There is another nuance that is easy to lose. If no PROM yet has enough evidence for a firm conclusion, COSMIN permits reviewers to identify one or more candidates that have potential for use while further research is conducted. The manual states that such a PROM should have at least very-low-quality evidence for sufficient content validity and should be accompanied by a research agenda.2 This does not create a fourth formal recommendation category. It creates a responsible way to act when the evidence is incomplete.
Content validity can end the argument early, but certainty matters
COSMIN calls content validity the most important measurement property because the items must be relevant to the construct and population, cover the important content, and be comprehensible to respondents.24 A PROM can be internally consistent and reliable while measuring an incomplete or inappropriate slice of the intended construct. Numerical stability does not repair conceptual mismatch.
The current stopping rule is narrower than “any content-validity concern means reject the PROM.” COSMIN says that when there is high-quality evidence that content validity is insufficient, the remaining measurement properties need no longer be evaluated because the PROM should not be recommended.2 That statement gives content validity a genuine non-compensatory role, but only under the specified evidentiary condition.
Low-certainty evidence of insufficiency, indeterminate content evidence and complete absence of studies are not the same state. In an emerging field, reviewers may have only the review team’s ratings or weak development evidence. COSMIN’s current methodology grades content-validity certainty separately for relevance, comprehensiveness and comprehensibility, and allows a total summarized rating when appropriate.2 If reviewers believe one of those aspects should count more heavily than another, the manual recommends not creating the total content-validity summary rather than silently changing the weights.
The 2026 methodological work by Chambers and colleagues is useful here because it shows that applying the content-validity method is still difficult in real reviews. Problems include finding development material, defining the correct PROM version, dealing with poor reporting and separating development evidence from later content-validity studies.5 Selection should therefore respect the content-validity conclusion without pretending that producing that conclusion was always simple or perfectly objective.
Structural validity has a different stopping status
For reflective measures, COSMIN moves from content validity to the internal structure of the PROM. High-quality evidence of insufficient structural validity creates a basic interpretive problem: it may be unclear what the score represents. The current manual says reviewers may decide not to evaluate the later measurement properties in that situation.2
That wording is deliberately different from the content-validity rule. A selection guide should preserve the distinction. High-certainty insufficient structural validity deserves prominent attention, especially when the intended decision depends on that exact total score or subscale, but it should not be converted into an automatic exclusion rule that COSMIN does not specify.
Score identity matters just as much. COSMIN treats subscales separately, and a multidimensional PROM can have one subscale that is well supported and another that is not. The manual recommends making recommendations for each subconstruct separately where relevant.2 Choosing “the questionnaire” without naming the score can therefore be methodologically meaningless.
The raw coefficient is usually the wrong comparison unit
Once measurement-property evidence has been synthesized, selection should work from those summarized conclusions rather than compare raw numbers across PROMs. A reliability estimate of 0.91 in a weak study is not automatically better selection evidence than an estimate of 0.78 supported by a stronger body of evidence. Nor does an alpha of 0.94 make one scale inherently preferable to another with an alpha of 0.82.
The COSMIN criteria provide thresholds for rating individual or summarized measurement-property results, but those criteria belong to the evaluation stage.3 Selection is downstream. Reopening the raw coefficients and ranking their magnitudes discards the methodological-quality and certainty work that the systematic review has already done.
This is especially important for measurement error. COSMIN evaluates whether SDC or limits of agreement are smaller than the minimal important change when that information is available; when MIC is not defined, the result can be indeterminate.3 An indeterminate measurement-error conclusion is not evidence that the PROM performs badly. It tells you that an important link in the interpretation chain is missing.
“Not evaluated” should never masquerade as “insufficient”
A selection matrix often becomes misleading through a simple formatting decision: all empty cells are shown in red. That turns ignorance into negative evidence.
COSMIN already distinguishes insufficient results from indeterminate findings.3 A practical comparison needs to go one step further and distinguish a property that has never been evaluated from a property that was evaluated but could not be interpreted. The former is a research gap. The latter may reflect missing information, a method that does not support a clear result, or conflicting evidence.
This distinction also affects newer versus older instruments. A newer PROM may have less quantitative validation simply because it has existed for fewer years. That should not give it the same label as a long-established PROM with convincing evidence of an inadequate property. Conversely, a long publication history does not protect an older instrument from modern evidence that its content is outdated or insufficient.
Feasibility is not a consolation prize
A common simplification says measurement quality is evaluated first and feasibility is merely a tie-breaker. COSMIN’s current wording is more useful than that. Feasibility and interpretability are not measurement properties, but they are important when selecting a PROM for a particular application and can be decisive in the final choice.2
The manual explicitly mentions time, number of items, access to the PROM, licensing, cost, administration mode, available translations and popularity. It also notes that interpretability, including availability of MIC values, may influence the choice among more than one PROM that can be recommended for use.2
Two principles can coexist. A practical advantage cannot make high-quality insufficient measurement evidence disappear. At the same time, a theoretically elegant PROM that cannot be obtained, cannot be administered in the required language or imposes unrealistic burden may be a poor practical choice for a particular implementation.
Even access is not a simple eligibility rule. COSMIN discusses the situation in which a PROM or manual cannot be obtained and says exclusion is one possible option but explicitly does not recommend that option as the default; alternatives include obtaining the PROM or reviewing what can be established from available information.2 “Unavailable today” should therefore be recorded as a practical constraint, not rewritten as “psychometrically unsuitable.”
Popularity matters, just not in the way people think
“Everyone uses this questionnaire” is a poor argument for validity. Historical uptake does not establish content validity, structural validity, reliability or responsiveness.
Yet it would also be wrong to say popularity is methodologically irrelevant. COSMIN explicitly lists popularity among feasibility information, giving examples such as the number of studies using the PROM and recommendation in core outcome sets.2 Established use may also bring practical advantages: familiar interpretation, existing translations, accumulated reference data or easier comparison with earlier research.
The correct move is to keep this information in its own lane. Historical uptake can influence implementation and comparability. It does not receive a vote on whether an insufficient measurement property should be treated as sufficient.
Generic and condition-specific measures answer different comparison questions
There is no defensible general rule that condition-specific PROMs are always better, or that generic PROMs are preferable because they facilitate comparison. The choice depends on the construct and the purpose of the evidence synthesis.
A generic PROM may be attractive when the review needs broad comparability across conditions or when established population reference data are important. A condition-specific measure may have more relevant content for a narrow clinical question. Those are advantages only if they match the intended construct. A highly responsive condition-specific instrument that measures a different construct from the one required by the review is not a better choice.
The same caution applies to modern item banks and computer-adaptive testing. Their design may improve efficiency or measurement precision in particular contexts, but “newer” is not itself a measurement property. A CAT, static short form and legacy fixed questionnaire should be compared through the evidence for the exact score and administration approach being considered, not through assumptions about technological sophistication.
Multilingual use requires more than a list of translations
Language appears in PROM decisions in several different ways. A translation may exist. Patients may understand the translated items well. Studies may have evaluated cross-cultural validity. Formal measurement-invariance or DIF evidence may exist. These states should not be collapsed.
COSMIN treats comprehensibility as language dependent and recommends language-specific assessment when appropriate.2 Cross-cultural validity and measurement invariance are separately evaluated measurement properties.3 A translated questionnaire is therefore not automatically measurement-invariant.
The opposite overstatement is also risky: absence of a formal invariance study does not by itself establish that every multinational synthesis is invalid. The implication depends on what comparison is being made, how scores are used, and what evidence exists. The responsible selection statement should say which language versions are available, what evidence supports them, and what remains uncertain rather than converting “no invariance study located” into an unsupported universal prohibition.
No numerical weighting scheme is hiding the value judgments for you
PROM selection has the surface appearance of a multi-criteria decision problem. That can make a weighted matrix seem attractive: give reliability 20%, content validity 25%, cost 10%, responsiveness 20%, and let a spreadsheet calculate the winner.
The problem is not arithmetic. It is that the weights encode methodological and value judgments that are rarely justified by empirical PROM-selection evidence. More importantly, some COSMIN conditions are explicitly non-compensatory. High-quality evidence that a relevant property is insufficient supports a recommendation against the PROM in its current form. High-quality insufficient content validity has an even clearer stopping role.2 A cheaper licence or shorter questionnaire cannot logically “earn back” those points.
General multi-criteria decision methods can still teach a useful lesson: make criteria visible, make priorities visible, and examine whether a different legitimate priority would change the decision. That is very different from pretending there is a validated universal set of PROM weights.
A ranking can look more objective than the judgment that produced it. If a review team believes multilingual availability is decisive, it should say so. If respondent burden becomes decisive only after two PROMs have similarly strong evidence, say that too. A transparent sentence is often methodologically stronger than a composite score with two decimal places.
The decision may change when the application changes
Context dependence is not a weakness in the method. It is what prevents the review from claiming a universal “best PROM” when different users need different things.
Case 1: two PROMs can both be recommended, yet one fits the study better
Both can remain scientifically credible PROMs. For this multilingual application, River may be more suitable because feasibility and interpretability fit the project better. In a monolingual project with frequent repeated assessment, Stone might instead be preferred. Neither conclusion requires changing the psychometric evidence.
Case 2: “promising” is not the same as “recommended”
Case 3: different subscales from one PROM can lead to different conclusions
The PROM family cannot be assigned one undifferentiated quality label. The mobility score may remain a credible candidate while the participation score does not. COSMIN explicitly recommends separate conclusions for subconstructs when appropriate.2
Case 4: established use can matter without becoming evidence of validity
The historical instrument does not win because it is popular. The newer one does not win because it is modern. The review should ask which evidence is sufficiently certain, what remains unknown, what the intended application requires, and whether established implementation offers a genuine practical advantage.
Stakeholder input can clarify values without replacing measurement evidence
Core outcome set development provides a useful adjacent model. The COSMIN/COMET guideline combines conceptual considerations, identification of instruments, quality assessment and stakeholder consensus when selecting instruments for a core outcome set.6 That governance process should not be copied wholesale into every systematic review, but it demonstrates why patients, clinicians and trialists may legitimately contribute to decisions about burden, acceptability and implementation.
Ordinary PROM selection after an evidence synthesis does not require a universal Delphi panel. Patient or stakeholder input becomes particularly useful when two measurement-supported candidates differ in practical ways that the evidence review cannot decide for the target users: sensitive content, burden of repeated administration, ease of interpretation or accessibility.
The important separation is again between evidence and preference. Stakeholders should not vote an insufficient property into adequacy. They can help decide which of several defensible options better serves the intended population.
Regulatory “fit for purpose” is a useful comparison, not a COSMIN synonym
FDA’s final 2025 Patient-Focused Drug Development Guidance 3 uses a regulatory fit-for-purpose framework for clinical outcome assessments used in medical-product development.9 It links the clinical outcome assessment to a defined concept of interest, target population and context of use, supported by qualitative and quantitative evidence. That provides a useful parallel to the broader methodological idea that an instrument is suitable for a purpose rather than inherently “good.”
The regulatory framework should remain labelled as FDA guidance. It is not a COSMIN requirement for an ordinary evidence synthesis. The European Medicines Agency’s reflection paper on patient experience data is likewise not a finalized equivalent standard as of this evidence check; the EMA page continues to identify it as a draft reflection paper for which consultation has closed.10
This guide therefore uses “fit for purpose” in its ordinary methodological sense while making the regulatory provenance explicit whenever FDA concepts such as concept of interest or regulatory context of use are discussed.
AI can organize evidence, but it should not quietly choose the PROM
AI tools can help researchers retrieve candidate studies, structure evidence or draft descriptive comparison summaries. Those tasks are fundamentally different from assigning value to conflicting criteria. No validated PROM-specific LLM decision system is needed to justify a cautious rule here: final selection depends on context, relevance, certainty, feasibility and sometimes stakeholder priorities. Those are judgments that must remain reviewable by humans.
An AI system should not infer an unreported measurement property, convert missing evidence into a negative rating, choose numerical weights, or declare a universal “best PROM.” If AI contributes to the evidence-synthesis workflow, the method should make clear what the system did and what humans verified. PRISMA-COSMIN provides the reporting framework for automation used in systematic reviews of outcome measurement instruments.7
Write the final decision as an argument that another reviewer can inspect
A strong PROM-selection statement answers more than “Which instrument won?” It names the defined purpose, the candidate set, the exact version or score selected, the evidence that drove the choice, the main uncertainty, and the practical considerations that mattered.
If one PROM has high-quality evidence that all relevant properties are sufficient and the alternatives do not, the reasoning may be short. If several are recommended for use, the explanation should show why feasibility or interpretability favoured one for this application. If no candidate supports a firm conclusion, the report should not manufacture certainty. COSMIN explicitly encourages a research agenda for promising PROMs when further measurement-property work is needed.2
A useful conclusion might therefore say that a particular score is preferred for a defined population because the current evidence supports all relevant measurement properties with high certainty and the required language and administration mode are feasible; another candidate remains an acceptable alternative but has weaker interpretability for the planned application. Another review might conclude that no firm selection can yet be justified and identify the missing content-validity, reliability or responsiveness evidence required before that judgment should change.
That is not indecision. It is the difference between evidence-based selection and filling every comparison table with a winner.
Related methodological context
Questions about choosing a PROM
Is the PROM with the strongest reliability coefficient usually the best choice?
No. Selection should use the summarized measurement-property conclusion and certainty, not rank raw coefficients. Reliability is also only one part of the evidence needed for a defined PROM score and purpose.
Can good reliability and responsiveness compensate for poor content validity?
Not when there is high-quality evidence that content validity is insufficient. Current COSMIN guidance states that the remaining measurement properties then need no longer be evaluated and the PROM should not be recommended.
Is a newer PROM preferable to an older instrument?
Not automatically. A newer PROM may have stronger contemporary development evidence but a smaller validation base. An older PROM may offer reference data, translations and implementation experience while still needing to satisfy current measurement standards. Age of the instrument is not itself a quality rating.
Does missing cross-cultural validity evidence mean a translated PROM cannot be used?
It means the evidence gap should be visible. Translation availability, comprehensibility, cross-cultural validity and formal measurement invariance are different questions. Whether the gap is decisive depends on what comparisons the intended use requires.
Should I assign numerical weights to content validity, reliability, burden and cost?
There is no universal COSMIN weighting system for such a calculation. Fixed weights can create false precision and allow practical advantages to compensate mathematically for methodological problems that should remain visible. A transparent qualitative deliberation is safer unless a separately justified decision model has been established for the specific context.
What if none of the available PROMs has enough evidence for a firm conclusion?
Do not force a winner. COSMIN allows reviewers to identify PROMs with potential for use when firm conclusions are not yet possible, provided there is at least very-low-certainty evidence for sufficient content validity, and recommends defining a research agenda for the missing measurement evidence.
References
- Mokkink LB, Elsman EBM, Terwee CB. COSMIN guideline for systematic reviews of patient-reported outcome measures version 2.0. Qual Life Res. 2024;33(11):2929-2939. doi:10.1007/s11136-024-03761-6.
- COSMIN. COSMIN Manual for Systematic Reviews of Patient-Reported Outcome Measures. Version 2.0. Amsterdam UMC; 2024. Official COSMIN manual.
- COSMIN. COSMIN Criteria for Good Measurement Properties. Version 2.0. Amsterdam UMC; 2024. Official criteria.
- Terwee CB, Prinsen CAC, Chiarotto A, Westerman MJ, Patrick DL, Alonso J, Bouter LM, de Vet HCW, Mokkink LB. COSMIN methodology for evaluating the content validity of patient-reported outcome measures: a Delphi study. Qual Life Res. 2018;27(5):1159-1170. doi:10.1007/s11136-018-1829-0.
- Chambers C, et al. Assessing content validity: challenges of conducting systematic reviews of patient-reported outcome measures and recommendations to improve the application of COSMIN guidance. Qual Life Res. 2026. doi:10.1007/s11136-026-04261-5.
- Prinsen CAC, Vohra S, Rose MR, Boers M, Tugwell P, Clarke M, Williamson PR, Terwee CB. How to select outcome measurement instruments for outcomes included in a Core Outcome Set: a practical guideline. Trials. 2016;17:449. doi:10.1186/s13063-016-1555-2.
- Elsman EBM, Mokkink LB, Terwee CB, Beaton D, Gagnier JJ, Tricco AC, et al. Guideline for reporting systematic reviews of outcome measurement instruments (OMIs): PRISMA-COSMIN for OMIs 2024. J Clin Epidemiol. 2024;173:111422. doi:10.1016/j.jclinepi.2024.111422.
- Elsman EBM, Boers M, Terwee CB, Beaton D, Abma I, Aiyegbusi OL, Chiarotto A, Haywood K, Matvienko-Sikar K, Mehdipour A, Oosterveer DM, Mokkink LB, Offringa M. Systematic reviews of patient-reported outcome measures (PROMs): table templates for effective communication. Qual Life Res. 2025;34(12):3485-3495. doi:10.1007/s11136-025-04058-y.
- U.S. Food and Drug Administration. Patient-Focused Drug Development: Selecting, Developing, or Modifying Fit-for-Purpose Clinical Outcome Assessments. Guidance for Industry, FDA Staff, and Other Stakeholders. Final guidance. October 2025. FDA final guidance.
- European Medicines Agency. Reflection paper on patient experience data. Draft; consultation closed. EMA primary source.
- Gagnier JJ, de Arruda GT, Terwee CB, Mokkink LB; Consensus Group. COSMIN reporting guideline for studies on measurement properties of patient reported outcome measures: version 2.0. Qual Life Res. 2025;34(7):1901-1911. doi:10.1007/s11136-025-03950-x.