How to Use the COSMIN Risk of Bias Checklist

MetaSyn Academy guide to use the cosmin risk of bias checklist, illustrating the decisions documented by COSMIN risk of bias checklist navigation
By Dr. Esmaeel Saeedy Robat, Founder of MetaSyn Academy · Published meta-analyst (Nature Human Behaviour, 2026) Evidence checked: 10 August 2026 Critical-appraisal methodology guide
Method before score

Start with the study question, not the paper’s label

To use the COSMIN Risk of Bias Checklist correctly, first determine which measurement property the study actually assessed, then apply the corresponding current COSMIN standards to that study. Do not give the whole article one global rating, do not choose a box only because the authors used a familiar label, and do not mix the methodological-quality judgment with the later question of whether the PROM performed well. COSMIN’s method is modular because each measurement property has its own design requirements and preferred statistical methods.12

The shortest defensible workflow is this: identify the assessment of a measurement property; define the exact PROM, score and population to which it belongs; map the design and analysis to the COSMIN taxonomy; apply the relevant official box independently by two reviewers; use the lowest applicable standard to form the study-quality rating; resolve disagreement; then keep that risk-of-bias judgment separate from the later rating of the measurement-property result and the certainty of the body of evidence.

Why COSMIN risk of bias is not an article-quality score

An ordinary critical-appraisal habit is to think in papers: one paper, one quality judgment. COSMIN asks reviewers to think in measurement-property studies instead. The current manual defines an article as a published paper and a study as an assessment of a measurement property. A single article may contain a structural-validity study, an internal-consistency study and a reliability study, each with a different design and different statistical requirements.2

This distinction is more than terminology. Suppose an article reports a carefully specified confirmatory factor analysis and, elsewhere, a poorly described test-retest analysis in participants whose clinical stability is uncertain. The structural-validity component and the reliability component do not share one methodological fate merely because they appear in the same PDF. The relevant COSMIN boxes need to be applied to the appropriate studies separately. Conversely, when the same design and analysis support the same property across clearly comparable groups, the review team should decide whether separate appraisals are needed rather than multiplying ratings mechanically.

The principle is useful when designing a review-management table. The citation identifies the publication container; the measurement property identifies the appraisal question. The PROM version, scale or subscale and sample identify the evidence to which the judgment applies. That level of precision prevents an “adequate article” label from being carried into every result extracted from it.

Map the study from methods to COSMIN terminology The diagram shows an author’s terminology entering a classification step based on design and analysis before being mapped to a COSMIN measurement property and the corresponding risk-of-bias standards. What the paper calls it “reliability” “criterion validity” “sensitivity to change” Useful clue, not final route Read the methods What design was used? What was compared? Was change evaluated? What statistic/model was used? Which score/sample applies? Classify by methodological question COSMIN route Measurement property Corresponding box Applicable standards Defensible appraisal target Classification error comes before scoring error The wrong box can produce a polished but methodologically irrelevant appraisal.
Figure 1. COSMIN recommends mapping designs and analyses to its taxonomy because terminology used in primary articles may differ from COSMIN terminology.2

Read the methods before deciding which box applies

The current COSMIN checklist contains ten boxes: PROM development, content validity, structural validity, internal consistency, cross-cultural validity or measurement invariance, reliability, measurement error, criterion validity, hypotheses testing for construct validity, and responsiveness.1 The names make the system look simpler than it is. Primary studies frequently use overlapping or historically inconsistent terminology, and the paper’s heading can send a reviewer toward the wrong box.

Consider limits of agreement. Authors may place them under a broad “reliability” heading because they arose from repeated measurements. COSMIN classifies limits of agreement as evidence about measurement error, meaning the absolute error in scores, rather than the relative reliability question of how well people can be distinguished from one another despite error. The correct appraisal route therefore follows what the analysis estimates, not the section title in the paper.25

Criterion validity creates another recurring classification problem. Authors may compare their new questionnaire with an established questionnaire and call the result criterion validity. In most PROM settings, however, neither questionnaire is a true gold standard. COSMIN therefore treats such evidence as hypotheses testing for construct validity unless a defensible gold standard exists. This is why identifying the comparator and the logic behind it matters more than copying the validity label from the abstract.

Responsiveness also needs methodological reading. A paper may report that scores improved significantly after treatment and conclude that the measure is responsive. But change in mean score can result from the intervention or natural history. Responsiveness concerns the validity of the change score as a measure of change in the construct. The appraisal therefore asks how change was evaluated, what external evidence or hypotheses were used and whether the longitudinal design supports the claim.7

A wrong route is not rescued by careful scoring. If an analysis of measurement error is appraised with the reliability box, or a comparison with another PROM is treated as criterion validity without a gold standard, the reviewer can follow the selected standards perfectly and still answer the wrong methodological question.

The measurement model matters before structural validity and internal consistency

Structural validity and internal consistency carry an assumption that is easy to overlook: they are principally meaningful for reflective measurement models. In a reflective model, the construct is treated as causing the item responses, so items representing the same underlying construct are expected to relate to one another. A formative model works differently: distinct components combine to form the construct and do not have to behave as interchangeable indicators of one latent factor.2

That means a low inter-item correlation does not automatically expose a defective formative measure, and factor analysis does not necessarily answer a useful question for it. At the same time, current COSMIN guidance is more nuanced than a simple “formative equals not applicable” rule. If structural validity or internal consistency has been studied for a PROM based on a formative model, reviewers may choose to ignore those studies because their interpretation is problematic, or include and appraise them while explaining that those properties may not be relevant to the PROM.

When the measurement model is unclear, COSMIN recommends retaining available studies of structural validity and internal consistency and evaluating them on the basis that the model might be reflective. This is a good example of why a checklist does not remove methodological judgment. The reviewer may need to infer the intended relationship between items and construct from theory, item content and development work rather than from one explicit sentence in the paper.

Practical consequence: write down how you classified the measurement model and why. If two reviewers differ, the disagreement is substantive. Resolve the model before treating an internal-consistency or structural-validity appraisal as straightforward.

Five methodological families make the ten boxes easier to understand

Memorising ten box names is less useful than understanding the bias mechanisms the standards are trying to control. The boxes can be read as five families of methodological questions. This grouping is explanatory rather than a replacement for COSMIN’s official structure.

Content generation and evaluation

PROM development and content validity

These studies ask whether the content of the instrument was built and evaluated in a way that makes the intended construct, population and context visible. Appraisal therefore depends heavily on who contributed, how content was elicited or tested and whether the methods support judgments about relevance, comprehensiveness and comprehensibility.4

  • Do not replace patient input with statistical evidence.
  • Do not assume a later validation study repairs undocumented development.
  • Keep development evidence distinct from a later content-validity study.
Internal structure

Structural validity and internal consistency

These properties depend on how items represent the construct. Structural validity examines dimensional structure; internal consistency examines item inter-relatedness within the relevant reflective score. The second cannot be interpreted sensibly without knowing what dimensional structure the score is intended to have.

  • Identify the score or subscale actually analysed.
  • Check the measurement model before assuming factor analysis is relevant.
  • Keep model-fit or coefficient thresholds separate from risk-of-bias standards.
Group equivalence

Cross-cultural validity and measurement invariance

The methodological question is whether items or parameters behave equivalently across relevant groups. Translation quality and comprehensibility are important but answer a different question. An invariance study requires an analytical design capable of detecting group-related differences in measurement behaviour.

  • Identify which groups are compared and why.
  • Distinguish linguistic adaptation from empirical invariance testing.
  • Check that the model and sample are appropriate for the intended group comparison.
Repeated measurement

Reliability and measurement error

Both may use repeated observations, but the targets differ. Reliability concerns relative consistency among people; measurement error concerns the absolute error in their scores. Stability of the measured construct, time interval, measurement conditions and choice of analysis can therefore affect the trustworthiness of both.5

  • Do not let the broad term “test-retest” obscure which property is being estimated.
  • Distinguish association from agreement.
  • Ask whether true change could have occurred between assessments.
External relationships

Criterion validity and construct-validity hypotheses

Criterion validity depends on an adequate gold standard. Construct validity usually depends on explicit expectations about relationships with other measures or differences between groups. The credibility of the comparator and the timing of the hypotheses are central to the appraisal.

  • Do not call a familiar comparator a gold standard without justification.
  • Prefer hypotheses specified independently of the observed results.
  • Separate the quality of the study from whether the observed relationship confirms the hypothesis.
Longitudinal validity

Responsiveness

Responsiveness is the longitudinal aspect of validity. The design must allow the reviewer to judge whether the change score behaves as expected when the underlying construct changes, rather than simply showing that scores changed after an intervention.7

  • Distinguish treatment effect from measurement performance.
  • Examine the change hypothesis, comparator or anchor.
  • Match the analysis to the longitudinal validity question.

“Worst score counts” is non-compensatory, but it is not mindless

Each applicable COSMIN standard is rated on a four-point scale: Very good, Adequate, Doubtful or Inadequate. The overall methodological-quality rating for a study is based on the lowest rating among the applicable standards in the relevant box, using the “worst score counts” method.23

The rationale is non-compensatory. A fatal design flaw cannot be averaged away by several well-conducted aspects of the same study. If a reliability study cannot establish that patients were sufficiently stable on the construct, perfect reporting elsewhere does not restore the interpretation of repeated scores as evidence of reliability. COSMIN therefore treats some failures as capable of determining the overall study-quality rating.

But “worst score counts” should not be read as “every small imperfection produces an Inadequate study.” COSMIN explicitly designed the response options with this aggregation rule in mind. For some standards, the lowest available option may be Doubtful or Adequate rather than Inadequate because the flaw is not considered fatal. The strength of the overall downgrade therefore depends on the response options COSMIN assigns to the particular standard, not on an external rule invented by the reviewer.

Nor should reviewers average the item ratings or create their own numerical composite score. The COSMIN system is designed around the meaning of the weakest applicable methodological feature. Turning Very good, Adequate, Doubtful and Inadequate into arbitrary integers and averaging them would change the method into a compensatory score that COSMIN did not design.

COSMIN methodological standards, result criteria and certainty are separate layers The visual separates study methodological quality from the measurement-property result and from certainty of the synthesized evidence, showing that each layer answers a different question. LAYER 1 COSMIN Risk of Bias standards Question: Can the study’s result be trusted given its design and statistical methods? LAYER 2 Criteria for good measurement properties Question: Does the observed measurement-property result meet the relevant performance criterion? LAYER 3 Certainty of the summarized evidence Question: How confident are we in the property-level conclusion across the available studies? A high-quality study can show poor PROM performance; a favourable result can come from a high-risk study.
Figure 2. COSMIN distinguishes methodological standards from criteria for good measurement properties and from the later certainty judgment. The three layers should not be collapsed into a single “quality” score.2

Standards and criteria answer different questions

This distinction is one of the most important safeguards in COSMIN appraisal. Standards concern the design and statistical methods used to evaluate a measurement property. They are the basis of the Risk of Bias assessment. Criteria concern whether the measurement-property result itself is sufficiently good. They are applied after the result has been extracted.2

Consider reliability. A study could use a suitable repeated-measure design, establish stability, use a defensible interval and analyse agreement appropriately. It could therefore have strong methodological quality. The resulting reliability estimate could still be too low to meet COSMIN’s criterion for sufficient reliability. That is not a contradiction. A good study has produced credible evidence that the PROM performs poorly on that property.

The reverse is equally important. A study can report an apparently impressive coefficient while the design has serious risk of bias. The numerical value does not erase the methodological problem. In evidence synthesis, a favourable result from weak methodology should not carry the same evidentiary weight as the same result obtained in a methodologically sound study.

This separation also protects reviewers from misusing numerical thresholds. A model-fit threshold or reliability criterion belongs to the result-rating stage if COSMIN places it there; a sample-size or design requirement belongs to risk-of-bias assessment only when the current relevant standard includes it. The same number should not be penalized at several stages without methodological justification.

Sample size is not handled identically across properties

COSMIN’s current certainty guidance illustrates why property-specific thinking matters. Sample size is incorporated directly into risk-of-bias assessment for content validity, structural validity and cross-cultural validity or measurement invariance. For other measurement properties, total sample size is generally considered later as part of imprecision when certainty in the summarized result is graded.2

Structural and invariance analyses are exceptions because adequate sample size is tied to the stability and credibility of complex model estimates, and results from these studies are not simply made reliable by pooling several undersized analyses. Even here, the manual cautions that sample-size rules are rules of thumb and depend on model complexity, precision requirements and sampling characteristics. The correct lesson is therefore not “memorize one universal N.” It is “use the current property-specific standard and understand what role sample size plays in that design.”

Avoid double penalties. When sample size has already contributed to the risk-of-bias rating for a property, COSMIN says not to downgrade again for imprecision on the same basis during certainty assessment.

Reliability and measurement error look similar until you ask what the estimate means

Both properties often use repeated measurements in people whose underlying status should be sufficiently stable. They nevertheless answer different questions. Reliability is concerned with the proportion of total variation that reflects differences between people rather than measurement error. Measurement error is concerned with the absolute amount of error expressed in the instrument’s measurement units or an equivalent error metric.5

This distinction explains why Pearson or Spearman correlation is not an adequate substitute for an agreement-oriented reliability statistic. Correlation can remain high when all second measurements are systematically shifted upward. The relative ordering of individuals may remain similar even though agreement has changed. An intraclass correlation coefficient can address agreement or consistency depending on the selected model, but the reviewer still needs enough information to understand which model was used and whether it fits the intended reliability question.

For measurement error, the reviewer instead looks for a method that quantifies absolute error, such as a standard error of measurement, smallest detectable change or limits of agreement where appropriate. A paired test that merely asks whether the average score changed is not a direct estimate of measurement error. This is why the methods section, not the authors’ heading, is the safer basis for box selection.

Criterion validity should be used sparingly for PROMs

Criterion validity is conceptually straightforward only when a credible gold standard exists. For many patient-reported constructs there is no direct reference measure that can be treated as the truth. Pain interference, fatigue or perceived functioning cannot ordinarily be reduced to a laboratory value or another questionnaire simply because that comparator is well known.

COSMIN’s current manual gives particular situations in which a reference can reasonably function as a gold standard. For example, a full-length PROM when evaluating a shortened version, or the patient-completed PROM when evaluating a proxy-report version. Outside such situations, comparison with another instrument generally provides evidence for construct validity. Reviewers should therefore ask why the comparator is a gold standard before opening the criterion-validity box.

Construct validity depends on hypotheses, not a fishing expedition

Hypotheses testing evaluates whether scores behave as theory predicts in relation to other measures or known groups. The strongest design logic exists when the expected direction and, where appropriate, magnitude of associations or group differences are formulated before the results are examined. Post-hoc explanations are weaker because the observed data have already influenced the hypothesis.

That does not mean an entire study becomes worthless because a hypothesis is imprecisely written. The relevant COSMIN standards determine the methodological-quality rating. What matters for the reviewer is to distinguish an a priori test of a construct-validity prediction from a set of exploratory correlations followed by a favourable narrative.

The quality of the comparator also matters. A relationship with a measure that itself has uncertain relevance to the target construct is harder to interpret than a relationship with a well-characterized measure selected for a clear theoretical reason. Again, this is a study-design issue. Whether the observed correlation is ultimately large enough to confirm the hypothesis belongs to the later result-rating stage.

Responsiveness is validity of change, not proof that patients improved

Responsiveness is often confused with intervention effectiveness. A questionnaire can show a large pre-post improvement because the intervention works, because the disease changes naturally or because the sample was selected for high baseline scores. None of those facts, by themselves, establishes that the change score is a valid measure of change in the intended construct.

Contemporary COSMIN methodology treats responsiveness as the longitudinal aspect of validity.7 The design therefore needs a defensible expectation about change: a relationship with another measure of change, differences between groups expected to change differently, comparison with a suitable external criterion where one exists, or another design that tests the validity of the change score. A paired t-test or standardized response mean can describe observed change, but neither automatically proves responsiveness.

Missing reporting is uncertainty about the study, not permission to guess

Primary measurement-property studies are not always reported with enough detail to apply every COSMIN standard confidently. The problem may be an unstated ICC model, unclear patient stability, an unspecified factor-analysis procedure or absence of information about when hypotheses were formulated. Reviewers should separate “the method was inadequate” from “the report does not let us establish whether the method was adequate.”

The current COSMIN response categories include Doubtful for situations in which it is unclear whether a standard was met or a preferred method was used.2 That is an important methodological state. It prevents incomplete reporting from being silently promoted to Adequate, while avoiding the equally unjustified assumption that every missing detail proves a fatal flaw.

Author contact may resolve important uncertainty, particularly when one missing methodological fact would materially change the appraisal. If clarification is sought, keep the question, date and response with the review record. Do not retrospectively rewrite the publication as if the information had originally been reported; document the supplementary clarification separately.

Two independent reviewers reduce hidden reasoning

COSMIN recommends independent assessment by two reviewers, followed by consensus and third-reviewer consultation when necessary.2 Independent appraisal matters because many decisions are interpretive. Reviewers may disagree over the property being assessed, whether a comparator can function as a gold standard, the relevance of a measurement model or the seriousness of incomplete reporting.

COSMIN also notes that practising on a few studies and discussing ratings can be helpful before the main appraisal. That is sensible calibration, but the manual does not establish a universal rule that every review must pilot exactly three, five or another fixed number of papers. The number should be sufficient for the team to discover interpretive differences before they propagate across the full evidence base.

Earlier empirical evaluation of the original COSMIN checklist showed that percentage agreement could look relatively high while reliability coefficients for individual items remained modest, partly because of uneven distributions across response categories.6 The practical message is not to chase a particular kappa threshold during every review. It is to invest in shared interpretation, record disagreements and avoid hiding judgment behind a completed table.

Risk of bias contributes to certainty, but the two are not the same calculation

Once the results from relevant studies have been summarized for a measurement property, COSMIN grades certainty in that summarized result. Risk of bias is one of four downgrading factors, alongside unexplained inconsistency, imprecision and indirectness.2 The study-level “worst score counts” procedure therefore should not be confused with the later body-of-evidence certainty judgment.

If several studies contribute to a summarized result, the review team considers the methodological quality of the evidence contributing to that result. In some circumstances, very weak studies may be ignored when better-quality evidence determines the summary, provided the decision follows the COSMIN synthesis method and is transparent. The certainty judgment therefore depends on the evidence actually supporting the summarized result, not on mechanically choosing the single worst study in the entire review.

This difference matters when readers interpret a review. “One study was Inadequate” does not automatically translate into “the certainty is Very low.” Conversely, a property supported by one small study with otherwise good methods may still face imprecision. The reviewer needs to follow the current COSMIN certainty rules rather than invent a direct conversion table from study-quality ratings to GRADE levels.

Reporting the appraisal is a separate responsibility

PRISMA-COSMIN for OMIs 2024 governs reporting of systematic reviews of outcome measurement instruments. It asks authors to describe the methods used to assess risk of bias and to report the results transparently.8 Reporting the checklist name without explaining who applied it, how disagreements were handled or how the results entered the synthesis leaves readers unable to reconstruct the review process.

Reporting guidance should not be mistaken for conduct guidance. PRISMA-COSMIN can help authors report what they did; it does not replace the COSMIN systematic-review manual or the Risk of Bias Checklist. Likewise, good reporting does not prove that the underlying appraisal was methodologically correct. A review can describe a wrong procedure transparently.

At minimum, authors should make the COSMIN version clear, describe the reviewer process, report property-level study-quality judgments in a form readers can trace to the included studies and explain how those judgments influenced synthesis and certainty. If the review adapted the method, the adaptation and rationale should be visible.

Six difficult cases that expose common reasoning errors

Case 1

The authors report “criterion validity” against another questionnaire. Do not accept the label automatically. Ask whether the comparator is a defensible gold standard. If it is simply another measure of a related construct, the evidence will usually belong under hypotheses testing for construct validity.

Case 2

A multidimensional PROM has one Cronbach alpha for the total score. Before appraising the internal-consistency study, establish whether that total score is intended to be reflective and unidimensional. A large coefficient does not solve a dimensionality problem.

Case 3

A test-retest paper reports Pearson correlation and no agreement statistic. Recognize that correlation alone does not evaluate agreement adequately. Use the official reliability standards to judge the study; do not decide the final rating from this summary statement alone.

Case 4

A translated PROM underwent careful forward/back translation but no invariance analysis. The adaptation process may provide content or comprehensibility evidence, but it is not automatically a study of cross-cultural validity or measurement invariance in the COSMIN statistical sense.

Case 5

The instrument improved after surgery and the authors call it responsive. Ask what design tests validity of the change score. A pre-post effect demonstrates change in the sample; it does not alone demonstrate that the PROM detects change in the construct validly.

Case 6

The article never states whether the scale is reflective or formative. Do not guess silently. Review the conceptual model and item structure, record the uncertainty and apply the current COSMIN guidance for unclear measurement models.

Errors that can survive a polished review table

Appraisal error Why it matters Better practice
One quality rating for the whole article Different measurement-property studies can have different designs and different risk of bias. Appraise the relevant study of each measurement property with its corresponding COSMIN box.
Selecting the box from the terminology in the abstract Authors may use measurement-property terms differently from COSMIN. Classify the study from the design, comparator and analysis.
Averaging COSMIN item ratings This changes a non-compensatory system into an unsupported numerical score. Use the current COSMIN “worst score counts” approach.
Using performance thresholds as RoB standards Study methodological quality becomes confused with whether the PROM performed well. Separate standards from criteria for good measurement properties.
Treating every missing detail as proof of bad methods Poor reporting and poor conduct are related but not identical. Use the current response options and document unresolved information.
Calling another PROM a gold standard without justification The wrong validity box may be applied. Use criterion validity only when a defensible gold standard exists.
Assuming translation proves invariance Linguistic adaptation and statistical equivalence answer different questions. Distinguish content/adaptation evidence from invariance or DIF testing.
Turning a significant pre-post change into evidence of responsiveness Treatment effect is confused with validity of the change score. Evaluate the longitudinal validity design and hypotheses.

When specialist input is worth seeking

A checklist can make the appraisal structure explicit, but it cannot turn every review team into specialists in factor analysis, IRT, reliability theory, measurement invariance and longitudinal validation. Some studies are simple enough to appraise after careful reading of the COSMIN manual. Others contain complex models, unusual sampling designs or poorly reported analyses whose appropriateness cannot be judged from a statistical label alone.

Specialist input is especially useful when a structural model is complex, the reflective/formative classification is disputed, an IRT or Rasch analysis uses unfamiliar assumptions, an ICC model is not clearly interpretable, measurement-invariance methods are technically difficult or an unusual comparator is proposed as a gold standard. The purpose is not to outsource the reviewer’s judgment. It is to make sure the judgment is based on an accurate understanding of the method.

The same caution applies to automated appraisal. Administrative tools can help track studies, preserve quotations, flag missing fields and compare reviewer records. They should not be allowed to infer a final COSMIN rating merely from a handful of extracted statistics. The applicable standard may depend on study context, measurement model and methodological detail that cannot safely be reduced to a single threshold.

Questions reviewers often ask

What is the current COSMIN Risk of Bias Checklist version for PROM systematic reviews?

The current official checklist available from COSMIN is dated 27 August 2024 and instructs users to cite it as the COSMIN Risk of Bias Checklist version 3. It accompanies the 2024 COSMIN systematic-review guideline version 2.0.

Does “worst score counts” mean every minor flaw makes a study inadequate?

No. The overall rating is determined by the lowest applicable standard, but COSMIN designed the response options so that only sufficiently serious flaws can receive the strongest negative rating. Some standards do not offer Inadequate as the lowest response.

Can a Very good COSMIN study show that a PROM is poor?

Yes. Risk of bias concerns whether the study methods are trustworthy. A methodologically strong study can produce convincing evidence that the PROM has an insufficient measurement property.

Should structural validity always be appraised before internal consistency?

The interpretability of internal consistency depends on the dimensional structure of a reflective score. Reviewers therefore need evidence about structural validity or unidimensionality before treating an internal-consistency coefficient as meaningful for that score.

Is poor reporting automatically an Inadequate COSMIN rating?

No. COSMIN distinguishes situations in which a standard clearly was not met from situations in which the report does not allow the reviewer to establish whether it was met. Apply the response options in the relevant current box rather than converting every reporting gap into the same rating.

Does PRISMA-COSMIN replace the COSMIN Risk of Bias Checklist?

No. PRISMA-COSMIN for OMIs is a reporting guideline for systematic reviews of outcome measurement instruments. COSMIN’s Risk of Bias Checklist assesses methodological quality of included measurement-property studies.

References

  1. COSMIN. COSMIN Risk of Bias Checklist. Version 3. Dated 27 August 2024. Amsterdam UMC. Official COSMIN checklist.
  2. COSMIN. COSMIN Manual for Systematic Reviews of Patient-Reported Outcome Measures. Version 2.0. Amsterdam UMC; 2024. Official COSMIN manual.
  3. Mokkink LB, de Vet HCW, Prinsen CAC, Patrick DL, Alonso J, Bouter LM, Terwee CB. COSMIN Risk of Bias checklist for systematic reviews of Patient-Reported Outcome Measures. Qual Life Res. 2018;27(5):1171-1179. doi:10.1007/s11136-017-1765-4.
  4. Terwee CB, Prinsen CAC, Chiarotto A, Westerman MJ, Patrick DL, Alonso J, Bouter LM, de Vet HCW, Mokkink LB. COSMIN methodology for evaluating the content validity of patient-reported outcome measures: a Delphi study. Qual Life Res. 2018;27(5):1159-1170. doi:10.1007/s11136-018-1829-0.
  5. Mokkink LB, Boers M, van der Vleuten CPM, Bouter LM, Alonso J, Patrick DL, de Vet HCW, Terwee CB. COSMIN Risk of Bias tool to assess the quality of studies on reliability or measurement error of outcome measurement instruments: a Delphi study. BMC Med Res Methodol. 2020;20:293. doi:10.1186/s12874-020-01179-5.
  6. Mokkink LB, Terwee CB, Gibbons E, Stratford PW, Alonso J, Patrick DL, Knol DL, Bouter LM, de Vet HCW. Inter-rater agreement and reliability of the COSMIN checklist. BMC Med Res Methodol. 2010;10:82. doi:10.1186/1471-2288-10-82.
  7. Mokkink L, Terwee C, de Vet H. Key concepts in clinical epidemiology: responsiveness, the longitudinal aspect of validity. J Clin Epidemiol. 2021;140:159-162. PubMed record.
  8. Elsman EBM, Mokkink LB, Terwee CB, Beaton D, Gagnier JJ, Tricco AC, et al. Guideline for reporting systematic reviews of outcome measurement instruments (OMIs): PRISMA-COSMIN for OMIs 2024. J Clin Epidemiol. 2024;173:111422. doi:10.1016/j.jclinepi.2024.111422.
  9. Mokkink LB, Elsman EBM, Terwee CB. COSMIN guideline for systematic reviews of patient-reported outcome measures version 2.0. Qual Life Res. 2024;33(11):2929-2939. doi:10.1007/s11136-024-03761-6.

Similar Posts

Leave a Reply

Your email address will not be published. Required fields are marked *