COSMIN Risk of Vias Checklist Navigation Guide

By Dr. Esmaeel Saeedy Robat, Founder of MetaSyn Academy · Published meta-analyst (Nature Human Behaviour, 2026) Evidence checked: 10 August 2026 Current COSMIN Risk of Bias Checklist checked: version 3, dated 27 August 2024
Navigate first, judge second

Find the right COSMIN appraisal pathway before you judge the study

The difficult part of COSMIN risk-of-bias assessment often begins before a reviewer reaches the first standard. A paper may call an analysis “reliability” when it actually reports measurement error, describe comparison with another questionnaire as “criterion validity” when no accepted gold standard exists, or report several measurement properties in one article. The first job is therefore to identify what was actually assessed and decide which current COSMIN pathway applies. The structured navigation record below helps document that decision without reproducing the official checklist or calculating the final methodological-quality rating.12

A navigation record should not function as a COSMIN scoring engine. Its purpose is to define the appraisal unit, route the study to the relevant official COSMIN box, preserve the design and statistical facts needed for assessment, and keep reviewer reasoning visible. The final rating must be made by reviewers using the current official COSMIN Risk of Bias Checklist and manual.

Why a navigation layer is useful

COSMIN distinguishes an article from a study. An article is the published paper. A study, for COSMIN purposes, is an assessment of a measurement property. One article can therefore contain several separate studies: structural validity and internal consistency may be evaluated in the same sample, while reliability and measurement error may be evaluated in another part of the same paper. If a measurement property is reported separately for distinct subgroups and the design or evidence differs, separate appraisal may also be needed.2

That distinction changes the way a review record should be built. A single row labelled with the article citation and one global “quality” judgment is often too coarse. The reviewer needs to know which PROM version, scale or subscale, population or subgroup, measurement property and analysis the methodological judgment refers to. COSMIN’s current manual explicitly describes measurement-property assessment as modular: the properties actually evaluated determine which of the ten Risk of Bias boxes are relevant.12

The problem becomes more difficult because authors do not always use COSMIN terminology. A paper’s label should therefore be treated as a clue rather than the final classification. The current manual advises reviewers to examine the design and analysis and map them to the COSMIN taxonomy. It gives examples of limits-of-agreement analyses described by authors as reliability but treated by COSMIN as measurement error, and comparisons with another PROM described as criterion validity but usually treated as hypotheses testing for construct validity when the comparator is not a genuine gold standard.2

Article-level information

Keep the publication traceable

Record the citation, DOI or identifier, relevant supplement, exact PROM name and version, and the sample described in the article. These fields tell another reviewer where the evidence came from.

Study-level information

Keep the appraisal target precise

Record the measurement property, scale or subscale, population or subgroup and the design or analysis that produced the result. These fields tell another reviewer what is actually being judged.

One article can contain several COSMIN studies A published article branches into several measurement-property assessments. Each assessment is routed to a relevant COSMIN Risk of Bias box and receives its own methodological appraisal. Published article One citation does not imply one appraisal Assessment A Factor structure Scale or subscale + sample Structural validity route Assessment B Repeated measurements PROM version + stable sample Reliability route Assessment C Comparison with another measure Hypotheses + comparator Construct-validity route Separate RoB appraisal using the relevant official box Separate RoB appraisal using the relevant official box Separate RoB appraisal using the relevant official box
Figure 1. The article is the publication container. COSMIN’s methodological appraisal is attached to the study of a measurement property rather than to one undifferentiated article-level score.2

What a navigation record should capture, and what remains with COSMIN

A reproducible navigation record captures the facts needed to reconstruct a later judgment. These include the exact PROM version, relevant scale or subscale, study population, property under investigation, measurement model where relevant, study design, statistical approach, information that could not be verified, and the reasoning of independent reviewers. It should also identify the corresponding COSMIN box so that reviewers can open the current official checklist rather than rely on memory or an old secondary summary.

The navigation record should not reproduce the individual COSMIN standards. That boundary is both methodological and practical. COSMIN updates its materials, and its current checklist is the authority for the wording and available response options. Reproducing every standard locally would create a risk of using outdated wording and could encourage reviewers to treat an unofficial copy as authoritative. The official checklist dated 27 August 2024 identifies itself as version 3 and directs users to the current systematic-review manual for application guidance.1

The navigation record also should not judge whether an observed coefficient or model-fit statistic shows a good measurement property. COSMIN distinguishes standards, which concern study design and preferred statistical methods, from criteria, which are used later to judge whether a measurement-property result is sufficient, insufficient or indeterminate.2 Mixing those two stages is one of the easiest ways to turn critical appraisal into an invalid composite judgment.

Data boundary: this interactive form does not transmit or store the information you enter. Copy and print actions operate in your browser. Reloading or leaving the page can remove unsaved entries, so preserve your completed record in your review-management system.

COSMIN appraisal navigation record

Identify the appraisal unit, route it to the relevant current COSMIN box, document the evidence needed for judgment, then record the independent reviewer decisions after consulting the official checklist.

Use with the official COSMIN materials. The route shown here is an orientation aid based on the COSMIN taxonomy and current manual. It does not reproduce the box standards, calculate “worst score counts,” or determine the final risk-of-bias rating.

Define the publication and appraisal unit

Identify the publication clearly enough that another reviewer can locate the same evidence.
Record the sample included in the relevant analysis, not only the total recruited sample.

Identify what the study actually assessed

Structural validity and internal consistency require special attention to the reflective/formative distinction.
Do not assume the article’s label matches the COSMIN taxonomy. Use the design and analysis to confirm the route.
Choose a measurement-property pathway to display the navigation note.

Preserve the study facts needed for appraisal

Record independent judgments after using the official checklist

COSMIN recommends independent assessment by two reviewers and consensus, with a third reviewer consulted when needed.

How to route the ten current COSMIN pathways

The structured record above names the ten boxes in the current COSMIN Risk of Bias Checklist without reproducing the standards inside them. The notes below explain what kind of study question normally brings each pathway into scope. They are orientation rules, not substitutes for the official checklist.12

Box 1

PROM development

Route here when: the evidence concerns how the instrument was originally developed, including the conceptual basis, generation of content and early testing of the PROM.

Capture before appraisal: development population, construct definition, context, method used to elicit content and information about pilot or cognitive work.

Do not infer: a later psychometric validation paper is not automatically a PROM-development study.

Box 2

Content validity

Route here when: patients, professionals or other relevant participants directly evaluate whether the PROM content is relevant, comprehensive or comprehensible.

Capture before appraisal: who evaluated the content, which aspect was examined, methods used to obtain judgments and the population/context to which those judgments apply.

Do not infer: a high internal-consistency coefficient does not provide evidence of content validity.

Box 3

Structural validity

Route here when: the study evaluates dimensionality or internal structure using factor analysis, IRT, Rasch or another relevant structural model.

Capture before appraisal: measurement model, analysis type, model specification, sample included in that analysis and enough methodological detail to check the relevant official standards.

Model warning: structural-validity analysis is meaningful primarily for reflective measurement models; current COSMIN guidance treats formative-model evidence more cautiously.

Box 4

Internal consistency

Route here when: the study evaluates inter-relatedness among items in a scale or subscale.

Capture before appraisal: whether the score is intended to be reflective and unidimensional, the exact scale/subscale, analysis sample and reliability coefficient or model used.

Do not infer: Cronbach’s alpha for an undifferentiated multidimensional total score should not be treated as self-evidently meaningful.

Box 5

Cross-cultural validity or measurement invariance

Route here when: the study tests whether items or scores behave equivalently across language, cultural or other relevant groups using an invariance or DIF approach.

Capture before appraisal: groups compared, version/language, analytical model, sample per group and whether the design addresses item or parameter equivalence.

Do not infer: translation or cultural adaptation alone is not the same as empirical measurement-invariance testing.

Box 6

Reliability

Route here when: the study evaluates relative consistency or reproducibility across repeated occasions, raters or observers.

Capture before appraisal: evidence of stability where required, interval between assessments, comparability of measurement conditions, raters and the reliability statistic/model.

Do not infer: a simple correlation between repeated scores answers the same question as an agreement-based reliability analysis.

Box 7

Measurement error

Route here when: the study estimates the absolute error around repeated scores rather than only the relative ordering of people.

Capture before appraisal: repeated-measure design, stability, interval and conditions, and the method used to estimate error such as SEM, SDC or limits of agreement where applicable.

Do not infer: authors may call limits of agreement “reliability”; classify the study by what the analysis measures.

Box 8

Criterion validity

Route here when: the target PROM is compared with a defensible gold standard for the same construct.

Capture before appraisal: why the comparator can legitimately function as a gold standard and the analytical method used for the comparison.

Do not infer: another PROM is not automatically a gold standard. COSMIN notes that true gold standards are uncommon for PROMs.

Box 9

Hypotheses testing for construct validity

Route here when: construct validity is examined through relationships with other instruments or differences between groups expected to differ.

Capture before appraisal: comparator, pre-specified expectation where available, direction and magnitude of the expected relationship or group difference, and the analysis used.

Do not infer: statistical significance alone is not the same thing as a well-specified construct-validity hypothesis.

Box 10

Responsiveness

Route here when: the evidence addresses validity of change scores, specifically whether the PROM detects change in the construct over time.

Capture before appraisal: longitudinal design, timing, expected change, change comparator or hypothesis and analytical approach.

Do not infer: a statistically significant pre-post difference by itself does not demonstrate that the instrument measures change validly.

Three decisions deserve extra documentation

Some routes are straightforward. Others depend on judgments that should be written down because a second reviewer may reasonably reach a different conclusion. The first is the measurement model. COSMIN distinguishes reflective instruments, where items are treated as manifestations of an underlying construct, from formative instruments, where components collectively form the construct. Structural validity and internal consistency are primarily meaningful under a reflective model. When the model is unclear, the current manual recommends retaining the evidence and evaluating those properties on the possibility that the model is reflective rather than simply declaring the studies irrelevant.2

The second difficult decision is criterion validity. For most patient-reported constructs there is no unquestioned gold standard. The current COSMIN manual gives limited examples in which a comparator may be treated as a gold standard, such as using a long version when evaluating its shortened form or using the patient-completed PROM as the reference when evaluating a proxy version. A study that compares two different questionnaires measuring a similar construct will usually belong under hypotheses testing for construct validity instead.2

The third is the line between reliability and measurement error. Both can use repeated observations, and the same article may evaluate both. Reliability concerns how well people can be distinguished from one another despite measurement error; measurement error concerns the absolute amount of error in the score itself. The 2020 COSMIN work on reliability and measurement error reinforces the need to match design and analysis to the measurement question rather than to the terminology chosen by the authors.5

Navigation precedes COSMIN judgment A four-stage flow shows the reviewer identifying what the paper assessed, defining the exact appraisal unit, opening the corresponding official COSMIN standards and only then recording independent ratings and consensus. 1 Read the methods What question did the design and analysis actually address? 2 Define the study PROM, version, score, population and measurement property 3 Open COSMIN Use the corresponding current official box and manual guidance 4 Record judgment Rate independently, then document consensus Classification and documentation precede the final methodological judgment The official COSMIN checklist remains the authority for applying the standards and determining the final risk-of-bias rating.
Figure 2. Routing is a separate task from rating. The study should first be classified from its design and analysis; only then should reviewers apply the relevant official COSMIN standards.

Worked example: one paper, two different appraisal records

A repeated-measures validation paper

Imagine an article evaluating a PROM in 140 adults with a chronic condition. Participants complete the instrument twice. The authors report an intraclass correlation coefficient to describe how consistently people retain their relative position, and they also report limits of agreement to describe the absolute spread of score differences. The paper calls the whole section “test-retest reliability.”

A single article-level risk-of-bias row would hide an important distinction. Under the COSMIN taxonomy, the ICC contributes evidence about reliability, while limits of agreement contribute evidence about measurement error. The methodological requirements overlap because both depend on a defensible repeated-measure design, but they answer different measurement questions and correspond to different COSMIN boxes.25

Article record Same citation, PROM version, population and repeated-measure context.
Study record A Reliability pathway: preserve stability, interval, testing conditions and the relative-reliability analysis.
Study record B Measurement-error pathway: preserve the same repeated-measure context plus the method used to quantify absolute error.

This example therefore requires two structured appraisal records. Reviewers would then open the current official Boxes 6 and 7, assess the studies independently and record the final judgments. Neither judgment should be inferred from the ICC, limits of agreement or sample size alone.

How to handle missing information without manufacturing certainty

Risk-of-bias appraisal often reaches a point where the paper does not report enough detail to tell whether a methodological requirement was met. That is different from evidence that the method was definitely poor. The distinction should remain visible in the working record. Use the missing-information field to state exactly what cannot be verified: the interval may be reported without evidence that participants remained stable, an ICC may be given without enough information to understand the model, or an analysis may be named without describing how groups were defined.

Where clarification could materially change the appraisal, contacting study authors can be reasonable. Record the question and response rather than silently filling the gap with an assumption. If the information remains unavailable, use the response options and instructions in the current COSMIN box. The documentation step should not decide how missing reporting changes a particular standard because that decision belongs to the official appraisal method.

Independent reviewers are part of the method

Current COSMIN guidance recommends that two reviewers assess risk of bias independently and come to consensus, consulting a third reviewer when needed. It also suggests that reviewers may benefit from practising on a few studies and discussing their ratings before the full appraisal begins.2 The latter is a useful calibration step, not a fixed requirement that every review must pilot an arbitrary number of papers.

The reason for keeping two independent rationale fields is not bureaucratic. COSMIN appraisal includes judgments about study design, terminology and appropriateness of statistical methods. Earlier empirical work on the original COSMIN checklist found that agreement and reliability varied considerably across items, supporting the practical value of training, discussion and transparent consensus processes.6 A consensus note should therefore preserve why a judgment changed, not merely replace two different answers with a third answer.

Completion check before you close an appraisal record

  • The citation and exact PROM version are identifiable.
  • The score, scale or subscale being assessed is explicit.
  • The relevant population or subgroup is recorded.
  • The measurement property has been classified from the study design and analysis, not only the authors’ label.
  • The corresponding current COSMIN pathway has been checked against the official material.
  • Design and statistical facts are documented separately from the final methodological-quality judgment.
  • Missing information and assumptions are visible.
  • Independent reviewer judgments and the consensus rationale are preserved.
  • No measurement-property performance criterion has been mistaken for a risk-of-bias standard.
  • The record does not claim that completion certifies the study, PROM or review.

A completed navigation record is an audit trail, not a certificate

The value of a structured navigation record lies in traceability. It makes it possible to reconstruct why a paper was routed to a particular COSMIN box, which study facts were considered, where information was missing and how two reviewers arrived at consensus. That is different from automating critical appraisal. COSMIN’s four study-quality ratings (Very good, Adequate, Doubtful and Inadequate) are determined by applying the relevant standards, with the overall box rating based on the lowest applicable standard under the “worst score counts” approach. The current manual also explains that response options were designed so that serious flaws drive the strongest downgrades; methodological strengths in one part of a study do not simply compensate for a fatal weakness elsewhere.23

That final appraisal remains a human methodological judgment. The navigation record helps reviewers reach it in a controlled way and leaves enough information for another member of the review team to understand what was done. It also keeps the appraisal ready for the next COSMIN stages, where study quality is considered alongside the measurement-property results and later contributes to certainty assessment. Those later result-extraction and synthesis tasks are methodologically separate from the risk-of-bias navigation stage.

Get every MetaSyn template free, including this one.

Leave your email and I’ll send this resource as an editable Word file and a printable PDF, plus access to the smart online version. You’ll also get every new template as it’s finished. No noise, just the resources.

References

Scope note: This MetaSyn Academy navigation guide is an independent educational and documentation aid. It is not produced, certified or endorsed by COSMIN. It does not reproduce the individual COSMIN Risk of Bias standards, does not calculate the “worst score counts” result, and does not replace the current official COSMIN Risk of Bias Checklist or systematic-review manual. Users should check the current COSMIN materials before finalizing an appraisal.

  1. COSMIN. COSMIN Risk of Bias Checklist. Version 3. Dated 27 August 2024. Amsterdam UMC. Official COSMIN checklist.
  2. COSMIN. COSMIN Manual for Systematic Reviews of Patient-Reported Outcome Measures. Version 2.0. Amsterdam UMC; 2024. Official COSMIN manual.
  3. Mokkink LB, de Vet HCW, Prinsen CAC, Patrick DL, Alonso J, Bouter LM, Terwee CB. COSMIN Risk of Bias checklist for systematic reviews of Patient-Reported Outcome Measures. Qual Life Res. 2018;27(5):1171-1179. doi:10.1007/s11136-017-1765-4.
  4. Mokkink LB, Elsman EBM, Terwee CB. COSMIN guideline for systematic reviews of patient-reported outcome measures version 2.0. Qual Life Res. 2024;33(11):2929-2939. doi:10.1007/s11136-024-03761-6.
  5. Mokkink LB, Boers M, van der Vleuten CPM, Bouter LM, Alonso J, Patrick DL, de Vet HCW, Terwee CB. COSMIN Risk of Bias tool to assess the quality of studies on reliability or measurement error of outcome measurement instruments: a Delphi study. BMC Med Res Methodol. 2020;20:293. doi:10.1186/s12874-020-01179-5.
  6. Mokkink LB, Terwee CB, Gibbons E, Stratford PW, Alonso J, Patrick DL, Knol DL, Bouter LM, de Vet HCW. Inter-rater agreement and reliability of the COSMIN checklist. BMC Med Res Methodol. 2010;10:82. doi:10.1186/1471-2288-10-82.
  7. Terwee CB, Prinsen CAC, Chiarotto A, Westerman MJ, Patrick DL, Alonso J, Bouter LM, de Vet HCW, Mokkink LB. COSMIN methodology for evaluating the content validity of patient-reported outcome measures: a Delphi study. Qual Life Res. 2018;27(5):1159-1170. doi:10.1007/s11136-018-1829-0.
  8. Elsman EBM, Mokkink LB, Terwee CB, Beaton D, Gagnier JJ, Tricco AC, et al. Guideline for reporting systematic reviews of outcome measurement instruments (OMIs): PRISMA-COSMIN for OMIs 2024. J Clin Epidemiol. 2024;173:111422. doi:10.1016/j.jclinepi.2024.111422.

Questions about using the navigation guide

Does the navigator calculate a COSMIN risk-of-bias rating?

No. It deliberately does not calculate an overall rating. It helps identify the relevant COSMIN pathway and preserve the information needed for appraisal. Reviewers must apply the current official standards themselves and record their resulting judgments.

Should one article receive one COSMIN rating?

Not necessarily. COSMIN distinguishes an article from a study of a measurement property. A single article can contain several measurement-property studies, each requiring appraisal with the corresponding COSMIN box.

What if the terminology used by the authors does not match COSMIN terminology?

Examine the study design and statistical analysis and map the evidence to the COSMIN taxonomy. The current COSMIN manual explicitly warns that authors may use terms such as reliability or criterion validity differently from COSMIN.

What if I cannot tell whether a PROM is reflective or formative?

Record the model as unclear and document why. Current COSMIN guidance recommends retaining structural-validity or internal-consistency evidence when the model is uncertain and evaluating it on the possibility that the model is reflective, while making the uncertainty visible.

Is another PROM a gold standard for criterion validity?

Usually not. True gold standards are uncommon for PROM constructs. A comparison with another instrument often belongs under hypotheses testing for construct validity unless there is a defensible reason that the comparator functions as a gold standard.

Why does the resource ask for two reviewer rationales?

COSMIN recommends independent risk-of-bias assessment by two reviewers followed by consensus, with a third reviewer consulted when needed. Keeping the rationales separate before consensus makes disagreement and its resolution auditable.