Full-Text Exclusion Reasons in Systematic Reviews: How to Decide and Report

Guide to full-text exclusion reasons showing assessment evidence, one primary reason and a separate non-retrieval status

Design a usable reason hierarchy, explicitly separate workflow status from true eligibility, and log full-text exclusions without double counting.

Short Answer: Derive your exclusion reason categories directly from your protocol. Define them operationally, establish a consistent hierarchy for reports that fail multiple criteria, and preserve the exact text citation supporting each decision. Never merge reports not retrieved, ongoing studies, or records awaiting classification into completed full-text exclusion counts.

Full-text exclusion is a methodological claim

When you exclude a retrieved report, you assert that its evidence falls outside your review’s scientific boundaries. That assertion alters the final population available for synthesis and demands the exact same rigor applied to your statistical modeling. A label like “wrong population” is not self-validating. It is defensible only if the population rule was strictly defined in the protocol, the report was objectively assessed against that rule, the relevant text can be cited, and the outcome aligns with the broader selection ledger. PRISMA 2020 mandates transparent reporting of exclusions and reasons, but it does not dictate a universal taxonomy for every review.1,2

Weak exclusion logs damage a review in three critical ways. First, vague labels hide whether you applied the protocol consistently. Second, counting every failed criterion artificially inflates the exclusion total and breaks the flow diagram’s arithmetic. Third, treating a retrieval failure or an unresolved author query as an eligibility exclusion invents a judgment you never actually made. A defensible system separates the decision evidence, the reporting category, and the workflow status while preserving their relationships.3,4

Why the exclusion log matters

The full-text appraisal stage is where a potentially eligible record converts into a permanent, documented inclusion or exclusion. A brief label inside a PRISMA flow diagram box cannot preserve the report citation, the specific criterion, the text evidence, the reviewer identity, or the conflict history that produced the decision. The detailed exclusion log provides that audit trace and guarantees your category totals are accurate.

This log acts as the bridge between your protocol and the final manuscript. The protocol defines what you intended to include; the decision record proves how you applied that definition; the flow diagram reports the aggregate result; and the manuscript cites borderline studies whose exclusion might otherwise provoke reader skepticism. None of these components can substitute for the others. A simple summary table cannot prove which report was counted, while a massive internal spreadsheet without a controlled roll-up cannot prove the published totals are arithmetically coherent.

Build the taxonomy from eligibility criteria

Population or phenomenon

Use only when the assessed study population definitively fails the operational protocol rule.

Intervention or exposure

Define active component, dose, duration, or specific exposure boundaries before screening begins.

Comparator

Use when a specific comparator is a strict eligibility requirement, not simply because one is absent from the reporting.

Outcome or concept

Exclude based on outcomes only if the protocol made that outcome necessary for study eligibility (rather than just extraction).

Study design

Link the decision to your methodological design definition rather than relying on an imprecise journal publication label.

Report characteristics

Enforce report type, publication date, language, or peer-review status only when explicitly pre-specified.

Generate your exclusion categories directly from your operational eligibility criteria rather than importing a generic list from another paper. Each category requires a clear definition, inclusion/exclusion examples, defined boundaries with neighboring categories, and a direct link to the active criterion version. If you define a study design by its allocation method, reviewers must record the evidence regarding allocation rather than simply inferring the design from the article’s title. If your protocol permits data from a mixed population, do not apply a “wrong population” label until the specific subgroup extraction rule has been tested.3,5

Category granularity must serve interpretation. A single catch-all “Wrong PICO” category is too broad to demonstrate where the evidence boundary actually operated; conversely, dozens of hyper-specific labels produce sparse, unstable totals that readers cannot synthesize. A robust taxonomy reports the major dimensions publicly while retaining exact criterion clauses and free-text notes internally. Because a single report can legitimately fail several criteria, categories should be mutually interpretable even when they are not logically mutually exclusive.

From protocol criterion to reportable exclusion reason A criterion is operationalized, assessed against report evidence, retained as a decision record and rolled into one reportable primary reason. Protocol rule versioned criterion Report evidence page, table or source Decision event reviewer + resolution Primary reason one reportable count Secondary failed criteria remain in the audit record
Figure 1. Reporting compresses the decision, but the internal audit record retains its evidential basis.

Use one reportable primary reason without losing detail

A single report frequently fails several eligibility criteria. If every failure enters the flow diagram independently, the reason total will exceed the actual number of excluded reports. The standard, defensible solution is to assign one reportable primary reason according to a pre-specified hierarchy, while retaining all other failed criteria as secondary detail in your log. This is a practical Academy implementation recommendation; it is not a rigid PRISMA mandate.

  1. Order the criteria in a hierarchy that logically fits the review question and screening sequence.
  2. Apply the first failed criterion encountered in that hierarchy as the reportable primary reason.
  3. Record all other failed criteria, along with supporting text citations, in the detailed decision log.
  4. Compare the summed primary reasons with the total number of reports excluded after assessment to ensure perfect reconciliation.

The priority rule may follow the strict screening sequence or a scientifically meaningful hierarchy. A sequence-based rule assigns the first failed criterion encountered; this is easy for reviewers to reproduce but makes totals highly sensitive to the chosen question order. A substantive hierarchy privileges fundamental boundaries—such as study design or core population—but requires clearer judgment rules. Whichever approach you select, establish it before tackling the full-text workload, pilot it on multi-failure cases, and apply it consistently. Secondary failures must remain queryable so that reporting compression does not destroy valuable methodological data.

Reporting one primary reason per excluded report is simply an implementation mechanism designed to yield mutually exclusive aggregate counts. It does not imply the report failed only one condition, and the manuscript should avoid suggesting otherwise. If a team opts for a different reporting format, it must explicitly demonstrate how the reason counts map to the number of reports excluded. A reason table where column totals exceed the exclusion box can be analytically insightful, but it must be clearly labeled as multiple responses rather than presented as a reconciled PRISMA sum.

Separate decision outcomes from workflow status

Category Methodological Meaning Where it belongs
Excluded after full-text assessment The retrieved report was assessed against criteria and failed at least one. Full-text exclusion reasons.
Report not retrieved The document was sought but not obtained; eligibility was never assessed. Reports not retrieved; preserve retrieval attempts in logs.
Ongoing study The study is underway and cannot yet provide a completed eligible report. Separate ongoing-study record where relevant.
Awaiting classification Information or author clarification is still insufficient for a final decision. Separate unresolved-status record and follow-up plan.

“Not retrieved” is a critical distinction because the absence of a document is not evidence of scientific ineligibility. Record retrieval attempts, search routes, dates, and unresolved access barriers in your internal logs, while the PRISMA diagram correctly reports the paper as sought but not retrieved. Similarly, an ongoing study might be highly relevant to the review question but simply unable to contribute a completed report before the review’s cut-off date. “Awaiting classification” indicates the team deferred judgment pending further information; silently converting this status into an exclusion hides uncertainty and biases the reported evidence base.1,2

Workflow status must be time-stamped because it frequently changes. A report not retrieved at the first attempt may later be obtained via interlibrary loan; an ongoing study may publish its results; an unresolved report may be classified after an author responds. Preserve the earlier event and append the later event so both the final state and the pathway remain fully intelligible. This practice is particularly vital for living or updated reviews, where unexplained status replacements make successive flow diagrams appear inconsistent.

Do not confuse companion reports with duplicate studies

Multiple reports frequently describe a single study. Link them to a central study identifier and determine how each report contributes to eligibility and synthesis. A companion report is not automatically an ineligible duplicate study. If one report supplies essential baseline eligibility information while another supplies long-term outcomes, keep both linked to the included study and ensure you report both the study and report totals separately.

A companion report may lack relevance to a specific outcome yet still belong to an included study. Labeling it “duplicate” conceals valuable methodological details, adverse-event data, or follow-up timelines, ultimately producing incorrect report totals. The review database must preserve the report as an entity, its relationship to the study, and its specific role in the synthesis. If the report is ultimately unused, the record should distinguish its non-use from the complete exclusion of the underlying study. When linkage remains uncertain, document the evidence suggesting a relationship and keep the assignment open to revision.3

Prepare the exclusions readers are likely to question

PRISMA 2020 Item 16b explicitly asks reviewers to cite studies that might appear to meet inclusion criteria but were ultimately excluded, and to explain why. This requirement does not necessitate publishing the entire internal exclusion log. Candidate examples include high-profile trials, studies sitting precisely on the eligibility boundary, and well-known reports that readers reasonably expect to see included. Provide a concise, criterion-linked explanation for each.

Selection for this reportable list should be reader-centered rather than driven by reputation. Include a study when its apparent fit creates a realistic risk that readers will suspect an accidental omission or inconsistent rule application. The explanation should identify the decisive criterion and the relevant fact without reproducing confidential author correspondence or adding an argumentative critique of the study’s design. A citation followed by “wrong design” is significantly less informative than a short, direct statement noting that the allocation method did not meet the protocol’s definition of an eligible randomized trial.2

The internal log must remain complete even when only selected exclusions appear in the main article or supplementary file. Journal space limits never justify deleting the evidence, reviewer history, or secondary failed criteria. A durable, comprehensive archive allows peer reviewers, editors, and future update teams to trace questioned studies efficiently. Access to this log can be governed appropriately without exposing personal data or copyrighted full texts.

Reconcile logs before preparing the flow diagram

  • Verify every excluded report possesses exactly one reportable primary reason.
  • Verify the primary-reason total equals reports excluded after full-text assessment.
  • Verify not-retrieved reports are kept entirely outside the exclusion total.
  • Verify ongoing and awaiting-classification records are not silently treated as exclusions.
  • Verify companion reports successfully link to study identifiers.
  • Verify reason labels match the approved protocol version or a formally documented amendment.
Do not repair a mismatch by editing the final number first. Trace the underlying reports, status labels, and reason assignments, then correct the primary source record.

Reconciliation proceeds directly from report identifiers. For every report assessed for eligibility, confirm exactly one terminal state for the reporting period: included, excluded, ongoing, awaiting classification, or another explicitly defined state. For excluded reports, verify the presence of one reportable primary reason. Sum these states and compare them against the retrieval and assessment boxes, then group included reports into studies. This sequence detects duplicate rows, missing outcomes, and inappropriate study-level counting long before they are hidden inside an aggregate flow diagram.

Stratified checks reveal errors that grand totals conceal. Reconcile database and register routes separately from other identification methods when your selected PRISMA template reports them separately. Examine reason distributions by reviewer, time period, and criterion version for unusual shifts. A sudden spike in “other” exclusions likely indicates the taxonomy no longer reflects the protocol; a reason utilized by only one reviewer indicates an interpretive divergence. Such patterns serve as prompts for audit, not automatic evidence of misconduct. Dedicated diagram software certainly improves visual presentation and basic arithmetic, but it still relies entirely on valid source counts and correctly classified units.8

Manage changed reasons and eligibility amendments

Decisions change after conflict resolution, receipt of a companion report, or the formal amendment of an eligibility criterion. The log must preserve the original decision, original reason, reviewer, and timestamp, then append the resolved outcome and its basis. Overwriting history prevents the team from determining whether a discrepancy arose from reviewer misinterpretation, new evidence, or a changed rule. Versioned events also support transparent review updates, where a study excluded previously may suddenly become eligible under a new question or expanded time boundary.

When an amendment alters the taxonomy, create an explicit mapping between old and new categories and decide whether earlier full texts require reassessment. Relabeling old decisions solely to make the final table look uniform obscures a substantive methodological change. If categories are combined solely for publication simplicity, retain the original granular codes internally and document the aggregation step. The manuscript must disclose material post-protocol changes and their effect on study selection.

Use automation for consistency checks, not unsupported judgment

Software can force a reviewer to enter a reason when “excluded” is selected, prevent multiple primary reasons, flag reasons inconsistent with workflow status, and recalculate totals dynamically. These controls reduce clerical errors without usurping scientific eligibility decisions. Free-text classification algorithms or language-model suggestions may help reviewers locate a criterion faster, but the accountable human reviewer must verify the report evidence and lock in the final reason. Record any automated transformation applied to historical reason text and retain the source text so the classification remains auditable.

Automated quality checks should test invariants: an excluded report has one primary reason; a non-retrieved report has no full-text exclusion reason; every report has a stable identifier; every reason points to an active criterion version; and the primary-reason sum matches the excluded report count. Passing these checks establishes structural coherence, not scientific validity. Human review remains completely necessary to determine whether the criterion itself is appropriate and whether the evidence was interpreted correctly. The same principle applies upstream: evidence regarding single versus duplicate screening and large-review screening practices informs risk control, but it does not validate an exclusion reason after the fact.9,10

Write reasons that remain interpretable outside the review database

A stated reason should identify the failed boundary without requiring the reader to decode local project abbreviations. “P3 fail” may be efficient for an internal screening interface, but the exported record should map it directly to the full versioned criterion. Avoid evaluative, subjective language such as “poor study” unless methodological quality was an explicit eligibility condition and the review design explicitly justifies its use. Risk of bias is typically assessed after inclusion and should not be retrofitted as an exclusion reason. Likewise, “no usable data” must carefully distinguish the absence of an eligible outcome from eligible data that are merely difficult to extract or synthesize; those are fundamentally different methodological problems.

Concision remains compatible with specificity. A reportable statement can name the exact criterion and the decisive fact in a single sentence, while the internal record points to the specific page or table and preserves reviewer discussion. This layered approach supports transparent publication, efficient peer review, and future updating without turning the PRISMA flow diagram into an overloaded decision log.

Explore Related Methodology and Tools
Access practical worksheets and guides across the evidence synthesis workflow:

Limitations

Reason categories and their operational priority depend heavily on the specific review question and protocol. The examples provided here serve as starting points, not a universal taxonomy. Specialist review types may require additional status categories or modified reporting guidance.

A complete decision log cannot recover reports that were never identified or decisions made outside the recorded system. Nor does the consistent use of a reason code prove that the underlying eligibility criterion was unbiased. Taxonomy validation, reviewer calibration, retrieval practice, and protocol quality remain independent requirements. Scoping reviews, diagnostic-accuracy reviews, qualitative syntheses, and other designs may require entirely different decision concepts and should follow their applicable extensions and methodological guidance.6,7

Conclusion

Full-text exclusion becomes methodologically defensible only when the reason derives from an operational criterion, is supported by located evidence, attaches to a stable report identifier, and preserves the reviewer and resolution history. Utilizing one reportable primary reason creates mutually exclusive PRISMA totals while secondary failures retain essential analytical detail. Separating true eligibility decisions from non-retrieval, ongoing, and unresolved statuses prevents the review from claiming judgments it never actually made. The resulting record is not mere administrative residue; it is the definitive evidence that the review’s boundaries were applied consistently and can be rigorously appraised.

Advance Your Systematic Review & Meta-Analysis Methodology

Ensure your screening procedures, risk-of-bias appraisals, and meta-analytic models conform to international publishing standards. Strengthen your protocol foundation or join MetaSyn Academy’s comprehensive curriculum.

Frequently asked questions

How many reasons should be reported for each excluded full text?

Use one reportable primary reason per excluded report if that is the team’s chosen method for mutually exclusive totals, while preserving all additional failed criteria in the detailed log.

Is a report that could not be retrieved a full-text exclusion?

No. It was not assessed against the full eligibility criteria. Record it separately as a report not retrieved and preserve the retrieval attempts.

Should companion reports be excluded as duplicate studies?

Not automatically. Link every report to the underlying study and determine the role of each document. Several reports may legitimately contribute to one included study.

References and Methodological Sources

  1. Page MJ, McKenzie JE, Bossuyt PM, Boutron I, Hoffmann TC, Mulrow CD, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71. doi:10.1136/bmj.n71.
  2. Page MJ, Moher D, Bossuyt PM, Boutron I, Hoffmann TC, Mulrow CD, et al. PRISMA 2020 explanation and elaboration: updated guidance and exemplars for reporting systematic reviews. BMJ. 2021;372:n160. doi:10.1136/bmj.n160.
  3. Lefebvre C, Glanville J, Briscoe S, Featherstone R, Littlewood A, Metzendorf MI, et al. Chapter 4: Searching for and selecting studies. In: Higgins JPT, Thomas J, Chandler J, Cumpston M, Li T, Page MJ, Welch VA, editors. Cochrane Handbook for Systematic Reviews of Interventions. Version 6.5.1. London: Cochrane; 2025.
  4. Cochrane. Methodological expectations of Cochrane intervention reviews: standards for study selection, C39–C42 [Internet]. London: Cochrane; [cited 2026 Jul 27]. Available from: Cochrane MECIR C39-C42.
  5. Aromataris E, Lockwood C, Porritt K, Pilla B, Jordan Z, editors. JBI manual for evidence synthesis [Internet]. Adelaide: JBI; 2024 [cited 2026 Jul 27]. Available from: https://synthesismanual.jbi.global.
  6. Tricco AC, Lillie E, Zarin W, O’Brien KK, Colquhoun H, Levac D, et al. PRISMA extension for scoping reviews (PRISMA-ScR): checklist and explanation. Ann Intern Med. 2018;169(7):467-473. doi:10.7326/M18-0850.
  7. McInnes MDF, Moher D, Thombs BD, McGrath TA, Bossuyt PM, Clifford T, et al. Preferred Reporting Items for a Systematic Review and Meta-analysis of Diagnostic Test Accuracy Studies: the PRISMA-DTA statement. JAMA. 2018;319(4):388-396. doi:10.1001/jama.2017.19163.
  8. Haddaway NR, Page MJ, Pritchard CC, McGuinness LA. PRISMA2020: an R package and Shiny app for producing PRISMA 2020-compliant flow diagrams, with interactivity for optimised digital transparency and Open Synthesis. Campbell Syst Rev. 2022;18(2):e1230. doi:10.1002/cl2.1230.
  9. Waffenschmidt S, Knelangen M, Sieben W, Bühn S, Pieper D. Single screening versus conventional double screening for study selection in systematic reviews: a methodological systematic review. BMC Med Res Methodol. 2019;19(1):132. doi:10.1186/s12874-019-0782-0.
  10. Polanin JR, Pigott TD, Espelage DL, Grotpeter JK. Best practice guidelines for abstract screening large-evidence systematic reviews and meta-analyses. Res Synth Methods. 2019;10(3):330-342. doi:10.1002/jrsm.1354.

Similar Posts

Leave a Reply

Your email address will not be published. Required fields are marked *