How to Plan a Systematic Literature Review for HEOR and HTA

MetaSyn Academy guide to plan a systematic literature review for heor and hta, illustrating the decisions documented by HEOR systematic literature
Academic methodology guideEvidence checked: 10 August 2026

A systematic literature review can be rigorous, transparent and reproducible yet still be unusable for health technology assessment. The usual reason is not a failed Boolean search or an incorrect PRISMA diagram. It is that the review was designed around a scientific question while the eventual decision required a different population, comparator, evidence stream, data cut or analytical handoff. HTA adds a decision architecture to systematic-review methodology: evidence has to be relevant to an institutional choice, not merely eligible for a publication.1,2,13

This distinction matters for HEOR because the review is often only one component of a larger evidence program. The same project may need clinical comparative evidence, natural-history data, health-state utilities, resource use, published economic evaluations or RWE, each for a different purpose. The clinical SLR may later need to support an indirect comparison or network assessment, but feasibility cannot be declared before the evidence structure is known. A protocol built for decision use therefore has to do two things at once: preserve standard systematic-review rigor and anticipate the decision dependencies that will emerge after searching.3,5,16

The central idea: plan backward from the decision, but do not reason backward from the desired result. The decision tells you what evidence must remain usable. It does not justify changing eligibility or synthesis after seeing favorable findings.

1. The review question is not the whole decision problem

Academic systematic reviews usually begin by narrowing a question into eligibility criteria. HTA planning has one additional step above that question: define what decision the evidence is meant to support. NICE makes this explicit through the decision problem and scope, which link the population, technology, care pathway, comparators and outcomes to the appraisal. PBAC similarly organizes the clinical evaluation around the proposed patient indication and nominated comparator. Germany gives the comparator an even more formal role through the appropriate comparator in added-benefit assessment. These systems use different terminology, but the methodological function is the same: evidence relevance is defined by a decision context before the review is interpreted.1,12,13

That extra layer changes what a protocol must record. Jurisdiction is no longer administrative metadata. It can change the reimbursement population, treatment line, standard of care, comparator set, economic inputs and submission timetable. A trial population may be broader than the reimbursable population. A pivotal trial may use a comparator that is not the current local standard. An outcome important to regulators may be less informative to a payer than a different endpoint or time horizon. None of these differences makes the original trial “bad.” They make decision relevance a separate methodological question.

The decision lens applied to the same evidenceA central trial evidence box is surrounded by four different decision lenses: academic review, NICE-style reimbursement, EU JCA multi-PICO assessment, and economic model evidence use. Each lens asks a different question of the same trial.The same trial can answer different questions—or fail to answer themDecision relevance is an additional layer of appraisal, not a replacement for internal validity.PIVOTAL TRIALPopulation · comparator · outcomesdata cuts · treatment switchingrelative effect estimatesACADEMIC REVIEWDoes it meet the scientific PICO?May be a complete question in itselfREIMBURSEMENTIs this the relevant populationand comparator for this system?EU JCA SCOPEWhich PICO questions can this trialactually inform?ECONOMIC USEWhich effects, utilities or resourceinputs remain decision-relevant?
Figure 1. The decision lens. Internal validity does not change when the institutional question changes, but the evidence that is relevant, transferable and sufficient can change substantially.

The comparator is part of decision validity

Comparator selection is where this becomes concrete. A review may be beautifully executed around placebo while the payer needs an active comparator. A sponsor may assume that licensed alternatives define the comparator set while a national pathway uses a different standard. PBAC explicitly directs systematic searching toward direct randomized evidence against the main comparator and, when direct evidence is absent, toward evidence that can support an indirect comparison. NICE similarly expects relevant final-scope treatments to be represented in the clinical evidence identification process.2,13,14

This is not the same topic as assessing whether a common comparator makes an indirect comparison credible. That belongs to the dedicated indirect treatment comparison guide. The protocol-planning question occurs earlier: did the SLR identify the comparator evidence needed to let a later team make that judgment at all?

2. One technology can create several evidence questions

One intervention does not guarantee one decision PICO. Treatment line, subgroup, biomarker status, local standard of care or reimbursement population can split the decision into several questions. EU JCA makes this visible because the assessment scope is explicitly PICO-based at European level. Early operational reports show very different patterns: some assessments have a single research question, whereas the 2026 tovorafenib assessment contained several PICOs and large portions of the scope lacked usable comparative evidence.6,7,8,23

The lesson is not that every HTA now needs a dozen PICOs. The lesson is architectural: the protocol must be able to represent multiplicity without creating duplicate reviews. When several PICOs share the same disease, intervention and broad trial universe, a coordinated master identification strategy can be efficient. The evidence still needs PICO-specific mapping. If one question is a clinical-effectiveness comparison and another is a utility search, the evidence universe, study designs and source architecture differ enough that separate streams make more sense.

Example: same drug, two treatment lines. A single clinical search may retrieve both first- and later-line trials. That does not mean the studies belong in the same synthesis. The review needs explicit PICO tags so that eligibility, comparator relevance and downstream analyses remain separate. The efficiency is in shared identification; the rigor is in decision-specific mapping.

3. HEOR evidence is often a program, not one omnibus review

The phrase “HEOR SLR” can conceal several different tasks. Clinical comparative effectiveness, natural history, utilities, resource use and economic evaluations are all legitimate evidence questions, but they do not share one universal study design or source architecture. NICE’s evidence-submission structure separates these functions, and Cochrane’s economics methods likewise distinguish when economic evidence deserves a full systematic component versus a lighter approach.1,16

The useful unit of planning is therefore an evidence stream: a bounded evidence question with its own decision use, eligibility, sources, appraisal and outputs. Streams can be coordinated within one protocol governance system without pretending they are methodologically identical. A clinical RCT stream may support comparative effectiveness. A natural-history stream may establish prognosis or an external-control context. A utility stream may locate values for defined health states. A resource-use stream may combine published evidence with national administrative sources. The review is coherent because the decision links the streams, not because one search strategy can find them all.

Evidence streams use different methodological grammarsFour vertical evidence-stream columns show clinical effects, real-world context, utilities, and costs/resource use. Each column has different study designs, sources, appraisal priorities, and outputs, while all connect upward to one HTA decision.Different evidence streams need different methodological grammarsHTA / HEOR DECISIONThe common purpose connecting separate evidence streamsCLINICAL EFFECTSDesignsRCTs / comparative studiesSourcesdatabases · registries · CSRAppraisalcomparative validityOutputrelative effects / readinessRWE / NATURAL HISTORYDesignscohorts · registries · claimsSourcesRWD literature · data sourcesAppraisalfit for intended useOutputcontext · prognosis · controlUTILITIES / HRQoLDesignselicitation · mapping · trialsSourcesclinical + specialist evidenceAppraisalapplicability to parameter useOutputhealth-state valuesCOST / RESOURCE USEDesignsresource studies · economic evalsSourcesliterature + local tariffsAppraisalsetting, price year, relevanceOutputdecision-model evidence inputs
Figure 2. Evidence-stream grammar. The streams belong to one decision program, but their study designs, sources, appraisal questions and outputs differ enough that an omnibus eligibility rule can become methodologically incoherent.

RWE needs a declared use case

RWE is a good example of why “include observational studies” is not a protocol. NICE’s RWE framework distinguishes different purposes for real-world data and evidence. A registry may describe treatment patterns, supply a natural-history estimate, support an external control, characterize local applicability or estimate comparative effects. The validity requirements depend on that purpose. A study weak for causal treatment-effect estimation may still be fit for describing current care pathways. Appraisal therefore has to be attached to intended decision use rather than summarized as one global quality label.5

Economic evidence needs proportionality

HEOR planning also benefits from rejecting the idea that every model input needs a full systematic review. Some parameters are decision-critical and contestable; others are local administrative values. NICE DSU TSD 13 and Cochrane economics methods both support deliberate planning of evidence used for economic decisions without implying that one method fits every parameter. A full evidence stream is appropriate when competing estimates and uncertainty could materially change the decision. A targeted systematic search may be sufficient for a narrow secondary parameter. A national tariff or official price schedule may be the correct source for a jurisdiction-defined unit cost.3,16

This boundary keeps the guide separate from economic-model construction. The guide explains how literature evidence should be identified and preserved for later model use; it does not decide the state structure, extrapolation function or economic model itself.

4. Design backward from downstream analysis—without choosing the analysis too soon

A clinical review intended for HTA should preserve information that a later comparative-effectiveness team may need. That can include trial arms and treatment definitions, comparator identity, multi-arm structure, baseline population characteristics, subgroup definitions, outcome definitions and timepoints, analysis population, switching or subsequent therapy, study era, IPD availability, aggregate-data availability and the source/data-cut version. NICE and PBAC both make clear that evidence identification may need to extend beyond the new technology’s own trials when indirect comparison is relevant.2,13,14

The important boundary is between readiness and feasibility. Recording a potential common comparator does not prove that an anchored ITC is credible. Collecting candidate effect modifiers does not prove transitivity. Mapping treatment nodes does not prove that a network should be synthesized. Those judgments belong downstream and are handled in the dedicated network meta-analysis assumptions guide and ITC credibility guide.

A useful protocol sentence is conditional, not prophetic: after study identification and evidence mapping, assess whether direct pooling or indirect comparative synthesis is methodologically supportable. The protocol should prespecify the information and decision rule, not the answer to a feasibility assessment that has not happened yet.

5. Evidence has lineage and a clock

A publication-centric view of the evidence base becomes fragile in HTA. Cochrane explicitly treats the study, rather than the report, as the unit of interest and requires multiple reports of the same study to be linked. PBAC asks for a master list of relevant trials and their associated reports. HTA adds an additional complication: evidence for one trial can mature through registry entries, abstracts, journal papers, CSRs, regulatory reports and later endpoint-specific data cuts.14,15

There is no defensible universal rule that a journal paper must always replace a CSR, or that the newest source is always the best. The correct source can depend on the variable. Safety tables may be more complete in a CSR or regulatory assessment, while a later publication may contain a more mature survival cut. The protocol should therefore preserve record identity, data-cut date, endpoint maturity and the rationale for choosing one value over another.

The 2026 lurbinectedin JCA is a useful early operational example. Assessors distinguished a July 2024 clinical cutoff from a February 2025 cutoff and considered the later prespecified data cut more relevant for several HTA purposes because it had greater maturity, even though much of the developer’s dossier emphasized the earlier cut. The lesson is not that “latest always wins.” It is that data maturity is part of the evidence object and should be visible in the SLR dataset.23

Evidence lineage and decision clockA horizontal trial timeline shows registry, conference abstract, primary report, regulatory assessment, and later data cut. A second horizontal decision timeline below shows search execution, evidence cutoff, analysis freeze, and submission. Vertical connectors show that evidence records and decision dates are separate dimensions.The study evolves on one timeline; the HTA decision evolves on anotherA protocol should track both. Neither “publication date” nor “search date” is enough on its own.STUDY LINEAGERegistryAbstractPrimary paper / CSRRegulatory reviewLater endpoint cutMASTER STUDY IDLinks records; preserves data-cut and variable provenanceDECISION CLOCKSearch executedEvidence cutoffAnalysis freezeHTA submissionNew mature data may trigger reopening
Figure 3. Evidence lineage and decision clock. A trial can acquire new reports and data cuts while the HTA project moves toward its own evidence freeze and submission date. Protocol governance needs both dimensions.

Search currency should be trigger-based, not myth-based

There is no single cross-jurisdictional rule that every HTA search must be rerun within three, six or twelve months. Agency procedures differ, and the evidence velocity of the therapeutic area matters. Current CDA-AMC procedures, for example, are tied to reimbursement-review process milestones rather than a universal month-based literature-search rule. Cochrane’s own review-update timing standards are useful for Cochrane production, but they should not be mislabeled as reimbursement rules. For HTA planning, the stronger approach is to record search execution, evidence cutoff, analysis freeze and submission timing, then define material triggers such as a new pivotal trial, comparator change, mature data cut, safety signal, regulatory decision, submission delay or resubmission. Where the evidence program is intentionally maintained as a living systematic review, PRISMA-LSR provides a reporting framework for the living process; it should not be mistaken for an HTA-mandated update interval.10,11,19,22

6. Prespecification does not mean pretending the evidence base is already known

Protocols protect against result-driven decision making by making important methods explicit before the review unfolds. That principle remains essential in HTA. Yet a real evidence program can encounter a scope clarification, a newly required comparator, a new pivotal study or an evidence volume that is radically different from what was expected. The methodological question is not whether change is ever allowed. It is whether the change is governed in a way that makes its timing and potential bias visible.

NICE DSU TSD 27 is useful here because it addresses prioritization of studies and outcomes in NICE HealthTech literature reviews when evidence volume is higher or lower than anticipated. The lesson should be applied narrowly: controlled refinement can be defensible when the framework, triggers and safeguards are planned. It is not a general license to change eligibility after inspecting results.4

Controlled protocol adaptationA protocol starts with prespecified criteria. If a predefined trigger occurs, a governance gate checks whether refinement is permitted, whether decisions are blinded to results, and whether an amendment is documented. The process either keeps the original criteria or applies a controlled change.Prespecification can include rules for change without becoming post hocAPPROVED PROTOCOLeligibility · sources · outputsPREDEFINED TRIGGERscope change · unexpected evidence volumenew comparator · major new trialGOVERNANCE GATE• Is the trigger genuinely met?• Is the change independent of favorable results?• Is rationale, timing and impact documented?KEEP ORIGINALtrigger not met / change unsafeAMENDversioned, justified, auditableThe gate governs process. It does not guarantee that the revised review is free of bias.
Figure 4. Controlled adaptation. The protocol can define how legitimate changes are handled while preserving a clear boundary against result-driven post hoc selection.

7. Automation is part of protocol governance when it can change the evidence set

AI-assisted searching, screening and extraction have moved from a workflow convenience to a governance issue in at least one current HTA system. In July 2026, the HTACG adopted general principles for AI use in preparing JCA dossiers. They keep accountability with the health technology developer, require human oversight through AI-assisted steps, preserve the existing methodological standard, and expect transparent reporting of AI-assisted retrieval, screening, extraction, risk-of-bias, analysis or reporting. The guidance also asks for tool details and expects prompts to be retained and made available if requested.9

These are EU JCA-specific expectations and should be labeled that way. A broader Academy principle can still be derived: if automation can influence which studies enter the review, what data are extracted or how evidence is classified, the protocol should state the task, tool/version, human validation, audit trail and escalation rule. That is about reproducibility and responsibility. It is not an endorsement of a particular AI product, and it does not mean that using AI makes a review more rigorous.

8. What “fit for HTA” means

A decision-ready SLR does not need to predict the committee’s conclusion. It needs to make the evidence program traceable from decision scope to evidence use. The comparators should be justified before they become search exclusions. Multi-PICO scope should remain explicit. Evidence streams should be separated when their methods differ. Study reports should be linked to the underlying study and data cut. Downstream comparative or economic teams should receive the variables they need without forcing the SLR to decide their models prematurely. Evidence updates and amendments should follow a visible rule rather than informal memory.

That is why a conventional systematic-review protocol and an HEOR/HTA protocol are not competitors. The generic protocol supplies the foundation: transparent eligibility, reproducible searching, screening, extraction, appraisal and synthesis. The HEOR/HTA layer connects those methods to a real decision system. PRISMA-P remains useful for protocol reporting, PRISMA-S for reporting searches, and PROSPERO or OSF for appropriate registration. None of those tools, by itself, decides which comparator a payer needs or how the team should handle a new pivotal data cut three weeks before submission.17,18,20,21

The strongest warning example is simple. Imagine a PRISMA-compliant review with a published protocol, dual screening, a comprehensive search and appropriate risk-of-bias assessment. It concludes that a new therapy is effective versus placebo in a broad population. The actual reimbursement decision concerns a later-line subgroup versus an active local comparator. The review can be academically excellent and still fail the decision. More generic rigor does not repair a decision mismatch; the protocol has to identify that mismatch before the evidence is frozen.

Apply the reasoning with the protocol resource

The paired HEOR systematic literature review protocol template turns these concepts into operational fields. Use it to map the decision context, comparator/PICO architecture, evidence streams, evidence dates, study lineage, downstream-readiness requirements and amendment governance. The guide explains why those modules exist; the resource is where the team records the decisions.

Conclusion

The defining feature of an HEOR/HTA SLR is not a special database, a longer protocol or a new acronym. It is the explicit connection between review methods and the decision that will consume the evidence. Once that connection is made visible, several design choices become easier to justify: why the comparator landscape must be established early, why one evidence program may need several methodological streams, why multi-PICO mapping is different from duplicating searches, why study reports need lineage and data-cut control, why update triggers belong in protocol governance, and why downstream analysis should be prepared for without being prejudged.

That planning layer also creates a cleaner Academy architecture. Generic protocol pages can own generic conduct. The search hub can own detailed retrieval methods. Detailed NMA assumptions, ITC credibility, economic modelling and PRO measurement methods belong in their dedicated methodology pages. This guide owns the earlier planning point where those later methods must be anticipated without being prejudged.

Frequently asked questions

Is an HEOR systematic literature review a different review type?

Not necessarily. The distinctive feature is usually the decision-oriented planning layer rather than a separate taxonomy of review design. Standard systematic-review methods still apply, but the protocol must preserve jurisdiction, comparator, evidence-stream, timing and downstream-use requirements.

Does every HTA need multiple systematic literature reviews?

No. Evidence streams are conditional. A project should split streams when eligibility, source architecture, appraisal or decision use differs materially, not because an HEOR checklist says every submission needs a fixed number of reviews.

Can one search support several EU JCA PICOs?

Sometimes. When the disease, intervention and evidence universe substantially overlap, coordinated identification can support several PICO mappings. The individual decision questions still need explicit eligibility and evidence mapping. Separate searches are more appropriate when the evidence domains or study designs differ materially.

Why is study lineage more important than simply deduplicating references?

Deduplication removes repeated bibliographic records. Study lineage links different reports of the same underlying study, including abstracts, journal articles, regulatory documents and later data cuts. Those records can contain different information and should not be counted as separate trials.

Should the SLR protocol decide which economic model will be used?

No. It can define which evidence inputs and downstream uses must be supported, but model structure and economic analysis belong to the economic-modeling process.

Can a protocol change when the HTA scope changes?

Yes, if the change is transparent and governed. The amendment should record the trigger, timing, affected methods, rationale and impact. Result-driven eligibility changes without such governance remain a serious bias risk.

Does AI-assisted screening need to be reported?

Reporting expectations depend on context. For EU JCA dossiers, current HTACG principles specifically expect transparent documentation of AI-assisted evidence-synthesis steps. More generally, protocol-level documentation of tool/version, validation and human oversight improves auditability.

References

  1. National Institute for Health and Care Excellence. Single technology appraisal and highly specialised technologies evaluation: company evidence submission user guide (PMG24). Current web guidance, updated through March 2026. NICE.
  2. National Institute for Health and Care Excellence. Appendix B: Identification, selection and synthesis of clinical evidence. In: PMG24. NICE.
  3. Kaltenthaler E, Tappenden P, Paisley S, Squires H. NICE DSU Technical Support Document 13: Identifying, selecting and using evidence to inform the model structure and parameters. NICE Decision Support Unit. NICE DSU TSD index.
  4. NICE Decision Support Unit. Technical Support Document 27: Prioritising studies and outcomes for consideration in NICE HealthTech literature reviews. 2025. NICE DSU.
  5. National Institute for Health and Care Excellence. NICE real-world evidence framework. 2022. NICE.
  6. European Parliament and Council. Regulation (EU) 2021/2282 of 15 December 2021 on health technology assessment and amending Directive 2011/24/EU. Off J Eur Union. 2021. EUR-Lex.
  7. Health Technology Assessment Coordination Group. Guidance on filling in the joint clinical assessment (JCA) dossier template – Medicinal products. 2024. European Commission.
  8. Member State Coordination Group on Health Technology Assessment. Joint Clinical Assessment Summary Report of Tovorafenib, Version 1.0. European Union; 2026. European Commission.
  9. Health Technology Assessment Coordination Group. General principles on the use of Artificial Intelligence in the preparation of dossiers for Joint Clinical Assessments. Adopted 15 July 2026. European Commission.
  10. Canada’s Drug Agency. Pharmaceutical Reviews Update — Issue 61. 30 April 2026. CDA-AMC.
  11. Canada’s Drug Agency. Procedures for Reimbursement Reviews. Current procedures. CDA-AMC.
  12. Institute for Quality and Efficiency in Health Care. General Methods, Version 8.0. 1 July 2026. doi:10.60584/General-Methods_V8.0. IQWiG.
  13. Pharmaceutical Benefits Advisory Committee. Section 2.1: Literature search methods. PBAC Guidelines. Australian Government Department of Health.
  14. Pharmaceutical Benefits Advisory Committee. Appendix 3: Identify relevant trials. PBAC Guidelines. Australian Government Department of Health.
  15. Li T, Higgins JPT, Deeks JJ. Chapter 5: Collecting data. In: Higgins JPT, Thomas J, Chandler J, et al, eds. Cochrane Handbook for Systematic Reviews of Interventions. Version 6.5. Cochrane; 2024. Cochrane.
  16. Aluko P, Graybill E, Craig D, et al. Chapter 20: Economic evidence. In: Higgins JPT, Thomas J, Chandler J, et al, eds. Cochrane Handbook for Systematic Reviews of Interventions. Version 6.5. Cochrane; 2024. Cochrane.
  17. Shamseer L, Moher D, Clarke M, et al. Preferred reporting items for systematic review and meta-analysis protocols (PRISMA-P) 2015: elaboration and explanation. BMJ. 2015;350:g7647. doi:10.1136/bmj.g7647. Current PRISMA-P page.
  18. Rethlefsen ML, Kirtley S, Waffenschmidt S, et al. PRISMA-S: an extension to the PRISMA Statement for Reporting Literature Searches in Systematic Reviews. Syst Rev. 2021;10:39. doi:10.1186/s13643-020-01542-z.
  19. Akl EA, Khabsa J, Iannizzi C, et al. Extension of the PRISMA 2020 statement for living systematic reviews (PRISMA-LSR): checklist and explanation. BMJ. 2024;387:e079183. doi:10.1136/bmj-2024-079183.
  20. Centre for Reviews and Dissemination, University of York. PROSPERO: eligibility for inclusion. Current guidance. PROSPERO.
  21. Center for Open Science. OSF Registrations and Preregistrations. Current guidance. OSF.
  22. Cochrane. Chapter IV: Updating a review. In: Cochrane Handbook for Systematic Reviews of Interventions. Current version. Cochrane.
  23. Member State Coordination Group on Health Technology Assessment. Joint Clinical Assessment Report of Lurbinectedin, Version 1.0. European Union; 2026. European Commission.
Methodological scope note. This guide focuses on decision-oriented SLR planning for HEOR and HTA. It does not teach search syntax, NMA/ITC statistical assumptions, economic-model construction, PRO psychometrics or detailed jurisdictional dossier completion. EU JCA examples are used as current operational evidence and are distinguished from universal methodology. Evidence and guidance status were checked through 10 August 2026.

Similar Posts

Leave a Reply

Your email address will not be published. Required fields are marked *