Evidence assessment in a Joint Clinical Assessment (JCA) examines the relative clinical effectiveness and safety of a medicine against the PICO questions defined in the assessment scope. The suitability of available studies, the validity of direct and indirect comparisons, and uncertainty around the results determine how findings can be interpreted. Particular methodological challenges arise where direct comparative data are unavailable and the assessment relies on population-adjusted indirect comparisons or non-randomised evidence.
The first published JCA reports demonstrate the importance of prognostic factors and effect modifiers, between-study comparability, post hoc analyses and statistical precision. For health technology developers (HTDs), methodological readiness therefore begins during clinical development, before the JCA dossier is submitted.
Methodological framework for clinical evidence assessment in JCA
JCA assesses relative clinical effectiveness and safety using the available clinical evidence for the populations, interventions, comparators and outcomes specified in the assessment scope. Regulation (EU) 2021/2282 establishes the European framework, complemented by implementing provisions and methodological guidance. JCA does not determine national added benefit, reimbursement or price.
Assessing clinical evidence against PICO questions
Each PICO defines the comparison for which relevant studies and analyses must be identified. One study may inform several PICOs, while a single PICO may require evidence from several studies or an indirect comparison. Assessors consider whether study populations, interventions, comparators and outcomes match the scope. The development and consolidation of PICOs are covered separately in the JCA scoping and PICO guide.
Direct and indirect comparative evidence in JCA
Direct evidence compares an intervention and comparator within the same study. Indirect comparisons estimate relative effects across studies. Randomisation can balance known and unknown prognostic factors within a trial; between-study comparability must be justified separately for indirect comparisons. The choice of synthesis method therefore depends on the evidence network and the specific PICO.
Methodological uncertainty and limits of inference
An effect estimate is only as reliable as the studies and assumptions on which it rests. Risk of bias, missing data, heterogeneous populations, imprecise estimates and untestable modelling assumptions can all limit interpretation. Statistical significance does not remove these concerns. JCA must distinguish observed results from the uncertainty surrounding them.
Randomised controlled trials in JCA evidence assessment
Well-conducted randomised controlled trials (RCTs) can provide direct comparative evidence while reducing important sources of bias. Randomisation alone does not guarantee valid results: trial conduct, analysis populations, missing data and statistical planning also matter.
Randomisation, allocation concealment and blinding
Randomisation aims to distribute measured and unmeasured prognostic factors between arms. Allocation concealment prevents foreknowledge of assignment from influencing enrolment. Blinding may reduce differences in care and outcome assessment. The implications of an open-label design depend on the outcome: patient-reported symptoms may be more susceptible to assessment bias than objectively ascertained mortality.
Intention-to-treat principle and analysis populations
Under the intention-to-treat (ITT) principle, participants are generally analysed according to their randomised assignment. This preserves the original comparison. Post-randomisation exclusions may introduce bias, particularly where exclusions are related to treatment or prognosis. The population used for each analysis and deviations from the statistical analysis plan should be clear.
Prespecified analyses and control of statistical error
Prespecified analyses are defined in the protocol or statistical analysis plan before the relevant results are known. Prespecification reduces outcome-driven analytical choices. Primary endpoints, hypothesis tests and multiplicity control are particularly important. Hierarchical testing or other multiplicity procedures can address inflated false-positive risk. Post hoc analyses may be informative but require appropriately cautious interpretation.
Data cut-offs and evidence maturity
Overall survival and progression-free survival estimates may change as follow-up accumulates. Assessors should consider whether a cut-off was prespecified, the number of events and the maturity of the data. A later cut-off is not automatically the primary methodological reference: the planned analysis and its inferential status remain relevant.
Indirect comparisons in JCA: NMA, MAIC and other methods
Indirect treatment comparisons can estimate relative effects where no suitable head-to-head trial is available. Their validity depends on between-study comparability, the available data and the assumptions of the selected method. Multiple comparators in the assessment scope may increase the need for these analyses.
Anchored and unanchored indirect treatment comparisons
An anchored comparison uses a shared comparator: if X is compared with C and Y is compared with C, an indirect X-versus-Y estimate may be derived under suitable assumptions. An unanchored comparison has no common comparator and requires substantially stronger assumptions about population comparability. Relevant prognostic factors and effect modifiers must be addressed; unavailable or unmeasured variables may undermine validity.
Network meta-analysis in JCA
Network meta-analysis (NMA) combines direct and indirect evidence across a connected treatment network. Transitivity requires sufficient comparability of factors that could modify relative treatment effects. Where direct and indirect evidence overlap, consistency can also be examined. Differences in disease stage, outcome definitions or treatment settings can challenge these assumptions.
Matching-adjusted indirect comparisons
Matching-adjusted indirect comparison (MAIC) uses individual patient data from at least one study and often aggregate data from another. Patients are weighted to align selected baseline characteristics with the external population. The result depends on the covariates selected, available data and assumptions about remaining differences. MAIC cannot adjust for factors that were not measured.
Population-adjusted indirect comparisons
Other population-adjusted methods include simulated treatment comparison (STC) and multilevel network meta-regression (ML-NMR). They differ in modelling approach and data requirements; ML-NMR can integrate individual and aggregate data in a common model. Greater statistical sophistication does not compensate for inadequate comparability or unsupported assumptions.
Methodological requirements for indirect comparisons
|
Assessment area |
Methodological relevance |
|---|---|
|
Population comparability |
Between-study differences must not unduly distort treatment effects |
|
Comparator structure |
The evidence network must support the chosen method |
|
Prognostic factors |
Relevant determinants of outcomes need consideration |
|
Effect modifiers |
Differences in relative treatment effects must be addressed |
|
Outcome definitions |
Measures must be sufficiently comparable |
|
Model assumptions |
Assumptions must be justified and assessed |
|
Sensitivity analyses |
Plausible alternative assumptions should be examined |
Meeting these conditions matters more than the complexity of the statistical method.
Prognostic factors and effect modifiers in indirect comparisons
Prognostic factors influence outcomes irrespective of treatment; effect modifiers influence the relative effect of treatment. Both matter when assessing population comparability and the validity of adjusted indirect comparisons.
Distinguishing prognostic factors from effect modifiers
A prognostic factor predicts clinical outcome regardless of the assigned treatment. An effect modifier changes the relative treatment effect across patient groups. A characteristic may play both roles. The distinction affects which covariates must be adjusted for, particularly in unanchored comparisons.
Identifying relevant variables
Variable selection should be justified using clinical knowledge, systematic literature reviews, trials and registries. Statistical significance in a single dataset is not an adequate selection criterion. A non-significant association, especially in a small study, does not demonstrate clinical irrelevance.
Causal relationships and variable selection
Directed acyclic graphs (DAGs) can make assumptions about treatment, patient characteristics and outcomes explicit. They may help identify appropriate adjustment sets and avoid inappropriate adjustment. Their usefulness depends on the accuracy of the assumed causal relationships; they are not a substitute for evidence.
Unmeasured variables and residual confounding
If important prognostic factors or effect modifiers were not collected, statistical adjustment may be impossible. Differences in outcomes may then partly reflect population differences rather than treatment. Weighting cannot eliminate unmeasured confounding. Analyses should state which variables were included, which were unavailable and how residual uncertainty affects interpretation.
Statistical precision and effective sample size in MAIC
In weighted MAIC analyses, the effective sample size (ESS) indicates how far unequal patient weights reduce the amount of information the sample provides. A small ESS signals concentration of weights and may be associated with lower statistical precision. ESS is a diagnostic measure, not a comprehensive measure of uncertainty or risk of bias.
Weighting and effective sample size
MAIC assigns weights to align selected characteristics with an external population. Poor overlap may result in very high weights for a few participants. A commonly used calculation is:
ESS = (sum of weights)² / sum of squared weights.
The formula quantifies the distribution of weights; it does not capture all uncertainty in the treatment effect. A low ESS may indicate limited population overlap, but it cannot reveal unmeasured effect modifiers or residual confounding.
Covariate balance and standardised mean differences
Standardised mean differences (SMDs) describe covariate imbalance and can be examined before and after weighting. Good balance on measured variables does not establish overall comparability: unmeasured differences may persist.
Confidence intervals and statistical uncertainty
A low ESS may increase uncertainty and widen confidence intervals. A large point estimate can remain poorly supported when precision is low. ESS is not, however, a complete measure of MAIC validity: covariate selection, modelling assumptions and residual bias must also be evaluated.
Non-randomised evidence and external control groups in JCA
Non-randomised studies and external controls may supplement comparative evidence when suitable randomised comparisons are unavailable. Their interpretation depends on the control of between-group differences and other biases. They may be particularly relevant in rare diseases and small populations.
Single-arm studies and the absence of internal controls
Single-arm studies describe outcomes and safety under an intervention without a concurrently randomised control. They cannot, on their own, identify a causal relative treatment effect. Additional comparative data are needed.
Historical and external comparator cohorts
External cohorts may be derived from previous trials, registries or other datasets. Inclusion criteria, disease stage, prior therapy, outcome definitions and follow-up should be compared. Changes in clinical practice can further limit the relevance of historical controls.
Real-world data and registries
Real-world data (RWD) may come from disease registries, electronic health records or other care datasets. Suitability for comparative analysis depends on data quality, completeness of relevant covariates and bias control. Large sample size does not replace a sound study design.
Confounding and selection bias
Confounding occurs when variables are related to both treatment assignment and outcome. Selection bias can arise when inclusion processes create systematic differences. Adjustment can address measured factors but cannot automatically remove unmeasured confounding.
Risk of bias and internal validity of JCA evidence
Internal validity concerns whether a treatment effect has been estimated without material systematic distortion. Risk-of-bias assessment is therefore essential for both randomised and non-randomised evidence.
Risk of bias in randomised trials
Important domains include randomisation, deviations from intended interventions, missing outcome data, outcome measurement and selective reporting. Risks may differ across outcomes and analyses; a single blanket rating for a trial can obscure these differences.
Bias in non-randomised studies
Confounding, selection processes, inconsistent follow-up, incomplete records and differing outcome definitions require particular attention. The construction of comparison groups and the methods used to address systematic differences must be transparent.
Missing data and missingness
Missing data can bias estimates when missingness relates to treatment, prognosis or outcome. The impact depends on the amount and mechanism of missingness and the analytical approach. Complete-case analyses may be problematic if participants with missing data differ systematically from others. Assumptions and sensitivity analyses should be reported.
Outcomes and subgroup analyses in JCA evidence assessment
The interpretability of clinical outcomes depends on their definitions, measurement and statistical analysis. Results must address the outcomes and populations specified in the assessment scope.
Mortality, morbidity and health-related quality of life
Mortality outcomes include overall survival; morbidity outcomes can cover symptoms, functioning and clinical events. Health-related quality of life captures patients' physical, psychological and social experience. Clinical relevance, operational definitions and methodological quality all influence interpretation.
Patient-reported outcomes
Patient-reported outcomes (PROs) capture patients' own assessments. Validated instruments, appropriate timing and sufficiently complete data are important. Open-label designs and differential missingness may create bias, particularly for subjective measures.
Surrogate outcomes and clinical interpretation
Surrogate outcomes measure biological or clinical markers used in place of patient-relevant outcomes. Their validity depends on evidence linking changes in the surrogate to changes in outcomes that matter to patients. A statistically significant surrogate effect does not automatically establish a benefit in survival, morbidity or quality of life.
Prespecified and post hoc subgroup analyses
Subgroup analyses examine whether effects differ across patient groups. Prespecification, clinical rationale and suitable interaction tests are central. Significance in one subgroup but not another does not, by itself, demonstrate a difference between subgroups. Post hoc findings are more vulnerable to chance and require caution.
Sensitivity analyses and robustness of JCA results
Sensitivity analyses test how effect estimates respond to alternative assumptions and analytical choices. They are particularly valuable for indirect comparisons, external controls and complex models.
Sensitivity analyses for indirect comparisons
Analyses may vary study inclusion, modelling assumptions or adjustment variables. In MAIC, alternative covariate sets may change weights, ESS and effect estimates. In NMA, alternative models or restrictions to the network can be examined.
Robustness analyses for non-randomised data
Alternative cohort definitions, adjustment approaches and assumptions about unmeasured confounding can indicate how sensitive findings are to methodological decisions. Such analyses cannot prove that all bias has been removed.
Interpreting inconsistent analytical results
Substantial differences between sensitivity analyses indicate dependence on particular assumptions. Assessors should consider which assumptions are clinically and methodologically credible and communicate residual uncertainty, rather than selecting only favourable estimates.
Methodological lessons from the first published JCA reports
The first published JCA reports illustrate recurring challenges around indirect comparisons, prognostic variables, non-randomised data and analyses planned after data became available. Tovorafenib, lurbinectedin, tarlatamab and onasemnogene abeparvovec span different evidence settings. More elaborate statistical models cannot automatically repair limitations in the underlying data or comparability. Detailed case-by-case findings are covered in the separate insight, Joint Clinical Assessment: What the first four JCA reports reveal about evidence assessment.
Methodological preparation for JCA evidence assessment
JCA methodological readiness starts before dossier submission. Developers should examine seven areas: (1) which PICOs have direct randomised evidence; (2) which comparators require indirect synthesis; (3) whether relevant prognostic factors and effect modifiers are identified; (4) whether individual or aggregate data are available; (5) how weighting and heterogeneity affect precision; (6) which biases arise from study design, confounding and missing data; and (7) which sensitivity analyses are needed. Dossier structure, documentation and submission requirements are addressed in the separate JCA dossier guide. Timelines and deadlines are described in the JCA process guide.
How methodological differences continue in national procedures after the JCA is shown in the insight EU JCA and national HTA: Where methodological differences begin after the JCA. For Germany, the guide to the AMNOG dossier explains how evidence is prepared for the national benefit dossier.
Sources and regulatory framework
Regulation (EU) 2021/2282 on health technology assessment.
European Commission and HTA Coordination Group: methodological guidance on joint clinical assessments, direct and indirect comparisons, and clinical evidence assessment.
European Commission: 2026 JCA reports for tovorafenib, lurbinectedin, tarlatamab and onasemnogene abeparvovec.