On 11 December 2024, NICE published TA1023 [1], its guidance on elranatamab, sold as Elrexfio, a BCMA-directed bispecific for relapsed or refractory multiple myeloma. The drug had reportedly cleared the FDA more than a year before that, an accelerated approval in August 2023 on the strength of a single-arm trial, MagnetisMM-3 [2]. By the time it reached NICE's committee, whether it worked was not in dispute. What the committee could not accept was the comparator it had been measured against.
The company brought a matching-adjusted indirect comparison against panobinostat. NICE's committee found the crude adjustments had "nullified the MAIC, resulting in a naive comparison" and judged the whole thing "not credible" [1]. A second comparison, against selinexor, was called "informative but uncertain" [1]. The cost-effectiveness estimates those comparisons produced sat above the range NICE will pay. So elranatamab went to the Cancer Drugs Fund, not routine commissioning.
Read that back. The FDA and NICE did not disagree about whether elranatamab works. They disagreed about what it should have been compared to. And by December 2024, years after the pivotal trial was designed, sized and run, that was no longer a question anyone could answer.
This is the payer-side sequel to a problem regulatory teams already know well: you run a single-arm trial for good reasons, regulators accept it, and then the awkward questions start arriving. The regulator asks them first. The payer asks them last, and asks the expensive version.
Here is the received wisdom about building one evidence plan across regulator, payer and prescriber. Get the three functions in a room. Sync the timelines, and produce three well-coordinated packages. Regulatory affairs owns the label, market access owns the value dossier, medical affairs owns the clinician, and if everyone talks often enough the contradictions dissolve.
They don't dissolve. Coordination cannot reconcile a comparator that three teams chose three different ways at three different times.
Picture three survey crews sent to build on one plot, each measuring from its own benchmark peg. Every crew works to the millimetre, and every crew builds true to its own peg. Then the walls don't meet. The fault was never careless work on site — it was three reference points fixed independently before anyone broke ground, and the comparator is that reference point. You set it once, for all three audiences, or you spend launch year discovering the walls don't line up.
An integrated evidence plan is supposed to serve the regulator, the payer, the prescriber and, behind all three, the patient, from a single evidence base. That only holds if the shared factual core is fixed before each function starts optimising for its own audience. Fixed first, not coordinated afterwards. The existing case for an IEP, and the version of it written for regulatory interactions specifically, both assume this and rarely say it out loud: the plan's value is upstream, at the fork, not in the coordination that comes after.
We can be precise about which fork. Tafuri and colleagues looked at 31 parallel EU regulator–HTA scientific-advice procedures run between 2010 and 2015, and scored where the two audiences agreed [3]. On the patient population, full agreement ran to 77%. On endpoints, 60%. On the comparator: 44%, with outright disagreement at 30%, the worst of any evidentiary domain they measured, by a wide margin.
It survives the fix, too. In a follow-up, even after joint advice had been given, sponsors landed on a comparator that satisfied both the regulator and the HTA body in only 12 of 21 studies, against all of them for the primary endpoint [4]. Advice closes the endpoint gap, but the comparator gap walks straight through it.
The asymmetry is structural, and it is worth understanding rather than lamenting. A regulator can approve a drug against placebo, or on a surrogate that plausibly tracks benefit. A payer cannot build a reimbursement case on either. The payer needs the drug measured against what clinicians actually prescribe today, on an outcome patients actually feel, because that is the only comparison that tells them what they are buying instead of the current standard. Surrogate endpoints that satisfy a regulator are treated by payers with structural suspicion [5][6]. And when the head-to-head comparator simply isn't there, the indirect comparisons sponsors reach for fare badly: of 25 such comparisons reviewed in EU relative-effectiveness assessments, 24 were rated "unclear" in suitability by HTA assessors [7].
Lest this look like a regulator-versus-payer story, it happens inside a single system too. In Germany's AMNOG process, agreement between IQWiG's scientific assessment and the deciding committee's actual appraisal, on identical evidence, came out at a Cohen's kappa of 0.183 [8]. Two bodies, one country, the same data, barely agreeing. The comparator question is that unstable.
So, for Monday: tag the comparator as the single highest-divergence-risk decision in your plan, and treat choosing it as one joint decision, not three parallel ones.
Free download
The IEP Template Pack
The gap matrix, prioritisation grid and plan-on-a-page we use to build integrated evidence plans. Free to keep.
Get the template pack →Before the prescription, the proof it is achievable, because this is not a counsel of despair.
PARADIGM-HF was built as one trial [9]: sacubitril/valsartan was tested against enalapril (the genuine standard of care a cardiologist would otherwise reach for, not placebo) in 8,442 patients, with a validated patient-reported quality-of-life instrument, the KCCQ, built in as a secondary measure. One protocol answered three questions at once. The regulator's: does it beat an active control on hard outcomes. The payer's: is it better than what we already fund, and does it change how patients feel. The prescriber's: how does it perform head-to-head against the thing I currently write. The trial was reportedly stopped early for overwhelming benefit; the registry itself records only that it terminated, so treat the stoppage reason as publicly reported rather than a registry fact.
Contrast the PCSK9 inhibitors. Evolocumab and alirocumab both cleared regulators on registration trials built around a surrogate, LDL-C reduction [10][12]. Useful for approval. Insufficient for a payer, who wanted to know whether lowering LDL-C actually prevented heart attacks. So both programmes then ran enormous cardiovascular-outcomes trials to answer that question: FOURIER in 27,564 patients, ODYSSEY OUTCOMES in 18,924 [11][13]. Two efforts, sequentially, because the payer's question was never inside the pivotal. The divergence here is read off the registered design choices (comparator, endpoint, population), not off any stated intent to satisfy one audience or another.
The lesson is a cost comparison. Fixing the shared core once, at protocol stage, is a design decision. Running the second mega-trial to recover the answer you didn't build in the first time is a capital raise.
If you want the comparator settled before the pivotal, the institutional route now exists. The EU's Joint Scientific Consultation lets developers "obtain scientific consultation during the planning of the clinical studies ... on the information and evidence needs for a subsequent Joint Clinical Assessment" [14]. Read the timing on that: during planning, before the pivotal trial is designed. Paired with EMA parallel scientific advice, it puts the regulator and the HTA body at the same table while the comparator is still a choice rather than a fact.
This is not a fringe recommendation. In a set of 13 stakeholder interviews on JCA implementation, comparator choice was the single most-cited risk, and the explicit recommendation was to move the conversation earlier [15]. And it is live, not hypothetical. Under Regulation (EU) 2021/2282, the Joint Clinical Assessment phases in from 12 January 2025 for cancer medicines with new active substances and for ATMPs, from 13 January 2028 for orphan medicinal products, and from 13 January 2030 for everything else [16].
For Monday: if your asset is an oncology new active substance or an ATMP, the comparator conversation is now a formal, before-pivotal event you can book. Use it to fix the population, comparator and outcomes jointly, rather than discovering the divergence at assessment when nothing can be changed.
Fixing the comparator early does not make the divergence disappear. Analysis published by the Office of Health Economics argues that joint assessment consolidates the clinical-effectiveness question at EU level without eliminating national comparator mismatch, since pricing and reimbursement stay entirely national. Members of IQWiG, Germany's HTA institute, put a number on the underlying problem: roughly a third of requested comparators simply aren't available across every jurisdiction [19]. Worse, a single jointly-agreed comparator does not stay single. It fractures into national variants: one analysis found a single asset generating between 10 and 16 distinct PICOs and, at the extreme, hundreds of required analyses [17]. National assessment timelines diverge just as widely, from around 156 days in Germany to 529 in Poland on comparable assets [18]. You cannot pre-empt all of that from one advice meeting. And the recovery tools have their own critics: managed-entry agreements, the classic post-hoc fix, are argued in the literature to function sometimes as "an operational tool" for commercial price negotiation "rather than as a tool for managing the actual risk arising from immature data" [20].
All true. And every point on that list argues for fixing the comparator early, not against it. You cannot abolish 720 national analyses, but you can choose a comparator that is defensible as standard of care across your major markets, and you can build the patient-relevant endpoint into the pivotal instead of chasing it later. The point was never to collapse the divergence to zero. It is to shrink the part of it you will otherwise pay to reconcile after launch. Necessary, not sufficient, and still the highest-leverage move on the board.
Call what you pay when you skip that move the "reconciliation tax". It is the bill for buying back, at the assessment stage, an alignment you could have set at the design stage. Every instrument on it is more expensive than the decision it replaces.
And elranatamab was not the first drug through that particular door. Venetoclax was approved by the EMA and FDA on three single-arm trials in relapsed or refractory CLL; NICE's Evidence Review Group and committee found "the lack of an appropriate comparator(s)", not the cost-effectiveness model, to be the central defect, and its final guidance, published in November 2017, routed the drug to the Cancer Drugs Fund rather than routine use [23]. That was 2017. Elranatamab was 2024.
Same fork. Seven years apart.
Fixing the comparator early guarantees nothing. It will not carry you cleanly through three assessments in three systems that were never designed to agree. Nevertheless, it is the one variable where the evidence says the divergence is both the largest and the most preventable, and the one where the cost of getting it wrong compounds every year you leave it. One evidence plan means one comparator, chosen once, out loud, before the pivotal. This is the kind of decision an evidence plan exists to force early, and where landscaping the real standard of care across markets (the work Inovia's IEP and InovaCS are built to do) earns its place well before the pivotal is locked.
Get the monthly digest
The 5 things evidence leads need to know each month: regulatory moves, RWE developments and what they mean in practice. No pitch, one email a month.