Two companies. The same two drugs: secukinumab against adalimumab in ankylosing spondylitis. Novartis ran a matching-adjusted indirect comparison on its own patient-level data and concluded secukinumab came out ahead. AbbVie ran a matching-adjusted indirect comparison on its own patient-level data and concluded adalimumab held its ground on cost. Same category of method, same disease, opposite verdicts, because the two analyses balanced on different sets of covariates [1]. The same failure is documented in the peer-reviewed NICE methods literature [2].
The reflex is to file it under "the statisticians disagreed" and move on. The adjustment method feels like a downstream technicality, something the biostatistician settles after the trial reads out. The paradox says otherwise. The method and the covariate list are two separate levers, and either one, on its own, can flip your conclusion. Underneath sits a more useful fact: the method is fixed for you by one checkable feature of your data, decided before a single model is fitted, whether or not anyone on the team realises the decision has already been made. Get that wrong and you have bought a second, quieter error, one hiding behind the tidy covariate table everyone will actually audit.
This post is about that second error. It sits underneath the ECA design decisions in our external control arm checklist; assume here that the comparator source and design are already settled, and the only question left is which statistical adjustment to apply.
Ask one question before anything else: do you hold individual patient-level data (IPD), and for which arms? The answer sorts you into one of three conditions, and each condition licenses a different method.
One plain sentence on what each method actually does, because you have to defend the choice, not derive it. PSM pairs each treated patient with one or more comparator patients who had a similar modelled probability of being treated, then compares the matched sets. Match, then compare: that is the whole idea. Weighting keeps every patient but weights each by the inverse of their probability of being in the group they are in, building a balanced pseudo-population [6][7]. MAIC reweights your own IPD so its covariate averages line up with the comparator's published baseline characteristics, then compares. It exists because the comparator's IPD is not yours to have.
None of this is Inovia's invention. It is the operative logic of TSD 17, TSD 18 and the 2024 EU HTA CG guideline, three independent official documents drawing the same line in the same place [3][4][5]. Nor are these the only tools in the drawer: hybrid designs that blend a concurrent arm with static or dynamic borrowing are a fourth pattern, registered in real protocols [8]. But for the standard ECA comparison, the three-way gate is the decision.
This is the clean condition, with a textbook worked example in the peer-reviewed literature. A published analysis of the APT trial matched 406 treated patients 1:1 against 1,366 patients pooled from other trials, using propensity scores built on five covariates: age, tumour stage, oestrogen-receptor status, progesterone-receptor status and histological grade [9]. Full IPD on both sides, matched on named prognostic factors. That is the easy branch when the data actually supports it.
Here is the counter-intuitive part. PSM is used less than its data-condition fit would suggest. TSD 17's own review of NICE appraisals built on non-randomised comparative IPD found only 2 of 16 used propensity score matching; 6 used multivariate regression, and 7 applied no adjustment at all [3]. Having the gold-standard data condition does not mean teams reach for the gold-standard method.
That said, having it does not guarantee a better outcome either. Krüger, Cantoni and Van Engen's 2025 feasibility study of EU oncology approvals built on non-randomised data found that propensity-score-based methods did not lead to more positive HTA outcomes than the alternative unanchored MAICs [10]. Sit with that: the "preferred" data condition, the "preferred" method, and still no reliable payoff at the committee. So even on the easy branch, do not treat the method as settled by the data alone. IPTW is sensitive to a handful of extreme weights, matching quality has to be demonstrated, and the choice still has to be justified rather than assumed.
Free download
The External Control Arm Design Checklist
The 12 design points regulators probe first, in one checklist you can run against your protocol before database lock.
Get the checklist →MAIC is the method for the condition most biotechs actually face: your IPD, the competitor's published numbers, nothing more. The founding paper reweighted one trial's patients to match another's reported baseline characteristics, comparing adalimumab against etanercept without ever touching the etanercept IPD [6]. It is a genuine piece of engineering. It also hands you a bill, in three parts.
Part one: your sample quietly shrinks. Reweighting costs you effective sample size (ESS), and the cost is not small. TSD 18's own review of three reporting papers found an average ESS reduction of around 80% (range 57–98%) [4]. Separately, and from a different and broader sample, Phillippo et al.'s 2019 audit of 268 NICE technology appraisals found that among the nine MAIC-using appraisals that reported ESS, the median reduction was 74.2% (range 7.9–94.1%) [11]. Those are two different figures from two different samples. Do not average them into one; take them as two independent readings of the same uncomfortable thing.
Years ago I worked on an early-phase rare-disease programme where the only comparator in existence was a published aggregate cohort. No IPD to request, nobody to request it from. MAIC was the only lawful option we had. When we reweighted our patients to match the published baseline means, the effective sample size fell to a sliver of what we had enrolled. The trial had not got smaller. The confidence we were entitled to claim from it had. That is the phantom precision reweighting buys you: a tidy-looking interval drawn on a sample that has quietly collapsed underneath it.
Part two: anchored is necessary, not sufficient. An anchored comparison uses a common comparator to cancel out prognostic differences; an unanchored one cannot, and leans on the heroic assumption that you have measured and balanced every effect modifier. But the right category is not a safe harbour. In the pegcetacoplan NICE appraisal the company used a textbook-correct anchored MAIC, with eculizumab as the common comparator, and the committee still judged the results biased because key effect modifiers, including haemoglobin levels and transfusion history, could not be adjusted for across the two trials [12]. Right method category, wrong outcome, for a reason the anchored/unanchored choice does not fix.
Part three: the numbers can lie about their own certainty. A 2020 simulation of unanchored MAIC found 95% confidence-interval coverage of 11.2% unweighted, 93.8% when fully weighted with every covariate balanced, and 39.1% when only partially weighted. None reached the nominal 95%, even in the fully weighted best case [13]. And unanchored is not the exception in the real world; it describes the large majority of NICE population-adjustment applications. So if MAIC is forced on you, pre-specify your effect modifiers, report ESS honestly, and treat "we used an anchored MAIC" as the opening line of your defence, not the closing one.
Let's be frank about the case nobody wants to be in. Aggregate data on your side, aggregate data on the comparator's, no patient-level records anywhere. The 2024 EU HTA Coordination Group Methodological Guideline, adopted under Regulation (EU) 2021/2282, states it plainly: when non-randomised evidence is available only at the aggregated data level, "there are no adequate method available for reliable estimation of treatment effectiveness" [5].
There is no statistical rescue for a data condition that cannot carry one.
This is the branch the tool-first mindset skips straight past, because it assumes every problem is a modelling problem. It is not. If you are heading into aggregate-on-both-sides territory, that is a design red flag to fix while you still can, by securing comparator IPD or reconsidering the comparator source, not a puzzle to hand to the statistician after the fact.
One clarification, because the tracks get conflated. The EU HTA CG guideline sits on the HTA and Joint Clinical Assessment side, not the marketing-authorisation side. It is not EMA guidance. EMA's own 2025 reflection paper on real-world data (EMA/99865/2025, final) is silent on which adjustment method to use [14]. That silence is not evidence the question is open. It is resolved, on the HTA side, by the document above.
If you want the cost of treating the method as a downstream detail in plain currency, read the Dinutuximab Beta appraisal. Across changes to the comparison method and corrections to its implementation, the incremental cost-effectiveness ratio moved from the company's naive indirect comparison at £22,338 per QALY, to an ERG estimate of £111,858, to a committee-revised MAIC at £24,661, to a DSU-corrected model landing between £62,886 and £87,164 [15]. Same underlying clinical data throughout. Only the comparison method and how it was run changed.
A swing of that size, from the method alone, is the reimbursement decision.
Here is the strongest version of the objection, and it is not a strawman. A good methodologist will tell you the covariate list and the overlap between populations do the real work; the mechanical choice of matching versus weighting versus reweighting is second-order. And there is evidence for it. Remiro-Azócar and colleagues' 162-scenario simulation found MAIC accurate when its assumptions were met [16], and Signorovitch and colleagues' 2023 validation showed a MAIC reproducing the result of a real head-to-head RCT in psoriasis [17]. When the data condition holds and the assumptions hold, the method genuinely is the smaller worry.
Agreed. That is exactly the conditional, and it is the whole argument. The method is second-order only once the gate is cleared and the assumptions actually hold. The point of the gate is to check whether they do, before you rely on them. The large majority of real NICE population-adjustment applications are unanchored, precisely the setting where those assumptions are least likely to hold, and the Dinutuximab swing is what "second-order" looks like when they do not.
This is the layer of process rigour inside every serious ECA, alongside the fitness-for-purpose standard for regulatory-grade RWE and the regulatory context that drives you to build an external control at all. It is also the kind of comparator-data landscaping Inovia's RWE work and InovaSight are built to do early, while the data condition is still a decision rather than a discovery. If you have already run the single-arm trial and the questions are coming, the method is one of the first things worth pressure-testing.
The method is the cheapest error in the whole ECA to avoid, because it is settled by a fact you already know before you fit a single model. It is among the most expensive to discover in an HTA committee room, after the submission is filed. One parting oddity to leave you with: EMA's own 2025 reflection paper on real-world data runs to some length on confounding and effect modification, and never once names propensity scores, weighting or MAIC [14]. The operative method-choice logic lives on the HTA track. That is where to go looking for it.
Get the monthly digest
The 5 things evidence leads need to know each month: regulatory moves, RWE developments and what they mean in practice. No pitch, one email a month.
Article titles shown in quotation marks are given verbatim; entries without quotation marks carry a sourced descriptor pending title verification against the primary record.