Skip to content
Decision-tree diagram showing how IPD availability determines whether to use propensity score matching, weighting, or MAIC as the adjustment method for an external control arm.

PSM, weighting or MAIC? The one fact that decides which adjustment method you're allowed to use for an external control arm

Imi
Imi

Two companies. The same two drugs: secukinumab against adalimumab in ankylosing spondylitis. Novartis ran a matching-adjusted indirect comparison on its own patient-level data and concluded secukinumab came out ahead. AbbVie ran a matching-adjusted indirect comparison on its own patient-level data and concluded adalimumab held its ground on cost. Same category of method, same disease, opposite verdicts, because the two analyses balanced on different sets of covariates [1]. The same failure is documented in the peer-reviewed NICE methods literature [2].

The reflex is to file it under "the statisticians disagreed" and move on. The adjustment method feels like a downstream technicality, something the biostatistician settles after the trial reads out. The paradox says otherwise. The method and the covariate list are two separate levers, and either one, on its own, can flip your conclusion. Underneath sits a more useful fact: the method is fixed for you by one checkable feature of your data, decided before a single model is fitted, whether or not anyone on the team realises the decision has already been made. Get that wrong and you have bought a second, quieter error, one hiding behind the tidy covariate table everyone will actually audit.

This post is about that second error. It sits underneath the ECA design decisions in our external control arm checklist; assume here that the comparator source and design are already settled, and the only question left is which statistical adjustment to apply.

The one fact that picks the method

Ask one question before anything else: do you hold individual patient-level data (IPD), and for which arms? The answer sorts you into one of three conditions, and each condition licenses a different method.

  • IPD for both arms. You have patient-level records for your own trial and for the comparator cohort. Propensity score matching (PSM) and weighting (IPTW/IPSW) are on the table. This is the world NICE DSU Technical Support Document 17 (2015, final) was written for [3].
  • IPD for your own trial only, comparator available as published aggregate data. MAIC exists precisely for this. NICE DSU Technical Support Document 18 (2016, final) sets it out as the method for the case where a company holds IPD from its own trials but only "aggregate outcomes" from a competitor's [4].
  • IPD for neither side. Aggregate-only on both. Here the honest answer, per the 2024 EU HTA Coordination Group guideline, is that there is no adequate method [5].

One plain sentence on what each method actually does, because you have to defend the choice, not derive it. PSM pairs each treated patient with one or more comparator patients who had a similar modelled probability of being treated, then compares the matched sets. Match, then compare: that is the whole idea. Weighting keeps every patient but weights each by the inverse of their probability of being in the group they are in, building a balanced pseudo-population [6][7]. MAIC reweights your own IPD so its covariate averages line up with the comparator's published baseline characteristics, then compares. It exists because the comparator's IPD is not yours to have.

None of this is Inovia's invention. It is the operative logic of TSD 17, TSD 18 and the 2024 EU HTA CG guideline, three independent official documents drawing the same line in the same place [3][4][5]. Nor are these the only tools in the drawer: hybrid designs that blend a concurrent arm with static or dynamic borrowing are a fourth pattern, registered in real protocols [8]. But for the standard ECA comparison, the three-way gate is the decision.

When you hold both hands of cards: PSM and weighting

This is the clean condition, with a textbook worked example in the peer-reviewed literature. A published analysis of the APT trial matched 406 treated patients 1:1 against 1,366 patients pooled from other trials, using propensity scores built on five covariates: age, tumour stage, oestrogen-receptor status, progesterone-receptor status and histological grade [9]. Full IPD on both sides, matched on named prognostic factors. That is the easy branch when the data actually supports it.

Here is the counter-intuitive part. PSM is used less than its data-condition fit would suggest. TSD 17's own review of NICE appraisals built on non-randomised comparative IPD found only 2 of 16 used propensity score matching; 6 used multivariate regression, and 7 applied no adjustment at all [3]. Having the gold-standard data condition does not mean teams reach for the gold-standard method.

That said, having it does not guarantee a better outcome either. Krüger, Cantoni and Van Engen's 2025 feasibility study of EU oncology approvals built on non-randomised data found that propensity-score-based methods did not lead to more positive HTA outcomes than the alternative unanchored MAICs [10]. Sit with that: the "preferred" data condition, the "preferred" method, and still no reliable payoff at the committee. So even on the easy branch, do not treat the method as settled by the data alone. IPTW is sensitive to a handful of extreme weights, matching quality has to be demonstrated, and the choice still has to be justified rather than assumed.

Free download

The External Control Arm Design Checklist

The 12 design points regulators probe first, in one checklist you can run against your protocol before database lock.

Get the checklist →

When you hold only your own hand: MAIC and the bill it hands you

MAIC is the method for the condition most biotechs actually face: your IPD, the competitor's published numbers, nothing more. The founding paper reweighted one trial's patients to match another's reported baseline characteristics, comparing adalimumab against etanercept without ever touching the etanercept IPD [6]. It is a genuine piece of engineering. It also hands you a bill, in three parts.

Part one: your sample quietly shrinks. Reweighting costs you effective sample size (ESS), and the cost is not small. TSD 18's own review of three reporting papers found an average ESS reduction of around 80% (range 57–98%) [4]. Separately, and from a different and broader sample, Phillippo et al.'s 2019 audit of 268 NICE technology appraisals found that among the nine MAIC-using appraisals that reported ESS, the median reduction was 74.2% (range 7.9–94.1%) [11]. Those are two different figures from two different samples. Do not average them into one; take them as two independent readings of the same uncomfortable thing.

Years ago I worked on an early-phase rare-disease programme where the only comparator in existence was a published aggregate cohort. No IPD to request, nobody to request it from. MAIC was the only lawful option we had. When we reweighted our patients to match the published baseline means, the effective sample size fell to a sliver of what we had enrolled. The trial had not got smaller. The confidence we were entitled to claim from it had. That is the phantom precision reweighting buys you: a tidy-looking interval drawn on a sample that has quietly collapsed underneath it.

Part two: anchored is necessary, not sufficient. An anchored comparison uses a common comparator to cancel out prognostic differences; an unanchored one cannot, and leans on the heroic assumption that you have measured and balanced every effect modifier. But the right category is not a safe harbour. In the pegcetacoplan NICE appraisal the company used a textbook-correct anchored MAIC, with eculizumab as the common comparator, and the committee still judged the results biased because key effect modifiers, including haemoglobin levels and transfusion history, could not be adjusted for across the two trials [12]. Right method category, wrong outcome, for a reason the anchored/unanchored choice does not fix.

Part three: the numbers can lie about their own certainty. A 2020 simulation of unanchored MAIC found 95% confidence-interval coverage of 11.2% unweighted, 93.8% when fully weighted with every covariate balanced, and 39.1% when only partially weighted. None reached the nominal 95%, even in the fully weighted best case [13]. And unanchored is not the exception in the real world; it describes the large majority of NICE population-adjustment applications. So if MAIC is forced on you, pre-specify your effect modifiers, report ESS honestly, and treat "we used an anchored MAIC" as the opening line of your defence, not the closing one.

When you hold neither hand: the answer nobody wants

Let's be frank about the case nobody wants to be in. Aggregate data on your side, aggregate data on the comparator's, no patient-level records anywhere. The 2024 EU HTA Coordination Group Methodological Guideline, adopted under Regulation (EU) 2021/2282, states it plainly: when non-randomised evidence is available only at the aggregated data level, "there are no adequate method available for reliable estimation of treatment effectiveness" [5].

There is no statistical rescue for a data condition that cannot carry one.

This is the branch the tool-first mindset skips straight past, because it assumes every problem is a modelling problem. It is not. If you are heading into aggregate-on-both-sides territory, that is a design red flag to fix while you still can, by securing comparator IPD or reconsidering the comparator source, not a puzzle to hand to the statistician after the fact.

One clarification, because the tracks get conflated. The EU HTA CG guideline sits on the HTA and Joint Clinical Assessment side, not the marketing-authorisation side. It is not EMA guidance. EMA's own 2025 reflection paper on real-world data (EMA/99865/2025, final) is silent on which adjustment method to use [14]. That silence is not evidence the question is open. It is resolved, on the HTA side, by the document above.

The receipt: what the method cost Dinutuximab Beta

If you want the cost of treating the method as a downstream detail in plain currency, read the Dinutuximab Beta appraisal. Across changes to the comparison method and corrections to its implementation, the incremental cost-effectiveness ratio moved from the company's naive indirect comparison at £22,338 per QALY, to an ERG estimate of £111,858, to a committee-revised MAIC at £24,661, to a DSU-corrected model landing between £62,886 and £87,164 [15]. Same underlying clinical data throughout. Only the comparison method and how it was run changed.

A swing of that size, from the method alone, is the reimbursement decision.

"But isn't the method second-order?"

Here is the strongest version of the objection, and it is not a strawman. A good methodologist will tell you the covariate list and the overlap between populations do the real work; the mechanical choice of matching versus weighting versus reweighting is second-order. And there is evidence for it. Remiro-Azócar and colleagues' 162-scenario simulation found MAIC accurate when its assumptions were met [16], and Signorovitch and colleagues' 2023 validation showed a MAIC reproducing the result of a real head-to-head RCT in psoriasis [17]. When the data condition holds and the assumptions hold, the method genuinely is the smaller worry.

Agreed. That is exactly the conditional, and it is the whole argument. The method is second-order only once the gate is cleared and the assumptions actually hold. The point of the gate is to check whether they do, before you rely on them. The large majority of real NICE population-adjustment applications are unanchored, precisely the setting where those assumptions are least likely to hold, and the Dinutuximab swing is what "second-order" looks like when they do not.

What to do on Monday

  1. Establish comparator IPD availability first. It picks your method before any modelling starts. Do this before you fall in love with a comparator source, not after.
  2. If it is aggregate-only on both sides, treat it as a design constraint, not a method problem. Fix it upstream, where fixing is still possible. There is no adjustment that manufactures the data you do not have.
  3. If MAIC is forced on you, pre-specify effect modifiers and report ESS. Use an anchored comparison wherever a common comparator exists. Then remember it is necessary, not sufficient, and plan for the missing-effect-modifier challenge before the committee raises it.
  4. Budget for the shifted null. The 2024 EU HTA CG guideline requires MAIC- and PSM-based comparisons to clear a confidence-interval threshold shifted away from zero, not merely to clear zero, pricing in the adjustment's own added uncertainty. That mechanism is absent from the older NICE TSDs [5]. Power your study against that harder bar, or it will read underpowered when the committee does the sums.

This is the layer of process rigour inside every serious ECA, alongside the fitness-for-purpose standard for regulatory-grade RWE and the regulatory context that drives you to build an external control at all. It is also the kind of comparator-data landscaping Inovia's RWE work and InovaSight are built to do early, while the data condition is still a decision rather than a discovery. If you have already run the single-arm trial and the questions are coming, the method is one of the first things worth pressure-testing.

The method is the cheapest error in the whole ECA to avoid, because it is settled by a fact you already know before you fit a single model. It is among the most expensive to discover in an HTA committee room, after the submission is filed. One parting oddity to leave you with: EMA's own 2025 reflection paper on real-world data runs to some length on confounding and effect modification, and never once names propensity scores, weighting or MAIC [14]. The operative method-choice logic lives on the HTA track. That is where to go looking for it.

Get the monthly digest

The 5 things evidence leads need to know each month: regulatory moves, RWE developments and what they mean in practice. No pitch, one email a month.

References

Article titles shown in quotation marks are given verbatim; entries without quotation marks carry a sourced descriptor pending title verification against the primary record.

  1. Jiang et al. (2025). Named "MAIC paradox" analysis, secukinumab vs adalimumab in ankylosing spondylitis. Research Synthesis Methods. PMID: 41626938. https://pubmed.ncbi.nlm.nih.gov/41626938/
  2. Phillippo et al. (2018). NICE DSU TSD 18 companion review of population-adjustment methods in NICE technology appraisals; documents the ankylosing-spondylitis dueling-MAIC instance. Medical Decision Making. PMID: 28823204. https://pubmed.ncbi.nlm.nih.gov/28823204/
  3. Faria R, Hernández Alava M, Manca A, Wailoo AJ. "NICE DSU Technical Support Document 17: The use of observational data to inform estimates of treatment effectiveness in technology appraisal" (2015, final). NICE Decision Support Unit.
  4. Phillippo DM, Ades AE, Dias S, Palmer S, Abrams KR, Welton NJ. "NICE DSU Technical Support Document 18: Methods for population-adjusted indirect comparisons in submissions to NICE" (2016, final). NICE Decision Support Unit.
  5. EU Health Technology Assessment Coordination Group. Methodological guideline on direct and indirect comparisons for Joint Clinical Assessment (2024, final; adopted 8 March 2024 under Regulation (EU) 2021/2282).
  6. Signorovitch JE, et al. (2010). "Comparative effectiveness without head-to-head trials: a method for matching-adjusted indirect comparisons applied to psoriasis treatment with adalimumab or etanercept." PharmacoEconomics. PMID: 20831302. https://pubmed.ncbi.nlm.nih.gov/20831302/
  7. Austin PC. (2011). "An Introduction to Propensity Score Methods for Reducing the Effects of Confounding in Observational Studies." Multivariate Behavioral Research. PMID: 21818162. https://pubmed.ncbi.nlm.nih.gov/21818162/
  8. Externally controlled study registering a hybrid concurrent-plus-borrowing design (fourth pattern, distinct from PSM/weighting/MAIC). ClinicalTrials.gov: NCT07337447. https://clinicaltrials.gov/study/NCT07337447
  9. Amiri-Kordestani L, et al. (2020). FDA-authored external-control analysis of the APT trial (406 treated patients matched 1:1 to 1,366 pooled-trial patients on five covariates); cited here only as a published example of the IPD-on-both-arms PSM condition. Annals of Oncology. PMID: 32866625. https://pubmed.ncbi.nlm.nih.gov/32866625/
  10. Krüger, Cantoni & Van Engen (2025). Feasibility study of propensity-score methods vs unanchored MAIC across 2020–2023 EU oncology HTA outcomes. International Journal of Technology Assessment in Health Care. PMC11718718. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC11718718/
  11. Phillippo DM, et al. (2019). Audit of population-adjustment methods across 268 NICE technology appraisals; ESS-reduction figures. International Journal of Technology Assessment in Health Care. PMID: 31190671. https://pubmed.ncbi.nlm.nih.gov/31190671/
  12. Pegcetacoplan NICE appraisal / ERG analysis (2023). Anchored MAIC (common comparator eculizumab) judged biased on unmeasured effect modifiers. PharmacoEconomics Open. PMID: 37195551. https://pubmed.ncbi.nlm.nih.gov/37195551/
  13. Jiang Y, Ni W. (2020). Simulation of unanchored MAIC 95% CI coverage for single-arm time-to-event comparisons. BMC Medical Research Methodology. PMID: 32993519. https://pubmed.ncbi.nlm.nih.gov/32993519/
  14. European Medicines Agency. "Reflection paper on use of real-world data in non-interventional studies to generate real-world evidence for regulatory purposes" (EMA/99865/2025, final, published 3 April 2025).
  15. Pennington M, et al. (2019). NICE case study of adjustment-method impact on the Dinutuximab Beta ICER. PharmacoEconomics. PMID: 30465228. https://pubmed.ncbi.nlm.nih.gov/30465228/
  16. Remiro-Azócar A, et al. (2021). 162-scenario simulation of population-adjustment methods under limited IPD access; MAIC accurate when assumptions met. Research Synthesis Methods. PMID: 34196111. https://pubmed.ncbi.nlm.nih.gov/34196111/
  17. Signorovitch J, et al. (2023). Validation of a MAIC against a head-to-head randomised trial in psoriasis. Journal of Dermatological Treatment. PMID: 36724798. https://pubmed.ncbi.nlm.nih.gov/36724798/

Share this post