Skip to content
Inovia Bio case study on external control arm acceptance: blinatumomab's FDA and EMA regulatory pathway from single-arm trial to confirmatory randomised trial.
Strategy Real world evidence external control arms

The external control arm that got checked against reality: what blinatumomab's acceptance actually required

Imi
Imi

An external control arm is an estimate of something nobody randomised. You take patients treated with the old standard of care, in another study, in another decade sometimes, and from them you build a counterfactual: what would have happened to the patients in front of you, had they never received your drug. Here is the quiet problem at the centre of the whole method. You almost never find out whether the counterfactual told the truth, because the only way to check it properly would have been to run the randomised trial and if you could have run that, you would not have needed the external control in the first place.

It is dead reckoning: you fix your position without a landmark, plot a course, and trust that the arithmetic was honest.

Blinatumomab (Blincyto), in relapsed or refractory Philadelphia-negative B-precursor acute lymphoblastic leukaemia, is the rare case where a landmark eventually appeared. The FDA granted accelerated approval in 2014 on a single-arm trial contextualised against a historical control, and EMA followed in 2015 [1]. But acceptance came with a condition: run the randomised trial anyway. That trial, TOWER, reported a few years later. So for once, we can go back and mark the homework.

The punchline, stated honestly before the mechanism: the external-control estimate pointed the same direction the randomised trial later confirmed, with confidence intervals that overlap heavily without coinciding. The historical comparison put the hazard ratio for death at 0.536. The randomised trial landed at 0.71. The direction held. The magnitude was overstated.

That gap is the whole lesson, and I will come back to it. First, the chain of decisions that got the external control accepted, because that is what regulators actually respond to. They accept a sequence of individually defensible choices, each of which a reviewer can pull on separately without the whole thing unravelling. "A strong design," judged as a single gestalt, is not a thing anyone at the agency signs.

Why this external control arm case, and what it is not

This is not a checklist. We have already written the checklist. FDA Guidance on External Control Arms: A Checklist For Drug Developers walks through what the FDA's 2023 draft guidance on externally controlled trials asks of you, criterion by criterion [20]. This post does the thing the checklist cannot. It takes one real acceptance apart to show which criteria carried the regulatory weight in practice, at the level of the specific choice.

The timeline anchors it. Accelerated approval, FDA 2014 and EMA 2015, on a single-arm trial plus a historical control. Full approval, 2017 and 2018, on the confirmatory randomised trial [2]. Two regulatory decisions, years apart, resting on two different grades of evidence for the same drug in the same indication. That structure is the gift here. It is why blinatumomab, almost alone, lets us set the external control's estimate and the randomised answer side by side.

Five decisions carried the weight, so take them one at a time.

Decision 1: the data source, registered as its own trial

Nobody scraped this historical control together once the single-arm trial had read out. It was a defined cohort: 2,373 patients screened across "six national study groups and five large treatment centers" in the US and Europe, adults aged 15 or older at their initial ALL diagnosis from 1990 onward, with eligibility criteria written to mirror the interventional trial's own population (Gökbuget, Kelsh, Chia et al. 2016) [3]. Of those, 694 contributed complete-response data and 1,112 contributed overall-survival data.

The load-bearing move: Amgen gave that cohort its own ClinicalTrials.gov registration, NCT02003612, cross-linked to the pivotal single-arm study, NCT01466179 (MT103-211) [4]. Call it the "registered counterfactual": a comparator treated with the same pre-specification discipline as an interventional protocol, its eligibility locked and its existence on the public record before the comparison was ever drawn.

So what, on Monday? Pre-specification and traceability sound like the abstract virtues you gesture at in a cover letter. Here they took physical form: a separate protocol, matched eligibility, registered and dated before anyone knew how flattering the comparison would turn out to be.

Decision 2: matching on the trial's own eligibility

Mirroring the eligibility criteria did something more specific than make the two groups look alike. It meant the same patient would have qualified for either arm, a much stronger claim than a tidy baseline table, and the claim a reviewer actually interrogates.

We have a rare window into how the FDA's statistician saw it. Reported via Khachatryan et al. 2023 (a secondary channel, worth flagging, and not the review document itself), an FDA statistical reviewer wrote: "Although retrospective historical studies may not be directly comparable to prospective clinical trials, each of the historical studies provided was conducted in a large number of patients; accounted for differences in patient characteristics between studies; and independently derived a CR rate not exceeding 30% for patients receiving salvage therapies." [5]

Read what actually persuaded the reviewer, and it was nothing to do with design elegance. The size of the cohorts, the adjustment for patient differences, and an independently reproducible ceiling: salvage therapy in this population does not get you past a roughly 30% complete-response rate. That ceiling is the anchor, and it is why the trial's numbers, sitting well above it, could be believed.

So what: match on whether the same patient would be eligible for both arms, and hand the reviewer a benchmark they can reproduce from outside your file. A tidy baseline table does neither.

Decision 3: endpoints harmonised, numbers reported with their exact denominators

The endpoints, complete response and overall survival, were defined identically across the records. That sounds procedural. In practice it is the difference between a clean comparison and an apples-to-oranges embarrassment.

It also demands discipline about which number you quote. The trial's complete-response rate appears in the literature as two different figures, and both are correct. FDA's own accelerated-approval account (Przepiorka et al. 2015) reports 32% CR among 185 evaluable adults (95% CI 26-40%) [6]. The trial publication (Topp et al. 2015) reports 33% CR, rising to 43% when complete responses with partial haematological recovery (CRh) are included, among 189 treated patients (95% CI 36-50%) [7]. Different denominators, different endpoint definitions: a distinction to state precisely rather than a discrepancy to hide.

Translation: report every rate with its denominator and its exact endpoint definition attached. Quote "43%" without saying "CR plus CRh, among 189 treated" and you have handed a reviewer a reason to wonder what else you quietly rounded.

Free download

The External Control Arm Design Checklist

The 12 design points regulators probe first, in one checklist you can run against your protocol before database lock.

Get the checklist →

Decision 4: two independent methods for confounding

Then came the confounding, and rather than adjust once and hope, Amgen triangulated with two independent methods, both fully described.

The first was a weighted analysis, stratified into six prognostic strata, then pooled and reweighted to the trial's own stratum proportions. The result: a historical complete-response rate of 24% (95% CI 20-27%) against the trial's 43%, and median overall survival of 3.3 months against 6.1 [8].

The second was a propensity-score analysis using inverse-probability-of-treatment weighting across eight covariates. It put the historical complete-response rate at 26.7% against 49.3%, and the odds ratio at 2.68 (95% CI 1.67-4.31). And the number the whole file builds toward: an overall-survival hazard ratio of 0.536 (95% CI 0.394-0.730) favouring blinatumomab (Gökbuget, Kelsh, Chia et al. 2016) [9].

Two methods, same direction, similar magnitude. That is the defence against the single most obvious challenge to any adjusted comparison, which is that you went fishing and reported the one analysis that flattered you. One method invites the accusation; two, pre-specified and concordant, largely retire it.

A caution the record demands. This analysis is the Blood Cancer Journal paper by Gökbuget, Kelsh, Chia and colleagues, the historical-control study built for the submission. There is a second, larger Gökbuget 2016 paper in Haematologica, a broader international reference analysis published from the same registered data collection (NCT02003612); it reports a different, differently-selected population and different headline numbers, so conflating the two is an easy mistake to make [10]. You will also sometimes see a much lower historical CR rate quoted for this programme. I have left it out on purpose: it comes from a separate model-based meta-analysis, a "synthetic control arm" rather than the historical-control study (Khachatryan et al. 2023), so setting it beside this dataset would conflate two different methods. The 24% weighted figure is the one that belongs here.

So what: pre-specify at least two adjustment methods, then show they agree. To a reviewer, concordance across methods is worth more than precision within any single one.

Decision 5: the confirmatory trial, agreed as the price of approval

Here is the decision most teams underweight. Accelerated approval was not the end of the conversation. FDA's account states the condition plainly (Przepiorka et al. 2015): "A randomized trial is required in order to confirm clinical benefit." [11]

That sentence is why TOWER exists as a formal post-marketing commitment, written into the approval rather than volunteered after it. TOWER (NCT02013167) randomised 405 patients 2:1, 271 to blinatumomab and 134 to standard-of-care chemotherapy, using the same population definition and endpoint set as the two earlier studies. The primary endpoint, overall survival, came in at a hazard ratio of 0.71 (95% CI 0.55-0.93, p=.012): median survival of 7.7 months against 4.0, complete response 34% against 16%. The data monitoring committee recommended stopping early for benefit (Kantarjian et al. 2017, corroborated in the FDA-authored account, Pulte et al. 2018) [12].

What that means in practice: treat an accepted external control as a down payment on approval, with the balance still owed, and build the confirmatory randomised trial into your integrated evidence plan from the first meeting. More often than not the regulator will require it anyway, and you want to have chosen its population and endpoints yourself.

The homework, marked honestly

Now the part that makes this case rare. Let's be honest about what it does and does not show.

Put the two estimates next to each other. The external control, before the randomised trial existed, gave an overall-survival hazard ratio of 0.536 (95% CI 0.394-0.730). TOWER, the randomised trial, gave 0.71 (95% CI 0.55-0.93). The direction is the same, and the intervals overlap substantially: the range 0.55 to 0.73 sits comfortably inside both. But they do not coincide, and the external-control estimate was the more optimistic of the two. The method got the answer's direction and rough size right. It did not land on the number [13].

One thing has to be said plainly, because it is easy to imply otherwise and wrong to. This side-by-side is my own construction, assembled from two separately published papers for the sake of the lesson. It is not a comparison the FDA itself drew. The agency's retrospective account (Pulte et al. 2018) treats TOWER as the confirmatory answer on its own terms; it does not go back and re-score the historical estimate against it. The independent methodology literature has noted the two estimates are "quite close" (Seeger et al. 2020) [14], and that qualitative read is theirs. The precise numeric pairing is mine, offered transparently as exactly that.

Directionally right. Not exactly right. That is what "the external control worked" honestly looks like once you can finally check, and it is the reason you should never present an external-control estimate, to a regulator or to your own board, as the final number. It is a good bearing and nothing more.

A different route to the same outcome

Cerliponase alfa (Brineura), for CLN2 disease, reached a comparable acceptance by a visibly different path, and it complicates the tidy list above.

BioMarin never registered its comparator as a standalone trial. There was no "registered counterfactual" here. The historical comparison lived inside the pivotal and extension studies' own outcome definitions (NCT01907087 and NCT02485899), analysed by Cox proportional hazards (Schulz et al. 2018) [15]. So the separately registered comparator that looked so load-bearing for blinatumomab turns out not to be a precondition at all. It was one good way to prove traceability, among several.

That said, what both cases share is a natural-history dataset regulators had inspected and trusted. For Brineura that was DEM-CHILD, and the one independently confirmed regulator statement in this case is worth quoting directly (Nickel et al. 2018): "The DEM-CHILD core data were inspected and approved by the US Food and Drug Administration and European Medicines Agency." [16]

EMA authorised Brineura on 30 May 2017 under exceptional circumstances, a specific EU pathway distinct from a conditional marketing authorisation, used because, in the agency's words, "it has not been possible to obtain complete information about Brineura due to the rarity of the disease." EMA reported that 20 of the 23 treated children (87%) "did not experience the 2-point decline in movement and language skills seen historically in patients not receiving treatment." [17]

The granular matching detail here, with matched controls selected on exact motor-language score, age within three months, and genotype, reaches me through CADTH's synthesis of the FDA and EMA material (CADTH 2019) [18], not from FDA or EMA text I read directly. I flag it because provenance matters in a piece about provenance. And CADTH is worth quoting for a second reason: it carries the honest ceiling. Cerliponase alfa's effect estimates are not clean. The unmatched comparison of 23 treated against 42 historical controls produced a hazard ratio of 0.08 (95% CI 0.02-0.23), a very large apparent effect; the smaller matched subset of 17 showed a real but more modest difference. CADTH's own verdict on the exercise is one every external-control advocate should keep pinned up: "comparison with a historical control group cannot produce results that are as reliable as those within a randomized study." [19]

Two accepted external controls, two different construction methods, and the same sober ceiling on what either one could prove.

"This is survivorship bias with a case number"

The strongest objection to everything above deserves stating at full strength. Blinatumomab had a large effect in a desperate population with no good treatment. Of course the external control held up — an effect that big is hard to fake. A sponsor with a 1.3-fold effect in a crowded indication gets none of this grace, and dressing a winner's file up as a "chain of decisions" is a just-so story told with the benefit of hindsight. The decisions did not make the drug work. The drug worked, and the decisions came along for the ride.

Much of that is fair, and it is exactly the point. Magnitude did heavy lifting. The FDA reviewer's ≤30% complete-response ceiling matters precisely because the trial's 43% sits so far above it; move the trial down to 35% and that argument becomes a great deal harder to win. But notice that even here, with a large effect, the external estimate still drifted optimistic, 0.536 against a confirmed 0.71. The comfortable reading is that external controls work. The honest reading is narrower: where they are accepted, they are accepted decision by decision, and the smaller your effect relative to a well-characterised natural history, the more airtight every link in the chain has to be, and the more certain you should be that a randomised trial is coming whether you volunteer it or not. Our piece on the EMA's single-arm trial scrutiny makes the same case from the regulator's side of the table.

What to take to Monday's meeting

  • Register the comparator as its own protocol with matched eligibility, and lock it before the single-arm trial reads out. Brineura shows it isn't mandatory. It remains the cleanest way to prove pre-specification to a sceptical reviewer.
  • Match on eligibility, not on cosmetics. The test is whether the same patient would have been eligible for both arms, not whether the baseline tables look tidy.
  • Report every rate with its exact denominator and endpoint definition. "43%" is not a number a reviewer can use; "43% CR+CRh among 189 treated" is.
  • Triangulate confounding with at least two pre-specified adjustment methods, and show they agree. Concordance beats precision.
  • Assume the confirmatory randomised trial and design it in from day one. Accelerated approval on an external control is a down payment, and you want to have chosen the confirmatory trial's population and endpoints yourself.
  • Never present your external-control estimate as the final number. It is a bearing, not a fix.

The marked homework is the rarest thing in this field, and its verdict is worth having even though it sounds modest: a well-built external control can be trusted to point the right way. It will not land you on the exact figure, and you should never present it as if it could, to a regulator or to your own board. Being able to say even that much, with the receipts behind it, is not nothing. It is honest, and honest is what survives a reviewer. If you are building, or defending, one of these decision chains and want it to hold up under the kind of reading an FDA or EMA reviewer will give it, that is the work we do at Inovia Bio. It reads well alongside what to do when regulators start asking questions of a single-arm dataset.

Get the monthly digest

The 5 things evidence leads need to know each month: regulatory moves, RWE developments and what they mean in practice. No pitch, one email a month.

References

  1. Przepiorka D, Ko CW, Deisseroth A, et al. (2015). Clinical Cancer Research. PMID: 26374073. https://pubmed.ncbi.nlm.nih.gov/26374073/ — with Khachatryan A, Read SH, Madison T. (2023). Journal of Pharmacokinetics and Pharmacodynamics. PMID: 37095406. https://pubmed.ncbi.nlm.nih.gov/37095406/
  2. Pulte ED, Vallejo J, Przepiorka D, et al. (2018). The Oncologist. PMID: 30018129. https://pubmed.ncbi.nlm.nih.gov/30018129/ — with Khachatryan et al. (2023), PMID: 37095406.
  3. Gökbuget N, Kelsh M, Chia V, et al. (2016). Blood Cancer Journal. PMC5056974 (PMID: 27662202). https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5056974/ — historical control registered as ClinicalTrials.gov NCT02003612. https://clinicaltrials.gov/study/NCT02003612
  4. Amgen. Registered historical comparator, ClinicalTrials.gov: NCT02003612 (https://clinicaltrials.gov/study/NCT02003612), cross-linked to the pivotal single-arm study NCT01466179 / MT103-211 (https://clinicaltrials.gov/study/NCT01466179). Gökbuget et al. (2016), PMC5056974; Pulte et al. (2018), PMID: 30018129.
  5. Khachatryan A, Read SH, Madison T. (2023). Journal of Pharmacokinetics and Pharmacodynamics. PMID: 37095406. https://pubmed.ncbi.nlm.nih.gov/37095406/
  6. Przepiorka D, Ko CW, Deisseroth A, et al. (2015). Clinical Cancer Research. PMID: 26374073. https://pubmed.ncbi.nlm.nih.gov/26374073/
  7. Topp MS, Gökbuget N, Stein AS, et al. (2015). Lancet Oncology. PMID: 25524800. https://pubmed.ncbi.nlm.nih.gov/25524800/ — corroborated by Gökbuget et al. (2016), PMC5056974.
  8. Gökbuget N, Kelsh M, Chia V, et al. (2016). Blood Cancer Journal. PMC5056974. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5056974/
  9. Gökbuget N, Kelsh M, Chia V, et al. (2016). Blood Cancer Journal. PMC5056974. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5056974/
  10. Gökbuget N, Kelsh M, Chia V, et al. (2016). Blood Cancer Journal, PMC5056974 (https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5056974/) and Haematologica, PMID: 27587380 (https://pubmed.ncbi.nlm.nih.gov/27587380/) — companion analyses of the same registered data collection (NCT02003612); "synthetic control arm" figures reported via Khachatryan et al. (2023), PMID: 37095406.
  11. Przepiorka D, Ko CW, Deisseroth A, et al. (2015). Clinical Cancer Research. PMID: 26374073. https://pubmed.ncbi.nlm.nih.gov/26374073/
  12. Kantarjian H, Stein A, Gökbuget N, et al. (2017). "Blinatumomab versus Chemotherapy for Advanced Acute Lymphoblastic Leukemia" (TOWER). New England Journal of Medicine. PMID: 28249141. https://pubmed.ncbi.nlm.nih.gov/28249141/ — with Pulte et al. (2018), PMID: 30018129; ClinicalTrials.gov NCT02013167 (https://clinicaltrials.gov/study/NCT02013167).
  13. Author's own construction from separately published papers — Gökbuget et al. (2016), PMC5056974; Kantarjian et al. (2017), PMID: 28249141; Pulte et al. (2018), PMID: 30018129. Not a comparison drawn by the FDA.
  14. Seeger JD, et al. (2020). Pharmacoepidemiology and Drug Safety. PMID: 32964514. https://pubmed.ncbi.nlm.nih.gov/32964514/
  15. Schulz A, Ajayi T, Specchio N, et al. (2018). New England Journal of Medicine. PMID: 29688815. https://pubmed.ncbi.nlm.nih.gov/29688815/ — ClinicalTrials.gov NCT01907087 (https://clinicaltrials.gov/study/NCT01907087) and NCT02485899 (https://clinicaltrials.gov/study/NCT02485899).
  16. Nickel M, Simonati A, Jacoby D, et al. (2018). Lancet Child & Adolescent Health. PMID: 30119717. https://pubmed.ncbi.nlm.nih.gov/30119717/
  17. European Medicines Agency. Brineura (cerliponase alfa): EPAR overview. Authorised under exceptional circumstances, 30 May 2017. https://www.ema.europa.eu/en/medicines/human/EPAR/brineura
  18. CADTH. Clinical Review Report: Cerliponase Alfa (Brineura). Ottawa: Canadian Agency for Drugs and Technologies in Health; June 2019. NCBI Bookshelf NBK544815.
  19. Schulz A, Ajayi T, Specchio N, et al. (2018). New England Journal of Medicine. PMID: 29688815. https://pubmed.ncbi.nlm.nih.gov/29688815/ — with Khachatryan et al. (2023), PMID: 37095406, and CADTH (June 2019), NCBI Bookshelf NBK544815.
  20. US Food and Drug Administration. "Considerations for the Design and Conduct of Externally Controlled Trials for Drug and Biological Products." DRAFT guidance for industry, February 2023 (Docket FDA-2022-D-2983).

Trials referenced by ClinicalTrials.gov identifier: NCT01466179 (blinatumomab pivotal single-arm, MT103-211), NCT02003612 (the registered historical control), NCT02013167 (TOWER, confirmatory RCT), NCT01907087 and NCT02485899 (cerliponase alfa pivotal and extension studies).

Share this post