Skip to content
Twelve-point external control arm design checklist for FDA and EMA regulatory submissions
RWE Strategy Real world evidence

The external control arm design checklist: 12 points that survive scrutiny

Imi
Imi

An external control arm's verdict is written into its design long before anyone measures an outcome. Here is the cleanest demonstration I know.

In a blinatumomab analysis in relapsed Philadelphia-chromosome-negative acute lymphoblastic leukaemia, correcting a single design choice moved the hazard ratio for death from 0.56 (95% CI 0.47–0.67) to 0.98 (95% CI 0.83–1.15) [1]. The choice was how time zero, the moment of cohort entry, got defined: matched line of salvage treatment versus latest line. Same drug. Same 189 treated patients, same 1,112 historical controls, same data. The apparent halving of mortality lasted exactly as long as the clock was set wrong.

The clock did that, not the drug.

That is the argument of this whole post, so let me state it flat. External control arms are not rejected because they lack randomisation, and they are not accepted because someone built a persuasive case about cost and speed. They survive scrutiny, or they don't, on a small and stubbornly repeatable set of design decisions locked in before data collection starts: timing, population, how you measure the outcome, and whether you committed the analysis to writing before you saw the answer.

Those are checklist-able. The dice framing, the belief that some ECAs get through and some don't and you won't really know until review, is the story teams tell themselves so they can skip the checklist and act surprised later.

This is the layer beneath the guidance overview. If you want the map of what the FDA's draft ECA guidance says exists, we wrote that already. This post assumes it and goes to the mechanics: the choices that decide whether your comparison holds. And it is aimed squarely past the ROI pitch, because the cost-and-speed case gets you to want an ECA and has nothing to do with whether the one you build survives. If you have already run a single-arm trial and the regulators are now asking pointed questions, you are past wanting one. You are defending the design you have.

Cluster 1 Designing the clock and the cohort (who, and when)

Most ECA failures are decided here, in who sits in each arm and when their clock starts, before a single statistic is run.

1. Get time zero and concurrency right. Timing is the most destructive axis on this list, which is why it opens the checklist. The blinatumomab flip above was pure timing. Omburtamab is the cautionary tale with a vote attached. In October 2022 an FDA advisory committee voted 16–0 that there was insufficient evidence the drug improved overall survival. Trade coverage of that public meeting reported the FDA review's concerns: a control that ran 1990 to 2015 while the trial enrolled from 2004, and confounding by differential post-relapse radiotherapy. The same coverage quoted the agency calling the external control "not fit for purpose" [2]. Treat that FDA-attributed detail as reported, not as a line anyone here has read inside an FDA document. Sixteen to zero.

Monday: define time zero identically in both arms and pre-register the definition. If your comparator predates your trial by a treatment era, stop and find another one.

2. Baseline comparability, and the bar it has to clear. Zolgensma cleared this bar without a clean match, and it is worth understanding why. Reviewing the STR1VE-US trial (NCT03306277) against the PNCR natural-history cohort used as its comparator, CADTH documented "several differences" in age at symptom onset, CHOP INTEND score and feeding or ventilatory support that "could impact response to treatment or outcome regardless of treatment" [3]. Accepted anyway. The effect was large, untreated SMA type 1 is uniformly fatal and its course is well documented, and that combination is the ICH E10 threshold for tolerating an imperfect control.

Monday: tabulate every prognostic variable across both arms now, at protocol stage. Where you cannot match, you had better be able to argue the effect swamps the residual gap.

3. Align eligibility criteria, don't gather and hope. The ECA population has to be built to your trial's inclusion and exclusion criteria. Tafasitamab's RE-MIND comparator matched on nine covariates, which sounds thorough until you read CADTH's finding that other known confounders were not accounted for: ECOG performance status, IPI score, cell of origin. That left what the review called "substantial risk of bias" [4]. Arondekar et al.'s census of FDA critiques names inclusion and exclusion matching to the trial as a recurring complaint [5].

Monday: apply your trial's I/E criteria to the comparator source and document the attrition at each step. The patients you drop are as informative as the ones you keep.

4. Vet the data source before you commit to it. Not every dataset that exists is fit to be a control. Selinexor is the blunt lesson: the approval rested on the single-arm STORM trial, because, as Yap et al. report it, "the FDA identified methodological issues with the EHR data and results from the observational study were deemed inadequate to support regulatory decision making" [6]. A dataset's existence is not its suitability. Historical trial data and RWD both qualify as sources, with different strengths, and the work of proving a source fit-for-purpose against your endpoints is what regulatory-grade actually means.

Monday: landscape and vet the source against your endpoints first. Fit the dataset to the study, never the study to whatever dataset is convenient.

Free download

The External Control Arm Design Checklist

The 12 design points regulators probe first, in one checklist you can run against your protocol before database lock.

Get the checklist →

Cluster 2 — One ruler, both arms (what you measure, and how)

A comparison is only as honest as its measuring instrument, and the instrument has to be the same on both sides. You calibrate it once. You do not get to recalibrate after you have seen the result.

5. Ascertain the outcome the same way in both arms. Same instrument, same definition, same assessment method. Cerliponase alfa (NCT01907087) is the tidy example: the registry states the design intent up front, and the motor-language subscale of the Hamburg rating scale was collected the same way in the treated cohort and the natural-history comparator, a matched instrument you can point a reviewer to. The cost of getting this wrong is not small. Simulated outcome-misclassification produced a relative bias of 67% even at 0.9 sensitivity and 0.9 specificity, rising to 156% and 237% as those fell to 0.8 and 0.7 [7].

Monday: confirm the comparator's outcome was captured the same way yours will be. If it wasn't, don't use it.

6. Reach for objective endpoints where the indication allows. Investigator-assessed, subjective endpoints are where an unblinded ECA leaks bias, because the arm without blinding is the one assembled from records, and records don't argue back. ICH E10 warns about exactly this, and the EMA's single-arm trial reflection paper is pointedly cautious about time-to-event endpoints read without a control. Anchor to something hard: overall survival, event-free survival, a response defined unambiguously rather than by a clinician's gestalt. Subjectivity is where hope creeps in.

Monday: pick the most objective endpoint your indication supports, and blind outcome assessment to treatment status wherever you can.

7. Plan for missing data before it goes missing. Arondekar's census names "plans to handle missing data" as a recurring FDA critique [5], and the asymmetry is easy to build in without noticing. In the golodirsen programme (NCT02310906), the concurrent untreated comparator group followed an abbreviated assessment schedule and had no muscle biopsies, a data-provenance gap baked straight into the design that limits how far the comparator can carry inference.

Monday: write the missing-data strategy into the SAP, and check the comparator source isn't systematically missing the very things your trial arm will capture.

Cluster 3 Commit the analysis before you see the answer (how you compare)

Pre-specification is the single most repeated FDA critique in the whole evidence base, and it is entirely self-inflicted. Decide the analysis while you are still blind to what it will show. Decide it after, and you are not testing a hypothesis. You are shopping for one that fits.

8. Pre-specify the confounding-control strategy, and don't oversell it. Even a good adjustment leaves a residue. Gupta et al. reconstructed real randomised control arms as ECAs and measured the gap: a mean log-hazard-ratio difference of 0.247 unadjusted (a ratio of hazard ratios of 1.36), falling to 0.139 after adjusting for measured confounders (1.22), and 0.098 even after quantitative bias analysis for the unmeasured ones (1.17). The authors call that a best-case scenario [8]. The residue never reaches zero. That said, propensity-score methods are routinely accepted by FDA reviewers in practice, and just as routinely contested in the methodological literature, as Jahanshahi et al. note in citing Elze and colleagues and King and Nielsen [9]. So present them as regulator-accepted, not as a settled fix. Monday: name your adjustment method and covariate set in the SAP before the comparison, and pre-plan a sensitivity or bias analysis for the confounding you can't measure.

9. Lock the matching and statistical approach before database lock: the "retrofitted control" trap. Here is the failure family that recurs under case after criticised case. I call it the retrofitted control: a comparator assembled after the single-arm trial has read out, chosen once you already know the number you are trying to beat. Omburtamab, selinexor and the exon-skipping comparators all wear versions of it. A comparison specified after you have seen the trial result is a comparison a reviewer discounts on sight, and they are right to.

Years ago I worked on an early-phase rare-disease programme where we did the opposite on purpose. We landscaped the candidate data sources and wrote the matching approach and the SAP into locked documents before the trial database closed, precisely so that when the comparison came, nobody could accuse us of reverse-engineering it. It cost us weeks we felt we couldn't spare. It also meant the analysis we presented was the analysis we had committed to, in writing, while still blind to the result. That was worth far more than the weeks.

ICH E10's whole logic on external controls turns on specifying the comparison before you select the control, and Arondekar found the FDA naming the absence of "a prespecified study protocol" across real submissions [5].

Monday: lock the matching approach and SAP before database lock, and date-stamp them.

10. Frame the estimand and the intercurrent events, kept plain. What treatment effect are you actually estimating, and what happens when a patient switches treatment, discontinues, or moves to a subsequent therapy? Answer that before you build the comparator. Polito et al. worked a non-small-cell lung cancer case through an estimand and target-trial lens to show how much the framing moves the answer [10], and Gupta et al. mapped where ECA emulation strains against per-protocol effects and time-varying adherence [11]. You don't need the machinery to get the discipline: write down the population, the endpoint and the intercurrent-event strategy, then emulate the target trial you would have run if randomisation had been open to you.

Monday: draft the estimand on one page before you touch the comparator.

The honest objection

If a checklist genuinely predicted acceptance, reviewers would agree with each other. They demonstrably don't. Handed the identical seven oncology ECA submissions, "agreement in critiques between and among regulators and HTA bodies was low" [12]. Hold the evidence constant and the split gets starker: in one survey only 12.9% of regulatory-affairs respondents rated RWD studies favourably, against 47.6% of healthcare-provider respondents looking at the same scenarios [13]. So isn't acceptance a roll of the dice after all, and this checklist a comfort blanket?

No. The disagreement lives at the ceiling, in the marginal judgement calls about how good a strong design has to be. The checklist governs the floor. Blinatumomab's clock, omburtamab's era mismatch, selinexor's inadequate EHR source: no reviewer disagreed about those. They are the failures every reviewer flags. The checklist doesn't buy you a yes. It removes the reasons for an automatic no.

Cluster 4 Design for the reviewer who wants to say no (what bar you clear)

So build for that reader.

11. Respect the higher statistical bar. An ECA needs more. ICH E10 says it in language worth quoting in full, because sponsors forget it: the persuasiveness of an externally controlled trial "depends on obtaining much more extreme levels of statistical significance and much larger estimated differences between treatments than would be considered necessary in concurrently controlled trials" [14]. A modest effect that clears a randomised trial can fail an ECA on exactly this point. Zolgensma survived on the size of its effect, not the tidiness of its match. If your expected effect is modest, an ECA may simply be the wrong instrument, and it is cheaper to learn that at protocol stage than at review.

Monday: power and design for a larger, cleaner effect than a randomised trial would demand.

12. Design for several reviewers who disagree, not one known bar. There is no single ECA acceptance bar to teach to the test. Of EMA-approved cancer drugs in one 2016–2021 window that used an external control, "close to one-third of the external control submissions were ultimately not considered supportive evidence by the EMA," most often for heterogeneous populations, missing outcomes or inappropriate statistics [15][13]. That figure is EMA-specific and oncology-specific, not a number to hang on regulators in general. Nor is there a validated instrument to outsource the judgement to; the systematic reviewers who went looking concluded plainly that "no such instrument currently exists" [16]. So you design to the strictest plausible reader: FDA, EMA and the HTA bodies at once.

Monday: pressure-test the design against the most sceptical reviewer you can imagine, not the agency's known past preferences.

What is still missing

One last thing worth knowing, because it changes what you are designing against. There is still no EU guidance dedicated to external controls. The EMA's concept paper (EMA/125200/2026, adopted 21 May 2026) is only the first formal step toward one, and the reflection paper it promises is not expected until around Q2 2027 [17]. Until it lands, you design against ICH E10 (final, 2000), the FDA's 2023 draft guidance on externally controlled trials (FDA-2022-D-2983, still draft) [18], and the most sceptical reviewer you can conjure.

Two things I would want that EMA paper to settle. A common cross-agency comparability standard would be one, given how far reviewers demonstrably diverge when handed the same submission. The other is a clear position on hybrid and borrowing designs, the intervention-plus-supplementation constructions the FDA left out of scope and that keep surfacing in real programmes.

Until then, the checklist is the floor you build so that no reviewer can find a trapdoor. Set the clock before the race, never after.

Get the monthly digest

The 5 things evidence leads need to know each month: regulatory moves, RWE developments and what they mean in practice. No pitch, one email a month.

References

[1] Suissa S. (2021). "Single-arm Trials with Historical Controls: Study Designs to Avoid Time-related Biases." Epidemiology;32(1):94–100. PMID: 33009252. https://pubmed.ncbi.nlm.nih.gov/33009252/

[2] Pharmaphorum. (2022, October 30). "Trial bias concern scuppers Y-mAbs brain cancer drug in FDA vote." https://pharmaphorum.com/news/trial-bias-concern-scuppers-y-mabs-brain-cancer-drug-in-fda-vote (reporting the FDA ODAC briefing document for the 28 October 2022 meeting; the advisory-committee vote is public record; corroborated by CancerNetwork and Medscape coverage of the same meeting).

[3] CADTH. (2021, May). "Clinical Review Report: Onasemnogene Abeparvovec (Zolgensma)." Canadian Agency for Drugs and Technologies in Health, Ottawa. https://www.ncbi.nlm.nih.gov/books/NBK584042/

[4] CADTH. (2023, January). "Tafasitamab (Minjuvi): CADTH Reimbursement Review." Canadian Agency for Drugs and Technologies in Health. https://www.ncbi.nlm.nih.gov/books/NBK601733/

[5] Arondekar B, Duh MS, Bhak RH, DerSarkissian M, Huynh L, Wang K, Wojciehowski J, Wu M, Wornson B, Niyazov A, Demetri GD. (2022). "Real-World Evidence in Support of Oncology Product Registration: A Systematic Review of New Drug Application and Biologics License Application Approvals from 2015–2020." Clinical Cancer Research;28(1):27–35. PMID: 34667027. https://pmc.ncbi.nlm.nih.gov/articles/PMC9401526/

[6] Yap TA, Jacobs I, Baumfeld Andre E, Lee LJ, Beaupre D, Azoulay L. (2022). "Application of Real-World Data to External Control Groups in Oncology Clinical Trial Drug Development." Frontiers in Oncology;12. https://pmc.ncbi.nlm.nih.gov/articles/PMC8771908/

[7] Nourredine M, Gavoille A, Lepage C, Kassai-Koupai B, Cucherat M, Subtil F. (2025). "Accounting for Misclassification of Binary Outcomes in External Control Arm Studies for Unanchored Indirect Comparisons: Simulations and Applied Example." Statistics in Medicine;44(20–22):e70236. PMID: 40930536. https://pmc.ncbi.nlm.nih.gov/articles/PMC12422847/

[8] Gupta A, Hsu G, Kent S, Duffield SJ, Merinopoulou E, Lockhart A, Arora P, Ray J, Wilkinson S, Scheuer N, Ramagopalan SV, Groenwold RHH, Popat S, Hernán MA. (2025). "Quantitative Bias Analysis for Single-Arm Trials With External Control Arms." JAMA Network Open;8(3):e252152. PMID: 40136297. https://pmc.ncbi.nlm.nih.gov/articles/PMC11947839/

[9] Jahanshahi M, Gregg K, Davis G, Ndu A, Miller V, Vockley J, Ollivier C, Franolic T, Sakai S. (2021). "The Use of External Controls in FDA Regulatory Decision Making." Therapeutic Innovation & Regulatory Science;55(5):1019–1035. PMID: 34014439. https://pmc.ncbi.nlm.nih.gov/articles/PMC8332598/

[10] Polito L, Liang Q, Pal N, Mpofu P, Sawas A, Humblet O, Rufibach K, Heinzmann D. (2024). "Applying the estimand and target trial frameworks to external control analyses using observational data: a case study in the solid tumor setting." Frontiers in Pharmacology;15:1223858. PMID: 38344177. https://pmc.ncbi.nlm.nih.gov/articles/PMC10853363/

[11] Gupta A, Merinopoulou E, Duffield SJ, Gomes M, Kent S, Scheuer N, Machnicki G, Sanglier T. (2025). "Estimating per-protocol effects in external comparator analyses using real-world data." Journal of Comparative Effectiveness Research;14(11):e250029. PMID: 41047965. https://pmc.ncbi.nlm.nih.gov/articles/PMC12580923/

[12] Jaksa A, Louder A, Maksymiuk C, Vondeling GT, Martin L, Gatto N, Richards E, Yver A, Rosenlund M. (2022). "A Comparison of Seven Oncology External Control Arm Case Studies: Critiques From Regulatory and Health Technology Assessment Agencies." Value in Health;25(12):1967–1976. PMID: 35760714. https://pubmed.ncbi.nlm.nih.gov/35760714/

[13] Pignatti F, El-Galaly TC, Kaiser M, Porkka K, Doeswijk R, Mol P, Rivera DR, Lerro CC, Rohr UP, Verpillat P, Valachis A, Trapani D, Di Maio M, Latino N, Cordoba R, Cherny N, Koopman M, Martins-Branco D, Pentheroudakis G, Postmus D. (2026). "Assessing Overall Survival Benefits in Advanced Cancers: The Role of External Comparator Cohort Studies with Real-World Data." Clinical Pharmacology and Therapeutics;119(4):1080–1087. PMID: 41669939. https://pmc.ncbi.nlm.nih.gov/articles/PMC12997495/

[14] ICH E10. "Choice of Control Group and Related Issues in Clinical Trials" (Step 4, 20 July 2000; adopted as FDA guidance effective 14 May 2001), §2.5.2. Final. https://www.fda.gov/media/71349/download

[15] Wang X, Dormont F, Lorenzato C, Latouche A, Hernandez R, Rouzier R. (2023). "Current perspectives for external control arms in oncology clinical trials: Analysis of EMA approvals 2016–2021." Journal of Cancer Policy;35:100403. PMID: 36646208. https://pubmed.ncbi.nlm.nih.gov/36646208/

[16] Zayadi A, Edge R, Parker CE, Macdonald JK, Neustifter B, Chang J, Zhong G, Singh S, Feagan BG, Ma C, Jairath V. (2023). "Use of external control arms in immune-mediated inflammatory diseases: a systematic review." BMJ Open;13(12):e076677. PMID: 38070932. https://pmc.ncbi.nlm.nih.gov/articles/PMC10729249/

[17] EMA. "Concept Paper on the Development of a Reflection Paper on the Use of External Controls for Evidence Generation in Regulatory Decision-Making." EMA/125200/2026, adopted by CHMP 21 May 2026. Concept paper. https://www.ema.europa.eu/en/documents/scientific-guideline/concept-paper-development-reflection-paper-use-external-controls-evidence-generation-regulatory-decision-making_en.pdf

[18] FDA. "Considerations for the Design and Conduct of Externally Controlled Trials for Drug and Biological Products." Draft guidance for industry, 1 February 2023. Docket FDA-2022-D-2983. https://www.fda.gov/media/164960/download (existence, title, date and docket confirmed via the Federal Register notice; full text not independently read)

[T1] Novartis Gene Therapies. "Single-Arm, Open-Label, Single-Dose Gene Replacement Therapy Clinical Trial for SMA Type 1 (STR1VE)." ClinicalTrials.gov: NCT03306277. https://clinicaltrials.gov/study/NCT03306277

[T2] BioMarin Pharmaceutical. "A Study of Intracerebroventricular BMN 190 in Patients With CLN2 Disease (cerliponase alfa)." ClinicalTrials.gov: NCT01907087. https://clinicaltrials.gov/study/NCT01907087

[T3] Sarepta Therapeutics. "Study of SRP-4053 (golodirsen) in DMD Patients Amenable to Exon 53 Skipping." ClinicalTrials.gov: NCT02310906. https://clinicaltrials.gov/study/NCT02310906

Share this post