Inovia Bio Insights

External Control Arm Feasibility: A Self-Assessment

Written by Imi | 27-Jul-2026 13:58:04

The moment an external control arm gets floated in a strategy meeting, the first question in the room is almost always the same: is there a dataset? Someone volunteers a registry, an EHR extract, a historical cohort, and the conversation moves briskly on to matching methods and covariates as though the external control arm's feasibility had already been settled.

It usually hasn't.

Consider amyotrophic lateral sclerosis. ALS sits behind one of the largest natural-history repositories in neurology, roughly 5,000 patients (NCT05966038). By the data-availability logic it ought to be a slam-dunk external-control indication. It is not: the pivotal programmes overwhelmingly run concurrent, randomised controls anyway.
Between-patient variability in ALS progression is too large relative to the effect sizes most drugs can plausibly produce, so a single arm read against even a huge historical cohort cannot tell you what your drug actually did.

That is the availability trap: mistaking the existence of data for the feasibility of a comparison. ECA-readiness is a property of the disease, not the data warehouse. Whether your indication has it comes down to a handful of hard, mostly disqualifying questions you can work through by hand, before you commission a single line of analysis.

Data availability is necessary; it is nowhere near sufficient.

What actually decides ECA feasibility

So what does? The most useful evidence base available is Subramaniam and colleagues (2024), who went through 37 real single-arm approvals (20 at the FDA, 17 at the EMA) and catalogued which disease and design features were actually present [1]. Two were near-universal: a natural history that progressed without meaningful spontaneous improvement, in 90% of the FDA approvals and 88% of the EMA ones, and an objective primary endpoint, in the same 90% and 88%.

Those are the load-bearing factors.

A large expected effect size, the thing everyone assumes is the price of admission, was present in only 45% of the FDA approvals and 18% of the EMA ones. A big effect helps enormously, but on this evidence it is not a precondition.

Free download

The External Control Arm Design Checklist

The 12 design points regulators probe first, in one checklist you can run against your protocol before database lock.

Get the checklist →

The ten questions, disease first

Ten questions, grouped, and the grouping is the argument. The disease questions come first because a "no" there cannot be bought back with a better dataset. No number falls out of the end; each is a pass/fail read. I built this list myself based on what regulators have actually done and anchored every question to a named source rather than asking you to take my word for it.

Group A Does the disease hold still?

1. Is the natural history predictable, without spontaneous improvement? The EMA's single-arm reflection paper (final, 9 September 2024) names an episodic or waxing-and-waning course as a condition that undermines interpretability [5], and Subramaniam's ~90% makes a clean monotonic decline the near-universal feature of what gets approved.

Question 1 : Pull the natural-history literature and see whether you can honestly draw the untreated curve. If you can't, the ECA is very likely a non starter.

2. Do patients enrol at their worst? If your criteria select patients at a transient extreme, some improve without your drug, and a single arm cannot see it happening. That is regression to the mean, and only a concurrent control catches it, because it is screened the same way and improves in parallel. Kerr and colleagues (2026) show this baseline dependence inside randomised epilepsy trials, where placebo and active arms tracked together in some trials and not others [6]. An ECA cannot.

Question 2 : Does your entry criteria quietly select a transient low point?

3. Is the endpoint objective, and derivable identically on both sides? An objective primary endpoint appeared in ~90% and 88% of Subramaniam's approvals: a subjective or investigator-assessed one reintroduces exactly the bias an external comparison already carries, and it must be derivable the same way in the real-world source, not merely "captured" there.

Question 3: Confirm the endpoint exists, and is derivable in a real world setting 

Group B Is the effect big enough to clear the noise?

4. Is your expected effect larger than the outcome's natural variability? The MHRA's draft guideline puts it explains this point almost formulaically, at §56: "If the difference in efficacy outcomes between the treatment and control groups is small and in the range of general variability of the outcome it could be difficult to be confident in a comparison to a RWD ECA even if a statistically significant result is presented" [3]. Statistical significance is not the bar. The effect has to be large relative to how much the outcome moves on its own.

Question 4: Put your expected effect next to the endpoint's known variability before you put it near a p-value. Is your expected value a plausable treatment effect?

5. How big is "big enough"? Benchmark it, don't guess. No universal threshold exists, and be wary of anyone who offers one. IQWiG, the German HTA body, defines a "dramatic effect" as p≤0.01 with more than a tenfold change in relative risk (Mangla et al., 2026) [7] This is one agency's line, not a global rule, and Subramaniam's 45% and 18% put a large effect in a minority of approvals.

Question 5: locate your effect against real precedent, and treat a modest one as a red flag for an ECA specifically.

Group C Do you even need an ECA?

6. Could you randomise if you actually tried? The disqualifier sponsors skip. The MHRA is blunt, at §58: "If there are sufficient patients and it is possible to randomise then the decision to use a RWD ECA would be more difficult to justify... In such circumstances an RCT is preferred" [3]. Given regulatory precedent from the FDA, EMA and MHRA the bar is very simple, if randomisation is possible the expectation is that you must randomise.

Question 6: Run a randomisation scenario, weighing numbers, equipoise and ethics. Ask yourself if an ECA justified?

Group D Only now, the data

7. Do you have a source that fits, or only one that exists? This is where the availability trap does its real damage. A dataset existing for your disease says nothing about whether it captures your population, your prognostic factors, your time-zero and your endpoint. Lin and colleagues (2025) found 6 of 8 real-world-derived external controls matched their concurrent RCT arm [8], so even a fitted source is no guaranteed stand-in. The FDA's 2023 draft guidance reportedly stresses fitting the study to the source, not shoe-horning it into a convenient one [4].

Question 7: Design a protocol and then check if a data source contains the variables necessary to operationalise it (more on this in our target trial emulation package)

8. Is the source good enough on provenance, completeness and contemporaneity? Missingness, the collection era, whether the covariates you need to adjust for were even recorded. This gate fails silently: the dataset looks fine until a reviewer asks how a given variable was ascertained.

Question 8: Score potential RWE sources and assess whether they are appropriate to addressing your research question (regulatory-grade RWE package)

Group E Precedent and sequencing

9. Is there regulatory precedent in your space? Jahanshahi and colleagues (2021) catalogued 45 FDA external-control approvals [9]. The cleanest is onasemnogene abeparvovec (Zolgensma) in spinal muscular atrophy, where a predictable natural history, an objective endpoint and a very large effect converged: roughly 90% of treated infants alive without permanent ventilation against about 25% under natural history.

Question 9: find your closest approved precedent, or accept that you are the precedent-setter and price that in.

10. Can you pre-specify, and did you build the data early? Liu and colleagues (2025) found that of 180 published externally controlled trials, only 14 (7.8%) ran an upfront feasibility assessment of the data source, and only 16.1% pre-specified the external control [10]. The sponsors who win the options built the infrastructure years early. FA-COMS (NCT03090789), the Friedreich ataxia registry begun in 2003 and grown past 1,250 participants across 14 sites, later let omaveloxolone's developers run a propensity-matched comparison showing a 55% slowing of progression, 6.6 versus 3.0 mFARS points at Year 3, nominal p=0.0001 (Lynch et al., 2023) [11].

Question 10: Do you have infrastructure in place to support an ECA and have you run a landscaping study?

The gate is a judgement call

Now the honest objection: a by-hand read like this is false comfort, because feasibility judgements are genuinely contested. Heynemann and colleagues (2026) documented how value-laden they are [12]: the same package can read as sufficient to the sponsor and insufficient to the reviewer, and no checklist collapses that.

The best evidence for the objection is Biohaven's troriluzole in spinocerebellar ataxia, which drew a complete response letter in November 2025. By the sponsor's account, the FDA had said beforehand that "a large and robust treatment effect would be needed to overcome the biases of an externally controlled trial" [13]. Biohaven reported 50 to 70% slower progression on the SARA scale across the primary and eight secondary endpoints, and still received a CRL, reportedly citing potential bias, design flaws, a lack of pre-specification and unmeasured confounding [13]. All of that is sponsor-reported and corroborated by trade press, not confirmed from a primary FDA document. The sponsor believed it had cleared the stated bar. It hadn't. So what use is a checklist troriluzole would have passed?

Exactly the use it claims, and no more. The assessment is a triage instrument, not a verdict machine. It reliably kills the clear failures, the §58 case and the ALS-grade variability case, so you do not sink a year and a slice of runway into a comparator that was never going to fly. The genuinely ambiguous cases it flags honestly, and troriluzole is the example: strong on disease and effect size, exactly the profile that should engage the agency early and pre-specify hard, which is precisely where it came apart. That said, a triage tool is not a guarantee, and nothing here replaces an actual regulatory conversation. Use it to decide what you take to the agency, and when, not to wave through your own case.

Where each answer sends you

Which brings us back to the question everyone asks first and ought to ask last. The dataset is the final gate, not the opening one. A full data warehouse is a stocked pantry, not a dinner: it tells you a meal is possible, not that anyone in the building can cook.

Where your answers land tells you where you actually are. Get through the disease and effect-size questions clean, and the build-or-not decision is already made: what remains is design work, comparability, time-zero, pre-specification, the downstream checklist that starts where this one stops. The alternative is a disqualifier discovered late, after the single-arm data already exists, with the regulator asking the question you should have asked yourself. That is the reactive scenario this exercise exists to prevent. For the single-arm fundamentals beneath all of it, the EMA's reflection paper is still the clearest account of what a regulator will expect.

This upstream gate is where we spend much of our time at Inovia, with teams working out whether their indication clears the external control arm feasibility bar at all, before a penny goes into building one. What the ten questions buy you is risk reduction. It is the cheap, early read that stops you spending a year of runway proving something the disease was never going to let you prove.

Get the monthly digest

The 5 things evidence leads need to know each month: regulatory moves, RWE developments and what they mean in practice. No pitch, one email a month.

References

  1. Subramaniam et al. (2024). Ther Innov Regul Sci. PMID: 39285061. https://pubmed.ncbi.nlm.nih.gov/39285061/
  2. EMA. Concept paper on the use of external controls in clinical studies (EMA/125200/2026). Concept paper, adopted 21 May 2026 (first published 3 June 2026); reflection paper still in development. https://www.ema.europa.eu/en/development-reflection-paper-use-external-controls-evidence-generation-regulatory-decision-making-scientific-guideline
  3. MHRA. Draft guideline on the use of external control arms based on real-world data to support regulatory decisions. Draft; consultation 20 May – 14 July 2025. https://www.gov.uk/government/consultations/mhra-draft-guideline-on-the-use-of-external-control-arms-based-on-real-world-data-to-support-regulatory-decisions (PDF: https://assets.publishing.service.gov.uk/media/6825bab1a4c1a40fde4e63e5/Draft_MHRA_Guideline_on_Studies_with_RWD_ECA_May2025.pdf)
  4. FDA. Considerations for the Design and Conduct of Externally Controlled Trials for Drug and Biological Products. Draft guidance, February 2023 (reported via secondary sources; fda.gov not independently reachable this session). https://www.fda.gov/regulatory-information/search-fda-guidance-documents/considerations-design-and-conduct-externally-controlled-trials-drug-and-biological-products
  5. EMA. Reflection paper on establishing efficacy based on single-arm trials submitted as pivotal evidence in a marketing authorisation application (EMA/CHMP/458061/2024). Final; adopted 9 September 2024. https://www.ema.europa.eu/en/documents/scientific-guideline/reflection-paper-establishing-efficacy-based-single-arm-trials-submitted-pivotal-evidence-marketing-authorisation-application_en.pdf
  6. Kerr et al. (2026). Epilepsia. PMID: 41830605. https://pubmed.ncbi.nlm.nih.gov/41830605/
  7. Mangla et al. (2026). J Comp Eff Res. PMID: 41384576. https://pubmed.ncbi.nlm.nih.gov/41384576/
  8. Lin et al. (2025). Drug Discovery Today. PMID: 40054765. https://pubmed.ncbi.nlm.nih.gov/40054765/
  9. Jahanshahi et al. (2021). Ther Innov Regul Sci. PMID: 34014439. https://pubmed.ncbi.nlm.nih.gov/34014439/
  10. Liu et al. (2025). JAMA Netw Open. PMID: 40906478. https://pubmed.ncbi.nlm.nih.gov/40906478/
  11. Lynch DR, et al. (2023). Propensity matched comparison of omaveloxolone treatment to Friedreich ataxia natural history data. Ann Clin Transl Neurol (online 2023; issue 2024;11(1)). FACOMS / omaveloxolone propensity-matched natural-history comparison. PMID: 37691319. https://pubmed.ncbi.nlm.nih.gov/37691319/
  12. Heynemann et al. (2026). Asia-Pac J Clin Oncol. PMID: 42095345. https://pubmed.ncbi.nlm.nih.gov/42095345/
  13. Biohaven, Ltd. "FDA Issues Complete Response Letter for Biohaven's VYGLXIA (troriluzole) New Drug Application for Spinocerebellar Ataxia," 4 November 2025 (sponsor press release). https://ir.biohaven.com/news-releases/news-release-details/fda-issues-complete-response-letter-biohavens-vyglxia — corroborated by trade press (NeurologyLive: https://www.neurologylive.com/view/fda-issues-complete-response-letter-spinocerebellar-ataxia-agent-troriluzole; Neurology Advisor; FierceBiotech; pharmaphorum). Sponsor-reported FDA statement (attributed by Biohaven to FDA meeting minutes, 8 March 2024); corroborated by trade press, not confirmed from a primary FDA document.

Trials cited

  • [T2] Cerliponase alfa, late-infantile CLN2 Batten disease. ClinicalTrials.gov: NCT01907087. https://clinicaltrials.gov/study/NCT01907087
  • FA-COMS, Friedreich ataxia natural-history registry. ClinicalTrials.gov: NCT03090789. https://clinicaltrials.gov/study/NCT03090789
  • [T1] ALS natural-history repository (NCT05966038) and its pivotal trials, which use concurrent control despite it (NCT01281631, NCT02794857, NCT03280056).