Inovia Bio Insights

Natural History Study Go/No-Go: FDA External Controls

Written by Imi | 27-Jul-2026 20:20:44

Every rare-disease team feels the same pull, and it points one way: characterise the disease first. The FDA's 2019 draft guidance on natural history studies says roughly that, telling sponsors to start early, run the study prospectively, and lay the disease-course foundation before the pivotal is designed [6]. It reads like responsible due diligence, and it is genuinely hard to argue with when you are staring at an indication that has no published progression rate and no validated endpoint.

Then you look at what actually cleared the FDA.

Across the external-control approvals catalogued by Jahanshahi and colleagues, 45 of them spanning 2000 to 2019, not a single one was built on prospectively collected natural history data [2]. Zero of forty-five. The commonest source was retrospective natural history, at 44%. And in Vaghela and colleagues' 2024 review of 20 rare-disease approvals, the four applications that leaned on ad-hoc chart-review natural history data were all criticised, and "none of these applications reported RWD in their label claims" [1]. The data got collected; the regulatory function it delivered came to nothing.

So here is the uncomfortable version. A natural history study has no universal value; its worth is entirely function-specific. It earns its long timeline and its budget only when it is commissioned to carry a named regulatory job, an external control or a qualified endpoint or the disease-course definition a label will rest on, and only when that job can clear the narrow window regulators actually reward. Commissioned as generic "let's understand our disease" reassurance, it burns runway a lean biotech does not have.

The practical consequence is a change of question: stop asking "should we characterise our disease?" and start asking what named regulatory job this study would do, and whether that job can clear the window. That is a go/no-go you run before the first patient is enrolled, not an act of faith that more data always helps.

Name the job, or don't start it: the four regulatory functions of a natural history study

The go decision starts with one line written down before enrolment: which regulatory function is this study for? Liu and colleagues, cataloguing the FDA's draft natural history guidance, list the menu as "identification of patient population, identification or development of clinical outcome assessment and biomarkers, and design of external controls", plus trial-design optimisation [3][6]. Four jobs, concretely:

  • Define the patient population: who has this disease, how it is diagnosed, and what the eligible cohort actually looks like. This is what lets you write inclusion criteria a regulator will not later pick apart.
  • Develop or qualify a clinical outcome assessment or biomarker: a natural history study can show that a measure moves with the disease, and moves predictably, which is the precondition for using it as an endpoint at all.
  • Design the external control: a retrospective or contemporaneous cohort that stands in for the concurrent comparator a single-arm trial does not have. This is the headline use, and the one most teams mean when they say "natural history".
  • Optimise the trial design: progression rates, event timing and variance feed the sample-size calculation and the choice of timepoints. Get the disease course wrong here and the pivotal is under-powered before it opens.

If you cannot name which of the four before you enrol a patient, you do not have a study; you have a data-collection habit. And a data-collection habit should not be drawing on the regulatory budget.

The window the job has to clear: the FDA's external control arm credibility test

Naming a function is necessary but not sufficient: the function pays off only if it can clear the window regulators reward, and that window is old and well documented. ICH E10, the choice-of-control-group guideline finalised in 2000 and still the most quotable statement of the rule, holds that an external control is credible when "the study endpoint is objective, when the outcome on treatment is markedly different from that of the external control and a high level of statistical significance… is attained", and the disease has "a well-documented, highly predictable course" [8]. Its own caution sits one section up: "Inability to control bias is the major and well-recognized limitation of externally controlled trials."

Videnovic and colleagues distil the same into a three-condition admissibility test for a historical control: a serious disease with "an unmet medical need"; "a well-documented, highly predictable disease course that can be objectively measured and verified"; and "an expected drug effect that is large, self-evident, and temporally closely associated with the intervention" [5].

The precedent bears this out. Of Jahanshahi's 45 approvals, roughly one in three first-time rare-disease approvals leaned on an external control, 87% used an objective primary endpoint, and the agency, on their reading, "is less swayed by the size of natural history studies than by the rigor of data collection and clarity on the course of the disease" [2]. Vaghela's read is blunter: "The FDA generally accepted RWD studies demonstrating a large effect size despite the noted concerns and criticisms" [1]. What the agency actually rewards is an objective endpoint, a large self-evident effect and a predictable course — never sheer cohort size.

Which is why a great deal of natural history spend is what I would call "characterisation theatre": a study that generates disease data with no function capable of clearing this window, run because it looks responsible rather than because it will move a submission.

That said, when the window is genuinely cleared, the payoff is real, and the gallery is instructive:

  • Zolgensma, SMA type 1: 90% alive without ventilation versus 25% in the SMA natural-history cohorts, an objective endpoint against an enormous effect in a predictable and fatal course [2].
  • Nulibry, MoCD type A: a natural history cohort of 37 patients gave three-year survival of 53% (95% CI 28-73) untreated against 84% (95% CI 47-96) treated, and the FDA judged it "an adequate, well-controlled investigation" [4].
  • Crysvita, paediatric XLH: a "large effect size for reduction in RSS (50–59% versus 12% in the historical control)", drawn from a small retrospective cohort [2].
  • Myozyme, infantile Pompe: 18 treated infants against "a historical cohort of 61 untreated patients… (only 1 of whom was still alive)", on a ventilator-free survival endpoint [3].

Every one of those cleared the E10 window before it was ever submitted. So score your intended function against objective endpoint, large effect and predictable course before you commit the spend, and if the external-control function cannot clear it, do not fund a comparator study on hope.

Free download

The RWE Briefing Document Template

The section-by-section structure for the RWE part of a regulatory briefing, built around the questions reviewers actually ask.

Get the template →

When a named function dies anyway (1): the runway

Here the opening paradox resolves. FDA guidance wants prospective, protocol-driven natural history "initiated in the earliest drug development planning stages" [4][6], yet precedent says almost nobody manages it, zero of 45, because a prospective study rarely finishes in time.

How long does a natural history study take? There is no canonical number, and anyone quoting you a tidy figure is guessing. The honest anchors are these: enrolling even a small cohort, under 50 patients, "may take up to two or more years" in a very rare disease [5]; the CINRG Duchenne natural history study recruited "440 patients aged 2-28 years… from 20 centers in 9 countries and were followed up for up to 10 years" [3]; and the agency's own guidance concedes the tension directly, warning that initiating a prospective natural history study "should not delay interventional testing otherwise ready to commence" [2].

Then the asset-death exhibit. BioMarin ran a prospective Duchenne natural history study, NCT01753804, with a six-minute walk distance endpoint matched to its pivotal, 269 patients, started September 2012 [9]. The design was correct and fit for purpose. It was still terminated in October 2016 when the drug programme it supported failed. It died with the asset. A prospective natural history study run in parallel with your pivotal carries the pivotal's risk, and if the asset falls over, the natural history spend falls with it.

So decide the lead time honestly: start early enough to finish, or plan a retrospective or hybrid source that fits, and price in the chance that a parallel prospective study returns nothing at all.

When a named function dies anyway (2): the estimand-fit tax

Even a study that clears the window fails if its captured endpoint, population and follow-up do not line up with the pivotal's estimand, and this is the tax nobody prices in. Take Brineura in CLN2: the DEM-CHILD natural history cohort held 42 patients and the single-arm trial 24 treated, and after the responder definition was re-derived to match the natural history scale, only 17 matched pairs survived the exercise [3]. Most of the natural history cohort was unusable for the comparison, because the two datasets had not been captured against the same ruler.

Many years ago I sat on a rare-disease programme that did everything the orthodoxy asks. It commissioned a broad, registry-style natural history capture early, on the reasonable theory that more characterisation could only help. When the pivotal's estimand was finally locked, endpoint, population and follow-up window, the natural history data did not line up: the outcome had been captured on a different schedule, in a slightly different population, and the dataset could not carry the external control it had been bought to become. The money was spent. The function was not.

That is a fit failure, not a data-quality one, and it is designed in from the first protocol. Vaghela's markers make the point quantitatively: across the 20 approvals, 60% matched the real-world data duration to the pivotal, 65% ran to an a priori protocol, and 85% matched eligibility [1]. Ugoji and colleagues put the rule plainly, finalise the natural history protocol before the trial, not alongside it [4].

Duration mismatch alone can sink a read-out: in the FORT trial of fosmetpantotenate for PKAN, the 24-week PKAN-ADL endpoint was too short a window for a slow disease, placebo declined by only "1–2 points" and no significant treatment difference emerged [5]. A perfectly real disease, measured over the wrong interval, returns no signal. So design the natural history study to the pivotal's estimand, the same endpoint, population and follow-up window, and lock its protocol first. The comparability requirements are the ones the FDA lays out for external control arms, which we walked through in our FDA external control arm checklist; the pre-specification demands sit in the EMA single-arm reflection paper.

When the data exists and still does nothing

Registered is not resulted, and resulted is not accepted. This is where most wasted spend hides, because a natural history dataset that sits on a server feels like an asset long after the regulator has declined to use it. The four chart-review applications in Vaghela's review, triheptanoin, stiripentol, voretigene neparvovec and vestronidase alfa, were all criticised, and none reached a label claim on their real-world data [1]. Viltolarsen in Duchenne had its natural history data "excluded from its label claim" over disease heterogeneity and uncontrolled bias, and the drug was approved on the dystrophin biomarker instead, so the natural history work carried no function at all [1]. Defitelio in hepatic VOD saw its historical control cut "from 6867 to 123 and finally to 32 patients", with control and treatment data spanning "vastly different timeframes (2 years and 12 years, respectively)", and the US label carries no formal statistical comparison to that cohort [2][3].

Then the studies that returned nothing because they never finished: a SMA1 pilot natural history study terminated at four patients (NCT01547871); a Batten natural history study terminated at ten (NCT04644549); a CLN2 ocular chart-review withdrawn before it enrolled anyone (NCT04480476). Each is a real docket and real spend with nothing to show for it.

The pattern has not stopped: trade press through 2025 and 2026 has reported complete response letters across several rare-disease programmes leaning on external or natural history comparators, among them Regenxbio's RGX-121, Biohaven's troriluzole and Capricor's deramiocel, though the letters themselves are not public and should be read as reporting rather than adjudicated FDA findings [13].

Do not book a natural history dataset as regulatory evidence until a named function has cleared the window. Existence is not acceptance. It is the contested-comparator scenario that lands teams in front of regulators asking pointed questions, and the cause is almost always upstream: a study commissioned without a function it could clear.

The best argument against all this

You cannot define an endpoint, size an external control, or establish a disease course without characterising the disease first, so function-first looks like a false economy. Tell a lean team to collect narrowly against a named job and you all but guarantee the one variable they skipped is the one the FDA later wants. Broad characterisation is the foundation everything else stands on, so collect widely now, or fly blind.

It is a good argument, and it is also, mostly, answered by the evidence already on the table. Function-first still means collecting broadly enough to answer the job in front of you. The difference is collecting against a named target and the pivotal's estimand, and starting early enough to actually finish. The four chart-review failures and Brineura's 17-of-42 attrition are precisely what open-ended characterisation without a function-and-fit specification produces: data that existed and still carried nothing. And where the disease course genuinely is not yet understood well enough to clear the E10 window, that finding is itself the go/no-go answer for the external-control function, while the same study can still be a legitimate go for the endpoint-development function. Naming the function is what tells you which job the data can and cannot do.

The test to run before the first patient

So put it on one page and run it before enrolment: a go needs all five to hold for the intended function, and a no-go is any one of them failing.

  1. Name the function: external control, COA or biomarker qualification, disease-course definition, or trial-design optimisation. If none is named, this is not a regulatory study, and you should fund it from somewhere other than the regulatory budget, if at all.
  2. Clear the window: is the endpoint objective, is the expected effect large and self-evident, is the course well-documented and predictable? Fail this for the external-control function and the study may still be a go for endpoint development, but not as a comparator.
  3. Fit the estimand: will the captured endpoint, population and follow-up line up with the pivotal's, and is the natural history protocol finalised before the trial rather than beside it?
  4. Fund the runway: can you start early enough to finish, and if not, is there a retrospective or hybrid source that actually fits?
  5. Price the asset risk: if the asset dies, does the dataset keep independent value, or does it die with the drug?

Run this way, the go/no-go stops being an act of faith. It is the same ruthless prioritisation a lean evidence plan applies everywhere else, the minimum viable evidence discipline we set out in the 90/10 framework, applied to one expensive, slow, easy-to-rationalise line item. When resources are tight, the natural history go/no-go is exactly the sort of call an integrated evidence plan exists to force early, on paper, before the money moves. Anything less is characterisation theatre with a budget line.

One last thing worth naming

You run this decision with worse guidance than the stakes deserve. The FDA's most specific text on natural history studies, the 2019 draft under docket FDA-2019-D-0481, was never finalised [6]. The final rare-disease guidance that followed in December 2023 dropped the standalone natural history section entirely [7]. So the developer weighing a multi-year natural history commitment is navigating by ICH E10, a control-group guideline written in 2000 [8], and a handful of peer-reviewed retrospectives.

A finalised natural history guidance could have pinned down what "fit-for-purpose" means for the disease-course-definition function specifically: how much follow-up, matched to which estimand, at what level of endpoint objectivity, before a course counts as well-documented and highly predictable in the agency's eyes. It did not. Until it does, the decision sits with you, and it belongs before the first patient, not after.

Get the monthly digest

The 5 things evidence leads need to know each month: regulatory moves, RWE developments and what they mean in practice. No pitch, one email a month.

References

[1] Vaghela S, et al. (2024). "A systematic review of real-world evidence (RWE) supportive of new drug and biologic license application approvals in rare diseases." Orphanet Journal of Rare Diseases;19(1):117. PMID: 38475874. https://pubmed.ncbi.nlm.nih.gov/38475874/

[2] Jahanshahi M, et al. (2021). "The Use of External Controls in FDA Regulatory Decision Making." Therapeutic Innovation & Regulatory Science;55(5):1019-1035. PMID: 34014439. https://pubmed.ncbi.nlm.nih.gov/34014439/

[3] Liu J, et al. (2022). "Natural History and Real-World Data in Rare Diseases: Applications, Limitations, and Future Perspectives." Journal of Clinical Pharmacology;62(Suppl 2):S38-S55. PMID: 36461748. https://pubmed.ncbi.nlm.nih.gov/36461748/

[4] Ugoji C, et al. (2024). "Important tool in our rare disease toolbox: hybrid retrospective-prospective natural history studies serve well as external comparators for rare disease studies." Frontiers in Drug Safety and Regulation;4:1418050. PMID: 40979382. https://pubmed.ncbi.nlm.nih.gov/40979382/

[5] Videnovic A, et al. (2023). "Study design challenges and strategies in clinical trials for rare diseases: Lessons learned from pantothenate kinase-associated neurodegeneration." Frontiers in Neurology;14:1098454. PMID: 36970548. https://pubmed.ncbi.nlm.nih.gov/36970548/

[6] FDA. "Rare Diseases: Natural History Studies for Drug Development." Draft guidance, March 2019 (never finalised). Docket FDA-2019-D-0481; Federal Register 2019-05655.

[7] FDA. "Rare Diseases: Considerations for the Development of Drugs and Biological Products." Final guidance, December 2023; Federal Register 2023-28310. Finalises the 2019 "Common Issues" draft and does not carry a standalone natural-history section.

[8] ICH. "E10: Choice of Control Group and Related Issues in Clinical Trials." Step 4, 20 July 2000. §2.5.2, §2.5.4.

[9] BioMarin Pharmaceutical. "A Prospective Natural History Study of Duchenne Muscular Dystrophy." ClinicalTrials.gov: NCT01753804 (started September 2012, terminated October 2016). https://clinicaltrials.gov/study/NCT01753804

[10] NINDS. SMA type 1 natural history pilot. ClinicalTrials.gov: NCT01547871 (terminated). https://clinicaltrials.gov/study/NCT01547871

[11] Batten disease (CLN6/CLN3) natural history study. ClinicalTrials.gov: NCT04644549 (terminated 2022). https://clinicaltrials.gov/study/NCT04644549

[12] REGENXBIO. CLN2 ocular chart-review natural history study. ClinicalTrials.gov: NCT04480476 (withdrawn, no enrolment). https://clinicaltrials.gov/study/NCT04480476

[13] Trade-press and company reporting, 2025–2026 (complete response letters; letters not public, to be read as reporting rather than adjudicated FDA findings). REGENXBIO RGX-121 in MPS II (CRL 7 February 2026; FDA cited comparability of the natural-history external control); Biohaven troriluzole in spinocerebellar ataxia (CRL 4 November 2025; FDA cited real-world-data bias and design concerns); Capricor deramiocel in DMD cardiomyopathy (CRL July 2025; BLA compared to a natural-history dataset).