The pivotal programme you are costing out this quarter is almost certainly bigger than the decision you are actually asking the FDA to make. The gap is rarely a rounding error. It runs to a whole trial, a thousand extra patients, or two years of follow-up you will never need to win the approval you are aiming at.
First-time sponsors over-build on two levers: whether they really need a second pivotal trial at all, and how large and how long each one is. Get either wrong and you are buying evidence for a decision nobody asked you to make, at a price a lean biotech pays in patients, in months, and in cash it does not have spare.
So start with a real case, because the abstraction never lands on its own.
In 2007, ambrisentan reached the US market on two identically designed trials. ARIES-1 (NCT00091598) ran alongside ARIES-2, each lasting twelve weeks, each powered on a surrogate, six-minute walk distance, for a combined enrolment of 372 patients [9]. That was what a PAH approval looked like.
Six to eight years later, in the same indication, the picture had inverted. Macitentan's SERAPHIN (NCT00660179) reached approval on a single trial, as did selexipag's GRIPHON (NCT01106014) [10][11]. Each ran for years rather than weeks, up to roughly three or four, each was powered on a hard morbidity-and-mortality composite rather than a walk test, and each enrolled more patients on its own, 742 and 1,156 respectively, than ambrisentan's two trials managed between them.
Read the two eras side by side. The number of trials, their size and their duration all moved together, tracking the endpoint the sponsor chose to answer. Nothing about the biology of PAH changed between 2007 and 2015; what changed was the question each sponsor set out to ask. Weak surrogate: short trials, two of them. Hard clinical endpoint: one long trial, larger. There is no fixed quantity of trial that PAH "needs" only a quantity for a given decision. Which is the whole game.
Ask most first-time teams why they have two pivotal trials and the honest answer is that everyone has two pivotal trials. Nobody chose it; the statute has never required it.
Under the FD&C Act §505(d), as amended by the FDA Modernization Act of 1997, the agency may accept "one adequate and well-controlled clinical investigation and confirmatory evidence" as substantial evidence of effectiveness [4]. That provision is 28 years old. FDA's 1998 guidance implementing it (final) named the conditions: a single trial showing "a clinically meaningful and statistically very persuasive effect" on a serious outcome, where a confirmatory second trial "would be impracticable or unethical" [5]. A high bar, certainly. It is also a real, written door, and most sponsors walk past it without checking whether it is open to them.
A good many now do walk through. The share of approvals resting on two or more pivotal trials fell from 80.6% in 1995-1997, to 60.3% in 2005-2007, to 52.8% in 2015-2017 (Zhang et al., JAMA Network Open, P<.001) [1]. Downing et al. put single-trial approvals at 36.8% across 2005-2012 [2]; Godoy et al. at 42% across 2015-2023, with the FDA leaning on confirmatory evidence for 21.7% of those [3]. Roughly one in five in the late 1990s, roughly two in five today. One honest caveat, so nobody reads this as a licence to do less: over the same window the median trial grew, from 277 to 467 patients and from 11 to 24 weeks [1]. Fewer trials, each more rigorous. That is a substitution, not a discount.
Two named cases make it concrete: traditional approvals, not oncology, each one written up by the FDA's own reviewers. Pimavanserin (Nuplazid) went through on a single pivotal trial in which, per the agency's Division of Psychiatry Products, "80.5% of pimavanserin patients experienced at least some improvement in symptoms compared to 58.1% of patients taking placebo" [7]. Angiotensin II (Giapreza) rested on the single ATHOS-3 trial, approved, in the words of the Division of Cardiovascular and Renal Products, "based on the agreements emanating from the special protocol assessment" [8].
The newest signal I flag rather than lean on: FDA's revised draft guidance issued on 24 June 2026, under HHS's "Operation TrialBlazer", reportedly shifts the default posture away from two adequate and well-controlled investigations toward one plus confirmatory evidence [6]. If that holds, the exception becomes the starting point.
Monday's question: before you pencil in a second pivotal trial, put the one-plus-confirmatory route explicitly on the agenda of your next FDA interaction. Make the agency tell you no; do not tell yourself no on their behalf.
Free download
The TPP Template
Target and minimally-acceptable profiles side by side, every claim linked to the evidence that has to support it.
Get the template →The second lever is quieter and costs just as much. So what decides the number? Sample size and follow-up should be set by the specific quantity the approval decision turns on, an event count or a clinically meaningful difference or a defined observation window, and then they should stop. Everything past that point is evidence you are paying for that the decision is not using.
Vosoritide (Voxzogo, NCT03197766) enrolled 121 children over 52 weeks in achondroplasia, a condition affecting roughly 1 in 25,000 births [12]. Fifty-two weeks is long enough to read annualised growth velocity, the endpoint the approval turned on, and too short to read final adult height, which belongs to post-marketing follow-up. Inebilizumab (N-MOmentum, NCT02200770) capped its randomised, controlled period at 197 days in a lifelong disease [13], because the endpoint was time to an adjudicated attack: the window was sized to that event rather than to the illness behind it.
Vericiguat (VICTORIA, NCT02861534) ran the other way, enrolling 5,050 patients, the largest trial in this set by some distance [14]. The reason was mechanical rather than nervous: a hard mortality-and-hospitalisation composite in a lower-event-rate population needs that many enrolled to accrue enough events to read at all. Big was correct there — the discipline is identical to small being correct elsewhere.
Size to the question.
Now the price of getting it wrong. Moore et al., writing in BMJ Open, found single-trial programmes ran a median of US$28 million, two-trial programmes nearly double at US$45 million, and programmes of three to eleven trials US$91 million; costs, they wrote, "rise exponentially as more patients and clinic visits are required" [15]. Their earlier work put the range starker still, from under US$5 million for small uncontrolled orphan trials to US$346.8 million for a single non-inferiority trial needing a large population to detect a small effect [16]. So an unnecessary trial spends tens of millions of runway you do not get back, and an over-powered arm burns the same money in smaller notes.
Monday's question, again operational: for every parameter in your protocol synopsis, the N, the duration, the endpoint count, the visit schedule, name the decision it serves. If a number is buying certainty that a later decision needs, the confirmatory study, the post-marketing commitment, the payer conversation, then it is not sized to this approval, and this approval is the one on the clock.
So if the statute allows less, and the data show more teams taking it, why does the default still inflate? Let's be honest about the incentives around the design table: nobody sitting at it is paid to argue for less evidence. The statistician rounds the sample size up, because being under-powered is a career-defining mistake and being over-powered is merely a Tuesday; the board wants the most defensible-looking package to wave at the next investor; regulatory affairs pads the plan to dodge a second review cycle; and the CRO has not one earthly reason to propose itself a smaller contract. Every function, acting reasonably on its own terms, nudges the number up, and up compounds.
Call it the defensibility ratchet. It turns one way only, and the programme that comes out the far end is the biggest package anyone could defend, never the smallest that would suffice, because it is the vector sum of everyone's individual caution rather than a decision any single person made.
I have watched this up close. Years ago I sat with a small team that had carried two pivotal trials since the very first slide deck: a budget line, a rough timeline, an assumed number of sites. Nowhere in that history had anyone asked whether the expected effect size and endpoint might qualify them for one trial plus confirmatory evidence. The second trial was there because it had always been there. That is the ratchet, playing out in one room.
Now the strongest version of the other side, because it earns one. A single trial is a single point of failure, and a second is genuine replication: insurance against a fluke, a rogue site, an unlucky p-value on a day the effect happened to under-read. Longer follow-up insures against a safety signal that only surfaces at the advisory committee, or worse, after launch. A fuller package smooths both the review and the eventual payer conversation. It is a serious argument, and a good reviewer will make it to you.
That said, here is the answer. The claim is precise: size to the decision, and know which decision each parcel of evidence is actually for. Replication insurance is worth every penny when the effect is modest, the endpoint soft, or the trial leaning on a single dominant centre, which are precisely the conditions regulators name. It is a poor buy when the effect is large, persuasive, on a hard endpoint and consistent across sites. The EMA says as much in CPMP/EWP/2330/99 (2001, still in effect): the minimum requirement is generally "one controlled study with statistically compelling and clinically relevant results", though there are "many reasons why it is usually prudent to plan for more than one study" [17]. The same guideline closes with a line worth pinning above the design table: "there is no formal requirement to include two or more pivotal studies in the phase III program" [17].
One nuance before the close: severity and unmet need buy flexibility on both levers. FDA's rare-disease guidance (final, December 2023) signals the "broadest flexibility" for severely debilitating or life-threatening rare diseases, while leaving "what is enough" to case-by-case judgement [18]; Downing's data bear this out, with orphan-designated trials running a smaller median population than non-orphan ones, 98 patients against 294 [2]. The number moves with the decision. The discipline of asking never does.
None of this needs a methodology overhaul, only three questions run against the protocol synopsis already sitting on your desk.
This is the decision-first sizing an Integrated Evidence Plan is built to force, and it is exactly the synopsis pressure-testing that our clinical development work and InovaSight are built to do. You can start it yourself this week with a printout and a red pen, and it pairs with the minimum viable evidence framework once you know which decision you are sizing toward.
The evidence you build for a decision you are not making buys you nothing on the day it matters. That is runway you do not get back and, further down the line, a medicine that reaches the patient later than it needed to. Size the programme to the decision, then stop.
Get the monthly digest
The 5 things evidence leads need to know each month: regulatory moves, RWE developments and what they mean in practice. No pitch, one email a month.