Inovia Bio Insights

Sample Size for a First Efficacy Trial: Read the Curve

Written by Imi | 28-Jul-2026 12:02:20

You are three weeks from locking the protocol. The statistician has been patient, but now she needs the one input everything else hangs from: the treatment effect. How much better than control do you actually expect your asset to be? This is a first-in-indication asset, so there is no prior trial in humans to point at, no published Δ to anchor to, and no comparator readout you can borrow. You have a preclinical model, a handful of open-label responders if you are lucky, and a mechanism you believe in.

So you give her a number, it goes into the sample-size calculation, out comes an N, and that N drops into the protocol looking every bit as solid as the inclusion criteria sitting next to it. It is not.

That single number is the least defensible object in the entire trial, and your whole power calculation is balanced on top of it.

Everyone in the room knew it was a guess.

The number you are about to write down is one point on a curve you have almost certainly never plotted. This post is about plotting it.

The number is one point on a curve you never plotted

A power calculation does not generate a fact. It evaluates a function. Fix your alpha, fix the power you want, and the required sample size becomes a function of the one thing you do not know: the true treatment effect. Feed it an assumed Δ and it returns an N; feed it a smaller Δ and it returns a bigger one. Choosing a single effect size collapses that whole function down to one point and throws the rest of the curve away.

How much does the rest of the curve matter? In a 2026 worked example, the same trial with the same outcome required anywhere from 1,650 to 4,800 patients, depending purely on which of two experts' plausible belief about the effect was fed in [1]. Same design, same endpoint, a near three-fold swing in N driven by nothing but the assumption (it's from a cluster-randomised setting rather than a Phase 2, so read it as proof of the principle rather than a like-for-like analogue). The principle itself travels everywhere: "the" sample size does not exist, only a sample size as a function of what you assume.

So stop asking your statistician for "the N". Ask instead for N plotted across the full plausible range of effect sizes, because the shape of that curve is the real planning object and the single point you eventually lift off it is a downstream decision.

The single guess is wrong most of the time, and in a direction you can predict

If the effect size were an unbiased guess, this would be an academic quibble. It isn't one. And the problem is not confined to theory: the single point is empirically wrong, and it fails in a direction you can call in advance.

  • Cardiology: across 263 major cardiovascular trials, 82.1% overestimated the effect size they were designed to detect, with a median assumed effect of 0.72 against a median observed 0.91, a median relative overestimation of 23.1% [2].
  • Ethics oversight: in a separate review, 80% of trials overestimated their target effect, at a median relative overestimation of 59% (IQR 9 to 226%), while only 6% of ethics-committee reviews challenged the effect-size assumption at all [3]. The one checkpoint built to question an inflated number almost never looks at it.
  • Infectious disease: in malaria cluster-randomised trials, 73% of the effect sizes were overestimated by more than 10% [4], a third disease area pointing the same way.
  • Reproductive medicine: the median power to detect a clinically relevant five-point improvement was 13% [5], and that same paper prints its own table of required N across a range of plausible effects, a real published precedent for the range-based reporting I am arguing for.

Four independent reviews, one direction.

Why does the number always drift upward? Name the mechanism honestly. I call it the "fundable N": the sample size quietly reverse-engineered from your runway and your timeline, then justified backwards with the most optimistic effect the team can say out loud without laughing. You do not start from the evidence and arrive at an N; you start from the N you can afford to recruit and back-solve for the effect that makes it work. Even DELTA2, the consensus guidance on all this, says as much in politer language, as we will see.

Two adjacent problems get wrongly folded in here, so keep them out. This is not the claim that a preclinical signal shrinks by a third in humans, which describes publication bias inside the animal literature, not a translation rate. It is also distinct from enrolment-shortfall underpowering, where a trial fails to recruit the target it calculated. Both are real; neither is this. This cluster is about the assumption underneath the target, full stop.

The consequence is specific and cheap to act on. Because the bias has a known direction and a rough magnitude, haircut your point estimate on purpose and set the pessimistic end of your range deliberately. Say nothing, and the optimistic end becomes your anchor by default, which is how a field arrives at 82.1%.

Free download

The TPP Template

Target and minimally-acceptable profiles side by side, every claim linked to the evidence that has to support it.

Get the template →

Plot the cone, then read it

Weather forecasters stopped publishing a single track line for an approaching hurricane years ago, because a single line is precise, almost always wrong, and quietly convinces everyone downstream to plan for a landfall that will not happen. So they publish the cone: the spread of probable paths, prepared against at its edges rather than its centreline. Your sample size deserves the same honesty, because the single-point N is that discredited track line and the cone is the truth you already knew and chose not to draw.

Plotting N as an explicit function across an effect-size range is not something you will find sitting in registered protocols or published statistical analysis plans;  it is not standard practice. So treat what follows as a tool I use, illustrated by real designs, rather than an established convention you can point to.

  1. Plot the whole cone: compute N across the full plausible range of Δ, from the pessimistic end to the optimistic one, holding alpha and target power fixed, then plot it. It takes an afternoon, and that afternoon is the cheapest one in the entire programme.
  2. Find where the curve goes vertical: somewhere on the plot, a small drop in assumed effect produces a large jump in required N, and that steep region is where your programme's feasibility actually lives. If your chosen Δ sits on it, you are one modest disappointment away from a trial you cannot run.
  3. Invert the question: instead of asking what N you need for your assumed effect, ask what the smallest true effect is that a recruitable, fundable N could still detect, then ask whether an effect that small is even worth owning. For a resource-lean biotech this is the single most useful read of the whole curve.
  4. Anchor the pessimistic end in real data: the bottom of your range should not be a shrug, so where relevant real-world or natural-history data exists, a data-landscaping exercise can tighten the plausible range instead of leaving it to the loudest opinion in the room.

Years ago I sat on an early-phase, first-in-indication programme where the effect-size assumption had been set months earlier off a small open-label signal, and nobody had gone back to it since. When the team finally plotted N across the plausible range, the number already in the protocol sat squarely on the steep part of the curve, and a single point of disappointment in the true effect roughly doubled the trial. We had been one bad week of data away from an unrunnable study and never known it. We moved the committed N, pre-specified an interim, and briefed the board on the range rather than the point. The readout, when it finally came, landed below the original assumption and the trial survived it.

What changes is the object itself. You stop defending a single number sitting alone in the SAP and start handing over a curve the whole team has actually read, alongside a documented decision about which point you are committing to and why.

The regulators are already asking for the sample-size range

The obvious objection is that this is academic and that regulators want one clean number, so read the guidance closely. ICH E20, the new adaptive-designs guideline, reached draft status (Step 2b) in 2025, with the identical text issued as an FDA draft via the Federal Register. On sample size it frames the planning goal as ensuring sufficient power "under a range of plausible and clinically meaningful treatment effect sizes" [6]. That is a live regulatory document, still in consultation, endorsing precisely the framing in this post; it is draft rather than final, so treat it as direction of travel, but the direction is not ambiguous.

DELTA2, the 2018 consensus guidance on choosing the target difference, makes two useful moves. Its Recommendation 9 states that sensitivity analyses around the key sample-size inputs "should be carried out". Then it concedes, with rare candour for a guidance document, that in practice target differences are often "determined on convenience, the research budget, or some other informal basis" [7]. That is DELTA2 describing the "fundable N" back to you in its own words.

The EMA's reflection paper on adaptive designs, final since 2007, adds the guardrail that matters for the next section: sample-size reassessment "should not be seen as a substitute for careful planning" [8].

You now have named guidance language to put in front of a board, an IRB, or an agency to justify a range-based sizing rationale. None of it is your opinion. It came from the people who wrote the rules you are being asked to satisfy.

When the range is too wide to commit, build the update in

Sometimes you plot the cone and it is simply too wide, so that no single committed N is honest because the plausible effect straddles both the flat and the vertical parts of the curve. That is when you design so the data can update the number, rather than pretending you knew Δ on day one. Real trials show how.

  • Put the assumption in the open: KENDO (NCT03227328) used a group-sequential response-adaptive design and registered its effect-size assumption in plain sight, at 8 versus 12-month progression-free survival, with the simulated power that bought under two competing designs, 0.911 against 0.717 [9]. A sponsor showing both its assumption and its consequence where a reviewer can see them.
  • Stage the commitment: Simon's two-stage designs, used in trials such as surufatinib (NCT02966821) [10] and an adaptive pembrolizumab pilot that could switch regimen (NCT03136055) [11], let you run a small first stage before committing the full N, buying information before you bet the programme on it.
  • Pre-specify the look: AdrenOSS-2 (NCT03085758) built in an interim analysis after 50% of patients had completed the study [12], while the EMBOLD trial of ATA188 made its Part 2 dose and initiation contingent on Part 1 (NCT03283826) [13], both deferring a committed design decision until the data has had its say.

That said, re-estimation has an honest limit worth stating plainly. NaBen (NCT01908192) named sample-size re-estimation explicitly, after a 76-patient first part, and terminated anyway [14]. Re-estimation protects you against a bad point guess; it does not conjure a signal that was never there. This is exactly what the EMA meant when it called reassessment a safety valve rather than a plan, and if you find yourself needing to re-estimate again and again, that is itself a signal the assumptions were never understood. So if your effect range straddles the steep part of the curve, pre-specify an interim or a re-estimation rule now, but interrogate the range at the planning stage first rather than deferring the thinking to a future interim look.

"But a protocol needs one number"

Take that seriously; a lot of good statisticians would make it too. Regulators do not approve a curve, IRBs do not approve a curve, and investors do not write a cheque against a cone. The protocol carries a single N, the SAP carries a single Δ, and at some point you have to commit to a number and own the readout that follows. Pushed too far, a range starts to look like intellectual hedging: a way to dodge the hard commitment, and to a sharp reviewer it can read as a team that does not actually have a hypothesis.

That objection is right about the commitment and wrong about what the curve is for. You do commit to one N, and that never changes. What changes is how you get there: reading the whole curve first, knowing exactly what true effect it can and cannot detect, and pre-specifying what you do if you land at the ugly end. ICH E20 and DELTA2 do not ask you to submit a range in place of a number. They ask you to show that your number holds up across the plausible range. The range is what makes the point defensible, and a team that has read the cone commits harder, because it commits with its eyes open.

The one number you can actually defend

So, three weeks from protocol lock, when the statistician asks for the effect size, give them a number, because you always will and you always should. But hand it over having plotted the cone, having named the smallest effect your recruitable N can actually catch, and having decided in advance what you do if the truth turns out to sit at the pessimistic end.

Sizing a first efficacy trial under exactly this kind of uncertainty is a large part of the trial-design work we do with resource-lean teams at Inovia, and it is rarely the statistics that prove hard. It is the discipline to look at the whole curve before someone lifts a single point off it and calls it the plan.

Write one number in the protocol. Just make sure it is not the one number in the trial you cannot defend.

Get the monthly digest

The 5 things evidence leads need to know each month: regulatory moves, RWE developments and what they mean in practice. No pitch, one email a month.

References

[1] Aloufi A, Wilson KJ, Wilson N, Shaw L, Price C. "Bayesian design and analysis of two-arm cluster randomised trials using assurance: Extension to binary outcomes and comparison of Markov chain Monte Carlo and Integrated Nested Laplace Approximations." Clinical Trials 2026;23(3):336-346. PMID: 41776384. https://pubmed.ncbi.nlm.nih.gov/41776384/

[2] Olivier CB, et al. "Accuracy of Event Rate and Effect Size Estimation in Major Cardiovascular Trials: A Systematic Review." JAMA Network Open 2024;7(4):e248818. PMID: 38687478. https://pubmed.ncbi.nlm.nih.gov/38687478/

[3] Jansen MS, Groenwold RHH, Dekkers OM. "The role of research ethics committees in addressing optimism in sample size calculations." Research Integrity and Peer Review 2025. PMC12699925. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC12699925/

[4] Biggs J, et al. "A systematic review of sample size estimation accuracy on power in malaria cluster randomised trials." BMC Medical Research Methodology 2024;24(1):238. PMID: 39407101. https://pubmed.ncbi.nlm.nih.gov/39407101/

[5] Stocking K, et al. "Are interventions in reproductive medicine assessed for plausible and clinically relevant effects? A systematic review of power and precision in trials and meta-analyses." Human Reproduction 2019;34(4):659-665. PMID: 30838395. https://pubmed.ncbi.nlm.nih.gov/30838395/

[6] ICH. "E20 Adaptive Designs for Clinical Trials." Draft guideline, Step 2b, endorsed 25 June 2025; public consultation to 30 November 2025. Identical text issued as FDA draft guidance, Federal Register notice 2025-18897, 30 September 2025.

[7] Cook JA, Julious SA, Sones W, et al. "DELTA2 guidance on choosing the target difference and undertaking and reporting the sample size calculation for a randomised controlled trial." BMJ 2018;363:k3750. PMID: 30560792. https://pubmed.ncbi.nlm.nih.gov/30560792/

[8] EMA/CHMP. "Reflection Paper on Methodological Issues in Confirmatory Clinical Trials Planned with an Adaptive Design." CHMP/EWP/2459/02. Final, adopted 18 October 2007.

[9] "KENDO." ClinicalTrials.gov: NCT03227328. https://clinicaltrials.gov/study/NCT03227328

[10] "Surufatinib." ClinicalTrials.gov: NCT02966821. https://clinicaltrials.gov/study/NCT02966821

[11] "Pembrolizumab pilot (adaptive Simon's two-stage)." ClinicalTrials.gov: NCT03136055. https://clinicaltrials.gov/study/NCT03136055

[12] "AdrenOSS-2." ClinicalTrials.gov: NCT03085758. https://clinicaltrials.gov/study/NCT03085758

[13] "EMBOLD (ATA188)." ClinicalTrials.gov: NCT03283826. https://clinicaltrials.gov/study/NCT03283826

[14] "NaBen." ClinicalTrials.gov: NCT01908192. https://clinicaltrials.gov/study/NCT01908192