You are three weeks from locking the protocol. The statistician has been patient, but now she needs the one input everything else hangs from: the treatment effect. How much better than control do you actually expect your asset to be? This is a first-in-indication asset, so there is no prior trial in humans to point at, no published Δ to anchor to, and no comparator readout you can borrow. You have a preclinical model, a handful of open-label responders if you are lucky, and a mechanism you believe in.
So you give her a number, it goes into the sample-size calculation, out comes an N, and that N drops into the protocol looking every bit as solid as the inclusion criteria sitting next to it. It is not.
That single number is the least defensible object in the entire trial, and your whole power calculation is balanced on top of it.
Everyone in the room knew it was a guess.
The number you are about to write down is one point on a curve you have almost certainly never plotted. This post is about plotting it.
A power calculation does not generate a fact. It evaluates a function. Fix your alpha, fix the power you want, and the required sample size becomes a function of the one thing you do not know: the true treatment effect. Feed it an assumed Δ and it returns an N; feed it a smaller Δ and it returns a bigger one. Choosing a single effect size collapses that whole function down to one point and throws the rest of the curve away.
How much does the rest of the curve matter? In a 2026 worked example, the same trial with the same outcome required anywhere from 1,650 to 4,800 patients, depending purely on which of two experts' plausible belief about the effect was fed in [1]. Same design, same endpoint, a near three-fold swing in N driven by nothing but the assumption (it's from a cluster-randomised setting rather than a Phase 2, so read it as proof of the principle rather than a like-for-like analogue). The principle itself travels everywhere: "the" sample size does not exist, only a sample size as a function of what you assume.
So stop asking your statistician for "the N". Ask instead for N plotted across the full plausible range of effect sizes, because the shape of that curve is the real planning object and the single point you eventually lift off it is a downstream decision.
If the effect size were an unbiased guess, this would be an academic quibble. It isn't one. And the problem is not confined to theory: the single point is empirically wrong, and it fails in a direction you can call in advance.
Four independent reviews, one direction.
Why does the number always drift upward? Name the mechanism honestly. I call it the "fundable N": the sample size quietly reverse-engineered from your runway and your timeline, then justified backwards with the most optimistic effect the team can say out loud without laughing. You do not start from the evidence and arrive at an N; you start from the N you can afford to recruit and back-solve for the effect that makes it work. Even DELTA2, the consensus guidance on all this, says as much in politer language, as we will see.
Two adjacent problems get wrongly folded in here, so keep them out. This is not the claim that a preclinical signal shrinks by a third in humans, which describes publication bias inside the animal literature, not a translation rate. It is also distinct from enrolment-shortfall underpowering, where a trial fails to recruit the target it calculated. Both are real; neither is this. This cluster is about the assumption underneath the target, full stop.
The consequence is specific and cheap to act on. Because the bias has a known direction and a rough magnitude, haircut your point estimate on purpose and set the pessimistic end of your range deliberately. Say nothing, and the optimistic end becomes your anchor by default, which is how a field arrives at 82.1%.
Free download
The TPP Template
Target and minimally-acceptable profiles side by side, every claim linked to the evidence that has to support it.
Get the template →Weather forecasters stopped publishing a single track line for an approaching hurricane years ago, because a single line is precise, almost always wrong, and quietly convinces everyone downstream to plan for a landfall that will not happen. So they publish the cone: the spread of probable paths, prepared against at its edges rather than its centreline. Your sample size deserves the same honesty, because the single-point N is that discredited track line and the cone is the truth you already knew and chose not to draw.
Plotting N as an explicit function across an effect-size range is not something you will find sitting in registered protocols or published statistical analysis plans; it is not standard practice. So treat what follows as a tool I use, illustrated by real designs, rather than an established convention you can point to.
Years ago I sat on an early-phase, first-in-indication programme where the effect-size assumption had been set months earlier off a small open-label signal, and nobody had gone back to it since. When the team finally plotted N across the plausible range, the number already in the protocol sat squarely on the steep part of the curve, and a single point of disappointment in the true effect roughly doubled the trial. We had been one bad week of data away from an unrunnable study and never known it. We moved the committed N, pre-specified an interim, and briefed the board on the range rather than the point. The readout, when it finally came, landed below the original assumption and the trial survived it.
What changes is the object itself. You stop defending a single number sitting alone in the SAP and start handing over a curve the whole team has actually read, alongside a documented decision about which point you are committing to and why.
The obvious objection is that this is academic and that regulators want one clean number, so read the guidance closely. ICH E20, the new adaptive-designs guideline, reached draft status (Step 2b) in 2025, with the identical text issued as an FDA draft via the Federal Register. On sample size it frames the planning goal as ensuring sufficient power "under a range of plausible and clinically meaningful treatment effect sizes" [6]. That is a live regulatory document, still in consultation, endorsing precisely the framing in this post; it is draft rather than final, so treat it as direction of travel, but the direction is not ambiguous.
DELTA2, the 2018 consensus guidance on choosing the target difference, makes two useful moves. Its Recommendation 9 states that sensitivity analyses around the key sample-size inputs "should be carried out". Then it concedes, with rare candour for a guidance document, that in practice target differences are often "determined on convenience, the research budget, or some other informal basis" [7]. That is DELTA2 describing the "fundable N" back to you in its own words.
The EMA's reflection paper on adaptive designs, final since 2007, adds the guardrail that matters for the next section: sample-size reassessment "should not be seen as a substitute for careful planning" [8].
You now have named guidance language to put in front of a board, an IRB, or an agency to justify a range-based sizing rationale. None of it is your opinion. It came from the people who wrote the rules you are being asked to satisfy.
Sometimes you plot the cone and it is simply too wide, so that no single committed N is honest because the plausible effect straddles both the flat and the vertical parts of the curve. That is when you design so the data can update the number, rather than pretending you knew Δ on day one. Real trials show how.
That said, re-estimation has an honest limit worth stating plainly. NaBen (NCT01908192) named sample-size re-estimation explicitly, after a 76-patient first part, and terminated anyway [14]. Re-estimation protects you against a bad point guess; it does not conjure a signal that was never there. This is exactly what the EMA meant when it called reassessment a safety valve rather than a plan, and if you find yourself needing to re-estimate again and again, that is itself a signal the assumptions were never understood. So if your effect range straddles the steep part of the curve, pre-specify an interim or a re-estimation rule now, but interrogate the range at the planning stage first rather than deferring the thinking to a future interim look.
Take that seriously; a lot of good statisticians would make it too. Regulators do not approve a curve, IRBs do not approve a curve, and investors do not write a cheque against a cone. The protocol carries a single N, the SAP carries a single Δ, and at some point you have to commit to a number and own the readout that follows. Pushed too far, a range starts to look like intellectual hedging: a way to dodge the hard commitment, and to a sharp reviewer it can read as a team that does not actually have a hypothesis.
That objection is right about the commitment and wrong about what the curve is for. You do commit to one N, and that never changes. What changes is how you get there: reading the whole curve first, knowing exactly what true effect it can and cannot detect, and pre-specifying what you do if you land at the ugly end. ICH E20 and DELTA2 do not ask you to submit a range in place of a number. They ask you to show that your number holds up across the plausible range. The range is what makes the point defensible, and a team that has read the cone commits harder, because it commits with its eyes open.
So, three weeks from protocol lock, when the statistician asks for the effect size, give them a number, because you always will and you always should. But hand it over having plotted the cone, having named the smallest effect your recruitable N can actually catch, and having decided in advance what you do if the truth turns out to sit at the pessimistic end.
Sizing a first efficacy trial under exactly this kind of uncertainty is a large part of the trial-design work we do with resource-lean teams at Inovia, and it is rarely the statistics that prove hard. It is the discipline to look at the whole curve before someone lifts a single point off it and calls it the plan.
Write one number in the protocol. Just make sure it is not the one number in the trial you cannot defend.
Get the monthly digest
The 5 things evidence leads need to know each month: regulatory moves, RWE developments and what they mean in practice. No pitch, one email a month.
[1] Aloufi A, Wilson KJ, Wilson N, Shaw L, Price C. "Bayesian design and analysis of two-arm cluster randomised trials using assurance: Extension to binary outcomes and comparison of Markov chain Monte Carlo and Integrated Nested Laplace Approximations." Clinical Trials 2026;23(3):336-346. PMID: 41776384. https://pubmed.ncbi.nlm.nih.gov/41776384/
[2] Olivier CB, et al. "Accuracy of Event Rate and Effect Size Estimation in Major Cardiovascular Trials: A Systematic Review." JAMA Network Open 2024;7(4):e248818. PMID: 38687478. https://pubmed.ncbi.nlm.nih.gov/38687478/
[3] Jansen MS, Groenwold RHH, Dekkers OM. "The role of research ethics committees in addressing optimism in sample size calculations." Research Integrity and Peer Review 2025. PMC12699925. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC12699925/
[4] Biggs J, et al. "A systematic review of sample size estimation accuracy on power in malaria cluster randomised trials." BMC Medical Research Methodology 2024;24(1):238. PMID: 39407101. https://pubmed.ncbi.nlm.nih.gov/39407101/
[5] Stocking K, et al. "Are interventions in reproductive medicine assessed for plausible and clinically relevant effects? A systematic review of power and precision in trials and meta-analyses." Human Reproduction 2019;34(4):659-665. PMID: 30838395. https://pubmed.ncbi.nlm.nih.gov/30838395/
[6] ICH. "E20 Adaptive Designs for Clinical Trials." Draft guideline, Step 2b, endorsed 25 June 2025; public consultation to 30 November 2025. Identical text issued as FDA draft guidance, Federal Register notice 2025-18897, 30 September 2025.
[7] Cook JA, Julious SA, Sones W, et al. "DELTA2 guidance on choosing the target difference and undertaking and reporting the sample size calculation for a randomised controlled trial." BMJ 2018;363:k3750. PMID: 30560792. https://pubmed.ncbi.nlm.nih.gov/30560792/
[8] EMA/CHMP. "Reflection Paper on Methodological Issues in Confirmatory Clinical Trials Planned with an Adaptive Design." CHMP/EWP/2459/02. Final, adopted 18 October 2007.
[9] "KENDO." ClinicalTrials.gov: NCT03227328. https://clinicaltrials.gov/study/NCT03227328
[10] "Surufatinib." ClinicalTrials.gov: NCT02966821. https://clinicaltrials.gov/study/NCT02966821
[11] "Pembrolizumab pilot (adaptive Simon's two-stage)." ClinicalTrials.gov: NCT03136055. https://clinicaltrials.gov/study/NCT03136055
[12] "AdrenOSS-2." ClinicalTrials.gov: NCT03085758. https://clinicaltrials.gov/study/NCT03085758
[13] "EMBOLD (ATA188)." ClinicalTrials.gov: NCT03283826. https://clinicaltrials.gov/study/NCT03283826
[14] "NaBen." ClinicalTrials.gov: NCT01908192. https://clinicaltrials.gov/study/NCT01908192