In PIONEER 1 (Aroda et al., 2019), oral semaglutide was tested against placebo in type 2 diabetes. At the 14 mg dose, the placebo-adjusted reduction in HbA1c was 1.1 percentage points. It was also 1.4 percentage points. [1]
Same trial, same patients, same dataset. Two numbers, both pre-specified, both sitting in the published record, both correct. A 0.3-point gap that nobody miscalculated and nobody fudged.
If that sounds like a contradiction, it is only because of a habit most non-statistician teams have never had reason to unlearn. The two numbers answer two different clinical questions. And the thing that decides which question your trial answers is the estimand, the building block of what regulators now call the estimands framework, chosen ahead of the statistical method and settled, or so it should be, long before anyone opens a dataset.
Here is the uncomfortable part for anyone who has ever signed off a statistical analysis plan thinking the analysis was the statistician's department. The estimand is a clinical decision. It just happens to live in a statistics document. If you have ever waved through "we'll use the ITT population" as your answer to the analysis question, you delegated something you did not know you were delegating.
Let's be frank. For most clinical leads, the big analysis-population decision is ITT versus per-protocol, and everyone has a view. Intention-to-treat keeps the pragmatism and the randomisation intact. Per-protocol tells you what happens in the patients who actually took the drug as intended. Pick your side, brief the statistician, move on.
But both labels quietly duck the one thing that changes the answer.
Think about what happens after a patient is randomised but before their endpoint is measured. They need rescue medication. They stop the drug because of a side effect. They die. Statisticians call these intercurrent events, and every one of them forces a question: what does the outcome even mean now? ITT and per-protocol do not answer that question. They assume an answer, silently, and bury it inside a population label. Call it the silent default a strategy for handling intercurrent events that got chosen without anyone choosing it.
That is precisely the gap ICH E9(R1), the 2019 addendum to the E9 guideline on statistical principles, was written to close. [2] Instead of letting a population label imply how you treat post-randomisation events, you state the strategy outright, as part of the question itself.
And the old vocabulary was papering over more than most teams realise. When Rehal and colleagues reconstructed the REMoxTB tuberculosis trial, they found that 43% of patients (777 of 1,785) had missing or potentially contaminated results that could have shifted how they were classified under the old per-protocol and modified-ITT framing. [3] That was the analysis question hiding in plain sight, dressed up as a data-quality footnote.
An estimand is a precise description of the treatment effect a trial sets out to estimate. Under ICH E9(R1) it has five attributes, and you can hold your statistician to every one of them: [2]
Four of those five, a clinical development lead already reasons about instinctively. You argue about inclusion criteria, endpoints, comparators and effect measures every week of your working life. The fifth is the one nobody ever put in front of you, and it is the one carrying the argument. So it gets its own section.
Write the five down for your primary endpoint and you have a one-paragraph specification of exactly what your pivotal trial claims. Before the protocol locks, that costs you a meeting. Wait until database lock, and it can cost you the claim.
Free download
The TPP Template
Target and minimally-acceptable profiles side by side, every claim linked to the evidence that has to support it.
Get the template →It's the one that gets quietly delegated away, and it shouldn't be.
An intercurrent event is anything that happens after treatment starts and changes what your outcome means, or whether it exists at all: rescue medication, discontinuation, a switch to another therapy, death. ICH E9(R1) names five estimand strategies for dealing with them, and each one commits you to its own clinical claim. Swap the strategy and you swap what the trial actually asserts, not merely how a number gets crunched.
Now PIONEER 1 resolves. In the placebo arm, 14% of patients used rescue medication; in the semaglutide 14 mg arm, 1.1% did. [1] Treatment policy counts all of that real-world rescue and lands at 1.1 percentage points. The hypothetical (trial-product) estimand asks what the drug would have done had nobody reached for rescue, and lands at 1.4.
Same data, two questions, two answers, and both of them correct.
That 0.3-point gap is not a rounding error. On plenty of endpoints, a gap that size is the difference between a trial that hit and a trial that missed. Choose the estimand late, or by default, and you let the accidents of your dataset (who happened to need rescue, who happened to drop out) choose your clinical claim for you.
Once the events are in your dataset you cannot retrofit a strategy that needed you to collect things differently, or to have kept measuring outcomes after a patient discontinued, and no amount of statistical sophistication downstream will conjure back data you chose, months earlier and without realising it, not to gather.
You do not need a consultant to tell you to plan early. The guideline says it in its own words: the clinical questions and estimands "should be specified at the initial stages of planning any clinical trial" (A.3.4), and a late change "can reduce the credibility of the trial" (A.6). [2] Prospective is nearly free; retrospective is expensive, and sometimes impossible.
This is the same pre-specification discipline the FDA presses everywhere else. Its guidance on external control arms already insists that intercurrent events be assessed across both arms and that the protocol and SAP be locked up front rather than reverse-engineered afterwards. The estimand belongs in that same early conversation, inside your integrated evidence plan, not just in the SAP a statistician writes once the data are in.
A lean team hears "five attributes and five strategies" and reasonably concludes this is academic over-engineering that slows an already brutal process down. ITT served the industry for decades. The commonest complaint, captured in the title of a 2020 review by Mitroiu and colleagues that asked whether estimands were just "old wine in new barrels", is that the framework relabels problems statisticians already knew how to handle and charges you new jargon for the privilege. [5] And the killer version of the objection: even the statisticians cannot agree. Fleming and colleagues argue that treatment-policy-plus-composite should be the rigorous default and that the other strategies are routinely misapplied; Morris and colleagues, and Keene and colleagues, published direct rebuttals defending those strategies when they are properly justified. [4][6][7] If the experts are still fighting, why should a founder care?
Because the framework does not add the complexity. It names complexity that was always there, and that you were carrying whether you named it or not. PIONEER 1 did not become two numbers because someone applied E9(R1). It was always going to be two numbers; the only question was whether anyone chose which one on purpose. The expert fight makes the point rather than undermining it. Fleming and Morris are arguing about which strategy is most defensible, and that is a debate you can only have once you have specified an estimand at all. That part is not optional. ICH E9(R1) has been final adopted guidance since November 2019, and final FDA guidance since May 2021. [2][8] So it is already the standard a reviewer holds you to, rather than some proposal you might get asked about one day, and reviewers push hard on pre-specification, as the EMA's reflection paper on single-arm trials makes plain.
Treatment policy stays the most conservative starting point, the one that leans on the fewest assumptions. That said, whether to move off it, and how far, is exactly the judgement your statistician exists to make. Your job is to know that a judgement is being made at all.
Governing this well never meant deriving a hazard ratio over your statistician's shoulder. It takes four questions, and the nerve to ask them before the protocol locks, not after.
Ask those four and you catch the exact moment a "technical" choice quietly changes what your trial is allowed to claim. That is the whole exercise. No statistics degree required.
The estimand is a clinical decision that happens to be written in a statistics document. Which means a founder or a clinical lead does not merely get to have a view on it. They have to.
Get the monthly digest
The 5 things evidence leads need to know each month: regulatory moves, RWE developments and what they mean in practice. No pitch, one email a month.
[1] Aroda VR, et al. (2019). PIONEER 1 estimand analysis (oral semaglutide monotherapy vs placebo, type 2 diabetes). Diabetes, Obesity and Metabolism. PMID: 31168921. https://pubmed.ncbi.nlm.nih.gov/31168921/ — primary trial results corroborated in Diabetes Care (Aroda VR, et al., 2019). PMID: 31186300. https://pubmed.ncbi.nlm.nih.gov/31186300/
[2] ICH E9(R1): Addendum to Statistical Principles for Clinical Trials — Estimands and Sensitivity Analysis in Clinical Trials (EMA/CHMP/ICH/436221/2017). FINAL — Step 4 adoption 20 November 2019; adopted as final FDA guidance 12 May 2021. https://www.ema.europa.eu/en/ich-e9-statistical-principles-clinical-trials-scientific-guideline
[3] Rehal S, et al. (2023). REMoxTB estimand reconstruction (tuberculosis). Clinical Trials. PMID: 37277978. https://pubmed.ncbi.nlm.nih.gov/37277978/
[4] Fleming TR, et al. (2025). "Appropriate Implementation of ICH E9(R1) Estimand Strategies." Statistics in Medicine. PMID: 40394856. https://pubmed.ncbi.nlm.nih.gov/40394856/ (source of the EXSCEL post-discontinuation MACE figure)
[5] Mitroiu M, et al. (2020). "A narrative review of estimands in drug development and regulatory evaluation: old wine in new barrels?" Trials. PMID: 32703247. https://pubmed.ncbi.nlm.nih.gov/32703247/
[6] Morris TP, et al. (2026). Comment on Fleming et al. Statistics in Medicine. PMID: 41847726. https://pubmed.ncbi.nlm.nih.gov/41847726/
[7] Keene ON, Wright D, Fletcher C. (2026). Comment on Fleming et al. Statistics in Medicine. PMID: 41847717. https://pubmed.ncbi.nlm.nih.gov/41847717/
[8] FDA. Federal Register notice 2021-10066 (12 May 2021): adoption of ICH E9(R1) as final FDA guidance. https://www.federalregister.gov/documents/2021/05/12/2021-10066/e9r1-statistical-principles-for-clinical-trials-addendum-estimands-and-sensitivity-analysis-in
Trials referenced (ClinicalTrials.gov):
[9] Taiho Oncology. SUNLIGHT. NCT04737187. https://clinicaltrials.gov/study/NCT04737187
[10] Alnylam. KARDIA-2 (zilebesiran). NCT05103332. https://clinicaltrials.gov/study/NCT05103332
[11] UCB. BE HEARD I. NCT04242446. https://clinicaltrials.gov/study/NCT04242446
[12] Bond Avillion / AstraZeneca. BATURA. NCT05505734. https://clinicaltrials.gov/study/NCT05505734
[13] Novartis. VictORION-Mono. NCT05763875. https://clinicaltrials.gov/study/NCT05763875