When the PAOLA-1 investigators put bevacizumab in both arms of the trial, olaparib's label was already half-written. Olaparib went in as the add-on: 537 patients received it on top of bevacizumab, 269 got placebo on top of bevacizumab [6]. Whatever the efficacy curve showed, the only claim that architecture could earn was a combination one. And that is exactly what olaparib got: an indication for use in combination with bevacizumab. A sentence fixed the day the protocol locked, long before the first Kaplan-Meier curve.
Compare that to SOLO-1, the other pivotal olaparib trial. Olaparib monotherapy, 300 mg twice daily, 260 patients against 131 on placebo, in a BRCA-mutated population [5]. Monotherapy in, monotherapy out: a clean single-agent, BRCA-restricted indication. One drug, one label, two indication paragraphs, each a faithful print of the trial that produced it.
Here is the comfortable belief I want to take apart. Run a clean trial, generate strong data, and a good label follows; the wording is something the regulator sorts out at the end, once they have seen how well the drug works. It is a reassuring story and it is largely wrong.
Your label's ceiling is set the day you lock the protocol.
The regulator writes inside the box your design drew. They can trim that box. They will not reliably enlarge it, and the enlargement, when it comes, is the first thing pulled back under scrutiny.
So the protocol should be reverse-engineered from the exact label sentence you want. Three raw materials do most of the work: the population you enrol, the comparator and regimen you run, and the design that sets your evidentiary ceiling. Take each in turn.
The indication statement is not free prose. It is assembled, almost mechanically, from what your protocol supplies: the population you studied becomes the population descriptor, and what sat in each arm decides whether you earn a monotherapy claim, a combination claim, or a genuine head-to-head one. The regulator is a compositor setting type from the copy you handed over, not an author inventing it.
The agencies say as much. The EMA's assessor-facing guide on the wording of therapeutic indication (EMA/CHMP/483022/2019, adopted October 2019, final) tells reviewers the final wording "may be wider or more restricted compared to the therapeutic indication as initially proposed by the applicant as well as compared to the population studied" [1]. Hold on to "wider"; a lot of programmes quietly bank on it. On regimen the guide is just as direct: specify "monotherapy" in section 4.1 if that is how the product was studied, and specify the combination if it was studied in combination [1]. The arm structure writes the claim.
The US regulation is blunter still. 21 CFR 201.57(c)(2) says indications "must not be implied or suggested in other sections of the labeling if not included in this section" [2]. The studied ground is the ceiling. You cannot smuggle a broader use in through the back door of the clinical studies section.
One precise caveat, because it changes where you spend effort. Endpoints, unlike population and regimen, generally do not appear in the indication sentence at all; the EMA guide is explicit that "references to study endpoints should, in general, not be included in 4.1" [1]. Endpoint choice sets the evidentiary ceiling, not the wording. A brilliant endpoint cannot rescue a population you never enrolled.
Free download
The TPP Template
Target and minimally-acceptable profiles side by side, every claim linked to the evidence that has to support it.
Get the template →Population: Kalydeco (ivacaftor). Ivacaftor was first approved in 2012 in one narrow slice of cystic fibrosis: patients with the G551D mutation, aged six and up. That was no whim: it was the population two placebo-controlled trials had actually enrolled, STRIVE (NCT00909532, n=161, age 12 and over) and ENVISION (NCT00909727, n=52, ages six to eleven), both reading out FEV1 change at 24 weeks [4]. The label reached exactly as far as those trials reached, no further. The proof is Vertex's own later trial: KONNECTION (NCT01614470, n=39) extended the label to nine further non-G551D gating mutations, and its registry record names STRIVE and ENVISION as the reason for the original restriction [4]. Broadening was possible, it just cost more trials. Today the label reaches "patients aged 1 month and older who have at least one mutation in the CFTR gene that is responsive to ivacaftor based on clinical and/or in vitro assay data" [4], and that closing clause about in vitro data is itself a fossil of a later widening through laboratory bridging rather than a fresh trial for every mutation.
Comparator and regimen: Lynparza and the cardiovascular contrast. Back to olaparib, the cleanest exhibit there is. SOLO-1 and PAOLA-1 tested the same molecule and produced two different indication paragraphs, and the difference tracks the arm structure, not the effect size. Putting bevacizumab in both arms of PAOLA-1 did not make olaparib work better or worse. It made "in combination with bevacizumab" the only comparative sentence the trial could write [6]. If your commercial case rested on a monotherapy claim, that protocol foreclosed it on day one, silently, whatever the hazard ratio turned out to be.
Cardiovascular outcomes trials make the point at scale. PARADIGM-HF (NCT01035255, n=8,442) ran sacubitril/valsartan against enalapril, an active comparator and a real head-to-head [7]. That design can support a superiority claim over a named standard of care. Line up the placebo-controlled outcome trials of the same era against it, SHIFT, EMPA-REG OUTCOME, CANVAS, FOURIER, ODYSSEY OUTCOMES, LEADER, and every one added its drug on top of background therapy versus placebo on top of background therapy. Each can earn an "on top of standard care" additive claim and nothing more, however large the benefit. The ceiling was a structural consequence of the comparator, chosen at design, independent of how well any of these drugs worked.
ICH E10 has said this plainly since 2001. "The choice of control group is always a critical decision in designing a clinical trial," it reads; that choice "affects the inferences that can be drawn from the trial ... the acceptability of the results by regulatory authorities" [3]. And the line every founder eyeing a placebo arm for speed should pin above the desk: "Placebo-controlled trials lacking an active control give little useful information about comparative effectiveness" [3].
A few years ago I sat in a protocol review for an early-phase programme where the team had all but settled on an active comparator. The reasoning was reasonable on its own terms: an active-controlled design would clear the ethics committee faster and recruit more willingly than asking sick patients to accept placebo. Sensible instinct. What nobody had done was ask what claim that comparator could earn, or what it quietly ruled out. It had been chosen to solve an operational problem, and it was about to write a commercial one into the label. We caught it. Plenty of teams do not.
This is why comparator choice deserves more scrutiny than almost any other protocol line. If you are weighing a single-arm design at the far end of that spectrum, where there is no concurrent control at all, the trade-offs are sharper still, and the EMA's position on single-arm trials is worth reading alongside this.
Design: reaching for breadth without betting the narrow claim. Endpoint and design cannot write your indication sentence, but they decide how much of the population you can defensibly claim. Pembrolizumab is the worked example, and its broader development story is worth reading in full, so just the mechanism here. KEYNOTE-024 (NCT02142738, n=305) enrolled only PD-L1 "strong" NSCLC, TPS of 50% or more, and won the first monotherapy indication in that narrow band [8]. KEYNOTE-042 (NCT02220894, n=1,274) then reached for a broader PD-L1 "positive" population through a registered hierarchical testing sequence: overall survival was tested first in the TPS ≥50% group, then, only if that succeeded, in TPS ≥20%, and finally in TPS ≥1% [9]. The design could stretch towards the wide claim without ever risking the narrow one it already held. That is deliberate breadth: a staircase, not a wager on a single all-comers hypothesis.
The levers also compound. Liraglutide is one molecule: at the dose studied in type 2 diabetes, the setting of trials like LEADER (NCT01179048, n=9,341), it is Victoza; taken to 3.0 mg in a non-diabetic, weight-endpoint population through SCALE Obesity and Prediabetes (NCT01272219, n=3,731), the same molecule is Saxenda [10]. Dose, population and endpoint look like recruitment parameters. On the label, they turn out to be the raw material of the product you end up allowed to sell.
The largest, most recent dataset says the label is usually broader than the trial population, not narrower. Vokinger and colleagues (2025) examined 263 drugs and 278 indications approved across the US, EU and Switzerland between 2012 and 2023, and found approved label populations were, in aggregate, broader than the trial populations, more so in the United States than in Europe [11]. So if regulators routinely widen the studied population for you, why agonise over reverse-engineering the protocol? Why not enrol for speed and let the agency generalise?
Three reasons.
First, that aggregate broadening is a favour, not a plan. Call it the "extrapolation dividend" the discretionary breadth a regulator may grant on top of what you actually proved. It is real, and notably common enough to show up across a dataset of 263 drugs. But you cannot underwrite a Series B, a revenue forecast or a launch plan on a discretionary judgement call made after your data are in. A dividend you cannot predict is not an asset you get to spend in advance.
Second, it is granted unevenly, and the studied ground is what reliably holds. Sumi and colleagues (2020) found the granted indication differed from the trial population on biomarker status in 20 of 38 oncology approvals, just over half [12]. Feldman and colleagues (2022) identified extrapolation in only about 20% of 105 novel approvals, 23 instances across 21 drugs [13]. The dividend is the exception the data throws up, not the base case you design towards.
Third, and most telling: extrapolation is the first thing rolled back when anyone looks hard. Aducanumab was approved in 2021 with a label covering all patients with Alzheimer's disease, though it had been studied only in mild disease. Within weeks, under public scrutiny, the FDA narrowed it to the population actually treated in the trials: mild cognitive impairment and mild dementia [13][14]. The box the protocol drew reasserted itself almost immediately. There is a full post-mortem of that programme on this blog. The lesson holds: broad extrapolation is borrowed, and the loan can be called.
One honest caveat, so nobody accuses me of forcing every label change into a design story: not every narrowing is about design. Niraparib's label was broadened through an all-comers trial, then voluntarily narrowed in 2022, reportedly for a class-wide overall-survival safety signal across the PARP inhibitors, per trade coverage, rather than because its protocol had failed to support the claim [15]. Safety-driven changes are a different animal; keep them separate.
So, concretely, what do you do with a live protocol before it locks? Reverse-engineer it from the sentence you want on the label.
Write the exact target indication sentence first. Not a paragraph of aspiration, the single sentence you want printed on the approved label, wording and all. If you have not built the target product profile that sits behind it, that is the prior step, and there is a way to build it without the guesswork. This post assumes the sentence exists.
Underline every claim-bearing clause. The population descriptor. The monotherapy, combination or comparative claim. The line of therapy and treatment setting. Each underlined clause is a promise your trial has to be built to keep.
Name the single protocol decision that is each clause's raw material. Eligibility criteria supply the population descriptor. Arm structure supplies the monotherapy-versus-combination-versus-comparative claim. If a clause has no protocol decision feeding it, you have found a claim you are hoping for rather than designing for.
Pressure-test every feasibility-driven deviation against the claim it forecloses, before lock. Broadening eligibility to hit an enrolment timeline. Choosing an active comparator to clear an ethics committee. Adding a combination backbone to satisfy standard of care. Each can be the right call. Just make it knowingly, having priced the claim it costs you, rather than discovering the cost at approval. That is the whole PAOLA-1 lesson in one instruction.
Decide your broadening strategy on purpose. There are three honest routes to a broad label: prove the broad claim head-on; build a staged, hierarchical path to it, the way KEYNOTE-042 did; or accept a narrow first label and budget the later broadening studies into the plan and the runway, the way ivacaftor's programme did. Pick one deliberately. The failure mode is defaulting into a fourth non-strategy: enrolling for convenience and hoping the extrapolation dividend covers the gap.
Think of the protocol as a photographic negative and the label as the print. Everything in the print was captured in the negative first. No developing skill in the regulatory darkroom brings out detail the negative never held. Design for convenience and you get a convenient trial and a print to match: narrower and more caveated than the science deserved.
And someone pays for that gap the patients just outside the studied box pay first, in off-label uncertainty and reimbursement fights, and then your next raise pays, because the valuation is priced off the label you can defend, not the one you meant to earn. The trial you can run is always the cheaper option today. The label you want is the one that has to be standing at launch. Working backwards from that one sentence to a protocol that can earn it, before the design locks, is the work that pays for itself.
Get the monthly digest
The 5 things evidence leads need to know each month: regulatory moves, RWE developments and what they mean in practice. No pitch, one email a month.