Why promising assets die in Phase 2: five diagnosable design faults, and how to tell an avoidable death from a real one
The Phase 2 top-line email lands on a Tuesday and the primary endpoint missed. The board reads it as a coroner's report, the development lead starts drafting the wind-down memo and by Friday the asset has been quietly folded into the "lessons learned" deck. I understand the reflex, when you have one or two shots on goal and a runway measured in quarters, a missed p-value feels like permission to stop spending (and start lamenting).
But the reflex skips a question that has been sitting in regulatory guidance for a quarter of a century. ICH E10, final since 2000, puts it plainly: "If a treatment fails to show superiority to placebo, for example, it means either that the treatment was ineffective or that the study as designed and conducted was not capable of distinguishing an effective treatment from placebo." [1] Two possible parents for the same null. One is dead biology. The other is a trial that could never have detected a live effect, even if the drug worked.
A missed endpoint is a verdict on the trial before it is a verdict on the molecule.
Most drug programmes never reach patients, with some estimates putting attrition north of 90% [2], and the way we tally the casualties hides something important. Hwang and colleagues tracked 640 novel therapeutics through development: 344 (54%) failed, and 195 (57%) of those failures were charged to inadequate efficacy. [3] But "inadequate efficacy" is a bin, and into it we sweep two very different things drugs that genuinely do nothing, and drugs whose trials were built to look as if they do nothing.
So put the molecule in the dock, by all means. Just make sure the evidence assembled against it could ever have convicted a guilty drug, or acquitted an innocent one.
A null has two parents, and design error only fathers one of them
Assay sensitivity is the name for it: a trial's ability to tell an effective treatment apart from an ineffective one. Most protocols quietly bleed that sensitivity dry, and nobody notices until the top-line reads null.
Here is the asymmetry that should keep you up at night. Design error is not random noise that could push a result either way. It biases in one direction. ICH E10 again: "many trial imperfections increase the likelihood of failing to show a difference between treatments when one exists." [1] A sloppy trial is systematically more likely to bury a real effect than to conjure a fake one. The null you are staring at is exactly the result a badly built trial produces whether or not the drug works.
E10 even names the culprits that erode assay sensitivity: poor patient compliance, a population that responds poorly, concomitant non-protocol medication, spontaneous improvement in the people you enrolled, diagnostic criteria applied loosely, and endpoint assessment that drifts. [1] Read that list as a pre-flight checklist. Five of the recurring ways Phase 2 assets die map straight onto it, each carrying a question you can ask before the protocol locks. Ask it after the miss and it is an autopsy (a very expensive one at that). Ask it before, and it is something a great deal cheaper (and a lot less depressing).
Free download
The Rescue Readiness Checklist
A six-point diagnostic for a programme after an ambiguous readout: what's salvageable, and what to pull together before any rescue conversation.
Get the checklist →The five faults that kill Phase 2 assets, and the question each one leaves on your desk
Fault 1: you never found a dose that works. Dose-finding is where optimism is cheapest, and later most expensive. ICH E4, final since 1994, warned that titration-to-tolerability designs systematically push you towards the highest dose a patient can stand rather than the lowest dose that does the job. [4] Oncology spent thirty years proving the point, which is why the FDA's Project Optimus now presses sponsors to characterise exposure-response rather than crown the maximum tolerated dose.
Sotorasib is the clean illustration. Its early dosing, as The ASCO Post reported, chased a level "toxic in 25% to 30% of subjects, even though responses were seen at lower doses", while the pharmacology showed "administering more than 240 mg orally did not yield more drug entering the bloodstream." [5][6] Higher dose, same drug in the blood, more toxicity, no more benefit. (Aducanumab's discordant EMERGE and ENGAGE results are the other dose story everyone reaches for, and we have picked that one apart at length elsewhere. [7])
Question 1 : did you carry more than one active dose into Phase 2 and actually characterise exposure-response, or did you ride the MTD and hope?
Fault 2: the right drug in the wrong crowd. Average an effect across responders and non-responders and you can land on precisely zero. Gefitinib is the textbook case. In the ISEL trial it missed its primary survival endpoint in an unselected NSCLC population, with the benefit visibly concentrated in never-smokers and patients of Asian origin. [8] IPASS then enriched for EGFR mutation status and split the result inside a single trial: hazard ratio 0.48 (95% CI 0.36-0.64, P<0.001) favouring gefitinib in mutation-positive patients, and 2.85 (95% CI 2.05-3.98, P<0.001) favouring chemotherapy in the mutation-negative ones. [9] A pooled signal of nothing was a large benefit and a large harm cancelling each other out. (Mepolizumab tells a broadly similar story, moving from an unselected asthma population to an eosinophilic one, though I would read that through a secondary review rather than lean on it. [10])
Question 2 : is your effect being diluted across a population you could have split a priori on a marker the biology already handed you?
Fault 3: an endpoint too blunt to read the effect. A bathroom scale that reads only in whole stones will never register a genuine two-pound loss, and plenty of Phase 2 endpoints are exactly that blunt. Belimumab missed both co-primary endpoints at all three doses in its Phase II, with no dose response at all. [11] The post-mortem in the sponsor's own data was damning for the trial rather than the drug: the assumed annual flare rate was far too low (87% of patients flared by week 52), and 28% of the cohort was serologically inactive and unlikely to respond to anything. Notably, the team's statisticians built a new endpoint, the SLE Responder Index, out of that same failed dataset. [12] BLISS-52 then hit it, with SRI response of 51% and 58% on the two doses against 44% on placebo [13], and belimumab became the first new lupus therapy in roughly five decades. Same molecule, the same broad disease, and a ruler now fine enough to see the effect.
Question 3: has your endpoint ever actually detected the effect size you are powered to find, in this population, at this timepoint?
Fault 4: powered on a fantasy. Trial designers systematically assume larger effects than their drugs deliver, the winner's curse, documented by Gan and colleagues and independently replicated by Lord. [14][15] (The mechanism is human. You power on the number you are hoping for, not the one you would bet your own money on.) Mongersen is the invoice. Phase II produced a startling 55% and 65% remission on a two-week course against 10% on placebo; the properly powered Phase III, built on that effect size, found 22.8% against 25% (P=0.62) and stopped early for futility. [16][17]
Question 4 : is your power calculation anchored to your rosy Phase 2 point estimate, or to a deliberately discounted number you would actually stake the programme on?
Fault 5: an uncontrolled placebo response. Some indications are assay-sensitivity minefields, and E10 all but publishes the map: depression, anxiety, dementia, angina, symptomatic congestive heart failure, seasonal allergies, symptomatic GERD. [1] Placebo response in several of these is climbing over time, confirmed independently in psychiatry and in neuropathic pain. [18][19] Pimavanserin shows the fix. Two earlier Phase 3 trials on general psychiatric scales failed against what an FDA briefing document, quoted in a later review, called "a large placebo response"; the pivotal trial swapped in a purpose-built instrument (SAPS-PD) and delivered a least-squares mean change of −5.79 against −2.73 on placebo (P=0.0014), leading to approval in April 2016. [20] No new dose and no new population, just a sharper endpoint in a noisy indication.
Question 5 : did you defend assay sensitivity by design, with a placebo lead-in, centralised or blinded ratings, or enrichment for likely responders, or did you leave it to chance?
The honest half: not every miss is the trial's fault
A team that has just missed a primary will always prefer saying "the trial was flawed" to "the drug is dead". Hand that team a checklist of five reasons the trial might have been flawed and you have handed it a permission slip: five ways to keep burning runway on a corpse while calling it rigour. I have watched exactly this happen.
That said, here is why it holds up when used honestly. The same guidance that names the false-failure parent forces you to hunt for the true one, and three programmes show what that parent looks like when you find it.
- Solanezumab narrowed from mild-to-moderate Alzheimer's to mild-only disease with amyloid confirmation, and bolted on the more sensitive iADRS endpoint. The population fix from Fault 2 and the endpoint fix from Fault 3, both applied, and it still missed across EXPEDITION 1, 2 and 3. [21]
- Verubecestat shifted from mild-moderate to prodromal patients, and here is the tell: target engagement was never in doubt. Both doses cut CSF amyloid-beta by 63-81%. [22] The drug did everything it was designed to do at the molecular level and still failed to slow decline, stopping for futility with an unfavourable risk-benefit signal. [23]
- Dalcetrapib got the most rigorous redesign in this whole set: a company spun out specifically to test an ADCY9 rs1967309 AA pharmacogenomic hypothesis, enrolling only that genotype into dal-GenE, and it missed anyway. [24]
Confirmed target engagement alongside clinical failure is the signature of a true failure, and it is the finding that tells you to stop. That is the whole point of the diagnostic. Its job is to tell you which parent you are holding, so you spend the next tranche of cash on the right conclusion, even when that conclusion is that the biology is dead.
The takeaway: run the audit to find the truth, and be willing to let it tell you the biology is dead.
Rescue is a post-mortem you could have run as a pre-mortem
Look again at what the belimumab team actually did: They ran a diagnostic after the miss (wrong endpoint, wrong flare-rate assumption, a chunk of the population that could not respond), fixed it, and carried the fix into a prospective Phase III that won. The diagnostic worked. It just ran late, and it took a whole Phase II to trigger it.
Now set that against ivacaftor. Vertex enrolled only G551D-mutation carriers from the very first patient, roughly 4 to 5% of the cystic fibrosis population, because the mechanism was a gating potentiator and the biology dictated exactly who could respond. The same question the belimumab team answered in the wreckage, whether the population matched the mechanism, was asked and answered before the protocol ever locked. No rescue was needed. [25]
Same five questions. Two points in a programme's life, and two very different invoices.
This is the reframe I want you to keep. Clinical programme rescue and rigorous protocol design are the same diagnostic skill, applied at different moments and at wildly different cost. Call it the pre-mortem audit: the five-fault checklist a rescue runs over the wreckage, run instead over your protocol while it can still be changed. I was pulled into a Phase 2 programme once, after the miss, where the audit found the fault was fixable and had been sitting in the protocol all along. That kind of finding stings precisely because it was available months earlier, for the price of a design review.
We do the post-mortem version of this work under Clinical Program Rescue. The cheaper version asks the same five questions before your Phase 2 locks, the pre-specification discipline this blog keeps coming back to. Either way, when the top-line email lands on a Tuesday, the wind-down memo can wait. The first job is to work out which parent you are holding.
Get the monthly digest
The 5 things evidence leads need to know each month: regulatory moves, RWE developments and what they mean in practice. No pitch, one email a month.
References
[1] ICH E10 (2000), Choice of Control Group and Related Issues in Clinical Trials, Step 4 (FINAL). International Council for Harmonisation.
[2] Inovia Bio. How Biotechs Can Reduce Development Risk With RWE. https://blog.inovia.bio/inovia-bio-insights/how-biotechs-can-reduce-development-risk-with-rwe
[3] Hwang TJ, et al. (2016). "Failure of Investigational Drugs in Late-Stage Clinical Development and Publication of Trial Results." JAMA Intern Med. PMID: 27723879. https://pubmed.ncbi.nlm.nih.gov/27723879/
[4] ICH E4 (1994), Dose-Response Information to Support Drug Registration, Step 4 (FINAL). International Council for Harmonisation.
[5] The ASCO Post (February 2024). "Sotorasib, the Poster Child for Project Optimus: Truths and Fantasies."
[6] Moon H. "FDA initiatives to support dose optimization in oncology drug development: the less may be the better." PMC9253446. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC9253446/
[7] Inovia Bio. A Study In Failure: Learning From The Aducanumab Clinical Development Story. https://blog.inovia.bio/inovia-bio-insights/a-study-in-failure-learning-from-the-aducanumab-clinical-development-story
[8] Thatcher N, et al. (2005). ISEL trial. Lancet. PMID: 16257339. https://pubmed.ncbi.nlm.nih.gov/16257339/
[9] Mok TS, et al. (2009). IPASS trial. N Engl J Med. PMID: 19692680. https://pubmed.ncbi.nlm.nih.gov/19692680/
[10] Menzella F, et al. "Mepolizumab for severe refractory eosinophilic asthma: evidence to date and clinical potential." PMC5076744. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5076744/
[11] Wallace DJ, et al. (2009). Belimumab Phase II. Arthritis Rheum. PMID: 19714604. https://pubmed.ncbi.nlm.nih.gov/19714604/
[12] Furie RA, et al. (2009). SLE Responder Index. Arthritis Rheum. PMID: 19714615. https://pubmed.ncbi.nlm.nih.gov/19714615/
[13] Navarra SV, et al. (2011). BLISS-52. Lancet. PMID: 21296403. https://pubmed.ncbi.nlm.nih.gov/21296403/
[14] Gan HK, et al. (2012). J Natl Cancer Inst. PMID: 22491345. https://pubmed.ncbi.nlm.nih.gov/22491345/
[15] Lord SJ, et al. (2018). J Clin Epidemiol. PMID: 30297036. https://pubmed.ncbi.nlm.nih.gov/30297036/
[16] Monteleone G, et al. (2015). Mongersen Phase II. N Engl J Med. PMID: 25785968. https://pubmed.ncbi.nlm.nih.gov/25785968/
[17] Sands BE, et al. (2020). Mongersen Phase III. Am J Gastroenterol. PMID: 31850931. https://pubmed.ncbi.nlm.nih.gov/31850931/
[18] Agid O, et al. (2013). Am J Psychiatry. PMID: 23896810. https://pubmed.ncbi.nlm.nih.gov/23896810/
[19] Tuttle AH, et al. (2015). Pain. PMID: 26307858. https://pubmed.ncbi.nlm.nih.gov/26307858/
[20] Touma KTB, Touma DC. "Pimavanserin (Nuplazid) for the treatment of Parkinson disease psychosis: A review of the literature." PMC6007714. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6007714/
[21] Eli Lilly. Solanezumab EXPEDITION programme. ClinicalTrials.gov: NCT00905372, NCT00904683, NCT01900665. https://clinicaltrials.gov/study/NCT01900665
[22] Egan MF, et al. "Further analyses of the safety of verubecestat in the phase 3 EPOCH trial." PMC6685277. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6685277/
[23] Merck. Verubecestat EPOCH/APECS. ClinicalTrials.gov: NCT01739348, NCT01953601. https://clinicaltrials.gov/study/NCT01739348
[24] Roche/DalCor. Dalcetrapib dal-OUTCOMES and dal-GenE. ClinicalTrials.gov: NCT00658515, NCT02525939. https://clinicaltrials.gov/study/NCT02525939
[25] Vertex Pharmaceuticals. Ivacaftor STRIVE. ClinicalTrials.gov: NCT00909532. https://clinicaltrials.gov/study/NCT00909532
-1.png?width=1169&height=277&name=Inovia%20Logo%20Dark%20(1)-1.png)