Skip to content
Illustration of subgroup analysis credibility in clinical trials, weighing biological plausibility and replication against pre-specification alone

Subgroup rescue: "we pre-specified it" is the weakest defence in the room

Imi
Imi

Your pivotal trial missed. The overall primary endpoint came back short of significance, and yet there, sitting in the pre-specified subgroup, is a clean and convincing signal. Someone on the team says the reassuring thing: it was pre-specified, it's in the SAP, we're covered.

You are not covered...

The EMA's own subgroup guideline is blunt about it. "the absence of pre-specification cannot be taken as a direct argument that results in a particular subgroup lack credibility," it says, before going one step further: the arguments that decide credibility "should focus on biological plausibility and replication" [1]. Read that twice. The regulator is telling you, in its own words, that the line in your analysis plan is neither necessary nor sufficient. What moves a reviewer is whether the effect is mechanistically believable and whether it reproduces. That's the whole test for subgroup analysis credibility: mechanism, then replication. The rest is housekeeping.

None of that contradicts the pre-specification discipline we've argued elsewhere. For a single-arm trial, pre-specifying your endpoints and success thresholds is survival, and we've said exactly that ("Pre-Specify or Perish"). But subgroup credibility is a different question. Here, pre-specification is the floor you stand on. The case itself gets built from the mechanism and the replication.

Gefitinib: told no, then yes, on the same biology

The cleanest proof of what actually decides a subgroup case lives inside one drug's history, so let's start there rather than with a framework.

Gefitinib won accelerated US approval in 2003 on response rate alone [2]. The confirmatory ISEL trial then read out in 2004 and missed its overall survival endpoint in an unselected population. Buried in the analysis was a signal: never-smokers, and patients of Asian origin, appeared to benefit. This was not an arbitrary slice of the data. A real, biologically motivated story sat behind it. In June 2005 the FDA restricted access anyway [2], because a subgroup that looked plausible, with no defined predictive biomarker to anchor it, was simply not enough to attribute the effect to the drug.

Then the biology caught up. The prospective IPASS trial (NCT00322452) set out to test the thing directly, comparing gefitinib against chemotherapy and splitting patients by EGFR mutation status. The result was the kind a reviewer can believe: a treatment-by-subgroup interaction where the effect reversed direction between the groups. In mutation-positive patients, gefitinib cut the risk of progression sharply (PFS HR 0.48, 95% CI 0.36–0.64, P<0.001). In mutation-negative patients, chemotherapy won on response (objective response rate 1.1% versus 23.5%, P=0.001) [3]. Same drug, opposite outcomes, cleanly separated by a mechanism. In 2015 the FDA approved gefitinib for EGFR-mutation-positive metastatic NSCLC with a companion diagnostic [2].

Look hard at what changed and what didn't. The underlying hypothesis, that EGFR-driven lung cancer responds to EGFR blockade, was the same in 2005 and in 2015. What changed was that the claim stopped arriving as an unplanned slice of a failed trial and started arriving with a known oncogenic driver plus a formally tested, prospectively designed interaction. The story did not get more persuasive. The evidence did. Before you build a submission around your own subgroup, map it onto that axis and be honest about which end you are standing at.

What actually decides subgroup analysis credibility?

Strip the EMA subgroup guideline (final, EMA/CHMP/539146/2013, in effect since 1 August 2019) [1] and ICH E9 §5.7 (final, 1998) [4] down to what decides a case, and you get four things in a fairly stubborn order of importance.

Biological plausibility comes first. A mechanism a reviewer already recognises, such as a driver mutation or a known receptor, is the single strongest thing you can bring. What counts is a mechanism that was there in the biology before your trial existed. A story you assembled after the readout to explain who responded does not qualify, however well it explains the slide.

Replication and consistency come a close second. That means a second cohort, a confirmatory arm, consistent movement in a related endpoint, or fit-for-purpose real-world evidence pointing the same way. One dataset saying something surprising is a hypothesis; two independent datasets saying the same surprising thing is a finding.

The treatment-by-subgroup interaction test comes third, necessary but weaker than most sponsors think, and weak by design. The mechanics are worth spelling out, because the correct question is not whether the effect is significant inside your favoured subgroup; it is whether the effect differed significantly between subgroups (Kent et al. 2010) [5]. Those are not the same test, and the second one is brutally underpowered. Brookes et al. 2004 put a number on it: a trial with 80% power for the overall effect has only 29% power to detect an interaction of the same magnitude, and you would need to inflate the sample size roughly fourfold to close that gap [6]. So the subgroup's own flattering p-value settles the matter? It settles almost nothing. That said, even a significant interaction test is not a clean pass, because the EMA warns in the same guideline that "the sole reporting of an isolated p-value from a test for interaction is an inadequate basis for decision making" [1].

Pre-specification comes last. Worth doing, worth documenting, and, per the guideline you just read, not something whose absence counts against you. It is the floor.

Now count backwards: most sponsors attempting a subgroup rescue lead with the fourth item and barely touch the first two. The cause is rarely incompetence. It's usually the calendar: pre-specification is the one thing you can point to on the Monday after a miss, so it's the one that gets said out loud first.

Free download

The Rescue Readiness Checklist

A six-point diagnostic for a programme after an ambiguous readout: what's salvageable, and what to pull together before any rescue conversation.

Get the checklist →

The pre-specification alibi

I'll give the error a name, because it recurs: the pre-specification alibi. It is the belief that a line written into the SAP before the trial is itself the defence, that if you can prove you called your shot, the shot counts. It feels like evidence because it is documented and entirely within your control. It is none of the things a reviewer is trained to demand.

The data on how common it is are unkind. Sun et al. 2012 went through 469 randomised trials in the leading clinical journals and found 64 subgroup claims about a primary outcome. Of those, only 8% clearly pre-specified the hypothesis and showed a significant interaction test. 84% met four or fewer of ten credibility criteria. Their verdict: "the credibility of subgroup effects, even when claims are strong, is usually low" [7].

In fact, that pattern holds everywhere anyone has looked. In stroke trials, 65% of subgroup effects were pre-specified, but only 1% were rated high credibility (Ademola et al. 2023) [8]. In contemporary oncology trials from 2004 to 2020, credibility was low or very low in 93% of cases (94 of 101), and the trials that had missed their primary endpoint were four and a half times more likely to claim a differential subgroup effect with no interaction test at all (OR 4.47, 95% CI 1.42–15.55, P=.01) (Sherry et al. 2024) [9]. Pre-specification is common. Credibility is rare. Those two facts sit right next to each other, and the alibi lives in the gap between them.

A few years ago I sat with a small oncology team the week after their pivotal readout. The overall result was flat, but there was a subgroup slide, and it was beautiful: a clean separation, a tidy p-value, a patient group you could describe in a single sentence. The room had already decided this was the path forward. The first question was not about the p-value. It was: what is the mechanism, and who else has seen this effect? The silence that followed told everyone where the programme actually stood.

Why so rare? Because subgroups appear by chance, in numbers, in every large trial. The clearest demonstration is genuinely from the record. In ISIS-2, where aspirin's overall benefit after heart attack was overwhelming (P<0.00001), the investigators divided patients by astrological birth sign, and two signs, Gemini and Libra, showed a non-significantly adverse effect of aspirin (Sleight 2000) [10]. Aspirin works, and it does not stop working if you happen to have been born in June. Split any dataset finely enough and it will hand you a constellation. The sky is full of stars waiting to be joined into a shape.

"But regulators approve on post-hoc subgroups all the time"

The strongest version of the case against everything above deserves a fair hearing. Regulators do approve on subgroups that fail the tidy criteria. The EMA restricted durvalumab's PACIFIC label to PD-L1 ≥1% tumours on a post-hoc subgroup analysis. Atezolizumab plus nab-paclitaxel won its triple-negative breast cancer indication even though the trial's own hierarchical testing plan never formally validated the PD-L1-positive subgroup, because intention-to-treat overall survival was not significant (P=0.077) and the plan did not permit the subgroup to be tested [18]. BiDil reached the market on a lineage that began with a retrospective racial subgroup. So the criteria plainly are not mechanical. Doesn't that sink the thesis?

It doesn't: every one of those cases is contested on the record, and the contest runs along exactly the axis this post is about. The PACIFIC restriction was rejected by the FDA and by a panel of lung-cancer experts, who pointed to the subgroup's small size: 148 patients, around 60 survival events (Paz-Ares et al. 2020; Paratore et al. 2022) [11][12]. The atezolizumab signal? When IMpassion131 tested essentially the same biomarker hypothesis with a different chemotherapy backbone, it failed to replicate even the PD-L1-positive effect (Miles et al. 2021) [13]. Replication, again, doing the deciding. On CAPRIE, NICE and IQWiG examined the identical subgroup heterogeneity and reached opposite conclusions (Hasford et al. 2010) [14], which tells you these are judgement calls, not the output of a formula.

And where a subgroup approval has actually held up, replication is why. BiDil reached the market on the strength of A-HeFT, a new prospective trial in the target population that was stopped early for a 43% relative reduction in mortality. The old retrospective racial subgroup that started the story never carried the approval on its own. (The broader use of racial subgroups here remains criticised in the peer-reviewed literature Ellison et al. 2008 [15]; the FDA's own published defence is worth reading first-hand [16].) Where replication went the other way, the approval did not survive. The same fate met drotrecogin alfa (Xigris): approved in 2001 for the APACHE II ≥25 subgroup after a split 10–10 advisory committee vote; when PROWESS-SHOCK finally tested it in 2011, the benefit was not there, and the drug was withdrawn [17]. Replication killed it.

Run the pre-mortem before the agency does

So what do you actually do, staring at a subgroup slide after a miss? You interrogate your own claim the way a reviewer will, before you spend a submission finding out. A short pre-mortem:

  1. Mechanism, not narrative. Can you name a driver a reviewer already believes in, one that existed in the biology before your trial? If the honest answer is "it's a compelling story about who responds," you have a story. Stories do not get approved.

  2. Corroboration you can actually get. Decide now what would replicate the finding, and whether you can reach it: a second cohort, a confirmatory arm, or fit-for-purpose RWE. The cheapest time to plan replication is before the readout, not after the miss.

  3. Lead with the interaction test; do not bury it. If it is unimpressive, say so, and explain the power problem honestly. Leading with the flattering within-group number and hoping nobody asks for the interaction is precisely the move reviewers are trained to distrust, and they will see through post-hoc dredging.

  4. Pre-specify, then treat it as the floor. Do it, document it, and then do not mistake it for the argument.

There is an honest fork at the end of this. The EMA guideline is direct about the formally-failed-trial scenario (Section 5.4): "no further confirmatory conclusions are possible in a clinical trial where the primary null hypothesis cannot be rejected. One or more additional trials should usually be conducted" [1]. Sometimes the credible move is a confirmatory trial in the subgroup, designed properly and powered for the interaction. That is a strategy decision rather than a defeat, and it is far cheaper to reach when your evidence planning started before the trial read out rather than the week after. Pressure-testing a subgroup claim against these four tests, honestly, before it becomes a submission, is most of the actual work behind a lean minimum-viable-evidence plan or an integrated evidence plan; the paperwork comes after.

What the regulator is actually asking

None of this is new, which is the quietly damning part. Even in the last few years, sponsors have kept trying to carry whole programmes on a post-hoc subgroup and a good slide. The aducanumab saga is its own cautionary tale, and we will not retell it here.

The reviewer on the other side of the table couldn't care less whether your subgroup story is compelling. Compelling is easy; anyone can build a compelling slide from a failed trial. What they are checking is whether it's true, and they have a specific, unglamorous, decades-old pair of tests for that: is there a mechanism, and does it replicate. Bring those two, and pre-specification becomes a quiet footnote in your favour. Without them, that pre-specified line in your SAP is the weakest defence in the room.

Get the monthly digest

The 5 things evidence leads need to know each month: regulatory moves, RWE developments and what they mean in practice. No pitch, one email a month.

References

  1. European Medicines Agency / CHMP. "Guideline on the investigation of subgroups in confirmatory clinical trials." Final, EMA/CHMP/539146/2013, in effect 1 August 2019. https://www.ema.europa.eu/en/investigation-subgroups-confirmatory-clinical-trials-scientific-guideline
  2. Gefitinib regulatory history (dates verified against primary/FDA sources): FDA accelerated approval 5 May 2003 on response rate; ISEL confirmatory trial — overall-survival miss in an unselected population with a pre-planned never-smoker / Asian-origin subgroup signal (top-line readout December 2004; full report Thatcher N, et al. "Gefitinib plus best supportive care in previously treated patients with refractory advanced non-small-cell lung cancer (Iressa Survival Evaluation in Lung Cancer)." Lancet 2005;366:1527–37, PMID 16257339); June 2005 FDA labelling restriction following ISEL; 13 July 2015 FDA approval for EGFR-mutation-positive metastatic NSCLC with a companion diagnostic. FDA drug-approval summaries: gefitinib 2003 (PMID 14977817); gefitinib for EGFR-mutation-positive NSCLC 2015 (Kazandjian D, et al. Clin Cancer Res 2016;22:1307, PMID 26980062). ISEL predates mandatory ClinicalTrials.gov registration.
  3. IPASS trial — Mok TS, et al. "Gefitinib or carboplatin–paclitaxel in pulmonary adenocarcinoma." N Engl J Med 2009;361:947–57. PMID 19692680, DOI 10.1056/NEJMoa0810699. EGFR-mutation-positive PFS HR 0.48 (95% CI 0.36–0.64, P<0.001) confirmed verbatim in the published report; mutation-negative objective response rate 1.1% vs 23.5% (P=0.001) confirmed against the published trial data. Trial registration: ClinicalTrials.gov NCT00322452 (verified). https://clinicaltrials.gov/study/NCT00322452
  4. International Council for Harmonisation. "E9 Statistical Principles for Clinical Trials," §5.7. Final, 1998. https://database.ich.org/sites/default/files/E9_Guideline.pdf
  5. Kent DM, et al. (2010). PMID: 20704705. https://pubmed.ncbi.nlm.nih.gov/20704705/
  6. Brookes ST, et al. (2004). PMID: 15066682. https://pubmed.ncbi.nlm.nih.gov/15066682/
  7. Sun X, et al. (2012). BMJ. PMID: 22422832. https://pubmed.ncbi.nlm.nih.gov/22422832/
  8. Ademola A, et al. (2023). PMID: 36988330. https://pubmed.ncbi.nlm.nih.gov/36988330/
  9. Sherry AD, et al. (2024). PMID: 38546648. https://pubmed.ncbi.nlm.nih.gov/38546648/
  10. Sleight P. (2000). ISIS-2 astrological-subgroup discussion. PMID: 11714402; PMC59592. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC59592/
  11. Paz-Ares L, et al. (2020). PMID: 32209338. https://pubmed.ncbi.nlm.nih.gov/32209338/
  12. Paratore C, et al. (2022). PMID: 36228332. https://pubmed.ncbi.nlm.nih.gov/36228332/
  13. Miles D, et al. (2021). IMpassion131. PMID: 34219000. https://pubmed.ncbi.nlm.nih.gov/34219000/
  14. Hasford J, et al. (2010). CAPRIE / NICE vs IQWiG. PMID: 20172690. https://pubmed.ncbi.nlm.nih.gov/20172690/
  15. Ellison GTH, et al. (2008). BiDil critique. PMID: 18840235. https://pubmed.ncbi.nlm.nih.gov/18840235/
  16. Temple R, Stockbridge NL. "BiDil for heart failure in black patients: the U.S. Food and Drug Administration perspective." Ann Intern Med 2007;146(1):57–62. PMID 17200223, DOI 10.7326/0003-4819-146-1-200701020-00010. The FDA's stated justification for the BiDil decision; primary source verified (its authors were the FDA's Robert Temple and Norman Stockbridge, and it independently reports the A-HeFT 43% mortality reduction, 95% CI 11–63%). https://pubmed.ncbi.nlm.nih.gov/17200223/
  17. Drotrecogin alfa (Xigris): November 2001 FDA approval for severe sepsis at high risk of death (APACHE II ≥25) after a split 10–10 FDA Anti-Infective Drugs Advisory Committee vote (16 October 2001); the confirmatory PROWESS-SHOCK trial found no significant mortality benefit (28-day mortality 26.4% vs 24.2%, P=0.31) — Ranieri VM, et al. "Drotrecogin alfa (activated) in adults with septic shock." N Engl J Med 2012;366:2055–64, PMID 22616830, DOI 10.1056/NEJMoa1202290, NCT00604214; Eli Lilly announced worldwide voluntary market withdrawal of Xigris on 25 October 2011 (FDA Drug Safety Communication, 25 October 2011). https://pubmed.ncbi.nlm.nih.gov/22616830/
  18. Emens LA, et al. (2021). IMpassion130 (atezolizumab + nab-paclitaxel; ITT OS P=0.077). PMID: 34272041. https://pubmed.ncbi.nlm.nih.gov/34272041/

Share this post