Skip to content
Illustration of the sensitivity analysis menu for an external control arm study: alternative matching specifications, missing-data handling, and quantitative bias analysis, as reviewed by FDA and EMA assessors.

One robustness check is not a menu: the sensitivity analyses that make an external control credible

Imi
Imi

When a reviewer opens your external control arm study, the hazard ratio is not the thing on trial. Its survival is.

You ran the single-arm trial. You built the external control from real-world data, aligned the cohorts, produced a clean point estimate that clears your threshold. Good. Then a statistician at the FDA, or an Evidence Assessment Group at NICE, turns to the appendix and asks a different question altogether: does that number hold when the assumptions change? Or did the whole effect live and die on one modelling choice nobody stress-tested?

Here is the figure that should stop you. In a cross-sectional audit of 180 externally controlled trials published between 2010 and 2023, Liu and colleagues found that only 32 of them, 17.8%, ran any sensitivity analysis at all on the primary outcome. Only two, 1.1%, ran a formal quantitative bias analysis. [1] That is the broad published literature, not a hand-picked regulatory subset, and it is grim reading. The great majority of these studies handed a reviewer a single point estimate and, in effect, asked them to take it on trust.

The credibility question behaves like a contested decision in a match that hangs on one replay. A single camera angle, played once, almost never settles it; the officials want the same moment from every angle before the result stands. One robustness check is one camera angle. The reviewer is asking for the replay.

Call the failure mode the "one-and-done" robustness check: you re-run the analysis a single alternative way, watch the hazard ratio hold, and file the credibility question as closed. It is the commonest analytical shortcut I see in external control packages, and it misreads, badly, what is being scored. Our earlier checklist on the FDA's external control arm guidance covers the broad design questions; this piece goes one level down, into the specific menu of sensitivity analyses that decides whether your headline number survives contact with a sceptical reviewer.

The framework was always plural

Start with the document everyone defers to and few read closely. ICH E9(R1), the addendum on estimands and sensitivity analysis adopted on 20 November 2019, defines the term in its glossary as "A series of analyses conducted with the intent to explore the robustness of inferences from the main estimator to deviations from its underlying modelling assumptions and limitations in the data." [2] Read that again: a series, plural by construction, in which no single re-run qualifies.

The same addendum goes further at §A.5.1: "Estimation that relies on many or strong assumptions requires more extensive sensitivity analysis. Where the impact of deviations from assumptions cannot be comprehensively investigated through sensitivity analysis, that particular combination of estimand and method of analysis might not be acceptable for decision making." [2] An external control built on real-world data relies on many strong assumptions. By the framework's own logic, that is precisely the setting demanding the most extensive treatment, not the least.

The EMA says the same thing in plainer language. Its Guideline on registry-based studies, adopted 16 September 2021 (EMA/426390/2021), states at §3.8 that on confounding by indication "these do not provide a unique solution and several sensitivity analyses using different approaches should be performed." [3] Several analyses, using different approaches, and pointedly not one generic check.

So why does one-and-done persist? Because no single official document hands you the finished menu. The FDA's February 2023 guidance on externally controlled trials remains a draft as of mid-2026, and even a completed checklist for the broad design of an ECA is not the same as a numbered list of the sensitivity analyses your specific result needs. The menu is assembled rather than issued — synthesised from ICH's general framework and from what real submissions have actually been praised or penalised for. People run a single check and wait for a checklist that was never coming. If you want the context on why single-arm-plus-ECA designs attract this scrutiny in the first place, the EMA's reflection paper on single-arm trials is the companion reading; and the reminder that credibility is a property of process, not of a magic dataset, sits in what "regulatory-grade" RWE actually means.

What a failure looks like

The clearest specification of the menu is a rejection, itemised. Take the Selinexor (Xpovio) case, reproduced by the PSI Real-World Data Special Interest Group from the FDA's own public multidisciplinary review of a Flatiron-EHR-based external control submitted for a real NDA. [4] The agency's objections read as a checklist of the checks that were not run or not survived:

The aligned sample came to 13 eligible patients. Confounding was inadequately accounted for. Selection bias was visible in the numbers, with real-world overall survival of 3.5 to 3.7 months against a trial that had excluded patients with a life expectancy under four months. Immortal time bias followed from differing definitions of time zero across the arms. Twenty-seven of 64 Flatiron patients required exclusion for misclassification. Thirty-one per cent were missing ECOG data. And, fatally, the analysis was not demonstrably pre-specified.

On that last point the review is worth quoting as reproduced by the PSI group: "Without having reviewed and consented to a protocol and SAP, FDA cannot be certain that the protocol and SAP were pre-specified and unchanged... This uncertainty... could lead to overly optimistic conclusions." [4]

Read it again and it stops looking like bad luck. It reads as the menu, written out as a post-mortem. Every objection maps to a sensitivity analysis the sponsor could have specified and reported in advance. The reviewer built the menu anyway. They just built it after the fact, and named every gap while doing so.

Free download

The External Control Arm Design Checklist

The 12 design points regulators probe first, in one checklist you can run against your protocol before database lock.

Get the checklist →

What a pass looks like

Now the mirror image. Esnault and colleagues published a matching-adjusted indirect comparison of entrectinib against French HTA comparators in ROS1-positive NSCLC, and it shows the shape of a credible package. [5] They ran a named series: an unadjusted comparison, an alternative progression-free survival definition, an alternative comparator definition, a complete-case analysis, and a quantitative bias tipping-point analysis. Five angles, each reported, the conclusion holding across all of them. The tipping-point work returned an E-value of 2.673, and the authors could state plainly that "no tipping point can be reached." [5]

Notice the shape of the pass, though. The effect holds at the same size as the headline number across every angle the analysts named before they looked; nothing inflates, nothing shrinks. That is the standard being applied: the point estimate is the claim, and the series is the evidence that the claim is not an artefact of one convenient assumption. It's the same bar any RWE package brought to a regulator has to clear.

The menu itself: sensitivity analyses an external control arm needs

Strip it to the families a reviewer actually expects, and only the ones the evidence supports:

  • Alternative model, matching and comparator specifications: this is the core of the EMA's "different approaches." Vary the propensity model, the matching or weighting approach, and where a study rests on an indirect comparison, the comparator definition itself, as Esnault did with both the comparator and the PFS definition. Among NICE single technology appraisals that paired a single-arm trial with a real-world external control, Leahy and colleagues found 17 of 23 (roughly 74%) ran "broader sensitivity analyses... for example, sensitivity to different model specifications." [10] That is the NICE HTA population specifically, not the field at large. Reassuring: the direction and size of the effect stay stable as you change the specification. Concerning: the estimate drifts with each modelling choice, which tells the reviewer the result is a property of your model, not your drug.
  • Missing-data handling alternatives: in Liu's audit of the broad literature, only 7.2% of the 180 trials specified how missing data would be handled at all. [1] Selinexor's 31% missing ECOG was one of the review's cited weaknesses. [4] The discipline is to show the conclusion does not turn on the imputation assumption: Esnault carried a complete-case analysis alongside the primary; Thorlund and colleagues, comparing pralsetinib against real-world pembrolizumab controls, ran tipping-point hazard ratios of 0.36 to 0.41 across different missing-ECOG assumptions, all staying below 1. [8] Reassuring: the inference is invariant to how you treat what is missing. Concerning: it flips or evaporates under a plausible alternative.
  • Unmeasured-confounding quantification: the flagship, and rare enough to be worth a section of its own. This is the E-value or the quantitative bias tipping-point, and it answers the question a reviewer of any observational comparison is really asking. More below.

One caution on what does not belong on this list as a standard, named check. Alternative time-zero definitions and outlier exclusion are sometimes presented as reflex robustness items; the evidence does not support treating them that way for ECA credibility. Time zero is a design decision, handled up front in target trial emulation (Tran and colleagues set out the practicalities), and its failure mode shows up as the immortal time bias that sank the Selinexor comparison, not as a tidy before-and-after line item. [12] Specify it correctly at the design stage. Do not dress it up as a post-hoc sensitivity analysis.

The flagship, honestly

Quantitative bias analysis deserves the top billing and an honest caveat in the same breath.

What it does is direct: it estimates how strong an unmeasured confounder would have to be, in association with both treatment and outcome, to explain away the effect you observed. The E-value, formalised by VanderWeele and Ding in 2017, puts a single interpretable number on that threshold, and Zhang, Stamey and Mathur have argued it belongs as a first-line unmeasured-confounding method rather than a specialist afterthought. [6][7] The applied wins are real: Esnault's 2.673, Thorlund's 3.31 and 3.37 on the risk-ratio scale. [5][8] A confounder strong enough to breach those thresholds is usually one you would already know about.

That said, the caveat here is load-bearing. Gupta and colleagues put QBA through the most rigorous empirical test I have seen: 14 real randomised trials against 15 external control emulations, checking whether QBA actually moved the ECA estimate closer to the trial truth. It did, narrowing the mean log-hazard-ratio gap from 0.139 using measured confounders alone to 0.098 with QBA adjustment. [9] But their own verdict was sober: "though the improvement may be considered small." [9] QBA is a triangulating check that quantifies the threshold reviewers ask about. It is not a single number that ends the argument, and anyone selling it as a magic correction has not read Gupta.

It also remains rare in both populations worth naming: 1.1% in Liu's broad literature, and just one of 23 (about 4%) in Leahy's NICE submissions. [1][10] Which means running it well still sets your package apart from almost everyone else's.

Pre-specify the series, or it reads as dredging

Every point above collapses if the menu is assembled after a reviewer asks for it. ICH E9(R1)'s whole architecture assumes the series is planned against your stated assumptions in advance; the Selinexor review shows what the absence costs, in the agency's own words about being unable to be certain the protocol and SAP were pre-specified. A menu produced reactively looks, from the reviewer's chair, indistinguishable from a search for the specification that gives the nicest answer.

Many years ago I worked on an early-phase rare-disease programme where we were asked, weeks out from a submission, to show that a survival signal from an external comparison held up under a different index-date choice and a different matching model. We had not planned either analysis. So we were reverse-engineering reassurance under time pressure, and however the numbers landed, the exercise carried the whiff of the thing you least want it to: a result gone looking for its own confirmation. That is the position pre-specification exists to keep you out of.

What a fully pre-specified series looks like in a public protocol is easiest to see in two active-comparator drug-safety cohorts registered by London Health Sciences Centre and Lawson: NCT07179978 (lorazepam in chronic kidney disease) and NCT07098078 (domperidone in CKD). [13][14] These are real-world-data-versus-real-world-data safety studies, not single-arm-trial-plus-ECA efficacy submissions, so no regulator has weighed this exact menu in the context we are discussing; treat them strictly as an illustration of the shape. What they register up front is instructive: inverse-probability-of-treatment weighting with standardised-mean-difference balance checks, effect-measure modification, time-period restriction, a hazard-ratio-versus-risk-ratio cross-check, an E-value quantitative bias analysis, and cross-jurisdiction triangulation, with the second study adding a negative-control-exposure falsification test. [13][14] That is a named series, declared before the data. A feasibility-limited series is fine; an undeclared one is not.

The obvious objection

Fine, push back on all of this. For a 13-patient rare-disease external control, a five-item sensitivity menu looks like statistical theatre. You cannot meaningfully layer QBA plus four alternative specifications onto a handful of patients, and any reviewer knows it. Isn't the whole menu a luxury of well-powered oncology datasets?

No, and the small-n case is where it matters most. A point estimate from 13 patients is more fragile to a single assumption, not less, so the survival question weighs heavier, not lighter. The menu is not a fixed liturgy to be recited whole; it adapts to what the data can bear. Hashmi, Rassen and Schneeweiss have shown that the type of data source itself constrains which robustness checks are even feasible, which is an argument for choosing your checks deliberately, not for skipping them. [11] The discipline in a thin dataset is to run what the data support, and to declare transparently what you cannot run and why. And silence is not neutral: in Leahy's review, Evidence Assessment Groups "frequently highlighted" unquantified residual confounding as a source of uncertainty. [10] The gap gets penalised whether or not you name it. Better you name it first. It's the same triage logic behind minimum viable evidence everywhere else in a lean biotech's evidence plan.

The asymmetry

The menu gets run either way, and that is the whole argument compressed into a single line.

Either you specify the series in advance, run what your data support, and present the result as a package that already survived every angle you could name. Or the reviewer runs it for you, retrospectively, uninvited, during the assessment that decides whether your asset moves forward, and writes up each angle you missed as a source of doubt. Same menu, but very different authorship, and very different consequences.

Landscaping which sensitivity analyses a given precedent survived, and which drew objections, is exactly the kind of work worth doing before you lock a protocol rather than after you get the questions. It is part of how we approach external control and RWE strategy at Inovia Bio, and where InovaSight earns its place mapping your intended menu against what comparable submissions have actually been held to. Run the replay yourself, from every angle, while it is still yours to run.

Get the monthly digest

The 5 things evidence leads need to know each month: regulatory moves, RWE developments and what they mean in practice. No pitch, one email a month.

References

[1] Liu et al. (2025). "Design, Conduct, and Analysis of Externally Controlled Trials." JAMA Network Open. PMID: 40906478. https://pubmed.ncbi.nlm.nih.gov/40906478/

[2] ICH E9(R1) — Addendum on Estimands and Sensitivity Analysis in Clinical Trials, to the Guideline on Statistical Principles for Clinical Trials. FINAL, adopted 20 November 2019. International Council for Harmonisation.

[3] European Medicines Agency. Guideline on registry-based studies. FINAL, adopted 16 September 2021. EMA/426390/2021.

[4] Merrall, Izem & Wolfram (2023), PSI Real-World Data Special Interest Group, reproducing the FDA multidisciplinary review of the Selinexor (Xpovio) externally controlled analysis. Secondary source citing the public FDA review.

[5] Esnault et al. (2025). "Addressing challenges with Matching-Adjusted Indirect Comparisons to demonstrate the comparative effectiveness of entrectinib in metastatic ROS-1 positive Non-Small Cell Lung Cancer." BMC Medical Research Methodology. PMID: 40025432. https://pubmed.ncbi.nlm.nih.gov/40025432/

[6] VanderWeele & Ding (2017). "Sensitivity Analysis in Observational Research: Introducing the E-Value." Annals of Internal Medicine. PMID: 28693043. https://pubmed.ncbi.nlm.nih.gov/28693043/

[7] Zhang, Stamey & Mathur (2020). "Assessing the impact of unmeasured confounders for credible and reliable real-world evidence." Pharmacoepidemiology and Drug Safety. PMID: 32929830. https://pubmed.ncbi.nlm.nih.gov/32929830/

[8] Thorlund et al. (2024). "Quantitative bias analysis for external control arms using real-world data in clinical trials: a primer for clinical researchers." Journal of Comparative Effectiveness Research. PMID: 38205741. https://pubmed.ncbi.nlm.nih.gov/38205741/

[9] Gupta et al. (2025). "Quantitative Bias Analysis for Single-Arm Trials With External Control Arms." JAMA Network Open. PMID: 40136297. https://pubmed.ncbi.nlm.nih.gov/40136297/

[10] Leahy et al. (2026). "From Framework to Practice: Evaluating the Uptake of Sensitivity Analyses in The National Institute for Health and Care Excellence Real-World External Control Arm Evidence." Value in Health Regional Issues. PMID: 42249866. https://pubmed.ncbi.nlm.nih.gov/42249866/

[11] Hashmi, Rassen & Schneeweiss (2021). "Single-arm oncology trials and the nature of external controls arms." Journal of Comparative Effectiveness Research. PMID: 34156310. https://pubmed.ncbi.nlm.nih.gov/34156310/

[12] Tran et al. (2026). "Practical elements to consider when emulating a target trial." Journal of Clinical Epidemiology. PMID: 41765347. https://pubmed.ncbi.nlm.nih.gov/41765347/

[13] London Health Sciences Centre / Lawson. "Lorazepam and outcomes in chronic kidney disease (active-comparator safety cohort)." ClinicalTrials.gov: NCT07179978. https://clinicaltrials.gov/study/NCT07179978 (adjacent RWD-vs-RWD safety cohort, cited as an illustration of a pre-specified sensitivity menu only — not an ECA efficacy study.)

[14] London Health Sciences Centre / Lawson. "Domperidone and outcomes in chronic kidney disease (active-comparator safety cohort)." ClinicalTrials.gov: NCT07098078. https://clinicaltrials.gov/study/NCT07098078 (adjacent RWD-vs-RWD safety cohort, cited as an illustration of a pre-specified sensitivity menu only — not an ECA efficacy study.)

Share this post