Skip to content
External control arm causal inference: the exchangeability, positivity and consistency conditions that stand in for randomisation
RWE Real world evidence drug development

Randomisation has exactly one job. An external control arm has to do it by hand.

Imi
Imi

You did everything the guidance asks. You picked a data source that fit the study rather than shoe-horning the study into whatever registry happened to be lying around. You matched on every prognostic factor you could measure. You wrote the analysis down before you looked at a single outcome, which for an external control arm is a small act of discipline all its own, because there is no unblinding event to write it before. It was a clean run, and you took no shortcuts.

And none of that is what makes the comparison mean anything.

That is the uncomfortable part, and it is worth naming plainly, because a well-executed ECA feels causally valid in a way that has almost nothing to do with whether it is. Call it the "clean-run alibi": the quiet assumption that if the mechanics were sound, the causal claim behind them must be sound too. That comfort is misplaced. Execution was always necessary; it was never sufficient.

Here is what randomisation actually buys you, in one line. It does exactly one job: it makes the treated and untreated groups exchangeable by design, casting both from the same mould, so that any difference in outcome afterwards has nowhere left to hide except the treatment itself. An ECA has no coin toss. One act. One guarantee. Gone.

What randomisation bundled into that single act now arrives as an itemised invoice with three separate line items, each payable on its own: exchangeability, positivity and consistency. Banack's 2019 commentary calls them the three fundamental identifiability assumptions for causal inference [1], and the whole causal standing of an external control arm reduces to whether you can settle all three.

Three line items, three ways to fail, no two of the failures alike, and "we controlled for confounding" is an answer to none of them.

Emulating the trial sets the question. It does not answer it.

Target trial emulation is the framework that turns those three abstractions into something you can actually audit. Hernán and Robins framed observational analysis as "an attempt to emulate a randomized experiment... that would answer the question of interest" [2]. Time zero, eligibility, assignment: line them up against a trial you could in principle have run, and the identifiability conditions stop being philosophy and start being checkboxes.

But a filled-in emulation table is not a causal claim. The same authors are explicit that "even in the absence of residual confounding, the emulation of the target trial may fail when researchers deviate from simple principles" [3]. The reverse holds just as hard: you can get every mechanical principle right and still be answering a question the data cannot answer. Mechanical fidelity and identifiability are two different failures. A scoping review of 96 emulation studies found adherence to the framework patchy, with residual confounding flagged again and again [4]. The emulation is the invoice's formatting. The three conditions are whether you can pay it. Keep them separate in your own head, and keep them separate in your submission.

Exchangeability: the one job, and the line that decides everything

Start with something rare: an exchangeability check against known ground truth. INSIGhT was a randomised platform trial in glioblastoma. When its analysts lifted out the real, randomised control arm and dropped external controls in its place, the substituted comparison found no survival benefit for any of three drugs, with hazard ratios of 1.00, 0.93 and 0.88 (Rudra Gupta and colleagues, 2026) [5]. The external controls reproduced the trial's own answer. But the authors are careful about the condition attached: that result holds only given "comprehensive and accurate data on all potential confounders, in the absence of unmeasured confounding" [5]. Exchangeability was demonstrated, not assumed. Most of the time you cannot do that, because you do not have the randomised arm sitting beside you to check against.

So state it plainly. Exchangeability means the comparison groups would have had the same outcome under the same exposure, once measured differences are accounted for. Randomisation delivers this across everything, measured and unmeasured alike; an ECA cannot. As Thomas and colleagues' AMD emulation work puts it, trials can assume exchangeability across all confounders, measured and unmeasured, "due to random treatment assignment," whereas emulated trials are "limited to exchangeability only in measured variables" [6]. That gap, between all confounders and the ones you happened to measure, is the entire exposure of an external control arm.

You can hear a regulator draw that exact line. In its 2024 review of teclistamab, CADTH noted that "Cytogenetic risk was considered an important risk factor; however, it was not included in the main analyses due to a high level of missingness," a missingness of 37.1%, and concluded that "the magnitude and direction of any bias is unknown" [7]. The measured confounders were balanced and accepted. One important confounder went unmeasured, was named, and was left standing as a live threat. That is the measured-versus-unmeasured line in a real HTA body's own words.

That said, unmeasured does not mean exotic. RCT DUPLICATE, which reached regulatory-equivalent conclusions in 6 of 10 emulations, traced its single largest remaining discrepancy to "unmeasured frailty and lower socio-economic status" [8], not to a mis-specified model. Ordinary, clinically obvious variables the data simply did not hold.

So what do you do on Monday? List every confounder that matters. Mark which ones are actually present in your data source and which are not. For the ones that are not, do not write "no residual confounding." Quantify the doubt instead: an E-value tells you how strong an unmeasured confounder would have to be to overturn your result [9], and a tipping-point analysis tells you where it breaks. "Adjusted for" and "accounted for" are not the same sentence, and reviewers know the difference. Those pointed attribution questions a single-arm submission attracts are exchangeability questions in plain English.

Positivity: the condition that can flip the sign of your result

This one can reverse your answer, not merely widen the interval around it. In a stroke registry, tissue plasminogen activator was analysed two ways. Weighted to estimate the average treatment effect across everyone, the odds ratio came out at 10.77 (95% CI 2.47 to 47.04): apparent catastrophic harm. Weighted to estimate the effect in the treated, it was 1.11 (95% CI 0.67 to 1.84): essentially nothing [10]. The reason is positivity. Severe-stroke patients were almost never treated, so for that slice of the population there was no treated counterpart to compare against, and the average-treatment-effect estimate was extrapolating into a region with no data behind it. A separate re-analysis of the same registry two decades later reproduced the qualitative reversal: a risk ratio of 1.70 one way, apparent harm, against 0.82 the other, apparent protection [10]. Same patients. The effect changes sign depending on whether the estimand demands an overlap that does not exist.

Positivity means real overlap: a non-zero probability that a comparator patient could have received the treatment, and vice versa. Where there is no overlap, there is no comparison to make. What you have instead is extrapolation dressed up as an estimate.

The quiet version is far more common than the dramatic one. Take the Friedreich's ataxia natural-history registry FA-COMS (NCT03090789) and read its enrolment span against the eligibility band of the MOXIe trial (NCT02255435). The registry covers a much wider range of ages and disease severity than the trial's narrow entry criteria, which means large parts of the registry population have no trial-eligible counterpart at all. Comparing across that non-overlap is precisely the positivity problem, and it is visible from the two registered records alone, before anyone runs a model.

There is no single correct fix here, and this is exactly where honest practitioners diverge. The stroke analysts switched estimand, from the ATE to the ATT, because "the ATT requires positivity only among the treated group and so can be identified under weaker assumptions" [10]. The AMD group did the opposite: they trimmed the non-overlapping patients and kept an ATE-style target on the narrower population that remained [6]. Both are legitimate. What is not legitimate is weighting across a region where no comparable patient exists and reporting the number as though it meant something.

Monday: look at the propensity-score overlap before you look at the effect estimate. Decide your estimand from who could plausibly have been treated, not from which number you would rather report.

Consistency: are you even comparing the thing you think you are

Some years ago I worked on an early-phase programme in a rare metabolic disorder, and the external control arm was, on paper, immaculate. Matched on age, on baseline severity, on every prognostic marker the registry captured. It passed every balance check we threw at it. And it was close to meaningless, because the "untreated" state it encoded was not one thing. The registry ran more than a decade, supportive care had shifted underneath it more than once, and "no treatment" turned out to mean several different management regimens depending on which year and which centre a patient happened to land in. We had matched the patients beautifully and compared our drug against a moving target. No amount of propensity work fixes that, because the problem was never who sat in the comparator. It was what the comparator was.

That is consistency. VanderWeele's formal statement of the condition [11] comes down to this: the comparator's data has to represent a well-defined version of the specific treatment-versus-alternative contrast the trial is actually asking. It breaks in two ways, and both of them hide behind a clean match.

First, an active treatment masquerading as "no treatment." Read APPEX (NCT05842486) against APPOINT-PNH (NCT04820530): the external comparator there is a group already receiving an active alternative therapy. The contrast the analysis then reports is drug-versus-drug wearing the label drug-versus-control.

Second, a comparator that is secretly several comparators. Read the pralsetinib external-control record (NCT04697446) against ARROW (NCT03037385) and the group entered as "best available therapy" resolves, on inspection of the eligibility text, into at least four mechanistically distinct regimen classes. One label, four counterfactuals. The health-technology record echoes the same worry: the pralsetinib comparison was flagged for "differences in distribution and dosage of subsequent therapies vs Canadian practice" [12].

RCT DUPLICATE hit the same wall from the other side. Its authors found that "close emulation of placebo is impossible via RWD, and selection of an active placebo proxy may fundamentally change the study question" [8]: the sulfonylurea-as-placebo-proxy problem, a consistency failure sitting quietly inside what looked like a confounding fix.

Monday: before you defend the effect, write down in one sentence exactly what the comparator condition is. If you cannot, if it is "usual care" or "best available therapy" spanning several regimens across a decade, you have a consistency problem that no matching pass will remove.

Free download

The External Control Arm Design Checklist

The 12 design points regulators probe first, in one checklist you can run against your protocol before database lock.

Get the checklist →

Why naming beats lumping

Look back across those cases. Neurodegeneration, rare metabolic disease, haematology, oncology: unrelated conditions, different sponsors, and the same three failures recurring across all of them. Ultra-rare gene therapy does not have a monopoly on this. It is routine, and getting more so. Between 2011 and 2019, Patel and colleagues found that 52% of single-arm HTA submissions, 226 of 433, carried some form of external control data, with submission volume rising thirteen-fold over the period [13]. The question is now everywhere, which is exactly why sloppy vocabulary, the same confusion that swirls around "regulatory-grade RWE", is expensive.

Let's be frank: "we controlled for confounding" collapses three different failures into one phrase, and in doing so it throws away the repair path. Each condition repairs differently:

  • Positivity failure: change the estimand, or trim to the region of overlap.
  • Consistency failure: redefine the contrast, or abandon it. Matching cannot touch it.
  • Exchangeability failure: quantify the residual doubt and show it does not overturn you, or concede it.

Diagnosis dictates treatment. "Bias" diagnoses nothing. That is the entire case for three ungainly words over one comfortable one.

A caution against the opposite overreach, though. Three conditions satisfied is still not randomisation reclaimed. Meeting them is the strongest argument you have, and that is all it is. As Layton and colleagues put it, "terminology should reflect the observational nature of EC studies rather than suggest an equivalence to RCT control arms" [14]. You are building a defensible case, not reproducing a coin toss.

The obvious objection

Put the strongest version of the counter-case. This is epistemology for its own sake, and the reviewers do not even use the words. The FDA-facing literature on external controls frames the problem in terms of bias, propensity and covariate adjustment, not the formal vocabulary of exchangeability, positivity and consistency [15]. NICE's 2022 real-world evidence framework names a "target trial approach" and an observed-versus-unobserved confounding distinction, but not the formal triad [16]. The EMA's 2021 registry guideline covers how to run a registry, not the causal logic of comparing a registry cohort to a trial arm [17]. So why saddle a lean team with three Latinate conditions the people judging the dossier will never name back to them?

Because they ask all three anyway, in plain English, on every single-arm package. Exchangeability is just "what would have happened to these patients without the drug." Positivity asks whether these patients were even representative of the ones who could have been treated. And consistency is the blunter question: what, exactly, are you comparing against? The EMA's own reflection paper turns on the counterfactual, and single-arm reviewers press these questions relentlessly. They already ask them; the triad simply names the grammar sitting underneath. And naming it is what converts "we're worried about bias" into a specific position with a known remedy. The vocabulary gap is theirs. The diagnostic clarity is yours.

What is still missing

Two things, honestly, and neither is close to solved. There is no shared regulatory vocabulary for the three conditions, so every sponsor re-derives the argument from scratch and every reviewer receives it in a slightly different dialect. And there is no field-wide evidence on which of the three fails most often, so nobody can tell you where to spend your defensive effort first. Until both exist, the burden sits entirely with the sponsor.

Which leaves one operational demand, and the whole post reduces to it. For any external control arm you intend to defend, say out loud which of the three conditions you are standing on, and which one you are least sure of. Surfacing that answer in a data-landscaping exercise before you build the study, rather than after a reviewer forces it out of you, is the cheap version of this lesson. The expensive version arrives by post — on agency letterhead.

Because "we controlled for confounding" is not an answer to any of the three.

References

[1] Banack HR. (2019). "You Can't Drive a Car With Only Three Wheels." American Journal of Epidemiology;188(9). PMID: 31107525. https://pubmed.ncbi.nlm.nih.gov/31107525/

[2] Hernán MA, Robins JM. (2016). "Using Big Data to Emulate a Target Trial When a Randomized Trial Is Not Available." American Journal of Epidemiology;183(8):758–764. PMID: 26994063. https://pubmed.ncbi.nlm.nih.gov/26994063/

[3] Hernán MA, Sauer BC, Hernández-Díaz S, et al. (2016). "Specifying a target trial prevents immortal time bias and other self-inflicted injuries in observational analyses." Journal of Clinical Epidemiology;79:70–75. PMID: 27237061. https://pubmed.ncbi.nlm.nih.gov/27237061/

[4] Zuo H, et al. (2023). Scoping review of target-trial-emulation studies. Journal of Clinical Epidemiology. PMID: 37562726. https://pubmed.ncbi.nlm.nih.gov/37562726/

[5] Rudra Gupta T, et al. (2026). External-control reanalysis of the INSIGhT randomised platform trial. Journal of Clinical Oncology. PMID: 42172567. https://pubmed.ncbi.nlm.nih.gov/42172567/

[6] Thomas DS, et al. (2021). Target-trial emulation in age-related macular degeneration. Clinical and Translational Science. PMID: 33421321. https://pubmed.ncbi.nlm.nih.gov/33421321/

[7] CADTH / CDA-AMC. "Teclistamab Clinical Review Report" (PC0332), September 2024.

[8] Franklin JM, et al. (2021). "Emulating Randomized Clinical Trials With Nonrandomized Real-World Evidence Studies: First Results From the RCT DUPLICATE Initiative." Circulation;143(10):1002–1013. PMID: 33327727. https://pubmed.ncbi.nlm.nih.gov/33327727/

[9] VanderWeele TJ, Ding P. (2017). "Sensitivity Analysis in Observational Research: Introducing the E-Value." Annals of Internal Medicine;167(4):268–274. PMID: 28693043. https://pubmed.ncbi.nlm.nih.gov/28693043/

[10] Wiener C, et al. (2026). Positivity, estimand choice and effect-sign reversal in a stroke registry. Epidemiology. PMID: 41086382. https://pubmed.ncbi.nlm.nih.gov/41086382/

[11] VanderWeele TJ. (2009). "Concerning the consistency assumption in causal inference." Epidemiology;20(6):880–883. PMID: 19829187. https://pubmed.ncbi.nlm.nih.gov/19829187/

[12] Gupta A, et al. (2025). "Transportability of nonlocal real-world evidence and its relevance to health technology assessment: a primer." Journal of Comparative Effectiveness Research. DOI: 10.57264/cer-2025-0041. https://doi.org/10.57264/cer-2025-0041

[13] Patel D, et al. (2021). "Use of External Comparators for Health Technology Assessment Submissions Based on Single-Arm Trials." Value in Health;24(8):1118–1125. PMID: 34372977. https://pubmed.ncbi.nlm.nih.gov/34372977/

[14] Layton D, Hester L, Golozar A. (2025). "Editorial: External control arms for single-arm studies: methodological considerations and applications." Frontiers in Drug Safety and Regulation. DOI: 10.3389/fdsfr.2025.1579171. https://doi.org/10.3389/fdsfr.2025.1579171

[15] Jahanshahi M, et al. (2021). "The Use of External Controls in FDA Regulatory Decision Making." Therapeutic Innovation & Regulatory Science;55(5):1019–1035. DOI: 10.1007/s43441-021-00302-y. https://doi.org/10.1007/s43441-021-00302-y

[16] NICE. "NICE real-world evidence framework" (ECD9), 2022. Final.

[17] EMA. "Guideline on registry-based studies" (EMA/426390/2021). Final, adopted 26 October 2021.

Trial records referenced (ClinicalTrials.gov): FA-COMS NCT03090789; MOXIe NCT02255435; APPEX NCT05842486; APPOINT-PNH NCT04820530; pralsetinib external control NCT04697446; ARROW NCT03037385.

Share this post