Ask a team that has just committed to an external control arm what happens next, and you will almost always get the same first question: which database do the comparator patients come from? Claims or EHR? A disease registry, or a chart-abstraction exercise across a handful of academic centres? That call can absorb weeks. Vendor demos, feasibility counts, a licence negotiation that swallows half a quarter.
Here is the uncomfortable part.
The source you agonise over is rarely the thing a reviewer goes after.
Appiah and colleagues examined how NICE and Australia's PBAC responded to real-world comparator evidence, and found that HTA bodies seldom contest the choice of data source at all [1]. The fire lands elsewhere. Arondekar and colleagues, cataloguing the FDA's oncology RWE reviews from 2015 to 2020, found the agency litigating what sponsors had done with their data: whether eligibility matched, how missing data were handled, whether an analysis protocol existed before anyone saw an outcome [2]. Liu and colleagues, auditing 180 externally controlled trials, found the same fingerprints across the methods sections [3].
So the spine of this piece is a reversal. An external control arm's credibility is decided in the two decisions teams treat as clerical: how you curate the raw extract, and how you match it to your trial population. Those two matter far more than the source decision that keeps you up at night.
Three build-stage decisions sit in sequence: which source type, what curation standard, which matching method. This is the build-it layer, and it assumes you have already decided an ECA is the right instrument and already settled your comparator population; the upstream regulatory groundwork lives in our checklist on the FDA's ECA guidance. From here, we spend our attention where the reviewer spends theirs.
So which of the four is the safe default? None of them.
Every real-world source carries a signature weakness, and the FDA's 2024 final guidance on assessing electronic health record and medical claims data [4] names them without diplomacy. It treats neither source as inherently more trustworthy than the other. There is no gold-standard database waiting on a shelf to be licensed.
Take claims. The agency's own language is that "medical claims data can change during the run-off period and claims adjudication process... leading to variations in reported diagnoses and procedures over time" [4]. A diagnosis you saw in March may read differently once adjudication settles. The disciplined answer is a stability window: the asciminib comparator cohort registered under NCT06148493 requires at least six months of continuous pre-index pharmacy and provider coverage before a patient is allowed to count [5].
EHR data fails differently. Collection "is limited to the data captured within an EHR system or network, and may not represent comprehensive care", and the guidance warns it "may not accurately reflect the presence, characteristics, or severity of a particular disease" [4]. A patient who gets their bloods done across town, or dies outside the network, is a patient your extract quietly misreports. That is why the Flatiron-based comparator for nivolumab plus relatlimab in NCT07079644 demands at least two clinic encounters before a record qualifies [6] — a crude but real proxy for "this network actually saw this patient".
Registries feel safer, being disease-specific and clinician-curated. They are not automatically comparable. Hermans and colleagues rebuilt a HOVON-103 comparator from two registries covering the same disease and the same broad geography, and WHO performance status was missing for 66% of patients in one (the HARMONY Alliance database) versus 54% in the other (the Netherlands Cancer Registry) versus 1% among the trial's own controls [7].
Same disease, two registries, not interchangeable.
Chart abstraction sits at the opposite pole: bespoke and clinically granular, but small, and only as good as whoever is reading the chart at each site. The pralsetinib comparator in NCT04697446 is built this way, from records at named academic centres [8].
One case shows what disciplined sourcing can achieve. SCHOLAR-1 (Crump and colleagues, 2017) pooled four sources of genuinely different provenance (two randomised trials and two retrospective academic cohorts), abstracting 861 patients down to a 636-patient refractory DLBCL cohort with an overall response rate of 26% (95% CI 21–31%) and median overall survival of 6.3 months [9]. That mixed-provenance cohort went on to contextualise tisagenlecleucel and axicabtagene ciloleucel. The takeaway is about the handling, not the pooling. Heterogeneous provenance holds up when the curation is rigorous.
Before anyone concludes EHR "replicates" a trial control: it sometimes does. Schröder and colleagues rebuilt the IMblaze370 control arm from EHR data and landed close to the randomised result [10]. Test the same database more widely, though, and the picture greys. Tan and colleagues emulated 15 trials from one EHR source and found only partial fidelity, with real-world outcomes running systematically worse: a combined hazard ratio of 0.76 for overall survival and 0.84 for progression-free survival [11]. One flattering replication is a property of that comparison, not of the database.
The anti-pattern to name and kill is what I'd call "shelf-sourcing": licensing whatever data source is already in the building, or cheapest to access, and reverse-engineering a justification for it against the comparator claim later. Choose the source against the specific claim it has to answer, and write its failure mode into the protocol before you sign anything.
The same drug can legitimately use different sources for different questions. Iptacopan in PNH sits behind both APPEX (NCT05842486), a pre-approval efficacy-contextualisation study on hospital and real-world data [12], and a separate post-authorisation study on the IPIG PNH registry (NCT06903234) [13]. Nobody ran a controlled head-to-head of claims versus registry here. These are two different regulatory questions, each answered with a fit-for-purpose source.
Free download
The External Control Arm Design Checklist
The 12 design points regulators probe first, in one checklist you can run against your protocol before database lock.
Get the checklist →Fair challenge, and it deserves the strongest version. Some source failures cannot be repaired downstream. A claims database that structurally cannot see tumour response will never yield an oncology response endpoint, however you curate it. A registry that never captured a decisive prognostic covariate cannot be matched back into validity; no propensity score conjures a variable that was never recorded. Doesn't that make the source the foundational call after all?
For the disqualifying case, yes, and that is exactly why the first job is to match the source to the claim. But that is a screening step, not the contest. Appiah's finding holds precisely because sponsors rarely pick structurally impossible sources; they pick plausible ones and then under-curate them [1]. The disqualifying source is rare and obvious. The plausible-but-under-curated source is everywhere, and it is where credibility actually leaks away.
Let's be frank about where the meeting goes: it goes to curation.
And curation is not a tidy word for data-cleaning. The FDA defines it precisely: "data curation is the processing of source data through the application of standards for exchange, integration, sharing, and retrieval" [4]. It is a named discipline with expected evidence, which is why our earlier piece on what "regulatory-grade" RWE actually means puts the process, not the dataset, at the centre.
The guidance is specific about what that evidence looks like. It expects data elements that are "well-defined with consistent and known clinical meaning". It expects an "assessment of completeness of data elements including trends over time", not a one-off snapshot. Where you have mapped between coding systems, say SNOMED CT and ICD-10-CM, it expects you to demonstrate the mapping is accurate [4]. None of this is exotic. It is simply what a reviewer asks to see, in writing.
Common data models earn a mention here, and a caveat. The FDA points to its own networks as CDM examples: Sentinel, the Biologics Effectiveness and Safety Initiative, and PCORnet [4]. It does not endorse OMOP or any commercial model, so do not walk into a meeting citing "the FDA's OMOP guidance", because no such thing exists. And a CDM has a cost the guidance states plainly: "data in CDM-driven networks rarely contain all the source information present at the individual health care sites" [4]. Standardisation buys comparability and spends granularity.
The highest-value curation work happens before you have seen a single outcome, in harmonising eligibility to your trial. Two registered comparators show the texture. The odevixibat cohort drawn from the NAPPED registry (NCT07497724) imports a molecular exclusion straight from the interventional protocol, dropping patients with "known pathologic variations of the ABCB11 gene that predict complete absence of the BSEP protein" so the subtype mix matches [14]. The trastuzumab-deruxtecan comparator in ESPERANZA (NCT06973161) writes in a completeness rule aimed squarely at informative censoring: patients "known to have died must have a complete recorded date of death" [15].
Skip that work and the critiques write themselves. Arondekar's catalogue is unsparing: the FDA faulted the entrectinib, selinexor and tazemetostat applications, among others, because a protocol for the RWD analysis "was not submitted prior to the conduct of the study", and it flagged erdafitinib's RWD as "often incomplete", with missing data that "would not allow for the evaluation of comparability" [2]. Liu's audit quantifies how common the gap is: of 180 externally controlled trials, only 13 (7.2%) specified a plan for handling missing data, and 90.0% relied on non-contemporaneous historical data [3].
Some years ago I sat in on a feasibility review for an early-phase rare-disease programme where the registry had already been licensed (signed, paid, sitting in the building) before anyone had written down the comparator claim it was meant to support. The curation team inherited the fallout, spending months trying to harmonise an eligibility definition the data had never been built to capture. No amount of cleaning could manufacture the covariate the protocol needed. The source was rarely the problem in that meeting; what we had done with it, and failed to do before it, was the whole conversation.
Do it properly and it costs you. In a single-centre EHR pilot building an external control arm for ustekinumab in Crohn's disease, Rudrapatna and colleagues excluded 75% of 736 candidate records during curation, and 30% of those that survived the filter still had missing baseline data [16]. That is the honest trade, and for a resource-lean team it is the minimum-viable-evidence calculation in miniature: rigorous curation shrinks your cohort and surfaces what is missing rather than papering over it. And curation quality maps straight onto bias. Carrigan and colleagues showed that incomplete death capture skewed median overall survival by as little as 0.6 to 0.9 months when mortality sensitivity was high, around 91%, and by as much as 3.3 to 9.7 months when it was not [17]. The number a reviewer cares about moves by months on the strength of how well you captured deaths.
So pre-specify the curation plan (completeness definitions, coding-system mappings, eligibility harmonised to the trial) and lock it before a single outcome is in view. Pre-specification is the whole game. It runs through how you take an RWE package to regulators in the first place.
The method question comes last, and it is more constrained than most biostatisticians pretend. Both regulators deliberately fence it off from their data-quality guidance. The FDA states that "methods to achieve covariate balance are out of scope of this guidance" [4]; the EMA draws the identical line in its Data Quality Framework (finalised in 2023), excluding the "analytical methods to derive evidence" from a framework about the underlying data [18]. Read that correctly: the agencies set a comparability bar and let you choose the instrument. They will not bless a method for you. That EMA line sits downstream of an earlier fight over whether a single-arm design was defensible in the first place: ground we cover in our lessons from the EMA's reflection paper on single-arm trials.
Deterministic or probabilistic linkage: which do you use, and when?
The FDA does offer a plain on-ramp for linkage. Deterministic linkage "uses records that have an exact match to a unique or set of common identifiers"; probabilistic linkage "uses less restrictive" rules, and where you use it, "the analysis plan should include testing the impact of the degree of match" [4]. Exact when you hold the identifiers; probabilistic, with sensitivity testing, when you do not.
Beyond linkage, your data condition picks the method more than your statistical taste does. A study with the identifiers and the covariates can name its approach cleanly, as the belimumab lupus comparator in NCT07056621 does, using propensity-score matching and/or inverse-probability-of-treatment weighting [19]. Methods are not mutually exclusive: Oesterheld and colleagues matched an eflornithine neuroblastoma comparator on eleven covariates by propensity score while forcing an exact match on the one covariate you cannot fudge, MYCN amplification [20]. And when covariate overlap is poor, the method bills you for it in sample size. A MAIC of an axicabtagene-ciloleucel cohort saw its effective sample size collapse to 15.3 from 45 real patients once the weights were applied [21]. That is not a rounding error. That is most of your statistical power, gone to buy balance.
That said, one honest gap is worth stating: what performs best is not what gets used. Loiseau and colleagues found that outcome-prediction methods such as G-computation and doubly debiased machine learning can cut estimation error relative to propensity-score approaches in simulation [22]. Yet Liu's real-world audit found propensity methods still dominating practice by a wide margin, weighting in 57.1% of them and matching in 22.9%, and 43.8% of the matched studies that reported a balance assessment still carried unbalanced covariates [3]. The frontier and the field are not in the same place. The deep statistical comparison of these methods is its own treatment; here the operational point is enough. Overlap, sample size and identifier availability force the method long before preference does.
Play it back. You will spend weeks choosing a source, and a reviewer will barely blink at it. You will treat curation and matching as implementation detail, and a reviewer will spend the entire meeting there. That inversion is the thing to design around.
The Monday-morning version is short. Write the comparator claim down first, in a sentence. Choose the source against that claim, not against what is already licensed. Then pre-specify and lock the curation and matching plan before a single outcome is seen. Everything a reviewer will challenge is decided in that order, and most of it is decided after the source is chosen.
This is exactly the source-to-claim fit and pre-specified curation plan we pressure-test with clients at Inovia Bio, and the kind of evidence picture our tools are built to assemble. Get the sequence right and the database on your shelf stops being the argument. What you did to it, and what you wrote down before you touched it, is what a reviewer actually reads.
Get the monthly digest
The 5 things evidence leads need to know each month: regulatory moves, RWE developments and what they mean in practice. No pitch, one email a month.