Regulators never wrote one guidance document on how to generate real-world evidence. They wrote at least six: on externally controlled trials, on natural history studies, on disease registries, on selective safety-data collection under ICH E19, on validating a clinical outcome assessment from the patient's own perspective, and the 2018 framework document that started the split in the first place [1]. Six methodological answers, because six shapes of evidentiary uncertainty do not share one fix.
Say your gap analysis is done. The prioritised list sits in front of you a comparator gap here, a durability question there, a subgroup claim nobody can support yet. Most teams sort that list by convention, not by shape: reach for whichever evidence-generation activity the last programme ran, relabel it as this gap's answer, and cost it. Call it the "tactic reflex." It survives because "we need more RWE" sounds like a plan. Let's be blunt about what that phrase actually commits you to: nothing, until it's said out loud which of six things it means.
This is not the reconnaissance step: the data-landscaping exercise that tells you what evidence already exists belongs earlier, and the case for running an integrated evidence plan at all has been made elsewhere. This is the step after both: gap list in hand, tactic still unchosen. This is the decision framework for that choice.
Before an external control arm gets costed, check the FDA's own precondition: its 2023 draft guidance treats borrowing a past as appropriate specifically where a disease's natural history is well understood and does not improve without intervention: a gate on the tactic, not a menu item picked because a single-arm trial needs a comparator and this is cheapest [1].
Jahanshahi and colleagues' review of 45 FDA approvals (2000-2019) built on external control data shows how sponsors use that gate: retrospective natural history was the most common source (44%), ahead of baseline control (33%) and published or prior-study data (11% each); 80% were rare-disease approvals, 87% used an objective endpoint [2]. None of the 45 used prospectively collected natural-history data, despite the guidance naming prospective collection the preferred source [2]. That is not a coincidence: sponsors pick the class correctly and reach for the cheaper, retrospective version almost every time.
Zolgensma sits inside this family, not the durability family it resembles at a glance. Its May 2019 approval was standard, not accelerated, built on two open-label single-arm studies (n=15 and n=21) read against external natural-history data: "a large treatment effect (90% alive without ventilation versus 25% based on natural history)" [2]. Its comparator was not scavenged after the fact: Al-Zaidy and colleagues describe a prospective natural-history cohort built to serve as the external comparator before the pivotal trial needed one [3]. Engineered ahead of the gap, not assembled to answer it retrospectively.
The tactic for one comparator gap doesn't have to stay fixed. Blinatumomab closed the same gap twice: a single-arm trial paired with a historical remission-rate benchmark for accelerated approval, then a 405-patient randomised trial against investigator's-choice chemotherapy for full approval [4]. Cheap first, rigorous second, same gap.
Choosing the right class still doesn't guarantee it closes what it was chosen for. Only 8.3% of original oncology approvals and 0.8% of supplementals (2015-2020) used real-world evidence for efficacy at all [5], and entrectinib's real-world comparator used electronic health record data unlikely to be generalisable, on the FDA's own assessment, given the low rate of the relevant genomic testing in ordinary practice [5]. Right instrument, wrong population underneath it. Sharper still in the next branch. (FDA's ECA checklist covers the design mechanics this section only gates.)
A durability gap needs matching to the right kind of registry, not just "a registry." Which kind, though? The FDA's final registries guidance splits them into three kinds: disease, health-service and product [6]. The EMA's parallel guideline names four distinct jobs a registry can do: post-approval effectiveness follow-up, external-control-arm creation, post-approval safety collection, and pre-authorisation natural-history characterisation [7].
Some durability gaps don't need a bespoke build at all. The European Cystic Fibrosis Society Patient Registry (EMA-qualified 2018, 54,546 patients), the EBMT Registry (2019, 700,000+ patients) and Enroll-HD (2022, 21,561 Huntington's disease patients) are standing, qualified instruments a sponsor can use rather than commission [8]. The UK's RaDaR registry shows one registry serving two decisions at once: over 37,000 participants across more than 100 sites have supported proteinuria as a surrogate endpoint for regulatory acceptance, and fed NICE appraisals from the same dataset [9].
For a durability question inside a pivotal's own population, the tactic is an extension study: LIMMitless has followed risankizumab patients past five years, BE BRIGHT has tracked bimekizumab through 196 weeks, POETYK PSO's long-term extension carries four years of deucravacitinib data, among others [10].
Zolgensma reappears here only as a second illustration, not its approval story — that belongs above. Its 15-year LT-002 follow-up and 20-year RESTORE registry [11] show what a durability tactic looks like once a one-time gene-transfer product is approved, whatever pathway got it there. Elevidys runs the same pattern on a shorter clock: accelerated approval in June 2023 for ambulatory four- and five-year-olds rested on a micro-dystrophin surrogate, and its 10-year follow-up study, tracking 400 prior recipients toward a planned 2033 completion, exists to catch what a two-year pivotal cannot observe by construction [12].
Free download
The IEP Template Pack
The gap matrix, prioritisation grid and plan-on-a-page we use to build integrated evidence plans. Free to keep.
Get the template pack →Choosing the correct tactic class is necessary. It is not sufficient. Somewhere between the class and the result sits the grain the tactic was actually built to, and that is where Kaftrio's story earns its place as this piece's central case.
During the initial marketing authorisation of ivacaftor/tezacaftor/elexacaftor (Kaftrio), the CHMP asked for Cystic Fibrosis Foundation Patient Registry data to support a benefit claim in a specific genotype subpopulation: patients heterozygous for F508del and a gating mutation, and those heterozygous for F508del and a residual-function mutation. The tactic class was exactly right, a registry for a subgroup-generalisability gap. The first submission was rejected anyway: regulators judged the data insufficient as the sole evidence for that subgroup, citing too little detail on the modulator therapy used, its duration, the specific genotypes covered, and individual patient efficacy [13]. The class was right. The grain was wrong.
A second, better-specified pass succeeded: on the subsequent application extending use to patients aged 12 and older, updated registry data carrying genotype-level clinical-endpoint detail supported the claim [13]. Nothing about the tactic changed between submissions. Only whether the instrument was machined to the tolerance the gap required. A part cut to the wrong tolerance still fits the assembly, still turns, and still fails the moment real load goes through it.
I'd contend the mismatch is rarely diagnosed as a mismatch at all; it gets diagnosed as "the regulator moved the goalposts." The same paper's other named uses argue otherwise: post-authorisation registry studies covered Kymriah's long-term safety and under-3 efficacy; an existing haemophilia registry, EUHASS, was repurposed for Esperoct's PEG-accumulation question rather than a bespoke cohort; EURACAN collects long-term neurologic, hepatic and infection safety data for Vitrakvi [13].
One aside, reported rather than confirmed: commercial pressure can push a sponsor toward the faster tactic before the instrument is built to the gap's actual grain, and Elevidys's 2025 distribution-suspension request from the FDA, confirmed by the agency's own press release, is the case most often cited for that pattern, though the sharper trade-press claims about it aren't something this piece treats as settled [14].
Esperoct's EUHASS study makes the first point: check whether an existing registry already answers the question before building a new cohort. ICH E19 makes the second, harder point. When does a safety-signal gap actually need a full new dataset? Its entire premise runs the mismatch in the other direction: when a drug's common adverse-event profile is already well characterised, collecting the full standard safety dataset again is oversized relative to what is actually still uncertain [15]. Under-engineering is not the only failure mode here. Over-collecting is a mismatch too; it just looks responsible instead of careless.
A patient-experience gap is a validation programme, not a survey question bolted onto an existing endpoint list. PAH-SYMPACT was built because nothing adequately captured the lived symptom burden of pulmonary arterial hypertension; NSCLC-SAQ, LFA-REAL and a dedicated RSV patient-experience measure exist for the same reason in lung cancer, lupus and respiratory infection, with PRISM and the Activity Impairment in Migraine Diary further named instances of the same pattern [16]. The FDA's Patient-Focused Drug Development guidance series sets the standard each had to clear: qualitative concept elicitation first, then quantitative psychometric validation: not a bolt-on question [17].
A sixth gap type exists (does the drug work through the mechanism it claims to) and honesty requires saying the research behind this piece did not surface a clean non-oncology pairing for it; the closest available example is an oncology biomarker-directed umbrella trial [18]. Treat that as a thin branch, not a worked case: forcing one here would be the same mismatch this piece argues against, just committed by the writer instead of the sponsor.
If patients are already being treated outside your trial, structured capture of that exposure is a tactic waiting to happen, not an afterthought. Polak and colleagues found 49 drug-indication pairs approved by the FDA or EMA in part or whole on expanded-access data through 2018, 63% orphan-designated; in the 39 cases where expanded access reached the pivotal efficacy section for rare-disease medicines, 58% of all patients in that section had been treated under expanded access [19].
Ipilimumab is the clearest case in this whole taxonomy of an evidence-generation activity buying a specific decision. Its pivotal trial enrolled only 55 UK patients too few for the UK's reimbursement body to estimate real-world vial usage with confidence. Pooled data from 258 UK patients treated through expanded access, showing average use of 1.19 vials against the pivotal's 1.51, changed the cost-effectiveness calculation used at the reimbursement decision itself [19]. Twenty-one per cent of health technology assessments conducted for the NHS over the last decade have relied in part on expanded-access data [19].
Cholic acid went further: two expanded-access programmes, 63 and 22 patients, were the entire evidentiary basis, and the EMA approved it under exceptional circumstances [19].
Instrument that capture now, before the access programme is asked to double as your evidence base after the fact. Retrofitting it later is the harder version of the same tactic.
Once a team has more gaps than money, exactly the condition a tight-runway IEP already assumes, the six branches above tell you what each tactic looks like once resourced, not which gap gets built first. The CanREValue Collaboration's answer is the most concrete found in this research: a seven-criterion decision tool prioritising candidate real-world studies against uncertainties left over from an actual cancer-drug funding recommendation, validated when an eleven-member CADTH committee used it to sort nine real proposals into low, medium and high priority [20]. The organising variable was never scientific interest. It was which stakeholder's decision the gap was blocking, and how soon.
Timing isn't incidental to that. Murphy and colleagues document why it matters structurally: accelerated approval increasingly arrives with a thinner evidence package than HTA bodies need, so evidence generation timed only against the regulatory submission risks delaying the HTA decision behind it [21]. No source here quantifies what that lateness costs, and none should be invented.
Deloitte's two survey waves, kept to their own years: in 2022, 90% said their organisation tries to apply real-world evidence across the product lifecycle, half actively running integrated evidence planning for at least one asset [22]; by 2025, 96% planned to increase real-world data investment, yet most organisations' evidence planning remained fragmented: different teams and regions working in isolation [23]. Appetite rose. Coordination didn't keep pace.
There is no external benchmark to check your own sequencing against, either. DiMasi and colleagues, valuing an integrated evidence plan's components in the most recent peer-reviewed attempt, had to build their worked examples from hypothesised data, because no field-wide dataset exists on what fraction of prioritised gaps become funded studies [24]. Nobody, including the vendors selling you the one RWE study that will supposedly close whatever's missing, has a number to check your list against.
Pressure-test the gap list against the shape of each gap before it's costed, with the same minimum-viable-evidence discipline that governs any resource-lean biotech's evidence spend. A platform like InovaSight is useful for keeping that pressure-test current as new evidence lands, not for running it the first time. That part is still yours to do by hand.
Get the monthly digest
The 5 things evidence leads need to know each month: regulatory moves, RWE developments and what they mean in practice. No pitch, one email a month.