Target trial emulation for non-statisticians: why the protocol wins the regulator, not the maths
Say "target trial emulation" in a room of biotech founders and watch the shoulders drop. The phrase arrives wrapped in machinery: cloning, censoring, inverse-probability weighting, g-methods, propensity scores stacked three deep. It reads like a discipline you have to hire your way into. And so the clinical or regulatory lead who commissioned the real-world study in the first place quietly hands the whole thing to a biostatistician on day one and hopes for the best.
That instinct is where the credibility gets lost.
Because the part of a target trial emulation that actually earns a regulator's confidence is not the estimator. It is the protocol you write before anyone opens the dataset. The method has a founding paper (Hernán & Robins, 2016) [1] and it does eventually need real statistical work. But the decisions that make or break the emulation, the ones a reviewer will interrogate first, are design decisions. Clinical judgement. And they belong to you.
What is a target trial emulation, actually?
Let's be clear about what the phrase means, because the confusion is doing damage. A target trial is the randomised controlled trial you wish you could run to answer your question and cannot, because it would be too slow, too expensive, unethical, or already overtaken by the standard of care. So you specify that ideal trial on paper. Then you build your observational analysis to imitate it, component by component, and you stay honest about every place the imitation falls short.
Specifying it means writing down seven things before the data is touched:
- Eligibility criteria: who would enrol, judged only on information available at the start.
- Treatment strategies: the specific interventions being compared, including any grace period allowed to start treatment.
- Assignment procedure: how patients are allocated to strategies, and what you will do to stand in for the randomisation you do not have.
- Time zero: the single moment at which eligibility is met, a strategy is assigned, and the follow-up clock starts. Aligned across every arm.
- Outcome: defined and ascertained identically in each arm.
- Causal contrast: intention-to-treat analogue or per-protocol, chosen and stated up front.
- Analysis plan: and this is the one item you hand to the statistician.
Six of those seven are clinical and regulatory judgements — you do not need a g-formula to decide who your trial would have enrolled or when its clock should start. You need to know the disease.
Here is the trap I see teams fall into. Call it the "filter fallacy": the belief that if you take a database and filter it down to the trial's inclusion criteria, you have emulated the trial. You have not. Filtering copies the eligibility list and silently skips time zero and treatment assignment, which is precisely where emulations die. The TARGET Statement, a 21-item reporting checklist published in 2025 by an international panel of clinicians, statisticians and trialists, exists largely to drag this discipline into the open. It asks authors to identify the study as an observational emulation of a target trial and to clearly specify the target trial protocol [2]. The estimator is not what the checklist is anxious about. The protocol is.
Why does the protocol earn the regulator's confidence?
Think about the position a reviewer is in. An ordinary observational analysis lands on their desk saying, in effect, trust me, I adjusted for everything. There is nothing specific to argue with, so they argue with all of it. Every unmeasured confounder becomes a live objection, because you have handed them no structure against which to test one.
A pre-specified protocol changes that encounter completely. Now the reviewer is holding a witness statement, and a witness statement gets cross-examined clause by clause: whether the time zero is aligned, whether the grace period was defensible, whether the outcome was ascertained the same way in both arms. Those are specific, answerable questions, and specific objections are the ones you can win. Blanket suspicion is the one you cannot.
The confidence payoff is not hypothetical. In 2021, tacrolimus became the first drug approved by the FDA on the strength of real-world evidence; as reported by Hernandez and colleagues (2025), the FDA characterised the underlying study as showing how "a well-designed, non-interventional study relying on fit-for-purpose real-world data (RWD), when compared with a suitable control, can be considered adequate and well-controlled under FDA regulations" [3]. Read that phrase again. Adequate and well-controlled is the language of a registrational trial.
The benchmarking work points the same way. RCT DUPLICATE (Franklin et al., 2021) ran a pre-registered process built explicitly to mimic a regulatory submission, emulating a set of completed trials with real-world data. Regulatory agreement was reached for 6 of 10 emulations; estimate agreement for 8 of 10 [4]. Not a perfect record, and the authors did not pretend otherwise, but a striking one for a method its detractors still call speculative.
Reviewers are writing it down too. The UK's NICE, a health-technology-assessment body rather than a licensing regulator, tells sponsors in its Real-World Evidence Framework (ECD9, final, June 2022) to "design studies to emulate the preferred randomised controlled trial (target trial approach)" and to "avoid time-related biases due to differences between patient eligibility criteria being met, treatment assignment, and start of follow up" [5]. That last sentence is immortal time bias, named without the jargon. On the licensing side, the FDA Sentinel Innovation Center's own PRINCIPLED process guide (Desai et al., 2024), a methods paper rather than formal guidance but FDA-sponsored and FDA-branded, orders the work identically: step one is "specification of the target trial protocol," step two is "describing the emulation of each component" [6]. And the EMA is building capability rather than merely endorsing it, having commissioned the TARGET-EU catalogue study in September 2024; that work is ongoing and has concluded nothing yet, but the direction of travel is not ambiguous [7]. This is the same reviewability thread that runs through our earlier piece on what makes RWE regulatory-grade: a reviewer can interrogate a process. A dataset can't defend itself.
Free download
The RWE Briefing Document Template
The section-by-section structure for the RWE part of a regulatory briefing, built around the questions reviewers actually ask.
Get the template →Skip the protocol and no estimator saves you
Now the same argument from the wreckage. When emulations fail peer or regulatory scrutiny, they overwhelmingly fail on design, not maths, and almost always at time zero.
What is immortal time bias?
The failure has a name: immortal time bias. It is a race where you only start the stopwatch for the runners who have already reached the halfway mark. They cannot lose, because the clock ignores everyone who dropped out before the line. Misalign your time zero and you build exactly that race into your data, handing the treatment arm a stretch of guaranteed survival it never earned.
The numbers this produces are not subtle. Hansford and colleagues (2026) re-examined a set of published claims. A naive analysis of statins and lung cancer showed an odds ratio of 0.23, a headline 77% apparent reduction in risk; a proper sequential emulation put the hazard ratio at 1.02.
The effect had never existed.
The same paper corrected an ovarian-cancer result from a hazard ratio of 0.56 to 1.12, and a colon-cancer result from 3.33 to 0.96, the corrected figure landing squarely on the IDEA trial's own 0.96 [8]. In advanced cancer, Garcia-Albeniz and colleagues (2026) watched an apparent six-month risk ratio of 0.13 dissolve to 0.95 and 1.04 once the target trial was specified honestly [9].
Years ago I sat with a team who had spent the better part of three months polishing the weighting on an emulation. Elegant model. When we finally drew the follow-up diagram together, nobody in the room could say when follow-up was meant to start. The clever part had been perfected on top of a broken foundation.
Here is the contrast that should stay with you. Ren and colleagues (2026) audited 237 emulations in top-quartile journals. Methodologists were involved in 81.4% of them. Yet only 56.5% pre-specified a protocol, only 16.9% supplied a follow-up diagram defining time zero, and only 30.8% addressed unmeasured confounding [10]. The statisticians were in the room. The design still failed. Better maths was never going to fix a problem that sits upstream of the maths.
Time zero is where most of them die.
The strongest objection, and the honest limit
Let me give the other side its best shot. The design is the easy part, the argument runs; the hard, confidence-determining work is the estimation. Unmeasured confounding lives in the data, not in the protocol, and no amount of careful eligibility-writing tells you whether your weighting model actually closed the backdoor between treatment and outcome. That is statistician territory, and pretending a clinician can own it is how you end up with a beautifully specified study that is quietly wrong.
Every word of that is true, and it does not move me off the thesis. The estimation step is real, it needs a statistician, and it needs its own defences: cloning-censoring-weighting to handle time-zero-aligned strategies (Hernán, 2018) [11], and pre-registered negative-control outcomes to probe for residual confounding (Levintow et al., 2023) [12]. The Uppsala type-2 MI emulation (NCT06736353) does exactly this, registering two negative controls, readmission for bacterial pneumonia and readmission for hip fracture, explicitly to check for residual confounding [13]. Still, the audits answer the objection. Failures cluster on design and time zero even with methodologists present. The protocol is necessary first and it is the piece you own; the estimator is necessary second and the piece you commission. One comes first, the other comes second, and neither replaces the other.
That said, the protocol does not buy you everything. Chesang and colleagues (2025) reported a careful, protocol-aware emulation of the PR07 prostate-cancer trial that still landed on a hazard ratio of 0.48 against the trial's own 0.77, partly because of how time zero and the grace period had to be defined [14]. A faithful emulation is a defensible, reviewable attempt at the truth not a guarantee of the exact number. Anyone selling you the second thing is selling you something.
The non-statistician's pre-flight checklist
Before you brief a biostatistician, before you spend a penny on estimation, answer these on paper. If you cannot, the study is not ready.
- Eligibility, applied at time zero: who is in, using only what is knowable at the start? Eligibility confirmed by later events is how immortal time creeps in.
- Treatment strategies: what exactly is being compared, and what grace period do you allow for starting treatment?
- Time zero: can you draw the single moment where eligibility is met, a strategy is assigned, and the clock starts, identically in every arm? Draw the diagram. If the arms start their clocks at different events, stop and fix that before anything else. It is the same time-zero discipline the FDA leans on in its external-control-arm guidance.
- Outcome: is it defined and measured the same way across arms, with no arm advantaged by how the outcome is caught?
- Causal contrast: intention-to-treat analogue or per-protocol? Choose it, state it, justify it now, not after you have seen the results. (TECOS, NCT00790205, registered both contrasts separately; that is the discipline in practice [15].)
- Analysis plan: written and dated before the dataset is opened. This is your hand-off to the statistician.
- Negative-control outcomes: name at least one before you start, a result your treatment could not plausibly cause, as a tripwire for hidden confounding.
One meta-rule sits over all seven: write it down and date it before you open the data. In Ren's audit only 56.5% managed even that. It is the cheapest credibility you will ever buy, and the discipline sits squarely inside the minimum viable evidence a lean programme can actually afford.
The one hour you should not outsource
The checklist will not hand you the trial's exact hazard ratio. Chesang's team did everything right and still missed it. What it hands you is a study a reviewer can cross-examine and, often, believe, which is the only kind of real-world evidence worth submitting. If you have already run a single-arm trial and the regulators are now asking pointed questions, the target trial is the design your external comparison should have been built from in the first place.
The highest-leverage hour in the whole programme is the one your clinical or regulatory lead spends writing that protocol, before the statistician is briefed, before the data is pulled, before a single line of code is run. Skip that hour and no estimator, however elegant, buys it back. It is where most of a regulatory-grade RWE package is won or lost.
Get the monthly digest
The 5 things evidence leads need to know each month: regulatory moves, RWE developments and what they mean in practice. No pitch, one email a month.
References
- Hernán MA, Robins JM. (2016). "Using Big Data to Emulate a Target Trial When a Randomized Trial Is Not Available." Am J Epidemiol;183(8):758-764. PMID: 26994063. https://pubmed.ncbi.nlm.nih.gov/26994063/
- Cashin AG, et al. (2025). "Transparent Reporting of Observational Studies Emulating a Target Trial—The TARGET Statement." JAMA;334(12). PMID: 40899949. https://pubmed.ncbi.nlm.nih.gov/40899949/
- Hernandez LG, et al. (2025). Report naming the 2021 tacrolimus FDA real-world evidence approval and quoting the FDA's "adequate and well-controlled" characterisation. Clin Pharmacol Ther. PMID: 39807817. https://pubmed.ncbi.nlm.nih.gov/39807817/
- Franklin JM, et al. (2021). "Emulating Randomized Clinical Trials With Nonrandomized Real-World Evidence Studies: First Results From the RCT DUPLICATE Initiative." Circulation;143(10):1002-1013. PMID: 33327727. https://pubmed.ncbi.nlm.nih.gov/33327727/
- National Institute for Health and Care Excellence. "NICE real-world evidence framework (ECD9): Methods for real-world studies of comparative effects." Final, first published 23 June 2022. https://www.nice.org.uk/corporate/ecd9/chapter/methods-for-real-world-studies-of-comparative-effects
- Desai RJ, et al. (2024). "Process Guide for Inferential Studies Using Healthcare Data from Routine Clinical Practice to Evaluate Causal Effects of Drugs (PRINCIPLED): Considerations from the FDA Sentinel Innovation Center." BMJ, 12 February 2024. https://www.sentinelinitiative.org/news-events/publications-presentations/process-guide-inferential-studies-using-healthcare-data
- European Medicines Agency / HMA. "Comparative effectiveness and safety studies using the target trial emulation and estimand frameworks (TARGET-EU)." Catalogue study, contract signed 19 September 2024, ongoing. https://catalogues.ema.europa.eu/node/4440/administrative-details
- Hansford HJ, et al. (2026). Three quantified immortal-time-bias corrections (statins/lung cancer; ovarian-cancer chemotherapy duration; colon-cancer adjuvant duration). BMJ Medicine. PMID: 41737145. https://pubmed.ncbi.nlm.nih.gov/41737145/
- Garcia-Albeniz X, et al. (2026). Worked example of immortal-time bias in biologic-therapy timing in advanced cancer. Cancer Medicine. PMID: 42321955. https://pubmed.ncbi.nlm.nih.gov/42321955/
- Ren J, et al. (2026). Quality audit of 237 target trial emulations in top-quartile journals, 2017-2023. JAMA Netw Open. PMID: 41712213. https://pubmed.ncbi.nlm.nih.gov/41712213/
- Hernán MA. (2018). "How to estimate the effect of treatment duration on survival outcomes using observational data." BMJ;360:k182. PMID: 29419381. https://pubmed.ncbi.nlm.nih.gov/29419381/
- Levintow SN, et al. (2023). Rationale for negative control outcomes in real-world causal analyses. Pharmacoepidemiol Drug Saf. PMID: 36965103. https://pubmed.ncbi.nlm.nih.gov/36965103/
- Uppsala University. "Type 2 myocardial infarction target trial emulation study." ClinicalTrials.gov: NCT06736353. https://clinicaltrials.gov/study/NCT06736353
- Chesang J, et al. (2025). Target trial emulation of the PR07 prostate-cancer trial. J Clin Epidemiol. PMID: 40147703. https://pubmed.ncbi.nlm.nih.gov/40147703/
- Merck Sharp & Dohme. "Sitagliptin Cardiovascular Outcome Study (TECOS)." ClinicalTrials.gov: NCT00790205. https://clinicaltrials.gov/study/NCT00790205
-1.png?width=1169&height=277&name=Inovia%20Logo%20Dark%20(1)-1.png)