Skip to content
Statistical analysis plan pre-specification checklist for external control arm regulatory submissions

The honour-system SAP: pre-specifying an ECA analysis plan you can actually prove

Imi
Imi

You pre-specified, and you did the disciplined thing: you wrote the statistical analysis plan for your external control arm, locked it, and only then went near the external data.

And on its own, that will not save you.

In a 2026 systematic review of 54 oncology target-trial-emulation studies published in The Oncologist, every single one had, in the authors' words, prospectively specified the target trial protocol and emulation procedures. Exactly one had prospectively registered that specification anywhere a third party could check it: one in fifty-four, or 1.9% (Tan et al. 2026, PMID 42411774) [1].

That gap between what teams claim and what they can prove is where the "cherry picking" accusation walks in. It does not require a reviewer to show you cheated, only to notice that nothing stopped you. An unregistered plan is, to a sceptic, indistinguishable from one written after the answer was already known, which is exactly the vulnerability the EMA flagged when it warned how open single-arm trials are to post-hoc rationalisation. A locked SAP is your defence against that charge, but only if the lock leaves a mark someone else can read.

Why an ECA is exposed in a way a blinded trial never is

This problem belongs to external control arms specifically, for a structural reason most teams walk straight past. ICH E9(R1), the estimands addendum finalised at Step 5 in the EU on 30 July 2020, anchors pre-specification to a single event: the analysis is fixed prior to unblinding [2]. In a blinded trial that anchoring does real work, because the statistician physically cannot see which arm did better until the SAP is locked and the database is unblinded. That turns "we locked it first" into a checkable fact rather than a promise, because the blind is the witness.

An ECA has no unblinding event. None. The external data, whether a registry, a historical trial or a claims extract, is already sitting on a server the sponsor can reach long before the SAP is finalised, so you have run your single-arm trial, you know how your patients did, and by the time regulators arrive with pointed questions the comparator data has been in the room for months. The one mechanism that makes a blinded trial's pre-specification self-proving is simply absent. The wall a blinded trial builds for you, an ECA makes you build yourself, and that structural gap, rather than any failing of ECA teams, is why "I locked it early" has to become something a stranger can verify.

The SAP as a checklist of its own

Pre-specifying the analysis means writing a document, and that document has required contents of its own. The most useful template I have seen is HARPER (Wang et al. 2023), which builds a per-parameter field into the plan, so each study parameter carries its own flag for whether it was pre-specified [3]. The document then polices itself, telling a reviewer line by line what was locked rather than asking them to take it on trust.

What has to be in there before you go near the data:

  • Population, cohort and index date. The eligibility algorithm and the exact rule for time zero, fixed before any patient is selected. bluebird bio's ALD-103 shows the standard to aim at: an index rule, transplants performed "on or after January 1, 2013", that sits in the public registry record rather than an internal memo [4].
  • The covariate and effect-modifier list. Which prognostic variables you adjust for, and which effect modifiers the analysis conditions on, named explicitly rather than left as "covariates as appropriate", a blank cheque a sceptic reads as "whichever ones gave the cleanest result".
  • The estimand. State it in full, in the protocol, where it can actually be checked (Lynggaard et al. 2022) [5]. E9(R1) is the anchor here too: the estimand is a pre-specification commitment rather than a methods-section afterthought.
  • Primary and pre-specified sensitivity analyses, including a quantitative bias analysis for unmeasured confounding, specified in advance (Gupta et al. 2025) [6]. This has teeth: the FDA denied an Expanded Approval for ezetimibe and simvastatin on the strength of a tipping-point QBA (Thorlund et al. 2024, J Comp Eff Res, PMID 38205741) [7], so the bias-quantification method you lock matters as much as the primary model.
  • Missing-data and protocol-deviation rules, written before you know how much data is missing or which deviations occurred.

The FDA's own draft ECA guidance already tells you to pre-specify the ECA protocol and SAP rather than pick a comparator after your single-arm trial reads out [8], and our earlier checklist for that guidance walks through it. This is the document-level layer sitting underneath that instruction. On Monday, add a pre-specification column to your SAP template, HARPER-style, so every locked parameter is flagged as locked inside the document itself.

How you prove the lock: turning "trust me" into "check me"

Think of it as an alibi. You can be entirely innocent, having genuinely locked the plan before you saw a single outcome, and still be convicted of cherry-picking, because you gave the sceptic no witness to interview but yourself. Your word is not a timestamp. So build the witnesses in, deliberately, before you touch the data, and four mechanisms do the real work.

  • Dated registration before data collection. The strongest real precedent here is a law rather than a guideline: under Directive 2001/83/EC, Article 107n, a marketing authorisation holder running an imposed post-authorisation safety study shall submit the draft protocol before the study is conducted, which GVP Module VIII confirms means before data collection begins [9]. ENCePP's Code of Conduct (Rev 4, 2018) is the process analogue: the protocol is registered before the study starts, amendments are documented, and post-hoc analyses are confined to generating hypotheses [10]. The HMA-EMA Catalogues of RWD sources and studies have carried this infrastructure since February 2024 [11]. An ECA SAP is not a PASS in law. Even so, PASS shows what a checkable timestamp looks like once a regulator decides to require one — the date burned into the corner of the photograph, not written on the back in your own hand. And the discipline is routinely skipped where nobody enforces it: on EU-PAS, the longest-established registration site for observational studies, 57% of studies were registered without a protocol at all [3].
  • A documented amendments table. HARPER mandates one, recording what is changed, when it is changed, and why [3], and you can watch it work inside the public record, because GENEr8-1's registered exclusion criteria carry the verbatim marker "(effective as of Protocol Amendment 3)" [12]. A reviewer sees that a criterion changed, and which amendment changed it, without asking the sponsor for anything. Nobody objects to an amendment. What kills you is the one nobody wrote down.
  • Authorship separated from data access. As a matter of general good practice, keep the statistician who writes and locks the SAP firewalled from the outcome comparison until the lock is done. There is no ECA-specific rulebook that quantifies this, so treat it as ordinary hygiene borrowed from blinded-trial conduct: the person choosing the analysis should not be looking at the answer while they choose it.
  • Register the external control as its own study, and publish the matching rule. bluebird bio's ALD-103 was registered years before read-out as a purpose-built comparator for the single-arm ALD-102 [4], and EMBOLDEN spelled out its three-step external-control filtering rule directly in its public eligibility criteria [13], so both put the proof in the record instead of an unpublished file.

On Monday, choose your witnesses first: register the protocol on a dated registry, open the amendments table on day one, firewall the SAP author from the outcomes, and where you can, register the external control as its own study. Any one of these beats your word, and together they leave the accusation nothing to grip.

Free download

The External Control Arm Design Checklist

The 12 design points regulators probe first, in one checklist you can run against your protocol before database lock.

Get the checklist →

What it costs when the witness is missing

Some years ago I worked with a team that had done everything right and could prove none of it. They had pre-specified a subgroup, genuinely, in a plan drafted well before the comparator data was pulled, and when an HTA reviewer challenged that subgroup as a post-hoc convenience they had no dated registration, no version-controlled protocol, and no amendments log to reach for. Their only evidence was their own testimony. They were telling the truth — and they lost the argument anyway, because a reviewer cannot audit your intentions, only your records. That is the ordinary cost of an honour-system SAP.

Even study types that ought to be harder to fault get caught. Elevidys is a randomised, placebo-controlled trial, not an external control arm, yet its failure mode is precisely the one an ECA invites: FDA reviewers found that certain secondary-endpoint results were "neither prespecified nor statistically adjusted for the repeated analyses of the data", and that presenting them as evidence of effect was "misleading" (the review language quoted by Public Citizen; the internal dissent reported by STAT News) [14]. If that charge can land inside a blinded, randomised trial, an unregistered ECA SAP is standing in the open.

Registration alone is no magic ward, that said. In a six-year analysis of registered orthopaedic trials, the primary outcome still diverged from the registered protocol in 25% (78 of 309) of trials and the secondary outcome in 60%; the authors conclude that "registration is often treated as a procedural requirement rather than a rigorous commitment to a fixed study protocol" (Poursalehian et al. 2026, Clin Orthop Relat Res, PMID 41910627) [15]. Registration is necessary but not sufficient. And the reason all of this matters sits in a single number: a specification-curve analysis of one observational dataset found a total of 10 quadrillion possible unique analyses of the same data (Wang et al. 2024, J Clin Epidemiol, PMID 38354868) [16]. With that many forks in the road, "we didn't cherry-pick" is unfalsifiable unless you fixed the path in advance and can show precisely when you fixed it.

The honest objection

Here is the strongest case against everything above: this is bureaucratic theatre. A determined bad actor registers a deliberately vague protocol early, leaves himself room, and games the analysis anyway, so the timestamp proves nothing about intent; and even done properly, Jaksa et al. found that agreement in critiques between and among regulators and HTA bodies was low [17], so no paper trail buys you a predictable yes when two reviewers will fault two different things. Both halves are true, and neither is the point.

The paper trail was never going to buy a yes or stop a determined fraud. Its job is narrower and real: to move the argument off your credibility, which a sceptic will always discount, and onto the record, which they cannot. And the honest version has teeth. In the TBASEL emulation of the OAK trial, investigators found the source trial had never defined "protocol violation" and chose not to proceed with the emulation rather than retrofit a definition after the fact (Gupta et al. 2025) [18]. The paper trail is what makes that kind of discipline visible to a reviewer instead of invisible.

What the guidance still doesn't give you

Every document in this space tells you to pre-specify, and none tells you how to prove you did. The FDA's ECA guidance remains a draft, issued February 2023 and still not finalised [8]. The MHRA's ECA and RWD guideline is also still draft, its consultation closed 14 July 2025 with no final text published [19]. Both name the principle, and both leave the proof mechanism to you. The one place that turned pre-specification into a checkable, dated obligation did it by writing it into law, Article 107n [9], and it did so for post-authorisation safety studies, not for ECAs.

So until ECA guidance grows a proof requirement of its own, the witness is yours to build. Build it before the data is in the room, rather than after the questions arrive. That is the work our RWE and regulatory-strategy team does at Inovia Bio: building a SAP that can prove its own lock.

Write it first. Then make it impossible to argue you didn't.

References

[1] Tan et al. (2026). "Target trial emulation studies in oncology: a comprehensive analysis." The Oncologist. PMID: 42411774. doi:10.1093/oncolo/oyag258. https://pubmed.ncbi.nlm.nih.gov/42411774/ — abstract-only; 1.9% (1/54) prospective-registration figure quoted from the abstract.

[2] ICH E9(R1), Addendum on Estimands and Sensitivity Analysis in Clinical Trials. FINAL (Step 5; EU adoption 30 July 2020).

[3] Wang SV, et al. (2023). "HARmonized Protocol Template to Enhance Reproducibility of hypothesis evaluating real-world evidence studies on treatment effects: A good practices report of a joint ISPE/ISPOR task force (HARPER)." Pharmacoepidemiol Drug Saf. PMID: 36215113. https://pubmed.ncbi.nlm.nih.gov/36215113/

[4] bluebird bio. "ALD-102 (Starbeam)" ClinicalTrials.gov: NCT01896102, https://clinicaltrials.gov/study/NCT01896102 ; purpose-built external comparator "ALD-103" ClinicalTrials.gov: NCT02204904, https://clinicaltrials.gov/study/NCT02204904

[5] Lynggaard H, et al. (2022). "Principles and recommendations for incorporating estimands into clinical study protocol templates." Trials. PMID: 35986349. https://pubmed.ncbi.nlm.nih.gov/35986349/

[6] Gupta et al. (2025). "Quantitative Bias Analysis for Single-Arm Trials With External Control Arms." JAMA Netw Open. PMID: 40136297. doi:10.1001/jamanetworkopen.2025.2152. https://pubmed.ncbi.nlm.nih.gov/40136297/

[7] Thorlund et al. (2024). "Quantitative bias analysis for external control arms using real-world data in clinical trials: a primer for clinical researchers." J Comp Eff Res. PMID: 38205741. https://pubmed.ncbi.nlm.nih.gov/38205741/

[8] FDA. "Considerations for the Design and Conduct of Externally Controlled Trials for Drug and Biological Products." DRAFT guidance (February 2023; docket FDA-2022-D-2983; no finalisation identified as of this writing). https://www.fda.gov/regulatory-information/search-fda-guidance-documents/considerations-design-and-conduct-externally-controlled-trials-drug-and-biological-products

[9] Directive 2001/83/EC, Article 107n. FINAL, in force. Imposed non-interventional post-authorisation safety study: before the study is conducted, the marketing authorisation holder shall submit a draft protocol to the PRAC (or the national competent authority for a single-Member-State study); GVP Module VIII confirms protocol submission before data collection begins. https://eur-lex.europa.eu/eli/dir/2001/83/oj

[10] ENCePP Code of Conduct, Revision 4 (2018). FINAL.

[11] HMA-EMA. Catalogues of real-world data sources and studies (operative since February 2024).

[12] BioMarin. "GENEr8-1 (BMN 270-301)." ClinicalTrials.gov: NCT03370913. https://clinicaltrials.gov/study/NCT03370913 — registered exclusion criterion marked "(effective as of Protocol Amendment 3)".

[13] "EMBOLDEN." ClinicalTrials.gov: NCT03771898. https://clinicaltrials.gov/study/NCT03771898 — three-step external-control filtering rule stated in eligibility criteria.

[14] Public Citizen (26 February 2025) quoting FDA review language on Elevidys secondary endpoints; STAT News (2 July 2024) on the internal dissent. Elevidys/EMBARK is a randomised, placebo-controlled trial, not an ECA — cited as the same failure mode in a different study type.

[15] Poursalehian et al. (2026). "Do Published Orthopaedic RCTs Match Their Registered Protocols? A 6-year Analysis of Leading Orthopaedic Journals." Clin Orthop Relat Res. PMID: 41910627. doi:10.1097/CORR.0000000000003910. https://pubmed.ncbi.nlm.nih.gov/41910627/ — abstract-only; 25% (78/309) primary and 60% secondary discrepancy figures quoted from the abstract.

[16] Wang et al. (2024). "Grilling the data: application of specification curve analysis to red meat and all-cause mortality." J Clin Epidemiol. PMID: 38354868. https://pubmed.ncbi.nlm.nih.gov/38354868/

[17] Jaksa et al. (2022). "A Comparison of Seven Oncology External Control Arm Case Studies: Critiques From Regulatory and Health Technology Assessment Agencies." Value in Health. PMID: 35760714. doi:10.1016/j.jval.2022.05.016. https://pubmed.ncbi.nlm.nih.gov/35760714/ — abstract-only.

[18] Gupta et al. (2025). "Estimating per-protocol effects in external comparator analyses using real-world data." J Comp Eff Res. PMID: 41047965. https://pubmed.ncbi.nlm.nih.gov/41047965/

[19] MHRA. Draft guideline on external control arms / real-world data (public consultation closed 14 July 2025; no final text published). DRAFT.

Share this post