Skip to content
Illustration of a hybrid control arm design: a randomised trial borrowing external control data to shrink its concurrent control arm
single arm trials Methodology

Shrink the arm, keep the randomisation: the borrowing discipline that decides a hybrid trial

Imi
Imi

The pitch is genuinely attractive. You are running a randomised trial in a population that is too small, or recruits too slowly, to power a conventional two-arm RCT, so you keep the randomised control arm rather than abandoning randomisation the way a pure single-arm-plus-external-control design does. Instead you borrow statistical strength from historical or external control data and randomise fewer patients to control: call it a hybrid control arm design, a smaller trial on the same randomised backbone. What is not to like?

Here is the part the pitch skips over. Whether that trial is credible has almost nothing to do with the decision to borrow, and almost everything to do with how much weight the external data gets and what your design does when that external data disagrees with your own concurrently randomised patients. That second question is the design. Most teams treat it as a footnote.

When we wrote up the FDA's externally-controlled-trials guidance in our ECA checklist for drug developers [1], we flagged that the agency had left one thing conspicuously off the table: it "would have been nice to see the FDA thinking on hybrid control arms ... however this was determined to be out of scope." That gap has not closed, but it is finally being filled from another direction. ICH E20, the draft guideline on adaptive designs (Step 2b, endorsed 25 June 2025), carries the most directly on-topic regulatory text yet written on this exact design, in its Section 5.3 [2], and it is worth reading before your next design meeting.

One boundary first, because it matters. This hybrid control arm design is not the pure external-control-arm animal we covered in our posts on single-arm trials and regulator questions and the EMA reflection paper. There, the external data replaces the control group; here it supplements one. Randomisation stays the primary source of causal inference, and the external data's job is narrower, to shrink the size and variance of a comparison you are still running head to head.

What "borrowing strength" actually buys you

Strip the Bayesian vocabulary away and the mechanic is mundane. You enrol fewer patients into the concurrent control arm than a textbook power calculation demands. Then you let external control data, patients treated with the same comparator in a prior trial or a registry, stand in for some of the control patients you chose not to enrol. Formally, the external data forms a prior on the control-arm outcome, and your trial's own concurrent controls update it. The more the external data is allowed to count, the fewer concurrent controls you need.

Hobbs and colleagues put the design pattern into a single sentence back in 2011, describing the commensurate prior [3]: "if historical and concurrent controls emerge as commensurate, we might randomize fewer patients to the control group, thus enhancing the efficiency of the ongoing trial." That is the whole idea, and everything else is machinery for deciding what "commensurate" means and what to do when it fails.

For the arithmetic, take REWENEC-01 (NCT07337447), a Phase 2 trial in gastro-entero-pancreatic neuroendocrine carcinoma, a disease with an incidence under five per million, where a conventional two-arm trial is close to impossible [4]. It keeps a 4:1 randomisation ratio and blends patient-level data from a completed trial (BEVANEC, NCT02820857, a 153-patient two-arm trial whose FOLFIRI arm supplies the control data) with two French retrospective cohorts into a hybrid synthetic control [5]. The registry states the payoff plainly: 77 prospectively randomised patients are expected to "provide statistical power equivalent to that of a trial including 122 patients." Note the tense: the trial is not yet recruiting, with a 2026 start, so it gives us a clean, prospectively-designed example of the machinery, not an outcome.

The named methods are variations on one question: the power prior, the commensurate prior, the robust MAP or robust mixture prior (Schmidli et al.) [6], the multi-source exchangeability model, or MEM (Wei et al.) [7]. All of them answer the same thing: how much should the external data count, and should that weight be fixed, or should it move?

The dial, and the disagreement clause

There are two families, and the difference between them is the difference between a credible hybrid design and a liability.

Static borrowing fixes the weight in advance and uses the full strength of the external data regardless of how well it fits. It is simple, and it is dangerous precisely when you can least afford it, which is when the external control population turns out not to match your trial's. Viele and colleagues worked a plain binomial example every sponsor should sit with [8]: pool a historical control against a trial where the true rate differs, and "should the true control rate be 0.80, the type I error rate approaches 20%." A one-in-five false-positive rate, in a design that looks, on paper, perfectly powered.

Dynamic borrowing scales the weight to agreement. Line the external and concurrent controls up, and it borrows heavily, banking you the efficiency. Let them diverge and the same mechanism can "cease borrowing heavily" or, in the limit, "not borrow at all" [8]. The weight is a dial that responds to evidence, not a constant set once and forgotten.

So the design question is bigger than "do we borrow?" Two things actually decide it: how much, at most, the external data is allowed to count, and what happens, precisely, when it disagrees with the concurrent arm. Call that second commitment the disagreement clause: the pre-specified rule linking observed conflict to the degree of borrowing. Skip that clause and a hybrid design stops looking conservative and starts looking unexamined.

The dial is the design.

Everything downstream inherits whatever you set here, mistakes included, and it inherits them without ever raising a flag.

Free download

The External Control Arm Design Checklist

The 12 design points regulators probe first, in one checklist you can run against your protocol before database lock.

Get the checklist →

Three kinds of "real", and why the distinction is not pedantry

When someone tells you this pattern has "been done", ask which of three very different things they mean. Conflating them is how a design meeting talks itself into more confidence than the evidence supports. Three rungs, and they are not interchangeable. It is the same discipline behind asking what "regulatory-grade RWE" actually means before you accept someone else's claim of precedent.

Retrospective demonstrations. Take a completed, conventionally randomised trial and re-analyse it to show borrowing would have shrunk it. MILES [9], the 2015 FDA-approved sirolimus trial in lymphangioleiomyomatosis, was run as a straight 1:1 randomised trial with no borrowing at all. A 2023 re-analysis showed that dynamic borrowing "would have allowed the trial to enroll fewer concurrent controls while leading to the same conclusion", needing "7, 10, 15, and 22 fewer concurrent controls" at successive analyses, "without producing type I error inflation ... when concurrent and historical controls are comparable." For a trial that took "seven years to be completed at a cost of over $5M," that is not a trivial saving. COAST [7], a 180-patient randomised NSCLC platform trial, was similarly redesigned after the fact with MEM to show a hybrid control arm could "reduce the sample size of the IC arm by 50%." Both are compelling, but neither trial actually used the method: they tell you the pattern can work, not that it was used.

Real, regulator-engaged, but supportive or post-hoc. A rung up: DINAMO [10] was a genuine 26-week randomised, double-blind, placebo-controlled paediatric type 2 diabetes trial in 157 participants. Borrowing was not designed in to shrink enrolment; it was added after recruitment, when a variability problem surfaced (an observed standard deviation of 1.65% against an anticipated 0.9%), and the FDA capped the borrowed effective sample size at roughly 52 patients per group. The dynamic mechanism then earned its keep: for linagliptin, the prespecified superiority criterion "would not have been met with any choice of prior weight smaller than 0.542." The robust prior refused to over-trust a conflicting source, which is the disagreement clause doing its job. Belimumab in childhood SLE [11] is a second example, authored by FDA statisticians. A paediatric trial (NCT01649765) "could not be designed with adequate power for efficacy to make statistical inferences based solely on the pediatric data," so the agency ran a post-hoc Bayesian analysis borrowing from adult Phase 3 data with a robust mixture prior. "A prior weight of 55% or larger on the adult data provided 95% credible intervals that excluded an odds ratio of one," which "contributed to FDA's decision to approve belimumab as the first treatment for cSLE." Real, regulator-blessed use of the machinery, but supportive analysis, not a control arm deliberately shrunk at the design stage.

Genuinely prospective, or approved on the mechanism. The top rung, and the thinnest: REWENEC-01, above, is prospectively designed to reduce concurrent randomisation from day one, and its registry even states the hypotheses were "formulated a priori, as recommended by the FDA guidance document on trials with synthetic/external control arms" [4]. But it has no outcome yet, so the closest thing to a completed, approved, regulator-engaged example is PUNCH CD3 [12], the Phase III trial of RBX2660 (now REBYOTA) in recurrent C. difficile infection. It was 2:1 randomised, double-blind and placebo-controlled, randomisation fully intact, and it used a Bayesian hierarchical model to dynamically borrow treatment-effect information from the sponsor's own earlier PUNCH CD2 trial, at the FDA's own suggestion: the agency "acknowledged the increasing recruitment difficulties and recommended that innovative designs, such as formal borrowing of data in a Bayesian framework, could be pursued." Two superiority thresholds were pre-specified, both calibrated to control type I error without borrowing (posterior probability above 0.999 at one-sided 0.00125, above 0.975 at one-sided 0.025). The result held: 70.6% versus 57.5%, a treatment effect of 13.1%, posterior probability of superiority 0.991, and approval followed in November 2022.

That said, read PUNCH CD3 carefully before you lean on it. The borrowed data came from the same sponsor's own earlier-phase trial of the same drug, about the most defensible external source there is, and a long way from borrowing off an unrelated registry or a different sponsor's study. The borrowing also strengthened power; it did not shrink enrolment.

The pattern itself is not new; only its statistical sophistication is. Notably, Jahanshahi and colleagues, reviewing 45 FDA external-control approvals between 2000 and 2019, found that "a hybrid approach, where external control data were added to a concurrent randomized control arm ... was used for at least three products" [13]: velaglucerase alfa (Vpriv), corticotropin (Acthar) and centruroides scorpion anti-venom (Anascorp). None used Bayesian dynamic borrowing. They leaned on propensity adjustment and weighting, the older, static version of the same idea.

One honest gap, stated plainly. We did not find a completed drug approval that used Bayesian dynamic borrowing from a genuinely external source, a registry, an EHR or a different sponsor's trial, in a still-randomised confirmatory trial. PUNCH CD3 borrowed from its own sponsor's prior study; REWENEC-01 is designed the harder way but has not read out. If someone tells you this specific pattern is a well-trodden path to approval, it is not, yet. None of that argues for avoiding the design. It argues for over-preparing it.

The failure mode is quiet, not loud

Here is the strongest version of the case against this whole enterprise, and it deserves a fair hearing. A purist would say a hybrid design is statistical cosmetics a way to dress an underpowered trial up as a powered one. If you cannot power an honest randomised comparison, the argument runs, either run a bigger trial or do not make the claim. And the mathematics appears to back the purist. Kopp-Schneider and colleagues proved [15] that "strict control of type I error implies that no power gain is possible under any mechanism of incorporation of prior information, including dynamic borrowing," and later extended the proof explicitly to hybrid two-arm trials. No free lunch. If borrowing buys you power, it costs you type-I-error control somewhere.

That proof is correct, and it belongs on the wall of every team that designs one of these. But it argues for discipline, not abstinence. Dynamic borrowing does not eliminate the power-versus-error trade-off. What it does is bound it and make it visible, so you choose your position on that curve deliberately and disclose it, rather than stumbling into a hidden one. Solved? No bounded, and made visible. That is a real and defensible thing to bring to a regulator, and it is worlds away from naive pooling.

Because the danger in naive pooling is how well it hides. The failure lands quietly, inside a trial that otherwise looks clean, not as a loud flag at review. Lewis and colleagues showed exactly this on a real Roche/Genentech colorectal-cancer dataset of three trials [14]: "the naive combination of the controls results in a Bayesianly significant treatment effect for the novel Drug A," even though "no significant PFS benefit from Drug A exists when comparing with the concurrent control arm of standard of care." A false positive, manufactured entirely by over-trusting a mismatched external control. ICH E20 §5.3 states the same risk in regulatory language [2]: "Misspecification of the prior distribution can lead to lack of control of the probability of false positive conclusions."

Years ago I sat in a design discussion for a rare-disease programme where the team had picked a historical control cohort because it was the one they could get, not the one that matched. Nobody had written down what should happen if it disagreed with the concurrent controls. That is the meeting where the disagreement clause gets skipped, and the omission stays invisible until it is expensive.

Think of your concurrent control arm as a small panel of eyewitnesses to your trial's own events. Borrow well and you add a few more people who genuinely saw the same thing. Borrow badly, too heavily and from a poorly-matched source, and you pad the panel with confident witnesses to a different event entirely, who quietly outvote the handful who actually watched yours, and the verdict still reads as unanimous. It is just wrong.

What a credible hybrid control arm submission puts on the table

Strip the theory back to what you actually commit to, in writing, before the trial starts. ICH E20 §5.3 (draft, Step 2b, 25 June 2025) is the clearest current statement of the expectation, and it maps almost one to one onto the discipline above [2]. It calls for pre-specified justification of the amount of borrowing; a planning-stage discussion of "the maximum amount of borrowing and the relationship between observed conflict and the degree of borrowing", which is the disagreement clause in a regulator's own words; a preference for patient-level data from randomised trials as the external source; and evaluation of "the current trial data with no borrowing" as a sensitivity analysis.

The EMA has not written the equivalent for this design. But its reflection paper on extrapolation in paediatric medicines (EMA/189724/2018, adopted 17 October 2018) sets the tone plainly: randomised, controlled studies remain, in its word, "preferable" [16]. Borrowing supplements the randomised comparison; it does not earn you a pass on running one. (This is a different EMA document from the single-arm-trials reflection paper we covered separately, so do not conflate the two.)

On the FDA side, two documents are worth knowing exist. There is a draft guidance, "Use of Bayesian Methodology in Clinical Trials of Drug and Biological Products" (Docket FDA-2025-D-3217, dated January 2026) [17], and the finalised "Interacting with the FDA on Complex Innovative Trial Designs" guidance (Docket FDA-2019-D-3679, December 2020) [18], which is the formal engagement route for a design like this. One caution applies to both, though: at the time of writing, we were able to read only their Federal Register notices of availability, not the full guidance text, so do not take any specific numeric threshold as coming from either document without checking the source.

Turn all of that into what you walk into a pre-submission meeting holding:

  • The dial, named: state the maximum borrowed effective sample size, or the discount factor, up front. DINAMO's cap of roughly 52 patients per group is the shape of the thing.
  • The disagreement clause, written: specify what the design does as conflict between external and concurrent data grows, before you see the data, not after.
  • The no-borrowing analysis, committed: report the trial on its own concurrent data alone, as a pre-specified sensitivity check. ICH E20 asks for it, so bring it unprompted.
  • Source discipline: patient-level randomised-trial data beats a registry or EHR extract. Say why your source is what it is, and what you did about the gaps.
  • Early engagement: use the CID route to put the borrowing plan in front of the agency before it is locked. This is not a design to spring on a reviewer, and the same logic holds for any RWE-based regulatory strategy: show the agency your plan before you need it to be right.

Where the ground still moves

A hybrid control arm design can honestly do what the pitch promises. It can keep the credibility of randomisation while shrinking the number of patients you have to randomise concurrently, and REWENEC-01's 77 standing in for 122 is a real, if unproven, statement of the prize. But the prize is conditional. You collect it only if you name and cap the dial, write the disagreement clause before the data lands, and bring the no-borrowing analysis whether or not a reviewer asks for it. Skimp on any of that and the failure, when it comes, is the quiet kind that hides inside a clean-looking trial.

The FDA left hybrid control arms out of scope in its ECA guidance. ICH E20 §5.3 is now filling that gap, but its own footnote concedes the section "is not fully harmonized", and it remains under public consultation as of mid-2025. So the regulatory ground under this design is still moving, and anyone who tells you the rules are settled has not read the footnotes. The discipline, though, is the durable part. It will read the same wherever harmonisation lands.

If you are calibrating a borrowing plan now, the cheapest time to pressure-test the dial and the disagreement clause is while the numbers can still change, before the protocol locks and long before a reviewer runs the no-borrowing analysis for you. It is the same discipline behind a minimum viable evidence plan: spend the patients you can least afford to lose on the comparison that actually carries the claim.

Get the monthly digest

The 5 things evidence leads need to know each month: regulatory moves, RWE developments and what they mean in practice. No pitch, one email a month.

References

  1. Inovia Bio. "FDA Guidance on External Control Arms: A Checklist For Drug Developers." https://blog.inovia.bio/inovia-bio-insights/fda-guidance-on-external-control-arms-a-checklist-for-drug-developers
  2. ICH. "E20 Adaptive Designs for Clinical Trials." Draft guideline, Step 2b, endorsed 25 June 2025. Section 5.3. https://www.ich.org/page/efficacy-guidelines
  3. Hobbs BP, Carlin BP, Mandrekar SJ, Sargent DJ. (2011). "Hierarchical commensurate and power prior models for adaptive incorporation of historical information in clinical trials." Biometrics;67(3):1047-1056. PMID: 21361892. https://pubmed.ncbi.nlm.nih.gov/21361892/
  4. REWENEC-01. "A randomised Phase 2 trial in gastro-entero-pancreatic neuroendocrine carcinoma using a hybrid synthetic control arm." ClinicalTrials.gov: NCT07337447. https://clinicaltrials.gov/study/NCT07337447
  5. BEVANEC. "FOLFIRI +/- bevacizumab in neuroendocrine carcinoma." ClinicalTrials.gov: NCT02820857. https://clinicaltrials.gov/study/NCT02820857
  6. Schmidli H, Gsteiger S, Roychoudhury S, O'Hagan A, Spiegelhalter D, Neuenschwander B. (2014). "Robust meta-analytic-predictive priors in clinical trials with historical control information." Biometrics;70(4):1023-1032. PMID: 25355546. https://pubmed.ncbi.nlm.nih.gov/25355546/
  7. Wei W, et al. (2024). "A Bayesian platform trial design with hybrid control based on multisource exchangeability modelling." Statistics in Medicine;43(12). PMID: 38594809. https://pubmed.ncbi.nlm.nih.gov/38594809/
  8. Viele K, Berry S, Neuenschwander B, et al. (2014). "Use of historical control data for assessing treatment effects in clinical trials." Pharmaceutical Statistics;13(1):41-54. PMID: 23913901. https://pubmed.ncbi.nlm.nih.gov/23913901/
  9. Harun N, et al. (2023). "Dynamic use of historical controls in clinical trials for rare disease research: A re-evaluation of the MILES trial." Clinical Trials;20(3). PMID: 36927115. https://pubmed.ncbi.nlm.nih.gov/36927115/
  10. Sailer MO, et al. (2025). "Pharmacometrics-Enhanced Bayesian Borrowing for Pediatric Extrapolation - A Case Study of the DINAMO Trial." Therapeutic Innovation & Regulatory Science;59(1). PMID: 39373938. https://pubmed.ncbi.nlm.nih.gov/39373938/
  11. Pottackal G, et al. (2025). "Application of Bayesian statistics to support approval of intravenous belimumab in children with systemic lupus erythematosus in the United States." Lupus;34(9). PMID: 40577570. Trial: NCT01649765. https://pubmed.ncbi.nlm.nih.gov/40577570/
  12. Khanna S, et al. (2022). "Efficacy and Safety of RBX2660 in PUNCH CD3, a Phase III, Randomized, Double-Blind, Placebo-Controlled Trial with a Bayesian Primary Analysis for the Prevention of Recurrent Clostridioides difficile Infection." Drugs;82(15). PMID: 36287379. https://pubmed.ncbi.nlm.nih.gov/36287379/
  13. Jahanshahi M, Gregg K, Davis G, et al. (2021). "The use of external controls in FDA regulatory decision making." Therapeutic Innovation & Regulatory Science;55(5):1019-1035. PMID: 34014439. https://pubmed.ncbi.nlm.nih.gov/34014439/
  14. Lewis CJ, Sarkar S, Zhu J, Carlin BP. (2019). "Borrowing from historical control data in cancer drug development: a cautionary tale and practical guidelines." PMID: 31435458. https://pubmed.ncbi.nlm.nih.gov/31435458/
  15. Kopp-Schneider A, Calderazzo S, Wiesenfarth M. (2020). "Power gains by using external information in clinical trials are typically not possible when requiring strict type I error control." Biometrical Journal;62(2):361-374. PMID: 31265159. https://pubmed.ncbi.nlm.nih.gov/31265159/ (extended to hybrid two-arm trials: PMID: 37632266. https://pubmed.ncbi.nlm.nih.gov/37632266/)
  16. EMA. "Reflection paper on the use of extrapolation in the development of medicines for paediatrics." EMA/189724/2018, adopted 17 October 2018. https://www.ema.europa.eu/en/documents/scientific-guideline/adopted-reflection-paper-use-extrapolation-development-medicines-paediatrics-revision-1_en.pdf
  17. FDA. "Use of Bayesian Methodology in Clinical Trials of Drug and Biological Products." Draft guidance, Docket FDA-2025-D-3217, January 2026. (Existence, title, docket and DRAFT status confirmed via the Federal Register Notice of Availability, 12 January 2026; full guidance text not retrieved, so no content is attributed to it.) https://www.regulations.gov/docket/FDA-2025-D-3217
  18. FDA. "Interacting with the FDA on Complex Innovative Trial Designs for Drugs and Biological Products." Final guidance, Docket FDA-2019-D-3679, December 2020. (Existence, title, docket and FINAL status confirmed via the Federal Register Notice of Availability, 17 December 2020; full guidance text not retrieved, so no content is attributed to it.) https://www.regulations.gov/docket/FDA-2019-D-3679

Share this post