Making small N count: the rare-disease statistical toolkit, and the price each method charges
Two trials, same drug, same disease, same primary endpoint. Mexiletine for non-dystrophic myotonia. One of them (NCT00832000) enrolled 59 patients across seven sites in the US, Canada, Italy and the UK, and read out as a conventional multicentre crossover RCT [3]. The other (NCT02045667) enrolled 30 patients at a single site, analysed them as a combined series of n-of-1 experiments inside a Bayesian hierarchical model, and named the first trial as its own benchmark in the registry record [2].
Read that quickly and you reach the seductive conclusion that tempts every team designing a rare-disease clinical trial: small-N methods are a shortcut. Pick the clever design, halve your sample, and you've skipped six sites and banked two years before anyone has even seen your protocol. If it worked for mexiletine it will work for your asset.
That is the wrong lesson, and the EMA said so twenty years ago. Its 2006 guideline on clinical trials in small populations (CHMP/EWP/83561/2005, final) puts it flatly: "There are no special methods for designing, carrying out or analysing clinical trials in small populations. There are, however, approaches to increase the efficiency of clinical trials." [1] No magic. Just efficiency, bought at a price.
Carry this frame through the rest of the post. Every method in the small-N toolkit is an "efficiency loan": it lends you statistical power you did not have the patients to earn, and like any loan it demands "collateral", an assumption you post up front that the regulator can call in if it turns out worthless. Choose the method by name, ignore the collateral, and you have not de-risked your programme. You have moved the risk somewhere your statistician cannot see it. The lever is method choice and the collateral behind it, not data volume, and that choice is itself part of your regulatory case.
What runs out when the patients run out: the arithmetic of small-N trial design
Power is the obvious casualty. Shrink N and the trial's ability to detect a real effect collapses, so a drug that genuinely works can still read as a null. Less obvious, and more dangerous, is what happens to your estimates. At very small N a single patient swings the point estimate wildly, and in the extreme you meet what Kidwell and colleagues call the zero-numerator problem: a real effect can leave zero events in a cell, and the standard frequentist machinery, which divides by that count, simply breaks [4].
So before you argue about designs, do one piece of arithmetic. Work out which of your endpoints can actually reach significance at the N you can realistically enrol. If the honest answer is none of them, you do not need a bigger dream or a louder KOL. You need an efficiency technique. And every technique on the shelf charges for what it lends you.
Method one: borrow big, but post the collateral
Bayesian borrowing is the loan in its purest form: external or historical data, expressed as a prior, stands in for the patients you cannot enrol, then gets updated once your own patients arrive. What does that actually buy you? Power, or a smaller trial.
Kidwell and colleagues worked a progressive supranuclear palsy (PSP) example in full: a conventional 1:1 frequentist design needed 85 patients per arm, while a MAP-prior Bayesian design with 2:1 randomisation needed 85 in the experimental arm and just 43 on placebo, borrowing the rest. When the external and trial-placebo data agreed, the power gain ran as high as 14% [4].
That is not a toy. Harun and colleagues went back to the real MILES trial (sirolimus in lymphangioleiomyomatosis, a trial that took seven years to run) and re-analysed it with dynamic historical borrowing. With 1:1 propensity-matched historical controls, it could have reached its final analysis on 67 patients instead of 89, or 55 with 1:2 matching, and in the authors' words "without producing type I error inflation and preserving power" [5].
Then the collateral. Both gains assume the borrowed data is exchangeable with your trial population: that the historical patients would, on average, have behaved like your concurrent controls. Post that honestly and borrowing is genuine efficiency. Duck it, and the same machinery you leaned on for power turns on you. In Harun's own re-analysis, the scenario that used the US-only historical registry to stand in for the trial's Japanese concurrent controls (mismatched populations) produced wider confidence intervals and a larger required sample, not a smaller one [5]. Borrowing the wrong data cost patients rather than saving them.
But done properly, there is regulator-tested precedent for this. Empirical MAP priors were built precisely to detect and down-weight prior-data conflict, with a haemophilia A observational-safety application (Li et al. 2016) [6]. The external- or synthetic-control-arm version of borrowing, where the data forms a whole comparator rather than a prior, is a discipline in its own right; the FDA's thinking on it has its own checklist for external control arms, and this post will not re-litigate the ECA mechanics. One warning that saves grief later: a concurrent shared-placebo platform is not historical borrowing, whatever it looks like. Hold that thought for method three.
Monday payload: if you intend to borrow, the deliverable is not "we'll use a Bayesian design." It is a pre-specified exchangeability assessment, a robustified (mixture) prior that can discount a conflicting source, and a type-I-error simulation run under prior-data conflict, all agreed with the regulator before your SAP locks.
Method two: the patient as their own control, when the disease permits
The n-of-1 design is the most elegant efficiency in the box. Each patient is randomised repeatedly between treatment and control across several crossover periods, so each patient becomes their own comparator. Between-patient variability, the thing that eats your power in a small parallel trial, largely drops out. Aggregate a series of these under a Bayesian hierarchical model and you get a population read from a handful of people. That is how mexiletine reached a defensible answer on 30 patients at one site rather than 59 across seven [2][3].
It scales down astonishingly far. There is a formal single-patient n-of-1 of hydrocortisone in a lipomatosis-with-neuropathy syndrome that has around seven documented cases worldwide (NCT04821583, n=1) [14]. N really can equal one and still be a designed experiment.
When should you actually reach for an n-of-1 trial design? Only when the disease's physics genuinely cooperate, never for statistical convenience alone. The collateral here is physical, not statistical, and it is unforgiving. The n-of-1 design only works if the disease sits still between crossover periods and the drug washes out cleanly: a stable, reversible situation. The EMA said it plainly in 2006. N-of-1 trials are "most useful for fast-acting symptomatic treatments and in diseases that quickly return to stable baseline values after treatment" [1]. Read that against the reality of most rare diseases, which are progressive and irreversible, and the design rules itself out before you reach the statistics. Cheung and Mitsumoto make the quantitative version of the point in ALS: once the treatment effect is small, the number of periods and patients you need climbs even inside an n-of-1 framework [10].
There is genuine regulatory mileage when the physics cooperates. Weinreich and colleagues ran an aggregated n-of-1 of ephedrine in myasthenia gravis all the way through formal scientific advice, including the regulator's own limiting judgement on what it could and could not support [8]. A 2025 series of L-serine n-of-1 trials in GRIN2B-related neurodevelopmental disorder reported honestly mixed results [9]. And the method is not confined to academic single-centre curiosities: HYDRA (NCT02226198) ran a rosuvastatin crossover in paediatric homozygous familial hypercholesterolaemia across nine sites in seven countries, sponsored by AstraZeneca [13].
One expertise beat, take the most famous "n-of-1" story in rare disease: milasen, the antisense oligonucleotide designed and dosed for a single child with CLN7 Batten's disease [11]. It is not an n-of-1 trial in this sense at all. Kim-McManus and colleagues, some of them the same investigators, are explicit that individualised gene-targeted therapies "do not naturally lend themselves to crossover clinical trial designs": single-dose pharmacology and a long half-life leave nothing to wash out [12]. Milasen is individualised therapy development for one patient. A landmark, certainly, but a different animal from the crossover method in this section, and conflating the two will cost you at a scientific-advice meeting.
Monday payload: check the physics of your disease before you fall in love with the design. Progressive and irreversible? N-of-1 is off the table however few patients you have. Stable, reversible, fast-acting and washable? It may be the most efficient design available to you.
Free download
The RWE Briefing Document Template
The section-by-section structure for the RWE part of a regulatory briefing, built around the questions reviewers actually ask.
Get the template →Method three: efficiency from structure, not from strangers' data
The third family buys efficiency from the architecture of the trial itself, which means it carries no exchangeability risk at all. Nobody else's patients are involved. Three flavours are worth keeping separate.
Pre-specified stopping. You build in interim looks and stop early for success or futility, so a clear signal in either direction ends the trial before you have spent every patient. INHIBIT (156 haemophilia-A patients) stops for success at 75% enrolment on a high posterior probability of superiority; DIAN-TU, in autosomal dominant Alzheimer's (under 1% of all Alzheimer's), pre-specified a stopping threshold of posterior probability at or above 0.9952 [4]. A precision note, because it matters: those are Bayesian posterior-threshold stops. Classical frequentist group-sequential designs for rare disease are well developed in the simulation literature (Bayar et al. 2020; Kotalik et al. 2022) [18][17], but I found no named rare-disease registry trial that actually used frequentist interim stopping. Ask for the NCT number before you assume otherwise.
Hierarchical fixed-sequence testing. This one is routinely mislabelled, so be precise. When you have more than one primary endpoint, you can test them in a pre-set order and only proceed to the second if the first is significant, which controls your family-wise type-I error without splitting alpha across endpoints. STR1VE, the pivotal Zolgensma trial in SMA type 1 (NCT03306277, n=22), did exactly this: independent sitting first, then event-free survival, tested in fixed sequence [15]. That is multiplicity control by ordering. It is not a group-sequential interim-stopping design, and people who call it one misdescribe what actually protected the type-I error.
Platform and shared-placebo efficiency. A platform trial tests several treatments against one shared control, so no single arm has to fund its own placebo group. The HEALEY ALS Platform (NCT04297683 and NCT04436497) does this with Bayesian repeated-measures and shared-parametric survival models over a concurrent, contemporaneous shared placebo [16]. Note the word concurrent: that placebo runs alongside the treatment arms, in the same protocol, at the same time. This is emphatically not historical or external borrowing, and the efficiency it delivers costs you no exchangeability assumption at all; it is bought with logistics and a master protocol, the machinery the FDA's revised draft master-protocols guidance (24 June 2026, still unfinalised) is trying to standardise [20].
The collateral for this whole family is discipline, not an assumption about strangers. You pre-specify every boundary, every testing order and every stopping rule before patient one, and you live with them. There is one genuine interaction to flag: when sequential stopping is combined with borrowing, Kotalik and colleagues show the type-I-error behaviour flips depending on whether you sit under the global null or a local null, so the two techniques do not simply add up [17].
Honesty counterweight, DIAN-TU was a rigorous, elegantly pre-specified Bayesian platform, and its initial results showed neither investigational drug slowed cognitive decline [4]. Structure buys you efficiency. It does not buy you a signal that was never there.
Let's be honest about what borrowing does to your error rate
Here is the strongest case against everything I have just argued, put as strongly as it deserves. These methods are mature and regulators endorse them: the EMA's 2018 reflection paper on paediatric extrapolation (EMA/189724/2018, final) lays out Bayesian borrowing with guardrails and worked examples [7], and in January 2026 the FDA issued its first guidance dedicated to Bayesian methodology in drug and biological product trials (Docket FDA-2025-D-3217), naming the use of an informative prior to borrow external information in the primary analysis as an explicit topic [19]. Robustified and mixture priors are built to detect a conflicting source and down-weight it automatically. So, the argument runs, the exchangeability worry is overblown, the mathematics handles it, and warning lean teams off their most efficient tool just leaves patients waiting.
It is a fair challenge, and it is wrong in one specific, demonstrable way.
That said, even a mixture prior designed to guard against conflict does not fully neutralise it. In Kidwell's own PSP example the maximum type-I error under prior-data conflict was 6.3% with the robustified prior in place [4]. That is above the 5% you promised the regulator, from the very safeguard meant to protect it. Harun's MILES re-analysis makes the point from the other direction: non-exchangeable historical controls widened the confidence intervals and increased the required sample [5]. The safeguards reduce the damage; they do not abolish it. And the deeper problem is not fixable by a better prior, because exchangeability is unverifiable by construction: you are borrowing precisely because you lack the concurrent controls that would let you test whether the borrowed patients resemble yours. Kotalik puts it formally, the same method reducing type-I error under one hypothesis and inflating it under another [17].
Reduce or inflate? It depends entirely on one thing you cannot check when you design the trial.
Which is why regulators police the assumption, not the method. The EMA's 2018 paper (a different, later document from its single-arm-trials reflection paper, unpicked separately) is unambiguous: "It is important to quantify how much information will come from the prior relative to the actual data generated. The Type I Error properties of any Bayesian method should be investigated," and borrowing "to such an extent that data generated in the target population would not dominate cannot usually be supported" [7]. The clearest illustration is its paediatric Gaucher disease example, where extrapolation was accepted for the somatic manifestations of the disease and explicitly refused for the neurological manifestation [7]. Same disease, same programme, opposite decisions taken at the level of the specific assumption, not the method's name. The 2006 guideline had already said the same about priors: "a variety of reasonable prior distributions should be used ... to ensure that conclusions are not too heavily weighted on the prior beliefs" [1]. This is also why the fitness-for-purpose of the borrowed data matters as much as the model, a case argued in full in our piece on what regulatory-grade RWE actually means.
Two honesty guardrails before anyone reads this as a sales pitch for borrowing.
First: no method manufactures signal. DIAN-TU is the reminder above.
Second, and this one is easy to abuse, so be scrupulous. The elamipretide story in Barth syndrome gets reached for as a "regulators rejected small-N methods" example. It is not that. The TAZPOWER crossover (NCT03098797) enrolled around 12 patients; the programme drew a refusal-to-file and later a Complete Response Letter in May 2025 despite a positive advisory-committee vote, then found an agreed path to accelerated approval on a different endpoint, knee-extensor muscle strength, marketed as Forzinity [21]. The stated grounds were trial adequacy and endpoint choice in a very small trial, not prior-data conflict or exchangeability. To be plain about the boundary of the evidence: across all of this research I found no named FDA rejection that cited "prior-data conflict" or "exchangeability" by name. The elamipretide case shows small-trial and endpoint-adequacy risk in general. It is not a borrowing-specific rejection, and it should not be dressed up as one.
The honest answer to "does borrowing help or hurt" is therefore "it depends on an assumption you cannot confirm the day you design the trial." So make the assumption explicit, quantify the discounting, and simulate the error rate under conflict. Do not hide behind the method's name.
What to actually do on Monday
Strip it to a sequence you can run this week.
- Start with the physics of the disease, not the statistics. A stable, reversible, fast-acting treatment puts n-of-1 on the table. Progressive and irreversible takes it off, and your real choice is then between borrowing and structure.
- Prefer structural efficiency where the design allows it. Sequential stopping, hierarchical testing and shared-placebo platforms charge you pre-specification discipline, not an unverifiable assumption about other people's patients.
- If you borrow, name the collateral up front. A pre-specified exchangeability assessment. A robustified prior that can discount a conflicting source. A type-I-error simulation under prior-data conflict, with an agreed discounting rule. Lock all of it before the SAP locks, exactly as the EMA's 2018 paper asks [7].
- Take the method to the regulator early. The EMA's 2006 guideline actively encourages scientific advice for small-population designs, and Weinreich's n-of-1 went through precisely that route [1][8]. Method choice is part of the submission; get it pressure-tested before you build on it. There is a separate post on taking an RWE package to regulators.
- Don't oversell it, internally or to investors. The EMA's 2006 guideline is blunt that most orphan and paediatric approvals still rest on RCTs, and that there is no general paradigm change on offer [1]. The toolkit is an efficiency layer on a genuinely hard problem, the same minimum-viable-evidence trade-off every resource-lean biotech makes elsewhere, not a substitute for a control arm where one is feasible.
This is the kind of method-choice pressure-testing Inovia Bio does, and where InoviaCS tends to earn its keep is the unglamorous first step: landscaping which historical or external datasets even exist for your indication before you stake a design on borrowing from them.
The rule that hasn't been written yet
There is a reason this reads like an unsettled field. It is one.
The FDA's Bayesian guidance is a draft. It published on 12 January 2026, its comment period closed on 13 March 2026, and there is no final version as I write [19]. The revised master-protocols draft (24 June 2026) is also still open [20]. The framework for exactly these methods is being drafted right now, in public, while rare-disease teams design trials against it.
And the live question in that draft is one of degree: how much borrowing is too much, the same question the EMA answered case by case in its Gaucher example and that is now being generalised into guidance [7][19]. The answer will be written at the level of the assumption, the discounting and the error rate. It will not be settled by which method you name in your protocol.
Which brings it back to the patient at the floor of all this. For that lipomatosis-with-neuropathy syndrome with seven documented cases in the world, there will never be a 59-patient trial. There will never be a 30-patient one. For those seven people, a designed single-patient n-of-1 is the only thing standing between a defensible answer and no answer at all. Getting the method right, and posting its collateral honestly, is how you make their small N count.
Get the monthly digest
The 5 things evidence leads need to know each month: regulatory moves, RWE developments and what they mean in practice. No pitch, one email a month.
References
- European Medicines Agency (CHMP). "Guideline on Clinical Trials in Small Populations" (CHMP/EWP/83561/2005). FINAL; adopted 27 July 2006, in force 1 February 2007.
- Combined N-of-1 and Bayesian hierarchical analysis of mexiletine in non-dystrophic myotonia; n=30, single site. Registry title: "Combining N-of-1 Trials to Estimate Population Clinical and Cost-effectiveness of Drugs Using Bayesian Hierarchical Modeling. The Case of Mexiletine for Patients With Non-Dystrophic Myotonia." Sponsor: Radboud University Medical Center (confirmed). ClinicalTrials.gov: NCT02045667. https://clinicaltrials.gov/study/NCT02045667
- Multicentre crossover RCT of mexiletine in non-dystrophic myotonia; n=59, seven sites (US, Canada, Italy, UK). Registry title: "Phase II Therapeutic Trial of Mexiletine in Non-Dystrophic Myotonia." Sponsor: Richard Barohn, MD (University of Kansas Medical Center) (confirmed). ClinicalTrials.gov: NCT00832000. https://clinicaltrials.gov/study/NCT00832000
- Kidwell KM, et al. (2022). Application of Bayesian methods to accelerate rare disease drug development: scopes and hurdles. Orphanet J Rare Dis;17(1):186. PMID: 35526036. https://pubmed.ncbi.nlm.nih.gov/35526036/
- Harun N, et al. (2023). Dynamic use of historical controls in clinical trials for rare disease research: a re-evaluation of the MILES trial. Clin Trials;20(3):223-234. PMID: 36927115. https://pubmed.ncbi.nlm.nih.gov/36927115/
- Li JX, et al. (2016). Addressing prior-data conflict with empirical meta-analytic-predictive priors in clinical studies with historical information. J Biopharm Stat;26(6):1056-1066. PMID: 27541990. https://pubmed.ncbi.nlm.nih.gov/27541990/
- European Medicines Agency. "Reflection Paper on the Use of Extrapolation in the Development of Medicines for Paediatrics" (EMA/189724/2018). FINAL; adopted by CHMP 17 October 2018.
- Weinreich SS, et al. (2017). Aggregated N-of-1 trials for unlicensed medicines for small populations: an assessment of a trial with ephedrine for myasthenia gravis. Orphanet J Rare Dis;12(1):88. PMID: 28494776. https://pubmed.ncbi.nlm.nih.gov/28494776/
- den Hollander B, et al. (2025). Potential benefits of l-serine in children with GRIN2B loss-of-function variants: randomized n-of-1 trials. Mol Genet Metab;146(4):109268. PMID: 41265180. https://pubmed.ncbi.nlm.nih.gov/41265180/
- Cheung K, Mitsumoto H (2022). Evaluating Personalized (N-of-1) Trials in Rare Diseases: How Much Experimentation Is Enough? Harvard Data Sci Rev;2022(Spec Iss 3). PMID: 38283317. https://pubmed.ncbi.nlm.nih.gov/38283317/
- Kim J, et al. (2019). Patient-Customized Oligonucleotide Therapy for a Rare Genetic Disease (milasen; CLN7 Batten's disease). N Engl J Med;381(17):1644-1652. PMID: 31597037. https://pubmed.ncbi.nlm.nih.gov/31597037/
- Kim-McManus O, et al. (2024). A framework for N-of-1 trials of individualized gene-targeted therapies for genetic diseases. Nat Commun;15(1):9802. PMID: 39532857. https://pubmed.ncbi.nlm.nih.gov/39532857/
- HYDRA: randomised rosuvastatin crossover in paediatric homozygous familial hypercholesterolaemia; nine sites, seven countries; sponsored by AstraZeneca. ClinicalTrials.gov: NCT02226198. https://clinicaltrials.gov/study/NCT02226198
- Single-patient N-of-1 of hydrocortisone in symmetric lipomatosis with neuropathy; n=1. Sponsor: Erasmus Medical Center. ClinicalTrials.gov: NCT04821583. https://clinicaltrials.gov/study/NCT04821583
- STR1VE: onasemnogene abeparvovec (Zolgensma) in SMA type 1; n=22, single-arm; hierarchical fixed-sequence co-primary testing. Sponsor: Novartis Gene Therapies (confirmed). ClinicalTrials.gov: NCT03306277. https://clinicaltrials.gov/study/NCT03306277
- HEALEY ALS Platform: Bayesian repeated-measures and shared-parametric survival models over a concurrent shared placebo. Sponsor: Merit E. Cudkowicz, MD / Massachusetts General Hospital (confirmed). ClinicalTrials.gov: NCT04297683 and NCT04436497. https://clinicaltrials.gov/study/NCT04297683
- Kotalik A, et al. (2022). A group-sequential randomized trial design utilizing supplemental trial data. Stat Med;41(4):698-718. PMID: 34755388. https://pubmed.ncbi.nlm.nih.gov/34755388/
- Bayar MA, et al. (2020). Group sequential adaptive designs in series of time-to-event randomised trials in rare diseases: a simulation study. Stat Methods Med Res;29(6):1483-1498. PMID: 31354106. https://pubmed.ncbi.nlm.nih.gov/31354106/
- FDA. "Use of Bayesian Methodology in Clinical Trials of Drug and Biological Products." DRAFT; published 12 January 2026 (Docket FDA-2025-D-3217); comment period closed 13 March 2026; no final version as of writing. Federal Register 91(7), 12 January 2026 (FR Doc. 2026-00325).
- FDA. "Master Protocols for Drug and Biological Product Development." DRAFT (revised); published 24 June 2026; comment period open to 24 August 2026; unfinalised. Revises the 22 December 2023 draft. Federal Register, 24 June 2026 (FR Doc. 2026-12620; Docket FDA-2023-D-5259).
- FDA. "FDA Grants Accelerated Approval to First Treatment for Barth Syndrome" (press announcement, September 2025). Forzinity (elamipretide) granted accelerated approval on an improvement in knee-extensor muscle strength, an intermediate clinical endpoint, following a May 2025 Complete Response Letter and a 10–6 advisory-committee vote in favour (10 October 2024). Trial: TAZPOWER, ClinicalTrials.gov NCT03098797. https://www.fda.gov/news-events/press-announcements/fda-grants-accelerated-approval-first-treatment-barth-syndrome
-1.png?width=1169&height=277&name=Inovia%20Logo%20Dark%20(1)-1.png)