Skip to content
Inovia Bio featured image for the evidence gap audit: five axes for grading regulatory readiness in a drug development programme before FDA or EMA review.
Strategy IEP drug development

The evidence gap audit: five axes to grade your programme before a reviewer does

Imi
Imi

In July 2025 the FDA published something it had kept private for its entire history: the first tranche of its Complete Response Letter back-catalogue, the rejection notices that tell a sponsor precisely why an application failed upwards of 200 letters, though at that point limited to applications that had gone on to be approved [1]. Two months later, in September 2025, the agency released a second and different batch: 89 letters tied to applications that have still not been approved. The consultancy working through that September batch, reported that manufacturing and CMC problems now dominate the recent record, cited in more than half of the 89 letters they reviewed; equity analysts at Jefferies, reviewing the wider and still-growing repository, reported the same pattern [2].

That should give you pause, because it is not what the peer-reviewed literature says. The large academic studies of why drugs fail at the agency put efficacy and safety deficiencies at the centre of the story: Sacks and colleagues across 302 new molecular entities [3], Chahal on Refuse-to-File letters [4], Lurie on the gap between what sponsors say publicly and what the FDA actually wrote [5]. Two portraits of the same question, drawn a decade apart, and they do not match.

Sit with that for a second, because it is the point of this whole piece.

Here is the question every one of those studies is really answering. Is this evidence base strong enough to carry the weight of the decision the sponsor is about to put on it? A pivotal readout, a pre-BLA meeting, a term sheet. Think of your evidence base as a structure, and this evidence gap audit as a load test you run yourself, on your own schedule, before a reviewer or an investor runs it for you at the worst possible moment. Nobody books that test for your convenience.

So run your own first.

Nobody publishes this framework, and that is the point

Let me be plain about what follows. No FDA, EMA or ICH document sets out these five axes as a single readiness framework. I assembled it from five separate bodies of evidence, running from CRL and Refuse-to-File cause-coding studies, through safety-population benchmarking against ICH thresholds and the standing critique of single-arm trials, to the hard outcome data on regulatory advice. Naming the provenance is not a hedge. A framework honest about where it came from is one you can actually pressure-test, instead of a checklist handed down from nowhere.

The five axes:

  1. Mechanism and effect-size validity: does what you measured predict real benefit, and is the effect big enough?
  2. Natural history and endpoint characterisation: do you know what untreated looks like well enough to prove you beat it?
  3. Comparator and control strategy: you have to be able to defend how you know the drug, and not the drift, produced the result.
  4. Safety-database maturity and duration: is your safety database sized to find the thing that will hurt you?
  5. Regulatory precedent and engagement alignment: the mechanisms regulators built only help if you used them, and complied with what they told you.

This audit runs against an Integrated Evidence Plan you already have, or before you decide you need one. It does not replace the 90/10 framework, which tells you how much evidence is enough; this tells you where, specifically, you are short. If your runway is tight and you think you cannot afford the exercise, that is exactly when you can least afford to skip it.

Axis 1 Does the endpoint you moved actually predict benefit, and did it move far enough?

Ask anyone what sinks a first FDA cycle and most people will say safety. Sacks and colleagues' analysis of 302 new molecular entities submitted between 2000 and 2012 says otherwise: efficacy was the single most common reason drugs failed. Of 151 first-cycle failures, 58.9% had efficacy deficiencies: 13.2% on unsatisfactory endpoints, another 13.2% on poor efficacy against standard of care rather than placebo [3]. The European mirror is worse, not better. Tafuri and colleagues, cataloguing the grounds cited across 86 EU marketing applications that were withdrawn or refused, found efficacy deficiencies behind 67.9% of the major objections raised [6]. State the difference plainly and do not smooth it: the EU share is higher, and the two evidence bases are not the same size. Four independent analyses cover the FDA side across two decades; the EU side rests on two studies, one of them now well over a decade old. Do not read the FDA breakdown as if it describes Brussels.

Now the timing, woven in. Sacks found a median of 435 days from a first unsuccessful submission to eventual approval but the range ran from 47 to 2,374 days [3]. That spread is the whole argument for auditing early. A small effect-size gap costs you weeks. A wrong-endpoint gap can cost you the better part of seven years.

The named cases make it concrete. Eteplirsen's pivotal trial read out a dystrophin surrogate in 12 patients [7]; the confirmatory PROMOVI trial, 109 patients on a functional endpoint, ran for years afterward to close the gap between that surrogate and actual benefit [8]. Aducanumab is the sharper lesson. Its two pivotal trials, EMERGE and ENGAGE, were designed identically, at 1,643 and 1,653 patients [9][10], and returned discordant results on the same primary endpoint, later attributed in part to dose and exposure differences that surfaced too late to fix. The confirmatory ENVISION trial, 1,027 patients [11], was terminated before it could deliver a verdict. We told that story in full in A Study In Failure, so I will not re-narrate it here. The audit point is narrow. Check whether your primary endpoint has ever supported an approval in your indication, and whether your effect size beats standard of care, not whether it beats nothing.

Free download

The IEP Template Pack

The gap matrix, prioritisation grid and plan-on-a-page we use to build integrated evidence plans. Free to keep.

Get the template pack →

Axis 2 Do you know what "untreated" looks like?

You cannot prove you beat the natural history of a disease you have not characterised. Ataluren's unstratified Phase 2b trial enrolled 174 boys with Duchenne and missed its primary endpoint [12]. The redesigned ACT DMD trial, 230 patients, narrowed enrolment to a specific baseline-walking band, roughly 150 metres to 80% of predicted, precisely because the earlier trial had not understood where on the disease trajectory the endpoint could actually move [13]. That redesign is registry-documented, and it is an honest example in both directions. It helped, but it was not enough: ataluren has still never received full FDA approval.

Timing again. Makena reached accelerated approval on a surrogate; its confirmatory PROLONG trial took nine years from start to primary completion to test whether the assumed benefit was real [14]. In the event, it was not confirmed. That nine-year wait was bought at protocol design, long before anyone sat down for a pre-BLA meeting that could have done anything about it.

So ask yourself the flat question. Do you have natural-history or registry data characterising your endpoint's trajectory in the population you actually enrolled, or are you assuming it? If you are assuming it, that assumption is the gap.

Axis 3 The comparator question

Kept deliberately short, because this blog has covered it at depth elsewhere. Subramaniam and colleagues' 2024 catalogue of how single-arm and externally-controlled trials draw regulatory fire names three recurring faults: "non-contemporaneous ECAs, subjective endpoints, and baseline covariate imbalance between arms" [15]. If any of those describes your control strategy, you have an Axis 3 problem, and the fix lives in more detail than this axis will get here. For that detail, see our FDA ECA checklist and the EMA's position on single-arm trials. It is one axis among five, and I will move on.

Axis 4 Is your safety database sized to find the thing that will hurt you?

"We've dosed a lot of patients" is a feeling, not a threshold. The actual threshold has a name and a number. ICH E1, quoted by Bouwman and colleagues, sets it out: for chronic use "the total number of patients exposed should be at least 1,000, and data of 6 months of use should be available for at least 300 patients and 1 year exposure for at least 100 patients" [16]. The gap between that standard and reality is wide and uneven. Among 154 non-orphan chronic-use medicines, 54% met the six-month threshold. Among 101 orphan chronic-use medicines, 1% did [16].

Lorcaserin is the case that ought to haunt anyone tempted to under-size. Approved in 2012 without a cardiovascular and long-term cancer database of adequate size, it needed a postmarketing trial, CAMELLIA-TIMI 61, enrolling 14,673 patients, to supply what the pre-approval package never had [17]. That trial detected the cancer imbalance the original database was never sized to find, and the FDA's own review of that signal is what drove the drug from the market: cancer was diagnosed in 7.7% of lorcaserin patients versus 7.1% on placebo, spread across several tumour types, and the agency asked the sponsor to withdraw it. The pre-approval database was sound in design and too small in size to catch a rare signal.

Count your exposed-patient numbers against ICH E1 for your dosing duration, today. Orphan status may relax the number a regulator will accept. It does not relax the biology.

Axis 5 Did you use the door regulators built for you, and walk through it?

Regulators built formal mechanisms for exactly the conversation you are dreading. 21 CFR § 312.47 lists identifying "any additional information necessary to support a marketing application" as one of the End-of-Phase-2 meeting's stated purposes, alongside confirming the drug is safe to proceed to Phase 3 and evaluating the Phase 3 plan itself [18]. Using it is not the whole battle, though. Regnstrom and colleagues' EMA analysis found that obtaining scientific advice, on its own, was not associated with a positive outcome at all (odds ratio 0.96). Complying with the advice was, and powerfully so: compliant applications ran an odds ratio of 14.71 for success, non-compliant ones 0.17 (p<0.0001) [19]. Chahal's Refuse-to-File data tells the same story from the failure side. 26.2% of RTF letters cited presubmission advice the applicant had received and ignored, most often on trial design [4].

Here is a war-story, anonymised. Years ago I worked with an early-phase rare-disease team convinced their pivotal design was sound: they had built the whole programme around a patient-reported functional endpoint the steering committee loved, and on its own terms it was a defensible choice. What a scientific-advice meeting would have surfaced, had they sought it before locking the protocol rather than after, was that the agency had already watched that exact type of endpoint fail to persuade a review division elsewhere, and wanted a natural-history-anchored measure instead. They found the gap at the pre-submission stage, not the design stage, with the assay already validated and running at every site.

That gap did not get cheaper for the wait.

The cost of finding an Axis 5 gap late is measurable. Chahal found a median 182 days to resubmission after a Refuse-to-File letter, and a median 784 days from original submission to eventual approval for the applications that recovered [4]. The four sponsors who filed over protest, contesting the gap rather than fixing it, were never approved: all four [4]. Tibau's data on accelerated-approval cancer drugs is the sharpest single number in this whole piece. When a confirmatory trial was already running at approval, median time to withdrawal of the drugs that failed to confirm was 3.81 years; when no confirmatory trial was running, 7.31 years [20]. Later discovery, longer resolution, every single time.

One honest limit, in the regulator's own words. The EMA says of its own advice: "Complying with scientific advice therefore increases the chances of receiving marketing authorisation but it does not guarantee it" [21]. That is the ceiling on this entire framework, and I want it explicit. The audit surfaces named, addressable risk early. It does not predict approval. Anyone selling you a readiness score that promises the latter is selling you something.

So for every major design choice, can you name the specific agency advice it complies with? And where you departed, is the rationale written down before someone asks?

Isn't an evidence gap audit just IEP gap analysis with a scoring rubric bolted on?

Give the objection its full strength. A competent development team already knows its own weak spots. A formal five-axis audit is process theatre, rigour cosplay. A checklist that admits it cannot predict approval offers false comfort in exchange for real hours. If you are good, the argument runs, you do this in your head already.

Two answers. First, teams do not audit themselves evenly. They over-weight the axis they are strongest on and under-weight the one they have never been burned by: the immunologists grade the mechanism generously and forget the safety database is a year short. The value of a named five-axis pass is that it forces the uncomfortable axis into the room. Second, and this is the harder evidence, internal self-assessment drifts optimistic in a way you can measure. Lurie and colleagues compared what sponsors said publicly about their Complete Response Letters against what the FDA had actually written, and found the public accounts systematically omitted deficiencies the agency had cited [5]. Nobody is lying. It is the ordinary gravity of wanting your own programme to be ready, and it is precisely why the gut-call cannot be trusted. Grade the axis you are most confident about last, ideally with someone who did not build the programme grading it beside you.

The last-war checklist

Return to where we started. The documented pattern of what sinks applications is not fixed. The Sacks, Chahal and Lurie record put efficacy and safety at the centre [3][4][5]; the analyses of the newly disclosed CRL repository put manufacturing and CMC there instead [2]. Reasonable people can read that two ways. It might be a real shift in what fails, or it might be that manufacturing failures were always common and simply became visible the moment the letters went public. Probably it is some of both.

Either way, the operational lesson is identical. An audit built once, against last year's dominant failure mode, and then left on a shelf is what I would call a "last-war checklist": rigorously grading this year's programme against the pattern that sank the last generation of drugs. So the fix is a cadence: re-pull the currently disclosed failure patterns before each major milestone, and re-weight the axes against them, rather than trusting the version of the checklist you internalised three years ago.

Run the load test on your own timetable

Five axes, each with real criteria drawn from what has actually gone wrong for programmes like yours, not from a list of good intentions, graded honestly and early, and re-graded against current evidence rather than memory. That said, none of it works if you run it once and file it.

Run it by hand first, because that is the part that builds the judgement. Once you know the axes and can grade your own programme without flinching, the repetitive work of re-scoring each quarter against a refreshed pattern set is exactly the kind of thing worth systematising, and a platform like InovaSight, built to keep that evidence overview current without you re-running the whole evidence gap audit from memory, earns its place there, but only once the manual framework is understood. Never instead of it.

The gap you find yourself, early, is the cheap one. The gap a reviewer finds for you is not.


Get the monthly digest

The 5 things evidence leads need to know each month: regulatory moves, RWE developments and what they mean in practice. No pitch, one email a month.

References

  1. FDA. "Complete Response Letters" (transparency initiative). Initial release, July 2025: upwards of 200 letters, limited at that point to applications that were later approved. Second release, 4 September 2025: 89 letters tied to applications not yet approved. openFDA. https://open.fda.gov/
  2. The FDA Group. "Behind the Rejections: An Analysis of 89 FDA CRLs" (8 September 2025), analysing the FDA's September 2025 batch of 89 Complete Response Letters for not-yet-approved applications (manufacturing/facility-inspection issues in 56%, product-quality issues in 47%); and Jefferies equity-research analysis of the broader, still-growing CRL repository, reported via BioSpace. (Third-party industry analyses; directional, attributed.)
  3. Sacks LV, Shamsuddin HH, Yasinskaya YI, et al. (2014). "Scientific and regulatory reasons for delay and denial of FDA approval of initial applications for new drugs, 2000-2012." JAMA;311(4):378-384. PMID: 24449316. https://pubmed.ncbi.nlm.nih.gov/24449316/
  4. Chahal HS, Mukherjee S, Sigelman DW, Temple R. (2021). "Contents of US Food and Drug Administration Refuse-to-File Letters for New Drug Applications and Efficacy Supplements and Their Public Disclosure by Applicants." JAMA Internal Medicine;181(4). PMID: 33587091. https://pubmed.ncbi.nlm.nih.gov/33587091/
  5. Lurie P, Chahal HS, Sigelman DW, et al. (2015). "Comparison of content of FDA letters not approving applications for new drugs and associated public announcements from sponsors: cross sectional study." BMJ;350:h2758. PMID: 26063327. https://pubmed.ncbi.nlm.nih.gov/26063327/
  6. Tafuri G, Trotta F, Leufkens HGM, Pani L. (2013). "Disclosure of grounds of European withdrawn and refused applications: a step forward on regulatory transparency." British Journal of Clinical Pharmacology;75(4). PMID: 22891872. PMC3612734. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3612734/
  7. Sarepta Therapeutics. "Eteplirsen pivotal study (dystrophin surrogate, 12 patients)." ClinicalTrials.gov: NCT01396239. https://clinicaltrials.gov/study/NCT01396239
  8. Sarepta Therapeutics. "PROMOVI (eteplirsen confirmatory, functional endpoint, 109 patients)." ClinicalTrials.gov: NCT02255552. https://clinicaltrials.gov/study/NCT02255552
  9. Biogen. "EMERGE (aducanumab, 1,643 patients)." ClinicalTrials.gov: NCT02484547. https://clinicaltrials.gov/study/NCT02484547
  10. Biogen. "ENGAGE (aducanumab, 1,653 patients)." ClinicalTrials.gov: NCT02477800. https://clinicaltrials.gov/study/NCT02477800
  11. Biogen. "ENVISION (aducanumab confirmatory, 1,027 patients; terminated)." ClinicalTrials.gov: NCT05310071. https://clinicaltrials.gov/study/NCT05310071
  12. PTC Therapeutics. "Ataluren Phase 2b in DMD (174 patients; missed primary endpoint)." ClinicalTrials.gov: NCT00592553. https://clinicaltrials.gov/study/NCT00592553
  13. PTC Therapeutics. "ACT DMD (ataluren, 230 patients; stratified baseline-walking enrolment)." ClinicalTrials.gov: NCT01826487. https://clinicaltrials.gov/study/NCT01826487
  14. AMAG Pharmaceuticals. "PROLONG (hydroxyprogesterone caproate / Makena confirmatory)." ClinicalTrials.gov: NCT01004029. https://clinicaltrials.gov/study/NCT01004029
  15. Subramaniam D, Anderson-Smits C, Rubinstein R, Thai ST, Purcell R, Girman C. (2024). "A Framework for the Use and Likelihood of Regulatory Acceptance of Single-Arm Trials." Therapeutic Innovation & Regulatory Science;58(6). PMID: 39285061. https://pubmed.ncbi.nlm.nih.gov/39285061/
  16. Bouwman L, Leufkens H, Sepodes B, Torre C. (2026). "Safety population size and duration of exposure prior to approval of new medicines: A database analysis of medicines centralised approved in the European Union between 2011 and 2023." PLoS ONE;21(2). PMID: 41662378. https://pubmed.ncbi.nlm.nih.gov/41662378/
  17. Eisai Inc., in collaboration with the TIMI Study Group. "CAMELLIA-TIMI 61 (lorcaserin cardiovascular safety, 14,673 patients)." ClinicalTrials.gov: NCT02019264. https://clinicaltrials.gov/study/NCT02019264. Lorcaserin (Belviq) was originally developed and brought to approval in 2012 by Arena Pharmaceuticals, which later licensed US commercialisation to Eisai. FDA. "FDA requests the withdrawal of the weight-loss drug Belviq, Belviq XR (lorcaserin) from the market" (13 February 2020). https://www.fda.gov/drugs/drug-safety-communications/fda-requests-withdrawal-weight-loss-drug-belviq-belviq-xr-lorcaserin-market
  18. US Code of Federal Regulations. "21 CFR § 312.47 — Meetings (End-of-Phase-2)." Cornell Law School Legal Information Institute. https://www.law.cornell.edu/cfr/text/21/312.47
  19. Regnstrom J, Koenig F, Aronsson B, et al. (2010). "Factors associated with success of market authorisation applications for pharmaceutical drugs submitted to the European Medicines Agency." European Journal of Clinical Pharmacology;66(1):39-48. PMID: 19936724. https://pubmed.ncbi.nlm.nih.gov/19936724/
  20. Tibau A, Cliff ERS, Romano A, Borrell M, Molto C, Kesselheim AS. (2025). "Predictors of withdrawal of anticancer drug indications granted accelerated approval: a retrospective cohort study." EClinicalMedicine;84. PMID: 40687736. https://pubmed.ncbi.nlm.nih.gov/40687736/
  21. European Medicines Agency. "Scientific advice and protocol assistance." https://www.ema.europa.eu/en/human-regulatory-overview/research-development/scientific-advice-protocol-assistance

Share this post