The evidence roadmap: sequencing studies by what other studies structurally need
Your ranked gap list tells you what matters. It does not tell you what to run first.
If you have done the prioritisation work, whether that was a proper urgency/impact pass or the 90/10 filter applied to your evidence gaps, you already know which questions are worth spending on, and that was genuinely the hard part that plenty of teams never reach. But a priority order and a running order are two different lists, and the gap between them is where a surprising amount of scarce budget quietly dies.
Here is the move almost everyone makes without noticing it. They finish the ranking, look at the top of the list, and start the study sitting there, on the logic that highest priority ought to run first. Or they start whichever study is administratively easiest to stand up: the one with a CRO already lined up, the protocol half-written, the vendor who can begin next month. That is sequencing by convenience, dressed up as sequencing by priority.
Neither of those approaches asks the question that actually sets the order, which is whether this study's inputs even exist yet.
Because some studies structurally depend on others, in the sense that the later one consumes the earlier one's output as a design input. Run them out of order and you lose more than time. You spend the expensive study's budget generating an answer you cannot use, because the input it needed to be designed properly was not there when you designed it.
This post is about that ordering problem, evidence generation sequencing: given a fixed budget and a set of studies that cannot all run at once, what order the funded studies physically run in. It assumes you already have the plan, an IEP or a ranked gap list or the lean version you build when resources are tight, and it picks up exactly where prioritisation stops.
What does "dependency" actually mean when you're generating evidence?
Let me be clear about what I am not saying, because "sequence by dependency" can sound like a project manager importing a Gantt chart into a place it does not belong. This is not that.
The dependencies here are structural rather than administrative, because the later study genuinely cannot be designed credibly until the earlier one has produced something specific:
- Data-source feasibility comes before a matched external control arm. You cannot design a defensible ECA against a real-world data source you have not confirmed actually captures your endpoint, at the granularity and at the time-points your analysis needs. Gray et al. set this out as a four-step exchangeability process in which data-source feasibility sits before comparator-study design, not alongside it [1]. The FDA's draft ECA guidance rests on the same prerequisite, which we have broken down into a checklist for drug developers.
- Population characterisation comes before your effect-size assumption. If you do not know how the untreated disease actually behaves over time, then your powering is a guess wearing a calculation's clothes.
- Endpoint validation comes before the pivotal that rests on it. An endpoint nobody has shown to move with the disease is not really an endpoint yet, only a hypothesis about one.
Call the study at the bottom of one of these chains the "load-bearing study": it has no unmet inputs of its own, and other studies rest structurally on its output. Pour the upper storey before the foundation has cured, and it will not matter how carefully you pour it.
So for every study in your plan, write down what it consumes that must already exist: an endpoint definition, an eligibility band, a validated data source. Then look hard at each input. If any of them is itself the output of a study you have not run, that study cannot run yet, whatever its priority rank.
The cheap study that wrote the expensive study's protocol
The ataluren programme shows dependency done right, with the numbers publicly on the record.
Ataluren's Phase 2b in Duchenne muscular dystrophy (NCT00592553, 174 patients, 2008-2009) did something quietly essential [2]. It identified the baseline six-minute walk distance band, a 6MWD of at least 150 metres and no more than 80% of predicted for the patient's age and height, in which a drug-placebo separation was actually detectable. Outside that band, the walk test was simply too noisy to show anything real.
That finding went on to define the next trial. ACT DMD, the Phase 3 (NCT01826487, 230 patients, 2013-2015), restricted enrolment to exactly that ambulatory band [3] — roughly four years and two studies later, and the second one's entry criteria did not exist until the first one had produced them.
The Phase 2b was the load-bearing study, so it ran first. Nobody ran it for glamour; it was hardly the flashiest line on anyone's gap list. It ran first because the Phase 3 was uninterpretable without it. Rank those two studies by urgency and you would want the pivotal first, since it is the one that gets you approved, yet dependency says you cannot have it first. Your pivotal's eligibility band, or its primary endpoint definition, very often does not exist yet, and the study that produces it is your load-bearing study wherever it happens to sit on the priority list.
Free download
The IEP Template Pack
The gap matrix, prioritisation grid and plan-on-a-page we use to build integrated evidence plans. Free to keep.
Get the template pack →How often does the prerequisite get skipped?
If this were obvious in practice then the post would not need writing, and it is not obvious in practice.
Liu et al., in a 2025 audit in JAMA Network Open, went through 180 externally controlled trials published between 2010 and 2023 and checked what they had actually done [4]. Only 7.8% conducted a feasibility assessment of the data source before using it, only 16.1% had prespecified the use of external controls at all, and just 35.6% gave any stated reason for using them.
Read that first number again. Fewer than one in twelve externally controlled trials checked whether the data source was fit for the job before building on it, which means the prerequisite that everyone agrees is necessary gets skipped roughly nine times in ten.
That said, the feasibility step is no unsolved problem. Gray et al. describe exactly how to do it, and a well-resourced DMD external-control programme shows it done, and done well [5]. The knowledge plainly exists. What is missing is the discipline to run the load-bearing study first, because the expensive study is the one that feels like progress.
I watched this happen on an early-phase rare-disease programme, where the team had committed budget to a matched external control arm before anyone had confirmed, at the granularity the analysis needed, whether the registry captured the endpoint on a schedule that matched our assessment windows. It did not. The endpoint was in there and recorded, but on a cadence that missed the time-points that mattered, and coarsely enough to wash out the very effect the trial was powered to see. The team learned this after the money had been allocated, and the feasibility check that would have caught it would have cost a rounding error against the ECA it was meant to support. The order was the failure.
That is the general rule the ataluren chain and the Liu numbers point at from opposite directions. The cheap feasibility or landscape check is almost always the load-bearing study, and almost always the one that gets run last, if it gets run at all. Run it before you commit the expensive study's budget, not after.
One bet, placed twice: what no sequencing looks like
There is a failure mode worse than running things in the wrong order, and that is running the expensive things in no order at all, in parallel, with nothing positioned between them to tell you what a split result would mean.
Consider EMERGE (NCT02484547) and ENGAGE (NCT02477800), Biogen's two Phase 3 aducanumab trials [6][7]. These were two near-identically designed studies, started about a month apart and run side by side with no discriminating study between them, and both were halted for futility in March 2019. When a later reanalysis had one trial reading positive and the other flat, there was no instrument built to adjudicate the disagreement, and the years of argument that followed are well documented. Two very expensive shots were fired at one underlying bet, with nothing designed to explain a divergence if it came.
Contrast the EXPEDITION programme in Alzheimer's. EXPEDITION (NCT00905372, 1,000 patients) and EXPEDITION 2 (NCT00904683, 1,040 patients) both enrolled broad mild-to-moderate populations, and both missed their primary endpoint [8][9]. A post-hoc signal in the milder patients then shaped EXPEDITION 3 (NCT01900665, 2,129 patients), which restricted to biomarker-confirmed, mild-AD-only patients about a year later [10]. The population refinement that ideally comes before a Phase 3 arrived instead as the very expensive by-product of two Phase 3s that had already read out.
Here is the strongest objection to all of this, parallel is faster, and for a biotech burning cash, speed is survival, so sequencing adds calendar time you may not have. Two Phase 3s at once could have reached approval a year sooner than one-then-the-other, if they had hit, so is insisting on order not simply a luxury of the well-funded? No, and the distinction is precise. Parallelism is right when the studies are genuinely independent, meaning true shots on goal where neither one's design should be informed by the other's result, and you should absolutely run those at once. It turns to waste only when one study's design should have consumed the other's output and you ran them together so that it could not. The aducanumab pair was never two independent shots; Biogen effectively placed one bet twice, with nothing built to explain a disagreement, which is precisely the disagreement that arrived. Speed that front-loads spend before the discriminating information exists is not really speed. It buys you an expensive route to ambiguity.
The load-bearing study the regulator keeps flagging
One load-bearing study is so routinely under-sequenced that regulators themselves keep pointing at it, and that is natural history.
In rare disease especially, natural history sits upstream of nearly everything, because without it you cannot credibly set an endpoint, justify an effect size, or build an external control. And yet it is the classic deprioritised study, slow and unglamorous and easy to wave off as "not directly supporting the filing", right up until it becomes the thing holding the filing up.
The regulators have noticed. The FDA's guidance on rare-disease natural history studies has been sitting in draft since March 2019, and a November 2024 GAO report (GAO-25-106774) confirmed it was still draft, roughly six years unfinalised [11]. The six-year draft status is the hard fact here, and it holds because GAO confirms it directly.
The operational read is simple. If your programme needs natural-history data, and in rare disease it almost certainly does, then that study is almost certainly on your critical path and almost certainly under-sequenced. Surface it now, while it is a planning item, rather than later, when it is the reason your pivotal slips.
Draw the arrows before you draw the Gantt chart
So, here is the Monday-morning version. Before you set a single start date:
- List inputs, not just studies. For every planned study, write down what must already exist for it to be designed and run credibly: the endpoint definition, the eligibility band, the feasible data source, the natural-history comparator. Write down the inputs, not the output.
- Draw the arrows. Mark which study produces each input, and any study whose inputs come from a study you have not run is downstream, full stop. It cannot be your opening move, however high it ranks.
- Find the load-bearing studies. These are the ones with no unmet inputs of their own that unblock the most downstream work, and they run first, even the cheap unglamorous ones, even the ones sitting low on urgency/impact.
- Only now overlay budget and calendar. Dependency sets the order, and convenience and urgency become tie-breakers between studies at the same dependency level, rather than the primary sort key.
- Run the cheap feasibility or landscape check first. It is almost always load-bearing and almost always skipped, as the 7.8% shows.
The boundary is the whole point, so it is worth drawing hard. Priority decides which gaps you fund, which is the job your ranked list and your 90/10 pass already did, while dependency decides the order the funded studies can physically run in. Those are two different lists, and the people who conflate them are the ones drawing the Gantt chart before they have drawn the arrows.
The column almost no evidence plan has
Most evidence plans I see carry a priority column and a start-date column, and almost none carry a dependency column, which would be a plain field next to each study answering one question: what does this study consume, and does it exist yet?
That is the column I would add to every plan. It is not sophisticated and it does not need software — it needs someone to write, next to the expensive study, the name of the cheaper study that has to finish first, and to notice when that cell is empty.
At Inovia we spend as much time on that ordering problem as on the ranking one, because a well-prioritised plan sequenced in the wrong order still burns the budget just as effectively. If your gap list is sound but your running order was set by whatever happened to be easiest to start, that is the thing worth fixing before the next study opens. That is where the quiet money goes.
Get the monthly digest
The 5 things evidence leads need to know each month: regulatory moves, RWE developments and what they mean in practice. No pitch, one email a month.
References
[1] Gray et al. Four-step exchangeability framework for external comparators. PMID: 32440847. https://pubmed.ncbi.nlm.nih.gov/32440847/
[2] Ataluren Phase 2b in Duchenne muscular dystrophy (174 patients, 2008-2009). ClinicalTrials.gov: NCT00592553. https://clinicaltrials.gov/study/NCT00592553
[3] ACT DMD, ataluren Phase 3 in Duchenne muscular dystrophy (230 patients, 2013-2015). ClinicalTrials.gov: NCT01826487. https://clinicaltrials.gov/study/NCT01826487
[4] Liu et al. (2025). Methodological audit of 180 externally controlled trials, 2010-2023. JAMA Network Open. PMID: 40906478. https://pubmed.ncbi.nlm.nih.gov/40906478/
[5] Goemans et al. DMD external-control feasibility. PMID: 32611643. https://pubmed.ncbi.nlm.nih.gov/32611643/
[6] EMERGE, Biogen aducanumab Phase 3. ClinicalTrials.gov: NCT02484547. https://clinicaltrials.gov/study/NCT02484547
[7] ENGAGE, Biogen aducanumab Phase 3. ClinicalTrials.gov: NCT02477800. https://clinicaltrials.gov/study/NCT02477800
[8] EXPEDITION, Alzheimer's disease Phase 3 (1,000 patients). ClinicalTrials.gov: NCT00905372. https://clinicaltrials.gov/study/NCT00905372
[9] EXPEDITION 2, Alzheimer's disease Phase 3 (1,040 patients). ClinicalTrials.gov: NCT00904683. https://clinicaltrials.gov/study/NCT00904683
[10] EXPEDITION 3, Alzheimer's disease Phase 3 (2,129 patients). ClinicalTrials.gov: NCT01900665. https://clinicaltrials.gov/study/NCT01900665
[11] U.S. Government Accountability Office. Report GAO-25-106774 (November 2024), confirming the FDA's rare-disease natural history studies guidance (draft since March 2019) remains non-final. https://www.gao.gov/products/GAO-25-106774
-1.png?width=1169&height=277&name=Inovia%20Logo%20Dark%20(1)-1.png)