Skip to content
A federated real-world data network turning a clinical trial feasibility count into a live design tool for trial eligibility criteria.
Strategy drug development

Feasibility in minutes: how a fast patient-count query became a design tool, and what it still can't do

Imi
Imi

Ask a drug developer how many real-world patients would match a candidate set of eligibility criteria, and until recently you were asking for a wait. A CRO feasibility survey. A round of manual, site-by-site chart-review requests. Weeks of latency before a number came back, and by then the criteria had usually hardened into something close to a final protocol.

Years ago I worked on an early-phase rare-disease programme where exactly that happened. The feasibility question went out as a formal survey to a handful of candidate sites. The answer came back the better part of two months later, by which point the eligibility criteria had been through legal, been through the SAP, and were three revisions deep. The count confirmed the design was tight. It just could not tell us so early enough to change anything.

Now ask the same question against a federated real-world data network, and the feasibility count lands before the meeting is over.

Two peer-reviewed, named-mechanism numbers make the shift concrete. On the ACT network, built on i2b2/SHRINE architecture, a single query returned counts from 21 sites and identified 8,383 patients across the network in five minutes (Visweswaran et al. 2018) [1]. On TriNetX, a basic patient-count query (how many patients carry an arbitrary set of criteria) resolved in under half a second across more than 100 million patients (Palchuk et al. 2023) [2].

Read quickly, that is a vendor speed brag. Minutes instead of weeks, faster is better, next slide.

Count the hours saved, though, and you have missed the more interesting change. A feasibility count that used to arrive weeks after the criteria were locked now arrives while they are still on the whiteboard. The answer moved to before the decision, so the decision can finally respond to it.

One scoping note, because the entire argument rests on it. The minutes belong to the counting step. Running a full, submittable real-world evidence study on that same network is a different and far longer job, and we will come back to exactly how much longer. The count is fast. The study is not.

The practical consequence for Monday is small and specific: you can run the count yourself, at the design phase, before anyone drafts the eligibility section.

The query travels; the data stays put

The mechanism doing the work has a plain name: a federated real-world data network. The query travels out to each participating site, runs against data already mapped to a common model, and returns a count. The patient-level data never leaves the site. TriNetX works this way (Palchuk et al. 2023) [2], as does the ACT/i2b2-SHRINE network (Visweswaran et al. 2018) [1]. So does the FDA's own Sentinel Initiative, which runs the identical "query moves, data doesn't" principle at the agency's hand [3] — though Sentinel is a post-market safety-surveillance system, not a pre-protocol design tool, and it is worth not blurring the two.

Set that against a traditional CRO feasibility survey and the reason for weeks-versus-seconds becomes obvious. The survey commissions new information. Sites get contacted, someone at each site interrogates their own records or, more often, estimates from memory, and the responses are chased and aggregated by hand. Every site is a coordination round-trip. The federated count skips all of it, because the counting has in effect already happened; the query just asks for the tally.

The guidance, for its part, lags the practice. The FDA's November 2020 final guidance on enhancing the diversity of trial populations mentions real-world data exactly once (Section III.B), and only as an aid to recruitment and site identification, not as an iterative design tool [4]. Its December 2025 successor exists and supersedes it [5], though I am relying on trade-press summaries for its title and status and have not read the guidance itself. The self-service, design-phase use of a feasibility count is running well ahead of the documents that describe it.

Practically, that means the capability now sits in the design team's hands, not behind a six-week CRO round-trip. That relocation is the whole reason iteration becomes possible.

Free download

The RWE Briefing Document Template

The section-by-section structure for the RWE part of a regulatory briefing, built around the questions reviewers actually ask.

Get the template →

Eligibility criteria as a dial you can watch move

This is the part that changes the work rather than the calendar.

When the answer took weeks, eligibility criteria were something you committed to and then hoped about. The question in the room was "we'll find out in six weeks whether this protocol is enrollable." When the answer takes seconds, the question becomes "watch the count move as we loosen this one criterion." You are no longer designing the eligibility section with the lights off and flipping the switch only once the room is built.

The platform methodology papers make this claim directly, not as marketing copy. TriNetX's own peer-reviewed methods paper states that the network gives industry users rapid feedback on the impact of individual inclusion and exclusion criteria on cohort size at the trial-design phase, enabling real-time iteration and optimisation of the protocol before its release, with the potential for less need for time-consuming and costly amendments later (Palchuk et al. 2023) [2].

How much can a single criterion move a count? A retrospective, hypothetical simulation of Alzheimer's disease trial criteria against the OneFlorida database gives a sense of magnitude, and it needs its caveats stated plainly, because it is a peer-reviewed AMIA Annual Symposium proceedings paper modelling a hypothetical protocol, not a live sponsor decision and not a turnaround-time claim [6]. In that simulation, the eligible population grew from 373 patients under the original criteria to 503 when a cardiac-disease exclusion was dropped, to 603 when depression was also dropped, to 744 when uncontrolled hypertension was also dropped, and to 1,856 when all sixteen exclusion criteria were removed. Roughly a fivefold swing, from one disease area's criteria alone. Illustrative of scale, nothing more.

Now the discipline, because this is where the argument is easy to overrun. The lesson is not "loosen everything." Whether fewer criteria improve accrual is not a settled finding at all; one analysis of oncology trials found no association between the raw number of eligibility criteria and accrual success (Schroen et al. 2010) [7]. What is well supported is the narrower, harder claim: removing specific, named, high-impact criteria increases the eligible count. That is the ASCO and Friends of Cancer Research lineage. The 2017 consensus statement on broadening eligibility (Kim et al. 2017) [8]. The finding that removing recommended comorbidity restrictions could add up to 6,317 patient registrations a year (Unger et al. 2019) [9]. The more recent result that 13.6% of screen failures in pancreatic and biliary cancer trials fell on potentially modifiable criteria (Saj et al. 2026) [10]. The dial is worth turning. Which dial you turn is the entire skill.

There is even a formal analogue already running at scale, and it deserves naming so it is not confused with the fast version. NCT06314542, a study with AstraZeneca as collaborator, quantifies criterion by criterion how relaxing each NSCLC eligibility rule changes the eligible patient count, across a 50,000-patient real-world database [11]. That is exactly the analysis a self-service query performs, except it is retrospective formal research done after the protocols already existed, not a live count you run while the criteria are still soft. The self-service count does this on the fly, at the design phase. The trial does it properly, after the fact.

Monday version: put the count inside the protocol-design meeting. Test criteria variants live, so the eligibility section is written to a population you have already seen, not one you are hoping exists.

Why this matters when you have one shot

Keep this in proportion. Trials fail on recruitment often enough that the base rates are worth stating, and no more than stating.

Across a cohort of 1,017 randomised trials, 9.9% were discontinued specifically for poor recruitment, around 40% of all discontinuations in that sample (Kasenda et al. 2014) [12]. Among prematurely terminated cardiovascular trials, 53.6% cited lower-than-expected recruitment (Bernardez-Pereira et al.) [13]. Later cohorts from the same continuing research programme as Kasenda put poor recruitment at 37% and 45.4% of discontinuations respectively (Speich et al. 2022; Speich et al. 2025), though those papers share a consortium and a protocol and should be read as one programme, not three independent replications [14][15].

Here is the honest limit on all of it, though. None of this proves that skipping a feasibility count causes these failures. The closest evidence, an analysis of whether a documented pretrial accrual assessment predicted sufficient accrual, was not statistically significant (p=0.08) in a sample of just 82 oncology trials (Schroen et al. 2010) [7]. So the recruitment base rates are real and well documented; the causal chain from "no feasibility count" to "trial dies" is not. Treat the fast count as cheap insurance against one avoidable slice of recruitment risk, not as a proven cure for it. For a biotech with one or two assets and a finite runway, that slice is still worth buying, because discovering an unenrollable design after protocol lock is the expensive kind of mistake.

Speed is not fitness: the count is not the cohort

Now the strongest objection, stated at full strength. If a patient-count query answers the population question in seconds, why bother running a slow, expensive real-world evidence study at all? Take the count as the evidence and move on.

That said, the fast count and the regulatory-grade study are different operations that happen to share an architecture, and the gap between them is measured in months. EMA's fourth DARWIN EU experience report describes the same federated, data-never-moves model at genuinely regulatory scale: more than 250 million patients, 40 data partners, 18 countries [16]. And yet a full study on it ran a median of 4.8 months (IQR 4 to 6) from protocol approval to results, with the report noting that some regulatory requesters need answers in weeks, which the current model cannot meet [16]. That is the same network as the minutes-fast count, doing the submittable job, taking two-thirds of a year.

The report draws the exact line this post draws. Feasibility assessment sits as a distinct, faster category inside that pipeline, and it has grown from 4% to 17% of DARWIN EU's output year-on-year, precisely so that teams can clarify population size, exposure and outcome frequency before committing to more resource-intensive studies [16]. Its feasibility success rate, whether the question could be answered with the available data at all, was 77% of the 108 topics assessed [16]. Fast feasibility work and slow evidence generation are two jobs, and the regulators themselves file them separately.

This is where a specific failure mode lives, and it earns a name: the count-cohort gap. A federated feasibility count is a tally of matching records in EHR and claims data. It is not a consented, individually verified, enrollable cohort. The PCORnet opioid surveillance demonstration returned an aggregate count of 15,438,284 matching records across 24 sites (NCT03743493), plainly a query result and not fifteen million enrolled patients [17]. And matching on paper is not the same as being reachable and enrolled: the SAGE study set out to test whether eligibility criteria explain under-representation in trials and found they explain only part of it (NCT01754636) [18]. A feasibility count cannot see the rest.

So the count sizes the opportunity; it does not validate the evidence. What a submittable, regulatory-grade study actually demands is a separate discipline, one this blog has covered in how to take RWE to the regulators. If the cohort is destined to become an external control, the bar is higher still, and the FDA's external-control-arm guidance checklist is the place to start. Use the fast count to design. Never use it to skip.

What the speed is actually for

The speed lives in the counting step, and nowhere else. Used well, it means you walk into protocol lock having already watched which criteria your real-world population can bear, rather than discovering it weeks later, when changing a criterion means a formal amendment. The count buys you a better-designed trial. The cohort still costs you months.

That is the honest shape of it. Bringing a fast feasibility count into the room where eligibility criteria actually get written, which is one of the things InovaCS is built to support, is a cheap way to stop hoping about a population you could simply have looked at. The looking is quick now. The evidence still has to be earned.

Get the monthly digest

The 5 things evidence leads need to know each month: regulatory moves, RWE developments and what they mean in practice. No pitch, one email a month.

References

[1] Visweswaran S, et al. (2018). "Accrual to Clinical Trials (ACT): A Clinical and Translational Science Award Consortium Network." JAMIA Open;1(2):147-152. PMID: 30474072. https://pubmed.ncbi.nlm.nih.gov/30474072/

[2] Palchuk MB, et al. (2023). "A global federated real-world data and analytics platform for research." JAMIA Open;6(2):ooad035. PMID: 37193038. https://pubmed.ncbi.nlm.nih.gov/37193038/

[3] FDA Sentinel Initiative. Program documentation. https://www.sentinelinitiative.org/

[4] FDA. "Enhancing the Diversity of Clinical Trial Populations — Eligibility Criteria, Enrollment Practices, and Trial Designs." Guidance for Industry (FINAL), November 2020. https://www.fda.gov/regulatory-information/search-fda-guidance-documents/enhancing-diversity-clinical-trial-populations-eligibility-criteria-enrollment-practices-and-trial

[5] FDA. "Enhancing Participation in Clinical Trials — Eligibility Criteria, Enrollment Practices, and Trial Designs." Guidance for Industry (FINAL), December 2025. Title and status confirmed via trade-press summaries; substantive content not independently read.

[6] Li Q, Guo Y, He Z, Zhang H, George TJ, Bian J. (2020). "Using Real-World Data to Rationalize Clinical Trials Eligibility Criteria Design: A Case Study of Alzheimer's Disease Trials." AMIA Annu Symp Proc;2020:717-726. PMID: 33936446; PMCID: PMC8075542. https://pubmed.ncbi.nlm.nih.gov/33936446/

[7] Schroen AT, et al. (2010). "Preliminary evaluation of factors associated with premature trial closure and feasibility of accrual benchmarks in phase III oncology trials." Clin Trials;7(4):312-321. PMID: 20595245. https://pubmed.ncbi.nlm.nih.gov/20595245/

[8] Kim ES, et al. (2017). "Broadening eligibility criteria to make clinical trials more representative: ASCO and Friends of Cancer Research joint research statement." J Clin Oncol;35(33):3737-3744. PMID: 28968170. https://pubmed.ncbi.nlm.nih.gov/28968170/

[9] Unger JM, et al. (2019). "Association of Patient Comorbid Conditions With Cancer Clinical Trial Participation." JAMA Oncol;5(3):326-333. PMID: 30629092. https://pubmed.ncbi.nlm.nih.gov/30629092/

[10] Saj F, et al. (2026). "Broadening the gates: Analysis of potentially modifiable study entry criteria in pancreatic and biliary tract cancer trials." Cancer;132(1):e70227. PMID: 41417596. https://pubmed.ncbi.nlm.nih.gov/41417596/

[11] Cancer Institute/Hospital, Chinese Academy of Medical Sciences (AstraZeneca collaborator). Criterion-by-criterion analysis of NSCLC eligibility relaxation on a real-world database. ClinicalTrials.gov: NCT06314542. https://clinicaltrials.gov/study/NCT06314542

[12] Kasenda B, et al. (2014). "Prevalence, characteristics, and publication of discontinued randomized trials." JAMA;311(10):1045-1051. PMID: 24618966. https://pubmed.ncbi.nlm.nih.gov/24618966/

[13] Bernardez-Pereira S, et al. (2014). "Prevalence, characteristics, and predictors of early termination of cardiovascular clinical trials due to low recruitment: insights from the ClinicalTrials.gov registry." Am Heart J;168(2):213-219.e1. PMID: 25066561. https://pubmed.ncbi.nlm.nih.gov/25066561/

[14] Speich B, et al. (2022). "Nonregistration, discontinuation, and nonpublication of randomized trials: A repeated metaresearch analysis." PLoS Med;19(4):e1003980. PMID: 35476675. https://pubmed.ncbi.nlm.nih.gov/35476675/

[15] Speich B, et al. (2025). "Nonregistration, discontinuation, and nonpublication of randomized trials: A systematic review." JAMA Netw Open;8(9):e2524440. PMID: 40899933. https://pubmed.ncbi.nlm.nih.gov/40899933/

[16] EMA/HMA. "Real-world evidence framework to support EU regulatory decision-making: 4th report on the experience gained with regulator-led studies from February 2025 to February 2026." EMA/66577/2026. https://www.ema.europa.eu/en/documents/report/real-world-evidence-framework-support-eu-regulatory-decision-making-4th-report-experience-gained-regulator-led-studies-february-2025-february-2026_en.pdf

[17] PCORnet Opioid Surveillance Demonstration. Feasibility examination; aggregate query count across 24 sites. ClinicalTrials.gov: NCT03743493. https://clinicaltrials.gov/study/NCT03743493

[18] Assistance Publique - Hôpitaux de Paris. "SAGE — under-representation in clinical trials and the role of eligibility criteria." ClinicalTrials.gov: NCT01754636. https://clinicaltrials.gov/study/NCT01754636

Share this post