Skip to content
Atlas comparing real-world data source selection across claims, EHR, registries and wearables for regulatory-grade evidence
RWE Strategy Real world evidence

The RWD source atlas: choose your data source by its blind spot, not its brochure

Imi
Imi

The budget finally lands. After months of building the internal case, you have money to license a real-world dataset, and a sensible instinct takes over: find the biggest, cleanest, most-cited source you can afford, and buy the impressive one. Understandable. It is also the wrong first move, and it is expensive to get wrong.

Fitness lives in the pairing of dataset and question, not in the dataset alone. That is the whole of real-world data source selection: a source that is gold for sizing a disease can be worthless for an efficacy endpoint, and the ranking flips the instant the question changes.

So the "best" source is a category error.

Many years ago I worked on an early-phase rare-disease programme where the team licensed a large, well-regarded claims extract because it was available, affordable and already cleaned. It sized the population beautifully. Then the question moved to a treatment response defined by a laboratory value, and that value simply was not in the data, because nobody bills for it. The money was already spent. The gap was found last, when it should have been found first.

Call it the disqualifying gap: for any given question, every source carries one blind spot that, if it lands on the exact variable your endpoint depends on, disqualifies the source no matter how strong everything else about it is. The brochure sells you the strengths. Only the disqualifying gap tells you whether the source can actually answer your question. (We have made the case before that fitness for purpose is about process, not the brand of the dataset; see what regulatory-grade RWE really means. This is the source-choice half of that same discipline, and licensing the wrong source is one of the most avoidable ways to pile risk onto a lean programme.)

The rule before the tour

What actually decides whether a real-world data source is fit for regulatory use? Regulators judge real-world data on two axes: reliability and relevancy (the SPIFD framing, Gatto et al. 2022 [1]). Reliability you can mostly buy back with process. Relevancy you cannot, because it is decided far down at the level of the single variable you need, and that is where the disqualifying gap lives. The headline strength of a source tells you what it is good at. It tells you nothing about whether it can see the one thing you came for.

Think of each source as a witness you might one day put in front of the FDA or a payer to support a regulatory submission. The job is to know, before you call one, exactly what it cannot testify to. Not to call the most impressive witness.

And "assess the source" is now the regulator's stated expectation, not a courtesy. The FDA has finalised dedicated guidance on assessing EHR and medical claims data (FINAL, July 2024 [2]), on assessing registries (FINAL, December 2023 [3]), on digital health technologies for remote data acquisition (FINAL, December 2023 [4]), and on RWD/RWE considerations overall (FINAL, August 2023 [5]); the EMA's registry-based-studies guideline has been final since October 2021 [6]. Two agencies, dedicated documents, one message: characterise what your source cannot do before you rely on it.

Let's be honest about what you are actually buying. Four witnesses, four blind spots.

Claims: the accountant who logged every transaction but never examined the patient

What administrative claims data testifies to is scale and continuity. Because a claim is a payment record, the denominator is clean and enrolment is continuous, which makes claims the natural home for disease sizing, drug-utilisation patterns and safety surveillance at population scale. They can even reach effectiveness when the trial is emulable. In the RCT-DUPLICATE programme, a claims-based emulation reproduced the trial's regulatory conclusion in 6 of 10 cases and put the real-world hazard ratio inside the RCT's 95% confidence interval in 8 of 10 (Franklin et al. 2021 [7]); across 32 emulations, agreement reached a Pearson correlation of 0.82 (Wang et al. 2023 [8]). Read Wang's own caveat, though: the concordance held only "when design and measurements can be closely emulated, but this may be difficult to achieve."

Where claims go silent is everything clinical. The code exists to justify a bill, not to describe a patient. Claims lack "reasons for treatment selection … results of diagnostic testing, vital statistics … and symptomatology" (Stamas et al. 2024 [9]). Ask a billing code to identify a disease on its own and it can badly miss: for pulmonary arterial hypertension, ICD codes alone returned a positive predictive value anywhere from 3.3% to 66.7% (Gillmeyer et al. 2019 [10])  a PAH-specific figure, but a sobering one. Continuity has an edge too. Up to 22% of US commercial members disenrol each year, and imposing a two-year continuous-enrolment rule cost one cohort 58% of its patients (Stamas et al. 2024 [9]). Even Sentinel, the largest US distributed claims-and-EHR safety network, runs into the wall: its own literature flags an inability to identify many conditions of interest to a satisfactory level of accuracy (Brown et al. 2020 [11]).

Makes or breaks: buy claims for sizing, utilisation, hard events they can see (a hospitalisation, a dispensing) and safety signals. Do not buy claims to power an endpoint that lives in a lab value, a symptom score, disease severity or an in-hospital medication. The FDA's July 2024 guidance [2] is, in effect, a checklist for interrogating this witness before you trust it.

That endpoint is exactly what the accountant never wrote down.

Free download

The RWE Briefing Document Template

The section-by-section structure for the RWE part of a regulatory briefing, built around the questions reviewers actually ask.

Get the template →

EHR: the treating clinician who knows the patient intimately, but only inside the building

The electronic health record is the opposite witness. It carries the clinical depth claims cannot: labs, vitals, clinician notes, biomarkers, staging, severity. It is the raw material for oncology natural history and for most EHR-derived external controls.

Its disqualifying gap is the front door. One health system sees only the care delivered inside it, and patients get care elsewhere. Merola et al. 2022 put a number on it: a mean continuity ratio of roughly 0.18 in the general population and about 0.45 even in oncology [12]. A single EHR routinely sees a minority of a patient's encounters, and "EHR discontinuity (missing out-of-network encounters) can lead to information bias." Watch what that does to an endpoint. Even purpose-built oncology EHR data captured only about two-thirds of deaths in its structured fields; mortality sensitivity climbed from 66% to 91% only after amalgamating EHR with commercial and Social Security death data plus manual abstraction (Curtis et al. 2018 [13]). Trustworthy? Only if the deaths are actually sitting in the fields you read. Depth has holes of its own: across 55 profiled paediatric databases, just 22% held any genetics or biomarker data (Wharton et al. 2024 [14]). That said, a representative, primary-care-anchored resource such as CPRD softens the catchment problem, though its own profile concedes that positive predictive value tends to run high while sensitivity runs lower, and secondary-care detail can be incomplete (Herrett et al. 2015 [15]).

Makes or breaks: EHR is your source for clinical characterisation, natural history and endpoint definition where the network is complete. It breaks on any event that walks out of the door: an out-of-network death, a progression scanned elsewhere, the total cost of care. Before you trust an EHR endpoint, ask what fraction of the patient's care the system actually sees.

Registries: the hand-picked panel, and who got picked shapes the testimony

When is a registry actually the strongest real-world data source to build on? A disease or product registry is curated on purpose, and that curation is the strength. Registries are the natural-history backbone of rare disease and the usual base for an external control. The exhibits are real. Ultragenyx built the GNEM-DMP natural-history registry (NCT01784679, 319 patients enrolled) explicitly to "identify and validate biomarkers … inform … design and interpretation of clinical studies"; the UK's RaDaR rare-kidney registry (NCT06065852) is targeting roughly 35,000 participants.

The disqualifying gap has two parts: who got in, and what got recorded. Selection bias first. When Jarada et al. 2023 compared a limited-catchment cohort with a population-based one, the catchment cohort overstated systemic-therapy use (67.4% versus 40.8%) and reported a median survival 2.07 times greater, because "non-referral is an important source of selection bias" [16]. That is the Alberta case, and the magnitudes are illustrative rather than universal, but the mechanism is not optional. Then the variables. A registry records what its original clinical purpose required, which is rarely what a comparative-effectiveness or HTA analysis needs: comorbidity, concomitant therapy, PROMs, resource use and safety fields were "lacking partly or completely in all evaluated registries" in one recent review (Wilpshaar et al. 2026 [17]), which is exactly the confounder set an external control depends on.

This is the one source family where both agencies now hand you a dedicated rulebook: the FDA's registry guidance (FINAL, December 2023 [3]) and the EMA's registry-based-studies guideline (EMA/426390/2021, FINAL, October 2021 [6]), and the FDA document explicitly anticipates linkage to other sources. If you are building an external control on a registry, that is precisely the source-selection decision our external-control-arm checklist walks through.

Makes or breaks: registries are the go-to for natural history and external controls in rare disease. They break when the catchment is referral-skewed, or when the confounders you must adjust for were never collected. You cannot adjust for what nobody wrote down. Audit the entry route and the variable list against your confounder set before you build anything on top.

Wearables and sensors: the witness that never blinks, as long as someone keeps it worn

Can a wearable device actually carry a primary regulatory endpoint? Start with the proof, because it is genuinely new. In July 2023 the EMA qualified stride velocity 95th centile (SV95C), captured from a wearable, as a primary endpoint for Duchenne muscular dystrophy trials  "the first digitally derived outcome measure to receive formal regulatory qualification as a primary endpoint … of any indication" (Servais et al. 2024 [18]). The analytical validity was strong: 98.7% of strides detected against motion capture.

So a sensor can, in principle, reach regulatory-grade.

Now the same paper's disqualifying gap. That qualification needed a minimum wear rule of 50 hours, because below it "intra-patient variability increased dramatically"; and the reassuring compliance was self-flagged as a best case, "positively affected by well-trained specialists and exclusion of patients who refused to wear the device." Do not carry that adherence into a pragmatic or decentralised setting. But the deeper trap is validation. Fit-for-purpose here means three separate tests: verification, analytical validation and clinical validation (the DiMe V3 framework, Goldsack et al. 2020 [19]). A consumer device can be verified, meaning the sensor works, while never being clinically validated for the endpoint you actually want to claim. That gap is what the FDA's digital-health-technologies guidance (FINAL, December 2023 [4]) is built around.

Look at where wearable endpoints sit in real protocols. In EMBARK (NCT05096221, Sarepta and Roche, Phase 3, 126 patients), SV95C from an ankle-worn device is registered as a secondary endpoint while clinician-assessed NSAA stays primary. ActiLiège Next (NCT05982119, SYSNAV) is a dedicated study to validate the wearable endpoint itself. The most mature case, continuous-glucose-monitor time-in-range, is a primary endpoint in a Phase 2 insulin study (Lilly's LY900014, NCT04585776) in a common disease, with a sensor that has years of validation behind it. Note what is missing from that list: no wearable measure has yet carried a pivotal efficacy claim to approval on its own. In every exhibit it is a secondary, or the thing being validated.

Makes or breaks: wearables are unmatched for continuous, objective, real-life signal, and they are the only witness that can manufacture an entirely new endpoint. They break when the device is merely verified rather than clinically validated for your claim, or when your real-world population will not keep the thing charged and worn. Before you register a digital endpoint, know which of the three validation tiers you can actually evidence.

The atlas: choosing a real-world data source by what each witness can and cannot support

Read this grid down your column, not across a source's row. Your column is the question you actually have; the row is only the witness. Every cell is a verdict on fitness, and a strength in one cell tells you nothing about the cell directly below it.

Evidence purpose Claims EHR Registries Wearables
Disease sizing / epidemiology Strong (clean denominators) Biased by network catchment Only within catchment No
Natural history Coded events only, no severity Strong if network complete The backbone, if catchment fair Continuous signal if worn
External control Only events it can see Strong in oncology, in-network Canonical base, if entry route sound Not on its own
Novel / digital endpoint No Labs and biomarkers only Only what was collected Yes, with full validation
Safety surveillance Strong at scale, weak on many outcomes Deep but discontinuous Often missing safety fields Signal only
Effectiveness Only when the trial is emulable Strong when endpoint in-network Confounder-limited Secondary endpoint, so far
HTA / payer Resource use yes, outcomes no Partial Resource use and PROMs often absent Emerging

"Then just link everything" and the honest answer

Two fair objections deserve their strongest form. First: if every source has a blind spot, tokenise claims to EHR to registry and let the gaps cancel out. Second, sharper: registries are simply the gold standard for rare disease, so stop overcomplicating a solved problem. Both carry real weight. Linkage does close gaps; the Flatiron mortality figure climbing from 66% to 91% is itself a linkage result [13], and privacy-preserving tokenised matching can reach a precision of 97.0% and recall of 95.5% when several identifiers are available (Tyagi and Willis 2025 [20]). Registries are the natural-history engine of rare disease, and some are now built for triangulation from the start: MPN PROGRESSion (NCT07362225) deliberately fuses EHR, claims and patient-reported data in a single registry.

But here is why neither objection dissolves the rule. Linkage is powerful and lossy in the same breath. Drop to a single identifier and recall can fall to 64.8%, and in that same study fewer than 4% of eligible records carried an SSN available for matching [20]. You lose patients at the join, and when that loss is differential it becomes its own bias, a new blind spot manufactured by the fix. Linking also does not launder selection bias: a referral-skewed registry stays referral-skewed after you bolt claims onto it, and the variables nobody collected do not reappear because you joined two tables (Jarada et al. 2023 [16]; Wilpshaar et al. 2026 [17]). The agencies endorse linkage precisely because no single source suffices. The FDA's registry guidance anticipates it [3], and the EMA's DARWIN EU network standardises EHR, claims, registries and biobanks onto the common OMOP data model so they can be interrogated together (on the order of 250 million patients as of 2026, a figure that keeps moving) [28]. The literature says it plainly: "EHR systems often identify only fragments … A remedy is to link EHR data with … claims" (Schneeweiss and Desai 2024 [21]). Linkage is the honest move when no single witness can speak to the whole question. It is not a way to avoid choosing which witnesses to call, and you still pay at the join.

The decision rule you can run on Monday

Before you license anything, run three lines:

  1. Write the question and the exact endpoint variable first, in that order. Not the dataset. The variable.
  2. Name each candidate source's disqualifying gap for that specific variable, out loud, in writing. If you cannot name it, you do not know the source well enough to buy it.
  3. Keep only the sources whose gap does not touch your endpoint. If none survive alone, link  and pre-specify the join and the loss you expect at it before you assume two sources add up to one.

That is the whole atlas in three lines, and it is the same discipline behind the minimum viable evidence framework: buy the evidence the decision needs, not the most impressive dataset on the shelf. The most expensive real-world dataset is the one you licensed before you knew which variable your endpoint lived in. Name the gap first. Then call your witness.

Get the monthly digest

The 5 things evidence leads need to know each month: regulatory moves, RWE developments and what they mean in practice. No pitch, one email a month.

References

[1] Gatto NM, et al. (2022). SPIFD structured fit-for-purpose comparative-study design framework (reliability and relevancy). Clin Pharmacol Ther;111(1):122-134. PMID: 34716990. https://pubmed.ncbi.nlm.nih.gov/34716990/

[2] FDA. Real-World Data: Assessing Electronic Health Records and Medical Claims Data to Support Regulatory Decision-Making for Drug and Biological Products. FINAL, July 2024 (Federal Register notice 2024-16338; finalises the September 2021 draft).

[3] FDA. Real-World Data: Assessing Registries to Support Regulatory Decision-Making for Drug and Biological Products. FINAL, December 2023 (Federal Register notice 2023-28289; finalises the November 2021 draft).

[4] FDA. Digital Health Technologies for Remote Data Acquisition in Clinical Investigations. FINAL, December 2023 (Federal Register notice 2023-28262; finalises the December 2021 draft).

[5] FDA. Considerations for the Use of Real-World Data and Real-World Evidence to Support Regulatory Decision-Making for Drug and Biological Products. FINAL, August 2023 (Federal Register notice 2023-18841).

[6] EMA. Guideline on registry-based studies (EMA/426390/2021). FINAL, October 2021.

[7] Franklin JM, et al. (2021). RCT-DUPLICATE first results (10 trials). Circulation;143(10):1002-1013. PMID: 33327727. https://pubmed.ncbi.nlm.nih.gov/33327727/

[8] Wang SV, et al. (2023). RCT-DUPLICATE: emulation of 32 randomised trials with real-world data. JAMA;329(16):1376-1385. PMID: 37097356. https://pubmed.ncbi.nlm.nih.gov/37097356/

[9] Stamas N, et al. (2024). Practical considerations for the use of administrative claims data. J Health Econ Outcomes Res;11(1):57-66. PMID: 38425708. https://pubmed.ncbi.nlm.nih.gov/38425708/

[10] Gillmeyer KR, et al. (2019). Accuracy of claims-based algorithms in pulmonary arterial hypertension. Chest;155(4):680-688. PMID: 30471268. https://pubmed.ncbi.nlm.nih.gov/30471268/

[11] Brown JS, et al. (2020). The FDA Sentinel distributed data network. J Am Med Inform Assoc;27(5):793-797. PMID: 32279080. https://pubmed.ncbi.nlm.nih.gov/32279080/

[12] Merola D, et al. (2022). An algorithm to assess EHR data completeness / continuity. Ann Epidemiol;76:143-149. PMID: 35878784. https://pubmed.ncbi.nlm.nih.gov/35878784/

[13] Curtis MD, et al. (2018). Development and validation of a composite real-world mortality endpoint. Health Serv Res;53(6):4460-4476. PMID: 29756355. https://pubmed.ncbi.nlm.nih.gov/29756355/

[14] Wharton GT, et al. (2024). A global overview of real-world data sources. Pharmacoepidemiol Drug Saf;33(1):e5695. PMID: 37690792. https://pubmed.ncbi.nlm.nih.gov/37690792/

[15] Herrett E, et al. (2015). Data resource profile: Clinical Practice Research Datalink (CPRD). Int J Epidemiol;44(3):827-836. PMID: 26050254. https://pubmed.ncbi.nlm.nih.gov/26050254/

[16] Jarada TN, et al. (2023). Selection bias in real-world data used for HTA. Curr Oncol;30(2):1945-1953. PMID: 36826112. https://pubmed.ncbi.nlm.nih.gov/36826112/

[17] Wilpshaar MC, et al. (2026). Repurposing disease registries for regulatory and HTA use. Clin Pharmacol Ther (online). PMID: 42337928. https://pubmed.ncbi.nlm.nih.gov/42337928/

[18] Servais L, et al. (2024). Stride velocity 95th centile: the first qualified digital primary endpoint. Sci Rep;14:29681. PMID: 39613806. https://pubmed.ncbi.nlm.nih.gov/39613806/

[19] Goldsack JC, et al. (2020). Verification, analytical validation and clinical validation (V3) of digital measures. npj Digit Med;3:55. PMC7156507. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC7156507/

[20] Tyagi K, Willis SJ (2025). Privacy-preserving record linkage. JAMIA Open;8(1):ooaf002. PMID: 39845287. https://pubmed.ncbi.nlm.nih.gov/39845287/

[21] Schneeweiss S, Desai RJ, Ball R (2024). Linking electronic health records with claims data. Am J Epidemiol;194(2). PMID: 39013780. https://pubmed.ncbi.nlm.nih.gov/39013780/

[22] Ultragenyx. GNE Myopathy Disease Monitoring Program (GNEM-DMP), natural-history registry. ClinicalTrials.gov: NCT01784679. https://clinicaltrials.gov/study/NCT01784679

[23] National Registry of Rare Kidney Diseases (RaDaR). ClinicalTrials.gov: NCT06065852. https://clinicaltrials.gov/study/NCT06065852

[24] Sarepta Therapeutics / Roche. EMBARK. ClinicalTrials.gov: NCT05096221. https://clinicaltrials.gov/study/NCT05096221

[25] SYSNAV. ActiLiège Next (wearable-endpoint validation). ClinicalTrials.gov: NCT05982119. https://clinicaltrials.gov/study/NCT05982119

[26] Eli Lilly. LY900014, Phase 2 (CGM time-in-range primary endpoint). ClinicalTrials.gov: NCT04585776. https://clinicaltrials.gov/study/NCT04585776

[27] MPN PROGRESSion registry (EHR + claims + patient-reported linkage). ClinicalTrials.gov: NCT07362225. https://clinicaltrials.gov/study/NCT07362225

[28] EMA. Data Analysis and Real World Interrogation Network (DARWIN EU): data partners standardise sources (EHR, claims, registries, biobanks) into the OMOP common data model. European Medicines Agency. https://www.ema.europa.eu/en/about-us/how-we-work/data-regulation-big-data-other-sources/real-world-evidence/data-analysis-real-world-interrogation-network-darwin-eu

Share this post