Every vendor deck has the stamp on it now. "Regulatory-grade real-world data." "Regulatory-grade EHR data." It's in the webinar title, on the data-room slide, in the sales email, and it does one job: reassurance. Buy this, and the reliability problem becomes someone else's.
Here is the reverse. The FDA never defined a grade. Its July 2024 final guidance on electronic health records and medical claims data defines reliability as three things you have to be able to measure: accuracy, completeness and traceability. It pairs them with relevance, the data's fitness for your specific question. [1] The EMA's Data Quality Framework, final since October 2023, names the same dimensions. [2] And a review of 46 real-world data and evidence guidance documents worldwide found unanimous agreement that the data must be reliable and of good provenance, with no harmonised numeric minimum anywhere. [3]
So "regulatory grade" does not ship with the dataset. It is a property of what you can demonstrate about it, per submission, per variable. If you want the concept itself, the companion piece on regulatory-grade RWE covers what the term means and who profits from the confusion. This post is the other half: the operational bar.
One test cuts through all of it, and it earns a name: the reconstruction test. Can you get every load-bearing data point back to its source of truth? That is the whole question. The three axes below are just where it bites hardest.
Start where most buyers finish. Traceability is the axis the sales conversation never reaches, and the one a reviewer reaches first.
The FDA defines it plainly. In a footnote to the 2024 guidance, traceability is the method (an audit trail, say) that gives you a data element's provenance: where it came from and how it got into the record. [1] Provenance lives inside traceability, and you do not get one without the other.
This is where a lot of "quality documentation" quietly fails. A data dictionary describes your fields and a methods folder describes your process, but neither is an audit trail. 21 CFR Part 11 §11.10(e) asks for something stricter: a secure, computer-generated, time-stamped record of every action that creates, modifies or deletes a record, one that does not obscure what was there before. [4] (Part 11's scope is narrower than people quote, so don't oversell it.) A folder describing what you meant to do is not a record of what happened to each value. And you cannot backfill it later; provenance gaps do not close after the fact.
How badly does it hold up on real data? In one FDA-funded demonstration study, the Verantos-led TRUST work published in JAMA Network Open, investigators scored all three axes on 120,616 patients: traceability came in at 11.5% under a traditional approach, meaning barely one data element in ten could be tracked back to a source of truth, and reached 77.3% under an advanced approach. [5] One study, one method, so not a benchmark. But it shows which axis collapses hardest without a controlled pipeline.
Years ago I worked on an early-phase rare-disease programme where the summary tables were beautiful: clean, complete, every cell populated. Then a reviewer asked how one endpoint had been derived, we went to trace it, and could not. The value existed; the path from source event to the cell had not been kept.
A number with no chain of custody looks immaculate in a slide and evaporates the moment someone leans on it.
The engineered-in version looks different. In the STR1VE trial of onasemnogene abeparvovec (NCT03306277), the ventilatory-independence endpoint was derived solely from a single named device, the Philips Trilogy BiPAP. [6] No regulator stamped that dataset "regulatory grade"; a registry records design, not a verdict. But it is the right instinct: pin the source of a load-bearing measurement so the chain back stays short and auditable.
Monday test: ask for the audit trail and the linkage-validation record, not the quality claim. If a value cannot be walked back to a source event, it is not regulatory-grade for a submission, however clean the table looks.
Free download
The RWE Briefing Document Template
The section-by-section structure for the RWE part of a regulatory briefing, built around the questions reviewers actually ask.
Get the template →Completeness is the axis buyers think they understand, and usually measure wrong. It is not the size of the dataset. A dataset can carry ten million patients, sail through every headline quality check a vendor puts in front of you, and still be missing the single covariate your causal estimate turns on, which is the difference between a number a regulator trusts and one they quietly set aside. The EMA files this under extensiveness: information available against the total that could be available. [2] Volume flatters; the load-bearing fields are where you actually look.
Take a leading oncology dataset, characterised by its own developers. A haemoglobin measurement, a routine covariate, was observed for 91.7% of the advanced non-small-cell lung cancer cohort and 83% of the multiple myeloma cohort — 8 to 17% missing on a variable nobody would call exotic, reported by the people who built the data (Flatiron and Roche authors, not a third-party audit). [7]
But here's the thing. Everyone wants a single number: what percentage complete counts as "regulatory grade"? There isn't one, and anyone quoting you one is selling. What exists is endpoint-specific working bars. In a Flatiron validation, raw EHR data captured only 66% of deaths; amalgamating sources lifted mortality sensitivity to 91%, and the authors noted experts suggest a 90% sensitivity target for that mortality endpoint. [8] Read that for what it is: a bar set for one endpoint, in one use case, and reported. Not a percentage you clear once and carry everywhere.
You see the same logic written into eligibility. AstraZeneca and Daiichi Sankyo's ESPERANZA external-control study for trastuzumab deruxtecan (NCT06973161) requires that patients known to have died must have a complete recorded date of death. [9]
The completeness bar, at the front door.
Monday test: before you commit, write down the exact covariates and outcomes your estimand needs, then measure missingness on those fields, not the headline N. It's the same discipline as spending your evidence budget on the 10% that moves the decision: check the fields that carry the estimate, and let the rest be what it is.
Accuracy is the most intuitive of the three, so this stays short. It asks whether the recorded value matches the source truth. In the same FDA-funded demonstration study, accuracy measured 59.5% under the traditional approach and 93.4% under the advanced one, on the same 120,616 patients. [5] Same caveat: one study, one method, attributed, not a benchmark.
Then the leg that gets left off the deck entirely: relevance. And relevance is not really a property of the data at all. It lives in the fit between the data and your question: the exposures, outcomes, covariates and representative patient numbers your specific estimand demands. [1] Which is why no dataset is "regulatory grade" in the abstract. Validity, as Wang and Schneeweiss put it, must be judged on the triad of study question, design and data together. [10] Change the question, and a dataset that was fit yesterday is unfit today, without a single value changing.
Monday test: stop asking "is this dataset regulatory-grade?" Start asking "is it reliable and relevant for this estimand?" Then take that framing into the regulatory conversation early, because the answer moves with the question.
The counterargument deserves its best shot, and there are two.
First, the cost objection: this is gold-plating, full lineage reconstruction being a big-pharma luxury paid for by a data-management department a lean biotech simply does not have. Second, the sharper one: you have just admitted there is no pass/fail number, so if nobody will name the threshold, why pay to measure against a bar that does not exist?
Both are reasonable. That said, both fail in the same place. No pass/fail number does not mean nothing to measure. Reliability is quantified and reported per dataset and per variable; the TRUST investigators put it flatly, that there can be no pass-fail inflection point, and then measured every axis anyway. [5] What you owe a regulator was never "grade A". It is your traceability, your missingness on the fields that matter, your accuracy against source, and why that package is fit for this question.
And the cost runs the other way. When RWE has failed at the FDA, it failed on data reliability and relevancy concerns and a lack of pre-specification, not on being too small — the finding from a review of 34 RWE instances submitted between 1954 and 2020. [11] Studies get rejected for factors intrinsic to the source: missing endpoints, inappropriate follow-up time, a high proportion of missing data. [12] For a resource-lean team the reconstruction test is cheap insurance: weeks of due diligence up front against a rejected submission and a burned funding window.
Six questions to run against a dataset before you build a submission on it. If you cannot answer one of them with an artefact rather than an assurance, you have found your risk.
Run all six before budget is committed. The same checklist earns its keep twice, because this standard bites hardest in an external-control submission, where the control cohort has to survive exactly this scrutiny.
There is no universal threshold, and that is the right answer. The FDA declines to set one; the industry frameworks decline too, on the stated grounds that it depends upon the context and purpose. [14] The bar is reconstruction, not a score you clear once and frame on the wall.
If anything, the premium is rising. In December 2025 the FDA signalled it would accept de-identified and aggregated real-world data on a case-by-case basis for certain device submissions, with the expectation that similar thinking reaches drugs and biologics. [15] Lowering the identifiability bar does not lower the reliability bar; it raises it. The less a data point carries with it, the more the burden falls on your ability to trace what remains back to where it came from.
That is the work: reconstruction, evidenced, per question and per variable.
A grade was never on offer. If you want a second set of eyes on that test before a submission is built on a dataset, that is the kind of thing we do at Inovia.
Get the monthly digest
The 5 things evidence leads need to know each month: regulatory moves, RWE developments and what they mean in practice. No pitch, one email a month.
[1] FDA. "Real-World Data: Assessing Electronic Health Records and Medical Claims Data To Support Regulatory Decision-Making for Drug and Biological Products." Guidance for Industry, FINAL, July 2024.
[2] EMA. "Data Quality Framework for EU medicines regulation." EMA/326985/2023, FINAL, 30 October 2023.
[3] Sarri G, Hernandez L, et al. (2024). "Data quality frameworks for real-world data: a scoping review of guidance documents." J Comp Eff Res. PMID: 39132748. https://pubmed.ncbi.nlm.nih.gov/39132748/
[4] 21 CFR Part 11 §11.10(e) — Electronic Records; Electronic Signatures (audit-trail requirement). https://www.ecfr.gov/current/title-21/chapter-I/subchapter-A/part-11
[5] Riskin D, et al. (2025). "Data Reliability and Real-World Evidence: Accuracy, Completeness, and Traceability." JAMA Network Open. PMID: 40063029. https://pubmed.ncbi.nlm.nih.gov/40063029/ (FDA-funded TRUST demonstration study, Verantos-led; figures attributed as a single demonstration study, not a neutral benchmark.)
[6] Novartis Gene Therapies. "STR1VE: A single-arm study of onasemnogene abeparvovec in spinal muscular atrophy." ClinicalTrials.gov: NCT03306277. https://clinicaltrials.gov/study/NCT03306277
[7] Sondhi A, et al. (2023). CPT Pharmacometrics Syst Pharmacol. PMID: 37322818. https://pubmed.ncbi.nlm.nih.gov/37322818/ (Flatiron/Roche authors; haemoglobin observed for 91.7% aNSCLC, 83% MM.)
[8] Curtis MD, et al. (2018). "Development and Validation of a High-Quality Composite Real-World Mortality Endpoint." Health Serv Res. PMID: 29756355. https://pubmed.ncbi.nlm.nih.gov/29756355/ (Flatiron authors; 90% sensitivity is an endpoint-specific working bar, not a universal regulatory percentage.)
[9] AstraZeneca / Daiichi Sankyo / IQVIA. "ESPERANZA: An external control study for trastuzumab deruxtecan." ClinicalTrials.gov: NCT06973161. https://clinicaltrials.gov/study/NCT06973161
[10] Wang SV, Schneeweiss S. (2022). "A framework for visualizing study designs and data availability of real-world evidence studies." Clin Epidemiol. PMID: 35520277. https://pubmed.ncbi.nlm.nih.gov/35520277/
[11] Mahendraratnam N, et al. (2022). "Understanding Use of Real-World Data and Real-World Evidence to Support Regulatory Decisions." Clin Pharmacol Ther. PMID: 33891318. https://pubmed.ncbi.nlm.nih.gov/33891318/ (34 RWE instances, 1954–2020.)
[12] Zebachi S, et al. (2025). Clin Pharmacol Ther. PMID: 40601391. https://pubmed.ncbi.nlm.nih.gov/40601391/ (ATRaCTR framework; rejection factors intrinsic to the RWD source.)
[13] FDA. "Data Standards for Drug and Biological Product Submissions Containing Real-World Data." Guidance for Industry, FINAL, 22 December 2023.
[14] Castellanos E, et al. (2024). "Relevance and Reliability of Real-World Data." JCO Clinical Cancer Informatics. DOI: 10.1200/CCI.23.00046. (Declines to set numeric thresholds; depends on context and purpose.)
[15] FDA. Federal Register notice 2025-23252, 18 December 2025 — acceptance of de-identified/aggregated real-world data on a case-by-case basis for certain device submissions; the FDA's accompanying announcement signals the same consideration for drugs and biologics. Directional/paraphrase.