The real-world evidence field agrees on more than it lets on. Ask a regulator, a payer and a clinical operations lead what good RWE looks like and you get one answer: data that is robust, reliable, of good provenance, and transparently reported. Reviewing 46 guidance documents in 2024, Sarri and colleagues found unanimous agreement on exactly that [1]. The principles, at least, are settled.
The shared vocabulary is where it quietly falls apart. The US statute that created the FDA's RWE programme defines real-world evidence as data: "data regarding the usage, or the potential benefits or risks, of a drug derived from sources other than traditional clinical trials" [2]. The FDA's own 2018 Framework treats it as the opposite thing, the clinical evidence you are left with after analysing that data [3]. The EMA sides with the Framework, calling RWE "evidence derived from the analysis of RWD" [4]. One camp, three sentences, two incompatible readings of the field's central term.
That pattern runs right through the lexicon. The disagreement is rarely about principle and almost always about operational definitions and who owns each term. Linguists have a word for terms that look identical across two languages but mean different things: false friends. The RWE lexicon is full of them, and a team that assumes a shared word implies a shared definition will spend months talking past itself before anyone spots the bill.
So this glossary reads every loaded term three ways: how a regulator uses it (FDA, EMA, ICH), how a payer or HTA body uses it (NICE, ISPOR, ICER), and how it lands operationally under Good Clinical Practice. Some genuinely diverge. A few genuinely converge, and I have said which is which. A glossary that manufactured disagreement to prove a point would be worth nothing to the team relying on it.
Every entry runs the same way: a plain-English gloss, the regulator reading, the payer/HTA reading, the operational reading, and finally the line that matters, which is why the gap bites on Monday. Each entry is labelled DIVERGES, PARTIAL or CONVERGES.
Two honest caveats. The operational column is ICH-GCP and everyday trial usage, not a published CRO dictionary, because no contract research organisation keeps an official glossary to rival the regulators'. And where I point to a real trial, that shows how the term gets used in the wild on ClinicalTrials.gov, not how anyone formally defines it. Usage and definition are different things, and mistaking one for the other is half the problem.
Gloss: evidence about a medicine drawn from sources other than a conventional randomised trial.
Regulator: two readings sit inside one camp. The statute calls it data [2], while the 2018 Framework and the FDA's FINAL 2023 "Considerations for the Use of Real-World Data and Real-World Evidence" guidance treat RWE as the analysed evidence product [5]. The EMA lands with the latter, on evidence derived from analysing RWD [4].
Payer/HTA and operational: everyone downstream also means the evidence product, the analysed result that goes into a dossier.
Why the gap bites: tell a reviewer whose statutory frame reads RWE as raw input that "our RWE shows a benefit" and you have described your ingredients when they asked to see the finished meal. Internally it is worse: your HEOR lead's "RWE" and your regulatory lead's "RWE" can be different objects on the same slide. For the regulator's headline usage see our primer on RWE for regulatory strategy; for what "regulatory-grade" actually demands, we have argued that elsewhere.
Gloss: efficacy is how well a drug works under ideal, controlled conditions; effectiveness is how it performs out in the mess. Singal and colleagues put it cleanly: efficacy is "performance of an intervention under ideal and controlled circumstances, whereas effectiveness refers to its performance under 'real-world' conditions" [6]. A continuum, not a switch.
Regulator: the statutory approval standard is substantial evidence of effectiveness [7], and yet the FDA establishes it in controlled RCTs, which is efficacy in the sense above. The regulator says "effectiveness" and quietly means efficacy.
Payer/HTA: the payer means the other thing entirely, real-world performance in an unselected population. The GetReal consortium named the space between the two the efficacy-effectiveness gap [8].
Operational: the CRO powers the pivotal trial on the regulator's efficacy endpoint, and nobody measures effectiveness in the payer's sense unless someone specifically funds it.
Why it bites: promise a payer "effectiveness", back it with your pivotal RCT, and you have handed them efficacy wearing the wrong label. That is where launch value stories quietly come apart.
Gloss: data good enough for the specific question you are asking, which sounds like a universal test and is nothing of the sort. Sarri and colleagues document four bodies applying four different bars: the FDA frames fit-for-purpose around data reliability and relevance; the EMA also weighs extensiveness, coherence and timeliness; NICE routes it through data "quality", with reliability read as completeness and accuracy; and CDA-AMC adds data access and database version [1].
Regulator versus payer: the same phrase bolted onto materially different checklists.
Operational: whichever body you submit to sets the bar your data management has to clear, so "fit-for-purpose" on its own is never an actionable spec.
Why it bites: "our data are fit for purpose" can clear the FDA's bar and still fail the EMA's added tests, or NICE's. One claim, four examiners. This is why we argue regulatory-grade RWE is a process, not a dataset, and why choosing your evidence bar deliberately beats gold-plating everything.
Gloss: how a drug performs against the alternatives rather than against placebo. Sarri and colleagues found that RWE in comparative effectiveness research "was exclusively covered in HTA guidance" [1]. The regulators barely touch it.
Regulator: the agency generally approves against placebo or a defined endpoint and rarely requires head-to-head comparative data.
Payer/HTA: the payer makes it the centre of the decision. The IOM's 2009 definition frames comparative effectiveness research as generating and synthesising evidence that compares the benefits and harms of alternative options [9], which is the payer's whole question in one line.
Operational: rarely a design driver unless market access has a seat at the table early.
Why it bites: skip the comparative arm because the regulator did not ask for it, and you arrive at the payer missing the one thing their entire decision turns on.
Gloss: whether the benefit is worth the cost. ICER anchors value on long-term value for money and, as a cost-effectiveness benchmark, a range of roughly $100,000 to $150,000 per QALY, widened for ultra-rare conditions [10], and NICE reasons in cost per QALY too.
Regulator: there is no such concept. To the FDA and EMA the frame is benefit-risk, full stop, and ICH E9(R1), the estimand addendum that governs how they reason about a treatment effect, never once reaches for "value" in this sense [11].
Operational: irrelevant to trial conduct, entirely relevant to what you should have measured while you had the chance.
Why it bites: build a clean benefit-risk story for the regulator, assume it carries to the payer, and you find that "value" does not convert at par. Benefit-risk and cost-per-QALY are different currencies.
Free download
The RWE Briefing Document Template
The section-by-section structure for the RWE part of a regulatory briefing, built around the questions reviewers actually ask.
Get the template →Gloss: a comparator built from data outside your trial, whether historical, registry or natural history, instead of a concurrent randomised arm.
Regulator: caution comes first. ICH E10 sanctions external or historical controls only in unusual circumstances [12], and Cucherat and colleagues set the bar bluntly, calling them "only truly acceptable if the observed effect is dramatic" [13]. The FDA's DRAFT 2023 externally controlled trials guidance is sympathetic but heavily hedged [16].
Payer/HTA: the payer sees an opportunity. NICE recommends analytical frameworks to build up a comparator arm where none exists [1], and Siu and colleagues describe renewed interest in external controls as a way to supplement single-arm trials [14].
Operational: it becomes a data-engineering problem, sourcing, curating and matching a cohort that was never designed to be your control.
In the wild: ARASEC (NCT05059236, Bayer, registry enrolment count 223) runs single-arm darolutamide plus ADT and reads it against a previously conducted study [15]. Usage, not a definition.
Why it bites: your regulatory team treats an external control arm as a concession wrung from a wary agency; your access team treats it as the comparator that wins the HTA. Same artefact, opposite postures. We keep an FDA ECA checklist built on that DRAFT guidance, and the EMA's caution on single-arm trials sits in the same tension.
Gloss: a precise statement of exactly what treatment effect a trial is trying to measure. ICH E9(R1) defines it as "a precise description of the treatment effect reflecting the clinical question posed by the trial objective" [11], built from five attributes: treatment, population, endpoint, intercurrent-event strategy and population-level summary. Here the definition genuinely converges: every camp cites the very same addendum.
The divergence sits in which estimand each camp actually wants. Clark and colleagues say it exactly: "The actual estimand of interest varies by stakeholder" [17]. Regulators lean towards the hypothetical, more explanatory strategy, whereas the treatment-policy estimand is the pragmatic one a payer or operational lead is usually picturing.
Why it bites: agreeing on "the estimand" counts for little if regulatory means the hypothetical strategy while everyone else is thinking treatment-policy. You have agreed on a word, not on a number.
Gloss: the main outcome a trial is built to measure.
Regulator: to a regulatory statistician the primary endpoint is one of the five estimand attributes, "the variable (or endpoint) to be obtained for each patient that is required to address the clinical question" [11], rather than the headline result everyone quotes.
Operational: on the ground it is simply the outcome you power the study on.
The structural catch: ClinicalTrials.gov has no "real-world endpoint" field at all [18]. Every outcome is filed as a Primary or Secondary Outcome Measure, whether it came from an EHR, a claims feed or a bespoke CRF.
Why it bites: "real-world endpoint" is a marketing label with no registry category behind it, so do not assume it signals provenance or rigour to a reviewer.
Gloss: how well a trial's findings apply beyond the trial itself.
Payer/HTA: for the payer this is first-order. NICE defines generalisability as findings "intended to be applied to a target population from which the study sample was drawn" and external validity as "how well the findings ... apply to the target population of interest" [19]. For a payer buying for the whole NHS, this is the entire question.
Regulator: for the regulator it is a secondary caveat. In ICH E9(R1) external validity appears as something you might trade away while protecting internal validity, never the headline concern [11].
Operational: the stats and ops camp guards internal validity first, through randomisation, blinding and protocol fidelity.
Why it bites: the same trial can read as "rigorous" to your operational lead and "not generalisable" to the payer, and both are right. They are reading different validities first.
Gloss: a trial designed to test a treatment under everyday conditions rather than idealised ones. The PRECIS-2 tool frames "pragmatic" as a position on a nine-domain continuum running towards "explanatory", a methodologist's design dial rather than a binary label [20]. Singal's group treats effectiveness studies and pragmatic studies as near-synonyms [6], Berger notes that pragmatic RCTs carry "varying amounts of pragmatic (real world) elements" [22], and Gartlehner and colleagues found that "no validated definition of effectiveness studies exists" [21].
In the wild: PERFORM (NCT06372496, GSK) is filed under both "Pragmatic" and "Effectiveness Study", carries a "Usual Care" comparator, and yet sets its primary endpoint as trough FEV1, an idealised measure [23]. The Belfast trial (NCT02717689) bills itself as a "Pragmatic Trial ... Versus Standard Care" [24]. One label, the whole continuum.
Why it bites: "pragmatic" on a protocol tells you where on the dial the designer aimed. It tells you nothing about whether a payer will accept the result as effectiveness evidence. For one structured pragmatic design, see our piece on TwICS.
Gloss: the original records a trial result traces back to, and the oversight that checks them. Under ICH E6(R2) Good Clinical Practice (and note that E6(R3) reached Step 4 in January 2025 and is now the current revision), source data are the original records needed to reconstruct and evaluate the trial, and monitoring is the oversight of conduct against protocol, including source-data verification [25].
Regulator using RWE: lean on RWE and the "source" becomes an EHR or claims record never designed for research. Same word, different object.
Operational: a CRO's source-data-verification muscle memory assumes a monitored CRF. It does not transfer cleanly to a curated claims feed that no monitor ever visited a site to check.
Why it bites: "source data" quietly changes meaning the moment you move from a trial database to a real-world one, and the assurance model bundled with the old meaning does not follow it across for free.
This is not a loose synonym for "we used some real-world data". FULL REVASC (NCT02862119, Karolinska, registry enrolment count 1,542) randomised patients through an online module inside the SCAAR registry and followed them up via SWEDEHEART [26]. The registry is the trial's infrastructure, supplying both the randomisation and the follow-up, while a hard mortality/MI/revascularisation composite is still filed as an ordinary Primary Outcome Measure. "Registry-based" describes the plumbing and says nothing about data quality, so do not let it stand in for either.
This section is the argument's honesty test. If every term diverged, the thesis would be a conspiracy theory. Several plainly do not, and that matters as much as the ones that do.
"Real-world data" (CONVERGES on the gist, contested at the edges). The FDA (data routinely collected from a variety of sources [3]), the EMA (routine clinical practice [4]), NICE (data collected outside a highly controlled trial [19]) and ISPOR (data "not collected in conventional RCTs" [27]) broadly agree on what RWD is. That said, the edges are another matter: Makady and colleagues found 38 separate definitions of RWD in the literature and concluded that "consensus on the definition of RWD is lacking" [28], and Berger notes that some count single-arm trial data as RWD [22]. The gist is shared while the boundary stays a running argument nobody has closed.
Data becomes evidence (CONVERGES). Berger's line is the cleanest in the literature: "Evidence is shaped, while data simply are raw materials and alone are non-informative" [22]. Every camp accepts the direction of travel from RWD to RWE, even while they argue over what sits at each end.
Quality principles (CONVERGES). This is the Sarri "unanimous agreement" again [1], on data that is robust, reliable, of good provenance and transparently reported. Nobody in the field disputes those principles.
Target trial emulation (CONVERGES). Hernán and Robins framed the method as an attempt to "emulate a randomized experiment—the target experiment or target trial" [29], and per Sarri it is recommended alike by the FDA, NICE, CDA-AMC and IQWiG [1]. That is rare cross-camp unanimity on a method.
HETE versus exploratory (CONVERGES). The ISPOR/ISPE pairing of hypothesis-evaluating treatment effect against exploratory analysis is used consistently right across the field [22].
The pattern here is the whole argument in miniature: the principles converge, while the operational definitions and the ownership of terms do not.
Let us give the objection its best shot. Everyone in a development meeting is an expert, experts read context and self-correct, and they usually land on the right meaning without a glossary refereeing every noun. Real programmes fail on biology and on data, not on whether two clever people meant slightly different things by "effectiveness". A definitions page can feel like box-ticking that mainly makes the box-ticker feel rigorous.
Pedantry? No, and here are the receipts. The self-contradiction over "real-world evidence" runs all the way to the difference between data and evidence, written into the statute [2] and the Framework [3]. "Fit-for-purpose" resolves into four separate checklists at four separate bodies [1]. "Value" is missing from the regulator's vocabulary altogether even as it drives the payer's entire decision [11]. Those are different referents wearing a single spelling, and context does not reliably repair them across a regulatory-to-HEOR-to-operations handoff, because each function reads in good faith from its own dictionary. The failure is never loud. It is the quarter you lose when a plan built on one camp's definition meets another camp's expectation at the wrong moment.
Here is the version you can act on this Monday. Before the next joint regulatory, market-access and operations meeting, take the loaded terms off this list (RWE, effectiveness, fit-for-purpose, comparative effectiveness, value, external control, estimand, endpoint) and agree, out loud, which camp's definition you are using for this asset, this indication and this decision. Then write the answers down.
Put those answers on a living definitions page inside your integrated evidence plan, and update it as the programme moves. Unglamorous work. It is also most of what an IEP is for: a shared spine that stops three functions optimising against three private dictionaries. Building that translation layer with cross-functional teams is a good part of what we do at Inovia, and this glossary is where it starts. Pick your definitions before the meeting, not during the argument.
Get the monthly digest
The 5 things evidence leads need to know each month: regulatory moves, RWE developments and what they mean in practice. No pitch, one email a month.