Skip to content
Inovia Bio field guide to real-world evidence bias: confounding by indication, selection bias and immortal time bias explained in plain English.
RWE Real world evidence Methodology

The three biases that sink real-world evidence: a plain-English field guide to confounding, selection and immortal time

Imi
Imi

In 2010, a group of epidemiologists took a cohort of people with diabetes and asked a plain question. Did taking a statin for a year or more slow their progression to insulin? The first pass was reassuring. Statin users looked protected, with an adjusted hazard ratio of 0.74, which you can read as roughly a 26% lower risk. Then the same team reran the same data with one change to how they accounted for time. The number turned over completely, to a hazard ratio of 1.97, an apparent near-doubling of risk. Same drug, same patients, same outcome, same database. [1]

Nothing was wrong with the data. The original design had quietly credited patients with a stretch of survival they had not yet earned, and once that was corrected, the protective effect it had manufactured simply disappeared. Here is the part worth sitting with. You did not need a biostatistics degree to catch that error you needed to ask one blunt question about when the clock started.

Three biases are what actually sink a real-world-evidence study: confounding by indication, selection bias and immortal time bias. None of them is, at heart, a mathematical problem you hand to the statistician once the study is built. They are design problems. Each is detectable with a single plain-English question you can ask before the analysis is run, and a clean, expensive, "regulatory-grade" dataset does nothing to protect you from any of them, because none of them lives in the data. They live in how you set the study up.

Why you can't outsource the noticing

Let's be clear about the stakes, because this is not a theoretical worry. In 2022, Purpura and colleagues went through 88 FDA approvals from January 2019 to June 2021 where real-world evidence had been submitted to support a safety or effectiveness claim. In 65 of them (74%) the RWE actually made it into FDA's assessment. In 13 (15%) it was judged inadequate. [2]

Then there is the sentence every RWE lead should tape to their monitor. Reproducing real FDA reviewer language on one submission, the authors quote: "Due to major methodological issues (including immortal time bias, selection bias, misclassification, confounding, and missing data), the FDA does not consider these [RWE] results adequate to support regulatory decision making." [2]

Read that list again. Immortal time bias, selection bias, confounding. Three of the five named failure modes are the exact three this piece is about, and they appear not in a textbook but in the stated reason a real submission was set aside.

You might assume the people running the analysis will catch these for you. That assumption is shakier than it looks. When D'Andrea and colleagues reviewed 44 published tools for assessing the quality of non-randomised drug studies, only one of them (2%) addressed selection bias from the depletion of susceptibles, and only three (7%) addressed time-related bias at all. [3] The checklists the experts themselves reach for miss these routinely. So the founder or clinical lead who can ask the design question early is not meddling. On a lean programme, where a handful of evidence choices drives most of the value, they are often the last line that asks it at all.

Some years back I sat in a programme review for an early oncology asset, where a real-world survival comparison had been built to carry the next funding conversation. Someone asked, almost in passing, when the clock started for the treated group against the comparator. The answer from the room was that the statisticians would handle it in the analysis. They did handle it, it just landed three months and one awkward investor update later than it should have, after the comparison had already been shown around as though it were settled.

This is the sharp end of a point we have made before: what makes RWE genuinely "regulatory-grade" is the process, not the dataset. And these three biases are where that word, "process", stops being a slogan and turns into three questions you can say out loud in the room, before anyone has opened a statistics package.

Confounding by indication: would the comparison group ever have been given this drug?

Confounding by indication is the drug and the reason it was prescribed arriving through the door hand in hand. If sicker patients are the ones who get the treatment, and being sicker also drives the outcome, then the drug and the reason for it are so entangled you cannot tell which one moved the result.

One caution before the example. The phrase gets used loosely. Salas and colleagues showed back in 1999 that "confounding by indication" is used across the literature to mean at least three distinguishable things. [4] I am using it here in its strict sense: patients are channelled towards or away from a drug for reasons that also predict how they will do, whether because they are sicker, or because they are the "healthy user" type who takes their preventive medicine, eats well, and turns up to appointments.

The cleanest illustration is hormone replacement therapy. Through the 1990s, observational studies including the Nurses' Health Study found that women on combined HRT had roughly 40 to 50% lower risk of coronary heart disease. Then the Women's Health Initiative randomised trial reported the opposite, an increased risk, with a hazard ratio of 1.29 for the combined oestrogen–progestin arm. [5] For years this read like observational evidence and randomised evidence fundamentally disagreeing.

They were not disagreeing. In 2008, Hernán and colleagues re-analysed the same Nurses' Health Study data, this time emulating the WHI's eligibility rules and its intention-to-treat logic instead of the original ever-user-versus-never-user comparison, and reproduced the WHI's early increased-risk finding from the very same observational data. [6] The discrepancy had been a design artefact all along. Worth flagging for precision: a separate strand of this literature attributes much of the original gap to a "timing hypothesis" about the age at which women started HRT and their time since menopause, rather than to healthy-user confounding alone. [7] Both explanations are real, they are not the same explanation, and it would be a factual softening to let anyone collapse them into one.

The design that defeats confounding by indication is the active-comparator, new-user approach: instead of comparing users to non-users, you compare new starters on your drug to new starters on a rival drug prescribed for the same clinical reason. The RCT DUPLICATE initiative out of Brigham and Women's Hospital has now run more than 30 registered claims-data emulations of completed trials on exactly this principle, from ROCKET-AF (NCT04593056) to ARISTOTLE (NCT04593030) to LEADER (NCT03936049). [8] One of its registry records is admirably honest about the ceiling of the method, noting that randomisation "is also not replicable in healthcare claims data but was proxied through a statistical balancing of measured covariates." [8]

And when you genuinely cannot rule out that something unmeasured is driving your result, there is a sharper test than arguing plausibility. A four-continent emulation (NCT07677865) is asking whether GLP-1 receptor agonists prevent the onset of Alzheimer's, a protective signal that turned up across real-world cohorts of more than four million people but not in the EVOKE and EVOKE+ trials. Its check for hidden confounding is to look at negative-control outcomes, things the drug should not plausibly affect, such as accidental injury. [9] If your "protective" drug also appears to protect against being hit by a car, the problem is your comparison, not the drug.

The question to ask: would the comparison group plausibly have been prescribed this drug for the same reason, and if not, are the people who got it systematically healthier or sicker than the people who did not?

Free download

The RWE Briefing Document Template

The section-by-section structure for the RWE part of a regulatory briefing, built around the questions reviewers actually ask.

Get the template →

Selection bias: what had to happen before a patient could appear in the data at all?

Think of a bouncer working a guest list. Your dataset only contains the people who got past the door. If whatever the bouncer used to wave people through, be it a referral, surviving long enough to be diagnosed, or reaching a specialist centre, is also tied to the outcome you are measuring, then the comparison inside the room is skewed before you count a single person. (In technical terms you have conditioned on a collider, a variable pushed on by both the exposure side and the outcome side. [10] You do not need the term to smell the problem.)

Here is what the door does to the numbers. Jarada and colleagues followed 23,152 people newly diagnosed with metastatic cancer across 13 cancer sites in Alberta between 2010 and 2019. Look at the whole population, and 40.8% went on to start treatment. Look only at the subset who were referred to and seen by a medical oncologist, and that figure is 67.4%, about 1.65 times higher. Median overall survival tells the same story: 5.42 months across the whole population against 11.24 months in the referred subset, a little over twice as long. But the distortion was worst exactly where referral was rarest. In liver cancer, the selected subset overstated the treatment rate by 2.30 times and survival by 2.60 times. [11]

That said, one honest caveat on that example. It is a contrast between a whole population and a selected slice of it, not a "here is the biased number, here is the corrected number" reanalysis of a single study the way the statin case was. It shows you the size of the door effect. It does not hand you a tidy before-and-after.

You see the same mechanism in regulatory submissions. In the blinatumomab external-control case described by Seeger and colleagues, the trial participants were systematically older than the external comparison group, mean age 41 against mean age 38, with 28% aged 55 or over against 10%. [12] Different doors, different rooms, and any naive mortality comparison between them would have been measuring age as much as drug.

The question to ask: what had to happen to a patient before they could appear in this dataset at all, and does that same filter also predict the outcome?

Immortal time bias: is anyone being credited with time before they were exposed?

Go back to the statin reversal we opened with. That was immortal time bias, and its logic is worth seeing clearly. Suppose you time a group of runners, but you only count someone as a "runner" if they were still on the track at the one-kilometre mark. Everyone in that group survived to one kilometre by definition, because that is how you sorted them. The stretch before is immortal time: a period in which, by the study's own design, the outcome could not happen, wrongly counted as time at risk. [13]

The classic demonstration is Suissa's 2003 reanalysis of a Saskatchewan cohort of 979 people newly treated for COPD. The time-fixed analysis put inhaled corticosteroids at a mortality rate ratio of 0.69, a handsome benefit. Move to a time-dependent analysis that counted exposure only from when it actually began, and the rate ratio was 1.00. [14]

The benefit was bookkeeping.

This hands the non-statistician a genuinely useful tell, one I would name the "reward-for-waiting" effect. In the biased version of that COPD analysis, the apparent benefit grew mechanically stronger the longer the exposure window was stretched: a rate ratio of 0.98 at a 15-day window falling to 0.51 at 365 days. The corrected analysis barely moved regardless of window, sitting near 1.0 throughout. [14] So here is the plain rule. If a drug's apparent benefit gets bigger the longer you let the exposure window run, you are almost certainly looking at reward-for-waiting, not efficacy.

None of this is exotic. When Bykov and colleagues audited 155 published real-world studies in diabetes and cancer, they found the potential for immortal time bias in 63% of them, most often because treatment groups were defined by prescriptions dispensed during follow-up rather than at a fixed baseline. That is the majority of an entire therapeutic literature getting the same thing wrong. [15]

The fixes are well established, and increasingly stated in plain text in trial registries. A French national-claims emulation of tocilizumab versus methotrexate in giant cell arteritis (NCT07459335) records it outright: "A clone-censor-weight approach was used to account for treatment initiation timing and avoid immortal time bias." [16] Another indexes both arms at the same fixed point 90 days after treatment start (NCT07465926). [17] The common thread, as Her and colleagues put it, is that the active-comparator new-user design defuses confounding by indication and immortal time bias in one move, because both arms start follow-up at the same unambiguous point. [13]

One correction often uncovers the next

A tidy three-box checklist would be dishonest, so here is the complication. Fixing one bias frequently reveals another sitting underneath it. Return to the statin study: after the immortal-time correction flipped the hazard ratio from 0.74 to 1.97, the authors themselves attributed part of the residual signal to confounding by indication, because statin users carried higher cardiovascular risk to begin with. [1] One correction unmasked a second.

You can watch this happen in layers. Suissa and Dell'Aniello, looking at inhaled corticosteroids and lung cancer, reported a hazard ratio of 0.32 when immortal time and latency were both mishandled, 0.50 once immortal time was corrected, and 0.96 once latency was corrected too. [18] The biases peel back one at a time. So treat the three questions as a first pass that tells you where to look next, not a certificate of clearance. A clean answer on one does not exonerate the study.

Isn't this the biostatistician's job?

These three failure modes are precisely what propensity scores, time-varying exposure definitions and target-trial emulation exist to handle. A non-statistician lobbing half-understood questions about colliders and immortal time risks second-guessing the very experts they hired, slowing the work and eroding trust. Surely the responsible move is to hire good people and stay out of their way.

Fair. Now look again at what the three questions actually are. Who is the comparator? Who made it into the dataset? When does follow-up start? None of those is a statistical calculation. Every one is a design decision that has to be settled before a statistician can even choose a method, and each is far cheaper to fix on a whiteboard than in a reviewer response. Recall D'Andrea: the formal tools the professionals lean on caught time-related bias 7% of the time. [3] You are not being asked to run the analysis. You are being asked to make sure the design question reaches it early enough to still change something. Purpura's reviewer quote is what it costs when it does not. [2]

Three questions to take into your next real-world-evidence design meeting

You do not need to run a single correction to be useful in the room where an RWE study is designed. You need three questions.

  • Confounding by indication: would the comparison group plausibly have been prescribed this drug for the same reason, and if not, are they systematically sicker or healthier than the people who got it?
  • Selection bias: what had to happen to a patient before they could appear in this dataset at all, and does that same filter also predict the outcome?
  • Immortal time bias: is anyone being credited with surviving, or with not having the outcome, during time before they were actually exposed?

This is also where the regulators are heading. ICH M14, finalised in September 2025 and being adopted across FDA and EMA through 2026, is reported to treat selection bias, information bias, time-related bias and confounding as a structured part of how non-interventional studies should be assessed. [19] So the vocabulary is converging, and knowing to ask these three questions has stopped being optional.

Ask them before the reviewer does. And if you would value a second pair of eyes on an RWE study before it goes anywhere near a submission, that is exactly the kind of work we do.

Get the monthly digest

The 5 things evidence leads need to know each month: regulatory moves, RWE developments and what they mean in practice. No pitch, one email a month.

References

[1] Lévesque LE, Hanley JA, Kezouh A, Suissa S. (2010). "Problem of immortal time bias in cohort studies: example using statins for preventing progression of diabetes." BMJ;340:b5087. PMID: 20228141. https://pubmed.ncbi.nlm.nih.gov/20228141/

[2] Purpura CA, Garry EM, Honig N, Case A, Rassen JA. (2022). "The Role of Real-World Evidence in FDA-Approved New Drug and Biologics License Applications." Clin Pharmacol Ther;111(1):135–144. PMID: 34726771. https://pubmed.ncbi.nlm.nih.gov/34726771/

[3] D'Andrea E, et al. (2021). "How well can we assess the validity of non-randomised studies of medications? A systematic review of assessment tools." BMJ Open;11(3). PMID: 33762237. https://pubmed.ncbi.nlm.nih.gov/33762237/

[4] Salas M, Hofman A, Stricker BH. (1999). "Confounding by indication: an example of variation in the use of epidemiologic terminology." Am J Epidemiol;149(11):981–983. PMID: 10355372. https://pubmed.ncbi.nlm.nih.gov/10355372/

[5] Prentice RL, Langer R, Stefanick ML, et al. (2005). "Combined postmenopausal hormone therapy and cardiovascular disease: toward resolving the discrepancy between observational studies and the Women's Health Initiative clinical trial." Am J Epidemiol;162(5):404–414. PMID: 16033876. https://pubmed.ncbi.nlm.nih.gov/16033876/

[6] Hernán MA, Alonso A, Logan R, et al. (2008). "Observational studies analyzed like randomized experiments: an application to postmenopausal hormone therapy and coronary heart disease." Epidemiology;19(6):766–779. PMID: 18854702. https://pubmed.ncbi.nlm.nih.gov/18854702/

[7] Bhupathiraju SN, Grodstein F, Rosner BA, et al. (2017). "Hormone Therapy Use and Risk of Chronic Disease in the Nurses' Health Study: A Comparative Analysis With the Women's Health Initiative." Am J Epidemiol;186(6). PMID: 28938710. https://pubmed.ncbi.nlm.nih.gov/28938710/

[8] RCT DUPLICATE initiative (Brigham and Women's Hospital / Harvard), claims-data emulations of completed randomised trials via an active-comparator, new-user design. Representative registered records: ROCKET-AF emulation, ClinicalTrials.gov NCT04593056, https://clinicaltrials.gov/study/NCT04593056 ; ARISTOTLE emulation, NCT04593030, https://clinicaltrials.gov/study/NCT04593030 ; LEADER emulation, NCT03936049, https://clinicaltrials.gov/study/NCT03936049

[9] Multi-cohort target-trial emulation of GLP-1 receptor agonists and Alzheimer's onset, using negative-control outcomes. ClinicalTrials.gov: NCT07677865. https://clinicaltrials.gov/study/NCT07677865

[10] Lu H, Cole SR, Howe CJ, Westreich D. (2022). "Toward a Clearer Definition of Selection Bias When Estimating Causal Effects." Epidemiology;33(5):699–706. PMID: 35700187. https://pubmed.ncbi.nlm.nih.gov/35700187/

[11] Jarada TN, O'Sullivan DE, Brenner DR, Cheung WY, Boyne DJ. (2023). "Selection Bias in Real-World Data Studies Used to Support Health Technology Assessments: A Case Study in Metastatic Cancer." Curr Oncol;30(2). PMID: 36826112. https://pubmed.ncbi.nlm.nih.gov/36826112/

[12] Seeger JD, Davis KJ, Iannacone MR, et al. (2020). "Methods for external control groups for single arm trials or long-term uncontrolled extensions to randomized clinical trials." Pharmacoepidemiol Drug Saf;29(11). PMID: 32964514. https://pubmed.ncbi.nlm.nih.gov/32964514/

[13] Her QL, Rouette J, Young JC, Webster-Clark M, Tazare J. (2024). "Core Concepts in Pharmacoepidemiology: New-User Designs." Pharmacoepidemiol Drug Saf;33(12):e70048. PMID: 39586646. https://pubmed.ncbi.nlm.nih.gov/39586646/

[14] Suissa S. (2003). "Effectiveness of inhaled corticosteroids in chronic obstructive pulmonary disease: immortal time bias in observational studies." Am J Respir Crit Care Med;168(1):49–53. PMID: 12663327. https://pubmed.ncbi.nlm.nih.gov/12663327/

[15] Bykov K, He M, Franklin JM, Garry EM, Seeger JD, Patorno E. (2019). "Glucose-lowering medications and the risk of cancer: A methodological review of studies based on real-world data." Diabetes Obes Metab;21(9). PMID: 31062453. https://pubmed.ncbi.nlm.nih.gov/31062453/

[16] TACTIC-GCA: French national-claims (SNDS) target-trial emulation of tocilizumab versus methotrexate in giant cell arteritis, using a clone-censor-weight approach. ClinicalTrials.gov: NCT07459335. https://clinicaltrials.gov/study/NCT07459335

[17] Target-trial emulation of early GLP-1RA/SGLT2i add-on therapy, using a 90-day landmark for both arms. ClinicalTrials.gov: NCT07465926. https://clinicaltrials.gov/study/NCT07465926

[18] Suissa S, Dell'Aniello S. (2020). "Time-related biases in pharmacoepidemiology." Pharmacoepidemiol Drug Saf;29(9). PMID: 32783283. https://pubmed.ncbi.nlm.nih.gov/32783283/

[19] ICH M14: General Principles on Plan, Design and Analysis of Pharmacoepidemiological Studies That Utilize Real-World Data for Safety Assessment of Medicines. Step 4 final, 5 September 2025; adoption across FDA and EMA through 2026. https://database.ich.org/sites/default/files/ICH_M14_Step4_Final_Guideline_2025_0905.pdf

Share this post