Articles

Real-world evidence and AI: How EHR data is reshaping drug development decisions

AI is turning a decade of regulatory ambition around real-world evidence into a practical clinical development tool.
Written byErika Russell
| 7 min read
Clinical researchers reviewing anonymized patient data visualizations on a monitor in a modern research office.

Real-world evidence AI drug development is moving from pilot projects into routine use as EHR AI tools mature. Explore how sponsors and regulators are adapting.

GEMINI (2026)

Register for free to listen to this article
Listen with Speechify
0:00
7:00

Real-world evidence AI drug development has moved from a regulatory aspiration voiced for more than a decade into a working discipline inside sponsor organizations. Artificial intelligence (AI) now lets translational and clinical development teams mine electronic health record (EHR) data at a scale unmanageable by hand, feeding trial design, patient stratification, and safety surveillance. Data quality, bias, and regulatory acceptance remain open questions that determine how far this evidence can travel toward a label.

Key takeaways

  • Real-world evidence (RWE) draws on data generated outside controlled trials, including EHRs, claims, and registries, to describe how treatments perform in routine care.
  • AI and natural language processing (NLP) convert unstructured clinical notes into structured variables that support research at a scale manual chart review cannot match.
  • Sponsors increasingly use real-world data (RWD) to build external control arms and refine trial eligibility criteria before enrollment begins.
  • Machine learning (ML) models help identify patient subgroups likely to respond to a therapy, supporting stratification strategies in precision medicine.
  • The FDA and the European Medicines Agency (EMA) have each issued frameworks for evaluating RWE, though acceptance still depends on data quality and study design.

What RWE means in drug development

RWE is the clinical conclusion drawn from RWD, meaning information collected outside the controlled setting of a randomized trial. Sources include EHRs, insurance claims, patient registries, and increasingly, wearable devices and patient-reported outcomes tools. Unlike a randomized controlled trial, RWD captures how a therapy performs across the variability of everyday clinical practice, including comorbidities, off-label use, and adherence patterns that trial protocols typically exclude.

The distinction matters because RWD and RWE are not interchangeable. RWD is the raw material, often messy and unstructured, while RWE is the analysis-ready conclusion generated once researchers clean, structure, and interpret that data using an appropriate study design. The FDA has emphasized this distinction in its guidance on real-world evidence, noting that the strength of any RWE submission depends heavily on the underlying data's fitness for a specific regulatory question.

Continue reading below...
A scientist in a white lab coat looking into a microscope in a brightly lit modern laboratory.
WebinarsChoosing the right study for developmental and reproductive safety testing
Learn how developmental and reproductive toxicology study selection supports regulatory decision-making and generates meaningful nonclinical safety data.
Read More

Interest in RWE is not new. Regulators and sponsors have discussed its potential for over a decade, but the sheer volume and messiness of EHR data limited practical use until analytic tools caught up with the data itself.

AI methods turn EHR data into usable signal

AI methods, particularly NLP and ML, are what make EHR mining tractable at scale. The bulk of clinically meaningful information in an EHR sits in free-text notes, such as physician narratives, pathology reports, and discharge summaries, rather than in structured fields like diagnosis codes. A systematic review of NLP systems published in JAMIA Open examining how such tools extract activities-of-daily-living information from unstructured clinical notes found that studies most commonly used deep learning approaches, frequently combined with rule-based or classical ML methods, illustrating a broader pattern in clinical NLP research toward hybrid deep learning pipelines for this kind of extraction task.

Transformer-based language models trained on clinical text, a category that includes domain-adapted variants of BERT, have become a common backbone for this work because they capture context across long, jargon-dense passages that older keyword-matching tools missed. These models can flag diagnoses, medication changes, and adverse events buried in narrative text, then convert that information into structured variables researchers can analyze alongside coded data.

A general framework for AI-driven EHR mining in translational research typically follows a consistent sequence:

  1. Define the clinical question and the target variables needed to answer it, such as a diagnosis, treatment response, or adverse event.
  2. Assemble a representative EHR cohort and document its provenance, including the health systems and time periods covered.
  3. Apply NLP models to extract structured variables from clinical notes, pathology reports, and imaging summaries.
  4. Validate extracted variables against a manually curated subset to quantify accuracy before drawing conclusions.
  5. Combine structured and AI-extracted variables into an analysis-ready dataset suitable for the intended regulatory or research use.

Validation remains the step sponsors are most tempted to shortcut, yet it is the one regulators scrutinize most closely, since an unvalidated extraction pipeline can silently distort downstream conclusions.

RWE increasingly shapes clinical trial design

RWE increasingly informs trial design decisions long before sponsors finalize a protocol, rather than only appearing after a drug reaches the market. Sponsors use RWD to estimate disease incidence, model expected event rates, and refine eligibility criteria so that a trial enrolls a population that reflects the patients most likely to receive the drug after approval. This front-loaded use of RWE can shorten enrollment timelines and reduce the risk of a trial failing simply because it targeted the wrong population.

External control arms built from EHR-derived cohorts represent one of the more mature applications. In oncology, researchers have tested whether EHR-derived patient cohorts can emulate the control arms of published trials that supported prior approvals, an approach that matters most when a placebo or standard-of-care comparator is ethically or practically difficult to run, such as in rare diseases or serious unmet-need indications. This work, described in a peer-reviewed analysis of external control groups in oncology trial development, has shown that external control arms can approximate trial-based comparators under favorable conditions, though results vary with how closely the RWD cohort matches the trial population.

Continue reading below...
3D illustration of a membrane protein embedded within a lipid nanodisc, representing a native-like environment used for membrane protein stabilization and characterization.
Application NoteCharacterizing nanodisc-embedded membrane proteins
Mass photometry supports membrane protein characterization by providing rapid insights into sample composition, purity, and molecular assembly.
Read More

The comparison below illustrates how an RWE-informed approach differs from a trial-only development pathway at a practical level.

AspectTrial-only approachRWE-informed approach
Comparator dataConcurrent randomized control arm collected prospectivelyEHR-derived external control cohort matched retrospectively
Population definitionFixed eligibility criteria set before enrollmentCriteria refined using observed real-world treatment patterns
Timeline pressureFull comparator arm recruitment adds to trial durationComparator data often already exists, shortening timelines
Evidence continuityEvidence effectively ends at trial closeoutEvidence can extend into post-approval monitoring using the same data infrastructure

RWD supports more precise patient stratification

RWD helps identify which patient subgroups are most likely to benefit from a given therapy, refining stratification strategies that were once based largely on trial-derived assumptions. This extends the same logic behind biomarker-driven patient stratification into the post-trial setting: ML models can combine genomic, clinical, and treatment-history variables extracted from EHRs to detect patterns associated with response or resistance, patterns that are often too subtle or too rare to appear reliably within a single randomized trial's sample size.

In oncology specifically, researchers have used ML to extract real-world variables from EHRs at a scale that approximates manually abstracted data, according to a study on replicating real-world evidence in oncology using EHR data extracted by ML. That approach lets research teams study biomarker-defined subgroups across thousands of patients treated in routine practice, a scale few single-institution trials can match.

Stratification built on RWD still carries risk. Cohorts drawn from EHRs reflect whichever health systems captured the data, so subgroup findings can inherit demographic or geographic biases that do not generalize to the broader population a drug is meant to serve. Translational scientists building these models generally pair them with sensitivity analyses across data sources before treating any subgroup signal as robust enough to inform trial design.

AI strengthens post-market safety surveillance

AI-driven surveillance systems now scan EHR narratives for adverse event signals that structured coding alone would miss, strengthening pharmacovigilance after a drug reaches the market. Clinicians frequently document adverse events in free-text clinical notes well before coders code them, if they code them at all, so NLP and ML models trained to recognize adverse drug event language can surface safety signals earlier than claims-based or spontaneous reporting systems.

Researchers have tested several architectures for this purpose, including support vector machines, convolutional neural networks, and transformer-based clinical language models, with newer transformer approaches generally showing stronger generalizability when applied to EHR text from institutions outside their original training data, according to research on detecting adverse drug events from clinical narratives in EHRs. This generalizability question matters for regulators because a safety signal detection tool that performs well on one hospital system's notes but poorly on another's offers limited real-world reliability.

Large-scale distributed data networks add a second layer of surveillance capacity. The EMA's real-world data network, known as DARWIN EU, now draws on roughly 40 data partners covering approximately 250 million patients across Europe and has delivered around 110 studies since its 2022 launch, supporting pharmacovigilance activities alongside its broader role in regulatory review. Networks of this scale let regulators assess background incidence rates and monitor emerging safety questions without waiting for a formal post-marketing study to be designed and executed from scratch.

Continue reading below...
3D illustration of a protein complex composed of clustered spherical subunits arranged in a ring-like oligomeric structure, shown in shades of blue, cyan, and purple against a blue gradient background.
Application NoteUnderstanding protein oligomerization with mass photometry
Automated mass photometry helps reveal the complex dynamics of protein oligomerization and the factors that govern protein assembly.
Read More

The FDA and the EMA are adapting their views on AI-driven RWE

The FDA and the EMA have both built formal pathways for evaluating RWE, though neither treats it as a wholesale substitute for randomized controlled trial data. The FDA's draft guidance on non-interventional studies, issued in March 2024, gives sponsors recommendations for designing observational studies intended to contribute to a demonstration of effectiveness or safety, reflecting the agency's effort to formalize expectations that have accumulated informally over the past decade. The agency's broader RWE work sits within a dedicated program established after mandates in the 21st Century Cures Act pushed the agency to clarify how such evidence could support regulatory decisions.

The EMA has taken a parallel path centered on federated data infrastructure and a published data quality framework rather than a single guidance document, describing how the agency evaluates observational study proposals and data source quality across the DARWIN EU network, with an explicit focus on standardizing data before it ever reaches an evaluation stage. Neither agency has said AI-extracted variables face categorically different scrutiny than manually curated ones, but both have signaled that sponsors must document and validate extraction methods with the same rigor applied to any other analytical approach feeding a regulatory submission.

Translational and regulatory scientists working on programs that also involve AI in clinical trials and adaptive trial design increasingly treat RWE strategy as something to plan early rather than retrofit after pivotal data are in hand. That shift reflects a broader maturation across AI in drug discovery, one where evidence generated from routine care increasingly informs, rather than merely supplements, the decisions sponsors make from early development through post-market monitoring.

Real-world evidence AI drug development moves from pilot to practice

Real-world evidence AI drug development is settling into standard practice rather than remaining an experimental add-on to conventional trials. Sponsors that invest early in validated EHR mining pipelines, well-documented stratification models, and safety surveillance infrastructure put themselves in a stronger position to use RWE credibly across trial design, subgroup analysis, and pharmacovigilance rather than treating each application as a one-off exercise.

The open questions that have followed RWE for a decade, namely data quality, representativeness, and regulatory acceptance, have not disappeared, but AI has given translational and regulatory teams tools sophisticated enough to address them directly rather than defer them indefinitely. As frameworks from the FDA and the EMA continue to mature alongside the underlying AI methods, real-world evidence AI drug development is likely to become less a specialized capability and more a default expectation across the development lifecycle.

This article was produced under Drug Discovery News' AI Editorial Guidelines.

Frequently Asked Questions (FAQs)

  • What is real-world evidence in drug development?

    Real-world evidence is the clinical understanding derived from analyzing real-world data, such as EHRs, claims, and registries, rather than from a randomized controlled trial. It describes how a therapy performs across the variability of routine clinical practice.

  • How does AI use EHR data?

    AI methods, especially NLP and ML, extract structured clinical information from unstructured EHR text, including physician notes and pathology reports. That structured output supports research uses ranging from trial design to safety surveillance.

  • Does the FDA accept real-world evidence for drug approval?

    The FDA accepts real-world evidence for certain regulatory purposes, including some effectiveness and safety demonstrations, but the acceptance depends heavily on data quality and study design. The agency has issued draft guidance describing how sponsors should design non-interventional studies intended to support such submissions.

  • Can real-world data replace randomized clinical trials?

    Real-world data generally supplements rather than replaces randomized controlled trials, particularly for confirming safety signals, building external control arms in specific contexts, and supporting post-market monitoring. Regulators continue to treat randomized trial data as the primary evidentiary standard for most efficacy claims.

  • What are the main challenges in using EHR data for research?

    The main challenges include inconsistent data quality across health systems, missing or unstructured information in clinical notes, and biases introduced by which patient populations a given health system captures. Validating AI-extracted variables against curated reference data helps address these challenges before teams use the evidence in decision-making.

Add Drug Discovery News as a preferred source on Google

Add Drug Discovery News as a preferred Google source to see more of our trusted coverage.

About the Author

Here are some related topics that may interest you:

Loading Next Article...
Loading Next Article...
Subscribe to Newsletter

Subscribe to our eNewsletters

Stay connected with all of the latest from Drug Discovery News.

Subscribe

Sponsored

A scientist in a white lab coat looking into a microscope in a brightly lit modern laboratory.
Learn how developmental and reproductive toxicology study selection supports regulatory decision-making and generates meaningful nonclinical safety data.
Illustration of multiple three-dimensional patient-derived organoids suspended against a dark blue background, representing tumor models used in precision oncology research.
By combining organoid biology with precision automation, researchers developed a miniaturized organoid screening platform that could help speed personalized cancer treatment testing.
Illustration of multiple three-dimensional patient-derived organoids suspended against a dark blue background, representing tumor models used in precision oncology research.
By combining organoid biology with precision automation, researchers developed a miniaturized organoid screening platform that could help speed personalized cancer treatment testing.
Drug Discovery News December 2025 Issue
Latest IssueVolume 21 • Issue 4 • December 2025

December 2025

December 2025 Issue

Explore this issue