Articles

Biomarker-driven patient stratification: How AI is improving clinical trial enrollment

Better patient selection, not better molecules alone, decides many trial outcomes.
Written byErika Russell
| 7 min read
A researcher reviews genomic biomarker data on a screen in a modern clinical research laboratory.

Explore how AI patient stratification clinical trials use biomarker data to enroll the right patients and cut costly late-stage failures across development.

GEMINI (2026)

Register for free to listen to this article
Listen with Speechify
0:00
7:00

AI patient stratification clinical trials are emerging as a direct response to the single biggest driver of clinical trial failure: enrolling patients unlikely to respond to the investigational therapy. Sponsors now pair biomarker enrichment trials with AI models that combine genomic, proteomic, and imaging data to identify subpopulations most likely to benefit before the first patient is dosed. That approach only works when the underlying biological hypothesis is sound, because no algorithm can substitute for a validated mechanism of action.

Key takeaways

  • Enrolling the wrong patients, not a weak drug candidate, is a leading cause of costly late-stage trial failure.
  • Trials that use biomarkers for patient selection show meaningfully higher probability of success (POS) than trials that do not.
  • Machine learning (ML) models can combine genomic, proteomic, and clinical data to stratify patients beyond what a single biomarker allows.
  • Companion diagnostics require contemporaneous regulatory review with the therapeutic product they support under the FDA's framework.
  • Case studies such as the National Cancer Institute's Molecular Analysis for Therapy Choice (NCI-MATCH) trial demonstrate both the promise and the limits of biomarker-matched enrollment.

Why AI patient stratification matters in clinical trials

Patient stratification matters because it directly determines whether a clinical trial can detect a true treatment effect. When a trial enrolls a broad, biologically heterogeneous population, a drug that works well in a defined subgroup can appear ineffective simply because responders are diluted among non-responders. An analysis of more than 400,000 clinical trial records collected between January 2000 and October 2015 found that programs using biomarkers for patient selection in clinical trials achieved a POS nearly twice that of programs without biomarkers, at 10.3% versus 5.5% overall.

The effect is even more pronounced in oncology, historically the therapeutic area with the lowest approval odds. The same analysis of drug development success rates found that oncology programs using biomarkers for patient stratification reached a 10.7% probability of success, compared with just 1.6% for programs that did not stratify patients by a biomarker. That gap illustrates why translational scientists increasingly treat biomarker strategy as a design decision made at the protocol stage, not an afterthought layered on once enrollment stalls.

Continue reading below...
A scientist in a white lab coat looking into a microscope in a brightly lit modern laboratory.
WebinarsChoosing the right study for developmental and reproductive safety testing
Learn how developmental and reproductive toxicology study selection supports regulatory decision-making and generates meaningful nonclinical safety data.
Read More

Precision medicine trial design reframes the enrollment question from "how many patients can be recruited" to "which patients are biologically positioned to respond." This shift matters for every stakeholder in the ecosystem, from biostatisticians sizing the trial to clinical operations teams managing site burden. A trial designed around a validated biomarker hypothesis can be smaller, shorter, and more likely to produce an interpretable result, which is the underlying rationale behind AI patient stratification clinical trials programs now being piloted across sponsors.

Biomarker types and predictive value in trial enrollment

Biomarkers used for stratification fall into several categories, each with distinct predictive value and validation requirements. Prognostic biomarkers indicate the likely course of disease regardless of treatment, while predictive biomarkers specifically forecast response to a given therapy, a distinction the FDA emphasizes in its guidance on trial design. Genomic alterations, protein expression levels, circulating tumor DNA, and imaging-based radiomic signatures are all in active use, particularly in oncology and immunology programs.

The clinical significance of a biomarker depends on how tightly it links to the drug's mechanism of action. Research on biomarker-based lung cancer enrichment, focused on non-small-cell lung cancer (NSCLC) drug approvals from 2003 to 2021, found that biomarker-personalized trials showed improved progression-free survival compared with unselected trials, illustrating how predictive biomarkers tied to a drug target's pathophysiology can improve outcomes within a defined subgroup. Combining several biomarker categories, such as tumor mutational burden alongside protein expression and immune cell profiling, similarly tends to produce more reliable enrichment than any single marker used alone.

No single biomarker category is universally superior; the right choice depends on disease biology and the mechanism under study. Common categories include the following:

  • Genomic biomarkers, such as single-gene mutations or fusion events, that identify patients whose tumors or tissues carry an actionable molecular alteration.
  • Protein expression biomarkers, including receptor overexpression, that indicate whether a targeted therapy's binding site is present at a clinically relevant level.
  • Circulating biomarkers, such as cell-free DNA or specific proteins in blood, that allow less invasive, repeatable sampling during a trial.
  • Immune and inflammatory biomarkers, including cytokine panels and immune cell subsets, that are increasingly used to stratify patients for immunotherapies.
  • Imaging-derived biomarkers, such as radiomic or functional imaging signatures, that capture spatial or metabolic disease features not visible in a single tissue sample.

How machine learning enables multi-biomarker stratification

ML enables trial designers to combine multiple biomarkers into a single predictive signature rather than relying on one marker in isolation. Traditional single-biomarker cutoffs often fail to capture the biological complexity of diseases driven by multiple interacting pathways, which is why integrative approaches to ML-based patient selection have grown alongside the volume of multi-omic data available per patient. Algorithms trained on combined genomic, transcriptomic, and proteomic datasets can reveal patient subgroups that single-marker analysis would miss entirely, particularly in heterogeneous cancers and autoimmune conditions.

These models are also being applied directly to trial operations, not just biological discovery. Integrated models that combine laboratory values, imaging, and molecular data into a single stratification score have shown improved discrimination between patient subgroups compared with any individual biomarker used alone. Natural language processing (NLP) tools applied to unstructured clinical notes and structured eligibility criteria are likewise being tested to help screening teams identify eligible patients for early-phase oncology trials more consistently than manual chart review alone.

Continue reading below...
3D illustration of a membrane protein embedded within a lipid nanodisc, representing a native-like environment used for membrane protein stabilization and characterization.
Application NoteCharacterizing nanodisc-embedded membrane proteins
Mass photometry supports membrane protein characterization by providing rapid insights into sample composition, purity, and molecular assembly.
Read More

The governing constraint remains data quality. Multi-biomarker models are only as reliable as the harmonization of the underlying genomic, clinical, and laboratory datasets feeding them, and biomarker scientists still need to validate any model-derived signature against clinical outcomes before it can inform enrollment decisions. That validation burden is part of a larger shift toward adaptive trial design, in which sponsors adjust enrollment criteria and treatment arms as biomarker data accumulate during the study rather than fixing every parameter at protocol design.

Companion diagnostics and AI-driven trial enrichment

A companion diagnostic is an in vitro diagnostic device or test that is essential to the safe and effective use of a corresponding therapeutic product, and it is the regulatory mechanism that operationalizes biomarker-driven enrollment. The FDA's guidance on in vitro companion diagnostic devices clarifies that, in most circumstances, the FDA should approve or clear the diagnostic and its corresponding therapeutic product contemporaneously for the use indicated in the therapeutic product's labeling. This codevelopment requirement means diagnostic and clinical development timelines must be planned together from an early stage rather than sequenced afterward.

Companion diagnostic AI tools increasingly support this codevelopment process by helping classify complex biomarker signatures, including multi-gene panels and combined genomic-proteomic scores, into binary or tiered eligibility calls that a trial protocol can operationalize. This is a meaningful departure from field practice a decade ago, when most companion diagnostics relied on single-analyte assays such as immunohistochemistry staining for one protein. Multi-analyte, AI-classified diagnostics allow enrichment strategies that would be impractical to implement with manual interpretation of a single marker.

For a translational scientist building a biomarker enrichment plan, the diagnostic pathway is as consequential as the clinical protocol. Choosing an assay that has an existing regulatory pathway, or planning to codevelop one, can shorten site startup timelines because sites need diagnostic infrastructure to prescreen patients before enrollment even opens.

Regulatory requirements for biomarker-driven enrollment

Regulatory expectations for biomarker-driven enrollment are anchored in the FDA's formal guidance on enrichment strategies, which defines the acceptable design frameworks sponsors can use. The FDA's guidance on enrichment strategies for clinical trials describes several categories, including strategies that decrease population heterogeneity, identify patients more likely to have a disease-related endpoint event, and identify patients more likely to respond to treatment based on a pathophysiologic mechanism. Sponsors must justify which category applies and how they will measure and validate the enrichment criterion.

Biomarkers intended for regulatory use in enrollment decisions can also go through the FDA's formal Biomarker Qualification Program (BQP), a three-stage submission process established under the 21st Century Cures Act that lets sponsors qualify a biomarker for a specific context of use across drug development programs. The FDA's resources on biomarker qualification outline the letter of intent, qualification plan, and full qualification package stages sponsors must complete. This qualification pathway is distinct from, but complementary to, the companion diagnostic approval process, since a qualified biomarker still needs an analytically validated assay to measure it in a trial setting.

Continue reading below...
3D illustration of a protein complex composed of clustered spherical subunits arranged in a ring-like oligomeric structure, shown in shades of blue, cyan, and purple against a blue gradient background.
Application NoteUnderstanding protein oligomerization with mass photometry
Automated mass photometry helps reveal the complex dynamics of protein oligomerization and the factors that govern protein assembly.
Read More

The two main regulatory pathways relevant to biomarker-driven enrollment differ in scope and reuse across programs, as shown below.

PathwayPurposeRegulatory basis
Companion diagnostic (CDx) approvalApproves the specific assay used to identify eligible patients for a specific therapeutic product.The FDA's in vitro companion diagnostic devices guidance.
Biomarker Qualification Program (BQP)Qualifies a biomarker for a defined context of use (COU) across multiple drug development programs.The FDA's 21st Century Cures Act three-stage submission process.

Case studies in AI-driven patient stratification

Real-world precision oncology platform trials illustrate both the potential and the operational complexity of biomarker-matched enrollment at scale. The NCI-MATCH trial, sponsored by the National Cancer Institute (NCI) and the ECOG-ACRIN Cancer Research Group, screened nearly 6,000 patients using a sequencing panel covering 143 cancer-related genes and matched them to one of nearly 40 treatment arms based on tumor mutations rather than tumor site of origin. The trial found that 62.5% of its first 6,000 enrolled patients had tumor types outside the four most common cancers, demonstrating that genomic matching can surface actionable treatment opportunities in populations underrepresented in conventional site-specific trials.

Individual arms produced mixed but instructive results. In one substudy, the targeted therapy taselisib showed no objective tumor shrinkage in patients with PIK3CA mutations, yet 24% of those patients had progression-free survival beyond 6 months, a signal the NCI-MATCH trial results suggest warrants further investigation in specific tumor types. Another arm testing ado-trastuzumab emtansine in HER2-overexpressing tumors outside breast and gastric cancer produced partial responses only in patients with rare cancer types, underscoring that even a validated predictive biomarker does not guarantee uniform benefit across every histology that expresses it.

These outcomes reflect a broader lesson for biomarker scientists: a biomarker-matched trial design increases the odds of finding a signal, but it does not eliminate the need for careful hypothesis generation about mechanism and disease subtype. AI-assisted screening and NLP-based eligibility tools, as demonstrated in pilot studies of automated trial matching, can accelerate the operational side of enrollment once a validated biomarker hypothesis exists, but they cannot substitute for that hypothesis.

What AI patient stratification means for trial outcomes

AI patient stratification clinical trials succeed when sponsors pair algorithmic pattern recognition with a defensible biological rationale for why a given biomarker predicts response. The data are consistent on this point: programs that stratify patients by biomarker, whether through a single genomic marker or an ML-derived multi-marker signature, show substantially higher probabilities of success than unstratified programs, a pattern that shows up throughout AI in clinical trials design more broadly, and platform trials such as NCI-MATCH show how that principle scales across dozens of treatment arms simultaneously.

For translational and clinical development scientists, the practical takeaway is sequencing: define the biological hypothesis first, validate the biomarker or companion diagnostic against outcomes, and only then apply ML models to combine markers or automate screening at scale. Skipping that order risks building a sophisticated stratification model on an unvalidated premise, which does nothing to reduce the fundamental risk that continues to drive the majority of late-stage trial failures. Once a stratification model is validated, the same subgroup signals increasingly feed into real-world evidence and AI approaches for post-approval monitoring. This progression is part of a wider pattern across AI applications in drug discovery, where algorithmic tools consistently perform best when scientists layer them onto a validated biological premise rather than use them to generate one from scratch.

This article was produced under Drug Discovery News' AI Editorial Guidelines.

Frequently Asked Questions (FAQs)

  • How does AI improve patient stratification?

    AI improves patient stratification by combining genomic, proteomic, imaging, and clinical data into multi-marker models that identify likely responders more precisely than a single biomarker can, while also automating eligibility screening against unstructured clinical notes.

  • What is biomarker-driven trial enrichment?

    Biomarker-driven trial enrichment is a clinical trial design strategy that selects or restricts enrollment to patients most likely to respond to treatment or experience a disease-related endpoint, based on a genomic, proteomic, or clinical biomarker linked to the drug's mechanism of action.

  • How are companion diagnostics used for patient selection?

    Companion diagnostics are in vitro tests used to identify, before or during enrollment, which patients carry the biomarker required for a therapeutic product's safe and effective use, and regulators generally expect to review the diagnostic and the therapeutic product together.

  • What is precision medicine trial design?

    Precision medicine trial design structures a clinical trial around a predefined biological subgroup rather than a broad, unselected population, using biomarkers or companion diagnostics to align patients with the therapy most likely to benefit them.

Add Drug Discovery News as a preferred source on Google

Add Drug Discovery News as a preferred Google source to see more of our trusted coverage.

About the Author

Here are some related topics that may interest you:

Loading Next Article...
Loading Next Article...
Subscribe to Newsletter

Subscribe to our eNewsletters

Stay connected with all of the latest from Drug Discovery News.

Subscribe

Sponsored

A scientist in a white lab coat looking into a microscope in a brightly lit modern laboratory.
Learn how developmental and reproductive toxicology study selection supports regulatory decision-making and generates meaningful nonclinical safety data.
Illustration of multiple three-dimensional patient-derived organoids suspended against a dark blue background, representing tumor models used in precision oncology research.
By combining organoid biology with precision automation, researchers developed a miniaturized organoid screening platform that could help speed personalized cancer treatment testing.
Illustration of multiple three-dimensional patient-derived organoids suspended against a dark blue background, representing tumor models used in precision oncology research.
By combining organoid biology with precision automation, researchers developed a miniaturized organoid screening platform that could help speed personalized cancer treatment testing.
Drug Discovery News December 2025 Issue
Latest IssueVolume 21 • Issue 4 • December 2025

December 2025

December 2025 Issue

Explore this issue