When a Phase 3 clinical trial misses its primary endpoint, the outcome is often treated as definitive. The drug failed to demonstrate efficacy, development stops, and the therapy may never be tested again.
But a negative trial in a heterogeneous patient population does not necessarily mean that every participant failed to benefit. It may instead mean that any treatment effect was diluted when patients with very different disease biology and treatment responses were analyzed together.
That possibility has become increasingly important as drug development moves toward precision medicine. Many diseases once treated as single disorders are now understood to encompass biologically diverse patient populations. But pivotal Phase 3 trials are still generally designed to answer a binary question: does the drug work across the overall study population?
Now, advances in AI are making it possible to search clinical trial datasets for patterns that conventional analyses may miss. At the Alzheimer’s Association International Conference 2026, Jan Sedway and colleagues presented a reanalysis of the Anti-Amyloid Treatment in Asymptomatic Alzheimer’s Disease (A4) Study, a closely watched trial of Eli Lilly’s anti-amyloid antibody solanezumab. Although the trial failed to meet its primary endpoint, the AI-based analysis identified two clinically interpretable subgroups of patients who appeared to experience less cognitive decline with treatment than with placebo.
While the findings do not overturn the original trial, they do illustrate how new analytical approaches could change the way researchers interpret negative studies and potentially identify patient populations worth studying further.
Looking beyond the average patient
The conventional clinical trial framework was developed to produce robust evidence about whether a treatment is safe and effective across a broad population. Large Phase 3 studies enroll hundreds or thousands of participants to ensure adequate statistical power and to capture uncommon adverse events.
However, patients enrolled in the same study can differ substantially in their genetics, disease biology, rate of progression, comorbidities, and response to treatment. As a result, a drug that produces a meaningful benefit in one subset of participants may appear ineffective when its effect is averaged across an entire cohort. A study can therefore miss its primary endpoint even though a clinically meaningful response exists in a subset of participants.
“Clinical trials are designed to look at the average of all of these people in a very heterogeneous population,” Sedway, Senior Vice President of Clinical Science at NetraMark, told DDN. “Having visibility into who's responding and who's not … will help the industry in so many ways.”
This is not an entirely new concern. Researchers have long examined predefined subgroups based on characteristics such as age, sex, disease severity, or biomarkers. In some cases, post hoc analyses have revealed important differences in efficacy or safety that were not apparent in the overall trial population.
The challenge is that conventional subgroup analyses can struggle to detect more complex interactions, particularly in smaller datasets. A treatment response may depend not on sex or age alone, for example, but on a combination of cognitive performance, imaging findings, disease stage, and other baseline characteristics.
Identifying these higher-order interactions can require testing large numbers of possible combinations, increasing the risk of false-positive findings and making it harder to distinguish clinically meaningful patterns from statistical noise.
“What we are able to do is to look at everything altogether and therefore find the most salient aspects,” Sedway said. The Netra AI approach searches across multiple variables to identify combinations that may distinguish responders from nonresponders.
Importantly, the goal is not simply to find the most statistically distinctive subgroup, but to identify a population that could actually be used in a clinical trial. “We really want to be sure that we are identifying populations that are useful,” explained Sedway. To help achieve that, the model-derived subgroups are limited to between one and four variables. This prevents the model from defining an extremely narrow population using a long list of characteristics that may be statistically interesting but impractical to identify or recruit in a clinical trial.
Revisiting the A4 trial
The A4 study enrolled clinically asymptomatic older adults with elevated brain amyloid, a population thought to be at increased risk of developing Alzheimer's disease. Participants were randomized to receive solanezumab or placebo, with change in the Preclinical Alzheimer Cognitive Composite (PACC) over 240 weeks serving as the primary endpoint.
The trial generated significant interest because it targeted patients at a very early stage of disease, when many researchers believe interventions are most likely to succeed. However, the study failed to demonstrate a statistically significant benefit, and solanezumab was widely viewed as another disappointment in the long and challenging history of Alzheimer’s drug development.
For Sedway and colleagues, however, the dataset provided an opportunity to ask whether the average result concealed differences between patients. The researchers first created an enriched sample that included participants who had progressed to mild cognitive impairment and a matched group who remained cognitively unimpaired. They then applied the Netra AI process to search for baseline characteristics associated with differential treatment response.
The analysis went through multiple stages designed to identify stable and clinically meaningful patterns rather than simply selecting the strongest statistical association. Candidate subgroups were selected based on the consistency of the treatment effect, the stability of the identified pattern, and whether the defining features could be clinically interpreted.
The researchers then used counterfactual causal inference to assess how those characteristics might translate across the full randomized cohort of 824 participants. “We basically estimate what someone who was on placebo may have done if they were on the drug and vice versa,” Sedway explained.
The analysis identified two patient profiles that appeared to derive greater benefit from solanezumab.
One subgroup, representing about 46 percent of the full trial cohort, was characterized by higher baseline right amygdala volume and stronger baseline performance on the Digit Symbol Substitution Test (DSST), a measure of processing speed and attention. These participants showed attenuated cognitive decline based on the PACC Total Sco with solanezumab compared with placebo.
The second subgroup, comprising about one-quarter of participants, was characterized by higher baseline right superior temporal volume and higher DSST performance. This group showed even greater separation between treatment and placebo arms in terms of PACC total score.
“Each of those [groups] essentially was saying the same thing,” Sedway said. “Both groups had better baseline cognitive scale performance and more volume in brain regions involved in processing social and emotional information.”
Those findings raise the possibility that anti-amyloid therapy may have been more effective in patients whose disease had not progressed as far, or whose brains retained greater structural integrity at the time treatment began. However, this analysis was retrospective and exploratory. The findings must be replicated in independent datasets and evaluated with appropriate corrections for multiple testing before they could support clinical or regulatory decisions.
From post hoc discovery to prospective trials
The A4 analysis was conducted after the trial had been completed. That means the findings cannot simply be used to conclude that solanezumab is effective in those populations. Post hoc findings can generate hypotheses, but they need to be replicated independently and prospectively before they can support clinical or regulatory decisions.
Sedway doesn’t necessarily see that as a limitation of the approach. “The goal is to identify something useful post hoc that can be prospectively applied to the next trial,” she said.
Phase 1 and 2 studies typically involve much smaller patient populations than Phase 3 trials, making conventional statistical analyses more challenging. But, if a meaningful treatment-response pattern can be identified in an early-stage study, developers could potentially use that information to stratify patients or otherwise inform the design of a subsequent trial.
“ That's really where NetraAI shines — its ability to identify meaningful subgroups even when using small datasets,” Sedway said.
The company is now seeking opportunities to demonstrate that process prospectively, ideally by using information from early-stage trials to identify patient populations and recruiting them in a subsequent Phase 3 trial. That would provide a much stronger test of whether AI-derived subgroups can improve clinical development than retrospectively finding a signal in a completed study.
Could failed drugs deserve a second look?
The A4 reanalysis does not change the original conclusion that solanezumab failed to meet its primary endpoint in the overall study population. Nor does it establish that the drug is effective for the two subgroups identified by the AI.
What it does suggest is that the binary labels applied to clinical trials — success or failure — may not capture the full biological complexity contained within a large dataset.
These implications extend far beyond Alzheimer's disease. Drug development is expensive, and a Phase 3 failure can result in years of research and hundreds of millions of dollars being written off. If a therapy genuinely benefits a well-defined subset of patients, identifying that population could provide a rationale for a new, more focused trial rather than abandoning the asset altogether.
That does not mean every negative trial contains a hidden success. Many drugs truly do not work. But as clinical datasets become richer and analytical methods improve, researchers may be able to distinguish between therapies that failed broadly and therapies that failed to show an average benefit across a heterogeneous population.
The idea also fits with a broader shift in clinical research. As therapies become more targeted, the traditional model of enrolling large, diverse populations and evaluating a single average treatment effect may become increasingly inefficient. Future trials could instead use model derived subgroups defined by biomarkers, imaging results or cognitive measures to enrich trials with patients most likely to respond.
As precision medicine advances, the question facing drug developers may increasingly shift from whether a therapy works for the average patient to understanding which patients are most likely to benefit. Rather than representing the end of a drug's story, negative trial results could provide the data needed to find the right patients for the drug.











