Articles

Target validation in the age of AI: What machine learning can and can't confirm

Machine learning ranks targets fast, but only genetics and bench experiments confirm they actually work.
Written byErika Russell
| 7 min read
Researcher in a genomics lab reviewing genetic data beside lab equipment used for target validation.

AI drug target validation speeds up disease association scoring, but lab and genetic evidence still decide which targets truly succeed. Explore the limits.

GEMINI (2026)

Register for free to listen to this article
Listen with Speechify
0:00
7:00

AI drug target validation has become very good at one narrow task: ranking thousands of candidate proteins by evidence weight in hours rather than months. Machine learning target validation tools integrate disease association data at a scale no human curator could match by hand. What these systems still cannot do is prove that hitting a target will change a patient's disease course, the point where most discovery programs still fail.

Key takeaways

  • Artificial intelligence (AI) shortens target ranking by scoring disease associations across large public datasets, but it does not substitute for laboratory proof of mechanism.
  • Genetic evidence, particularly from genome-wide association studies (GWAS) and Mendelian randomization (MR), remains the strongest available predictor of a target's clinical success.
  • Targets with human genetic support are roughly twice as likely to reach drug approval, based on historical analyses of development pipelines.
  • Experimental steps such as gene knockouts, biochemical binding assays, and animal disease models still supply causal proof that computational scores cannot generate on their own.
  • Programs that treat computational ranking as sufficient validation tend to encounter mechanistic failures later, in clinical trials, when the cost of being wrong is highest.

What target validation requires

Target validation asks a narrower and harder question than target identification: does modulating this specific protein or pathway actually change disease biology in a way a drug can exploit, and can that effect be shown, not just inferred. Identifying a plausible target from expression data or a genetic signal is comparatively easy; establishing that engaging it is both necessary and sufficient to alter a disease phenotype is not. That distinction separates computational target validation, which produces a ranked hypothesis, from experimental validation, which tests the hypothesis directly.

A validated target typically needs converging evidence from several independent lines: human genetics linking the gene to the disease, functional data showing that perturbing it changes relevant cell or tissue biology, expression patterns consistent with a disease role, and a safety profile that does not predict intolerable on-target toxicity. No single data type carries the whole burden. A drug discovery scientist weighing a new target is really weighing how many of these independent lines agree, and how strong each one is on its own.

Continue reading below...
3D illustration of a membrane protein embedded within a lipid nanodisc, representing a native-like environment used for membrane protein stabilization and characterization.
Application NoteCharacterizing nanodisc-embedded membrane proteins
Mass photometry supports membrane protein characterization by providing rapid insights into sample composition, purity, and molecular assembly.
Read More

This is also why so many programs stall well before a molecule reaches the clinic. A target that looks compelling in a single dataset, whether an expression signature or a disease association score, can still fail once tested against orthogonal genetic or functional evidence. Validation is the stage built to catch that failure early, before years of medicinal chemistry investment ride on an unproven premise.

Where AI adds value in target validation

Machine learning models excel at one thing in early discovery: integrating enormous, heterogeneous datasets into a single ranked score faster than any team of curators could manage by hand. AI preclinical validation tools pull together genetic association data, expression profiles, pathway annotations, known drug mechanisms, and animal model phenotypes, then weight and combine them into a target-disease association score. The Open Targets Platform, a public resource built for this purpose, now integrates evidence from 23 independent public data sources into ranked target-disease association scores and separately tracks more than 73 million literature-derived target-disease co-occurrences, giving researchers a single place to compare evidence across thousands of candidate genes at once.

This kind of computational target validation is genuinely useful for triage. It lets a discovery team compare a long list of candidates on consistent criteria, surface targets with unexpectedly strong genetic or genomic support, and deprioritize ones that look interesting in one dataset but lack corroborating evidence elsewhere. That is a real efficiency gain over manually cross-referencing separate databases and literature searches for each candidate.

What AI ranking cannot do is confirm causality. A high association score reflects correlation across existing datasets, not a demonstrated mechanistic link between the target and the disease process. Treating a strong computational score as validation, rather than as a prioritization signal, is the most common way AI is misapplied at this stage of drug discovery.

Genetic evidence: GWAS and Mendelian randomization

Human genetic evidence is the single strongest predictor of whether a drug target will hold up in clinical development, and it comes from methods that are older and more rigorously tested than most machine learning scoring systems. A landmark analysis of historical drug development data found that targets with genetic support were roughly twice as likely to result in an approved drug compared with targets lacking such support, a finding subsequently refined and largely reaffirmed as datasets have grown. GWAS supply the raw material for this kind of genetic target validation by identifying loci statistically linked to a disease across large populations, often complementing evidence gathered through broader multi-omics integration efforts that combine genetic, expression, and proteomic data.

Continue reading below...
3D illustration of a protein complex composed of clustered spherical subunits arranged in a ring-like oligomeric structure, shown in shades of blue, cyan, and purple against a blue gradient background.
Application NoteUnderstanding protein oligomerization with mass photometry
Automated mass photometry helps reveal the complex dynamics of protein oligomerization and the factors that govern protein assembly.
Read More

MR goes a step further than a GWAS association by testing whether a specific gene product plausibly causes a change in disease risk, rather than merely correlating with it. MR uses naturally occurring genetic variants as proxies for a protein's activity, exploiting the fact that variants are assigned essentially at random at conception, which limits the confounding that plagues most observational studies. A framework for applying Mendelian randomization to targets has helped formalize how researchers use this method to test whether altering a specific protein's level or activity would be expected to shift disease risk in the same direction a drug is meant to achieve.

Neither GWAS nor MR is infallible. Both depend on population-level data that may not generalize across ancestries, and MR carries its own statistical assumptions that can be violated in ways that are hard to detect. Even so, genetic evidence generated this way carries a kind of causal weight that association scores from expression or literature-mining data simply do not, which is why most rigorous target validation workflows treat it as a near-mandatory checkpoint rather than an optional data source.

The experimental validation AI cannot replace

No computational model, however comprehensive its training data, can substitute for direct laboratory perturbation of a target in a relevant biological system. Functional genomics tools built on clustered regularly interspaced short palindromic repeats (CRISPR) technology allow researchers to knock out or knock down a candidate gene and observe the phenotypic consequences directly, providing the kind of causal evidence that association-based scoring cannot generate. A large-scale CRISPR knockout screen in pancreatic cancer models illustrates the approach, identifying genes whose deletion changed drug sensitivity in ways that pointed to specific combination strategies.

Several categories of bench and in vivo work remain essential regardless of how sophisticated the upstream computational triage becomes:

  • Gene knockout and knockdown studies that test whether removing the target changes a disease-relevant phenotype in cells or model organisms.
  • Biochemical and biophysical binding assays confirming that a candidate molecule engages the target with the expected affinity and selectivity.
  • Animal disease models that test whether modulating the target changes outcomes in a system with intact physiology, not just isolated cells.
  • Orthogonal pharmacological tools, such as a chemical probe with a distinct mechanism from a genetic knockout, used to confirm that an observed effect is target-driven rather than an artifact of one method.

These experiments are slower and more expensive than a computational scoring run, which is precisely why AI is attractive for triage. But they answer a different question than any ranking algorithm can. A knockout study or an animal model result demonstrates that manipulating a specific target in living biology produces a specific outcome, closing a causal loop that no dataset integration, however large, can close on its own.

Derisking drug targets before chemistry starts

Derisking a target means systematically closing the gaps between a computational hypothesis and enough independent evidence to justify committing medicinal chemistry resources. Most experienced discovery teams follow some version of a staged framework before a target earns that commitment:

Continue reading below...
A 3D rendering illustrates a sandwich ELISA technique, where antigen detection is achieved between two layers of antibodies: a capture antibody and a detection antibody
EbooksThe four essentials of immunoassay quality
Learn the key characteristics that determine whether an immunoassay generates accurate and reproducible data.
Read More
  1. Use computational tools to triage and rank candidates across available disease association, genetic, and expression data.
  2. Check for independent human genetic support, ideally from both a GWAS signal and an MR result pointing in a consistent direction.
  3. Run orthogonal functional experiments, such as a genetic knockout paired with a distinct pharmacological tool, to confirm the target's causal role.
  4. Evaluate expression patterns and known biology for signals of on-target toxicity risk before investing further.
  5. Make an explicit go or no-go decision, documenting which evidence types agree and which remain unresolved, rather than allowing a strong computational score to substitute for that decision.

This staged approach reflects a shift already underway in how discovery organizations use AI: as a fast, transparent first filter rather than a final arbiter. Compared with fully manual triage, a computational front end lets teams evaluate far more candidates before committing bench resources, which is a genuine advance. The judgment about which targets clear the bar for chemistry, though, still rests on genetic and experimental evidence that machine learning systems assemble but do not generate themselves.

The table below synthesizes how different evidence types function across this process and where each one's limits sit.

Evidence typeWhat it establishesWhat it cannot establish on its own
AI-integrated association scoringRelative prioritization across large candidate sets using existing public dataCausal relevance of the target to disease biology
GWAS signalStatistical association between a genetic locus and disease risk in a populationWhich gene at the locus drives the effect, or the direction of causality
Mendelian randomizationA causal estimate of how altering a protein's activity is likely to shift disease riskDownstream pharmacology, tissue-specific effects, or dosing behavior of an actual drug
CRISPR knockout or knockdownDirect evidence that removing the target changes a cellular or organismal phenotypeWhether a drug-like molecule can engage the target safely and selectively
Animal disease modelsEvidence that modulating the target changes outcomes in intact physiologyDirect translation to human efficacy or safety at clinical doses

Building the case for AI drug target validation

AI drug target validation, used honestly, means treating machine learning as an evidence-integration and triage tool rather than as proof of mechanism. The programs most likely to avoid late-stage failure are the ones that pair computational ranking with genetic evidence from GWAS and Mendelian randomization, then close the loop with CRISPR-based functional experiments and animal model data before chemistry begins. That combination plays to each method's strength: speed and scale from computation, causal weight from genetics and the bench.

None of this makes AI dispensable. Faster, more comprehensive triage means discovery teams can evaluate more candidates and reach a well-supported shortlist sooner than manual curation would allow. But the final judgment about whether a target is worth years of chemistry and clinical investment still depends on the same genetic and experimental evidence it always has, just organized more efficiently than before. A broader look at how AI reshapes drug discovery, from target identification through clinical translation, situates this validation step within that larger pipeline, and a closer look at AI target identification methods fills in the steps this article covers in less depth.

This article was produced under Drug Discovery News' AI Editorial Guidelines.

Frequently Asked Questions (FAQs)

  • How is AI used in drug target validation?

    AI is mainly used to integrate and rank large volumes of disease association data, genetic evidence, and expression profiles so researchers can prioritize candidate targets faster. It functions as a triage and evidence-integration tool rather than as a substitute for laboratory proof of mechanism.

  • What is genetic target validation?

    Genetic target validation uses human genetic data, such as associations from GWAS, to test whether a candidate gene is plausibly linked to a disease. It is valued because targets with genetic support have historically shown higher rates of clinical and regulatory success.

  • What is Mendelian randomization?

    Mendelian randomization is a statistical method that uses naturally occurring genetic variants as proxies for a protein's activity to estimate whether changing that activity would causally affect disease risk. It reduces the confounding common in standard observational studies because genetic variants are assigned essentially at random at conception.

  • Can AI replace experimental target validation?

    No. AI can rank and prioritize targets using existing data, but only direct experiments, such as gene knockouts, biochemical assays, and animal models, can demonstrate that engaging a target actually changes disease-relevant biology.

Add Drug Discovery News as a preferred source on Google

Add Drug Discovery News as a preferred Google source to see more of our trusted coverage.

About the Author

Here are some related topics that may interest you:

Loading Next Article...
Loading Next Article...
Subscribe to Newsletter

Subscribe to our eNewsletters

Stay connected with all of the latest from Drug Discovery News.

Subscribe

Sponsored

Illustration of multiple three-dimensional patient-derived organoids suspended against a dark blue background, representing tumor models used in precision oncology research.
By combining organoid biology with precision automation, researchers developed a miniaturized organoid screening platform that could help speed personalized cancer treatment testing.
Illustration of multiple three-dimensional patient-derived organoids suspended against a dark blue background, representing tumor models used in precision oncology research.
By combining organoid biology with precision automation, researchers developed a miniaturized organoid screening platform that could help speed personalized cancer treatment testing.
3D illustration of a membrane protein embedded within a lipid nanodisc, representing a native-like environment used for membrane protein stabilization and characterization.
Mass photometry supports membrane protein characterization by providing rapid insights into sample composition, purity, and molecular assembly.
Drug Discovery News December 2025 Issue
Latest IssueVolume 21 • Issue 4 • December 2025

December 2025

December 2025 Issue

Explore this issue