Articles

Multi-omics integration for drug target discovery: What the data actually shows

What genomic, proteomic, and transcriptomic data integration actually delivers for drug target discovery.
Written byErika Russell
| 6 min read
Researcher reviewing layered omics data visualizations displayed on monitors in a translational research lab setting.

Multi-omics drug target discovery combines genomic, proteomic, and transcriptomic data. Explore what the integration actually delivers for target discovery.

GEMINI (2026)

Register for free to listen to this article
Listen with Speechify
0:00
6:00

Multi-omics drug target discovery promises that combining genomic, proteomic, transcriptomic, and metabolomic layers reveals targets no single dataset could surface on its own. The harder, more practical question is whether that combined signal, increasingly interpreted using artificial intelligence (AI) and machine learning models, is reliable enough to act on, given how noisy and disease-specific each layer can be. This article maps what the integration approach actually delivers, and where its practical limits still show up.

Key takeaways

  • Drug mechanisms with human genetic support are approximately 2.6 times more likely to reach approval, a central rationale for adding genetic evidence to multi-omics target discovery.
  • Large reference resources now profile thousands of proteins and hundreds of metabolic measures across tens of thousands of participants, giving multi-omics integration far more statistical power than a decade ago.
  • Statistical fusion methods such as Multi-Omics Factor Analysis (MOFA) can identify clinically relevant patient subgroups from combined omics layers that single-layer analysis would miss.
  • Batch effects, technical variation unrelated to the biological question being studied, remain a documented and only partially solved challenge across large multi-omics consortia.
  • A pan-cancer proteogenomics analysis spanning more than 1,000 tumor samples identified thousands of candidate druggable proteins, illustrating multi-omics integration's practical scale.

Why multi-omics integration improves target discovery

Single-omics analysis answers a narrow question: does this one data layer, genetic variation or gene expression or protein abundance, point toward a disease-relevant gene. Multi-omics integration asks whether several independent layers agree, and that convergence is a stronger signal than any one layer alone can provide, particularly for the early-stage target identification and validation work that determines whether a candidate advances toward experimental confirmation.

The clearest evidence for why this matters comes from outcomes data rather than methodology papers. Drug mechanisms with human genetic support are approximately 2.6 times more likely to reach approval than mechanisms without it, and that likelihood rises further when the causal gene assignment itself is made with high confidence, precisely the kind of confidence multi-omics integration is designed to build.

Continue reading below...
Researcher using a laptop with a digital DNA helix and molecular biology graphics overlaid, illustrating connected workflows for sequence design, data management, and therapeutic research.
ExplainersExplained: How can molecular biology teams scale therapeutic design with connected workflows?
To keep pace with modern drug discovery, researchers need molecular biology approaches that can support complexity without slowing down the science.
Read More

That said, genetic support alone does not resolve every target hypothesis. Combining genetic evidence with expression and protein-level data helps confirm not just that a gene is disease-associated, but that its protein product is active in the right tissue at the right time, a distinction that matters enormously for translating a statistical association into an actual drug program, and that gives multi-omics integration a second rationale beyond approval odds alone: the type of supporting evidence a target carries predicts not just whether a program succeeds, but how it is likely to fail if it does not.

The main data layers

Multi-omics integration draws on four principal data layers, each capturing a different level of biological information and each with its own strengths and blind spots. Genomic data identifies inherited variation linked to disease risk, while transcriptomic, proteomic, and metabolomic layers capture what a cell is actually doing at a given moment, information genomic data alone cannot supply.

Large reference datasets have transformed what is possible at each layer. Population-scale tissue expression atlases now link genetic variants to tissue-specific transcription, establishing the expression quantitative trait loci framework that underpins much of modern target validation. On the protein side, the UK Biobank's plasma proteomics initiative profiled 2,923 plasma proteins across more than 54,000 participants, identifying over 14,000 genetic associations with those proteins, the large majority previously undescribed.

Metabolomic layers round out the picture, capturing small-molecule metabolic signatures that bulk genomic or transcriptomic data cannot resolve on their own. A large UK Biobank metabolomics resource profiled 249 nuclear magnetic resonance-based metabolic measures, including lipoprotein lipids, fatty acids, and amino acids, across more than 118,000 participants and linked those measures to incidence and mortality across more than 700 diseases, giving multi-omics pipelines a metabolomic layer at a scale that was not available a decade ago. Each additional layer adds statistical power, but also adds a new set of technical and analytical challenges, which is precisely why fusion methods and harmonization practices matter as much as the raw data itself.

The table below summarizes what each layer captures and a representative large-scale resource for it.

Data layer

What it captures

Representative resource

Genomic

Inherited variation linked to disease risk

Large population biobanks with linked exome or genome sequencing

Transcriptomic

Tissue-specific gene expression activity

GTEx Consortium tissue expression atlas

Proteomic

Circulating and tissue protein abundance

UK Biobank plasma proteomics initiative

Metabolomic

Small-molecule metabolic signatures linked to disease

Large biobank metabolomics panels

Machine learning approaches for multi-omics fusion

Fusing several omics layers into a single analysis is a genuine statistical challenge, since each data type has its own scale, noise structure, and missingness pattern. Machine learning fusion methods have moved from simple concatenation of datasets toward purpose-built statistical frameworks designed specifically for heterogeneous, multi-layer biological data.

The MOFA framework, an unsupervised statistical method, was applied to 200 patient samples of chronic lymphocytic leukemia across four data modalities and identified latent factors, including a clinically meaningful patient-subgroup axis, that outperformed single-omics analysis at predicting time to treatment. Later extensions of the same framework adapted MOFA for single-cell, multi-modal data, broadening its use beyond bulk tissue samples.

Continue reading below...
A gloved laboratory technician selects a labeled blood sample tube from a rack containing multiple color-coded collection tubes.
Technology GuidesTechnology Guide: Sample preparation for modern analytical workflows
Analytical performance begins long before a sample reaches the instrument, making sample preparation one of the most important determinants of data quality.
Read More

Similarity-network approaches take a different mathematical route to the same problem, building a separate patient-similarity network for each omics layer and fusing those networks together rather than fusing the raw feature values directly; applied across multiple cancer datasets, this approach has substantially outperformed any single data type at identifying cancer subtypes with distinct survival outcomes. Graph-based deep learning methods represent a newer generation of fusion approaches, using neural networks that learn both within-layer and cross-layer patterns simultaneously rather than fusing pre-computed features, though they typically require larger training datasets than classical statistical methods to perform reliably.

None of these methods is universally superior. Statistical fusion approaches such as MOFA tend to be more interpretable and more forgiving of smaller sample sizes, while graph-based deep learning methods can capture more complex cross-omics relationships when enough training data is available. Choosing between them is a practical decision about dataset size and interpretability needs, not a question with one universally correct answer.

Data quality and harmonization challenges

Combining data layers only helps if the underlying data are actually comparable in the first place, and batch effects, systematic technical variation introduced by different instruments, reagents, or processing times rather than real biology, remain one of the field's most persistent problems. Left uncorrected, batch effects can produce misleading conclusions that look like biological signal but are actually artifacts of how and when samples were processed.

Correcting for this is harder at scale than it sounds. A recent review of batch-effect mitigation across large-scale omics studies notes that large-scale batch effect correction becomes more difficult as the number of separate batches in a consortium database grows, and describes a five-stage correction workflow, assessment, normalization, diagnostics, correction, and re-assessment, that still leaves some technical variation unresolved even when followed carefully. Algorithm choice matters more than intuition might suggest, and the lesson for discovery teams is not to default to the most sophisticated available method, but to validate whichever method is chosen against a known reference standard before trusting its output.

Missing data further compounds the harmonization problem, since not every sample is measured across every omics platform in most real-world cohorts. Researchers increasingly use cross-omics-informed imputation methods, which borrow information from one data layer to fill in gaps in another, though this approach itself introduces model-dependent assumptions that a careful analysis needs to account for rather than treat as free information.

Common harmonization hurdles fall into a few recurring categories:

  • Batch effects introduced by different instruments, reagent lots, or processing dates across a study.
  • Missing values where a sample was profiled on some omics platforms but not others.
  • Inconsistent naming and identifier conventions for genes and proteins across different databases and consortia.
  • Disagreement in the field over whether batch correction should be applied to each omics layer separately or jointly across layers at once.

Case studies: Targets found through multi-omics AI

The clearest demonstration of multi-omics integration's practical scale comes from cancer proteogenomics, where genomic, transcriptomic, proteomic, and phosphoproteomic data have been combined across large public tumor cohorts. A pan-cancer analysis integrating data from more than 1,000 tumor samples across 10 cancer types analyzed thousands of potentially druggable proteins; an independent commentary noted that only a small fraction of those proteins are currently targeted by cancer drugs approved by the FDA.

Continue reading below...
Gloved researcher transferring liquid into a microplate using a multichannel pipette.
EbooksA practical guide to better pipetting
Discover practical strategies to improve pipetting accuracy, reproducibility, ergonomics, and instrument performance across diverse laboratory workflows.
Read More

That same analysis illustrated how multi-omics data can connect a genomic event to an actionable protein-level finding. Researchers found that loss of the tumor suppressor TP53 was associated with a specific phosphorylation change on the protein TOP2A, and that uterine cancer cell lines carrying TP53 loss were more sensitive to a chemotherapy drug that inhibits TOP2A, a finding that ties a genomic marker directly to a therapeutic implication.

Multi-omics case studies like this one share a common structure: a genomic or transcriptomic signal nominates a candidate, a protein-level or functional dataset confirms that the candidate is actually active in the relevant tissue, and a cell-line or clinical dataset provides at least preliminary evidence that the finding has therapeutic relevance. That structure, more than any single algorithm, is what distinguishes a validated multi-omics finding from a purely computational hypothesis.

Cancer remains the disease area where this full structure, computational nomination through cell-line or clinical confirmation, is most consistently completed, largely because cancer has the deepest public multi-omics infrastructure of any disease area. Multi-omics case studies in other disease areas increasingly follow the same computational logic, but readers should note when a published "case study" stops at the computational-prediction stage rather than continuing through to functional or clinical confirmation, since the two are not equivalent evidence, a distinction covered in more depth in a broader look at machine learning's target validation limits.

What multi-omics drug target discovery still requires

Multi-omics drug target discovery has genuinely improved on single-layer analysis, and the evidence for that improvement is strongest where independent data types converge on the same biological conclusion. The technology's real value lies in raising confidence before a target enters costly experimental validation, feeding that confidence into the broader AI-driven drug discovery pipeline that carries a nominated target from initial identification through clinical translation, not in replacing that validation altogether.

Data quality and harmonization remain the field's most practical bottleneck today, more so than any shortage of fusion algorithms. Discovery teams that treat batch effects, missing data, and cross-layer consistency as first-class problems, rather than afterthoughts, are the ones most likely to turn a multi-omics signal into a target that survives the rest of the pipeline, and that discipline matters more to the outcome than which specific fusion method a team ultimately chooses for any given project.

This article was produced under Drug Discovery News' AI Editorial Guidelines.

Frequently Asked Questions (FAQs)

  • How is multi-omics data used in drug target discovery?

    Multi-omics data is used by combining genomic, transcriptomic, proteomic, and metabolomic layers so that a candidate drug target is supported by multiple independent biological signals rather than a single dataset.

  • What machine learning methods work best for multi-omics fusion?

    Statistical frameworks such as MOFA and newer graph-based deep learning methods are commonly used, with the right choice depending on dataset size, the number of omics layers, and whether the goal is patient subgrouping or biomarker identification.

  • What are examples of multi-omics-discovered drug targets?

    A widely cited example is a pan-cancer proteogenomics analysis that combined genomic, transcriptomic, and proteomic data from more than 1,000 tumors to identify thousands of candidate druggable proteins and specific mutation-linked drug sensitivities.

Add Drug Discovery News as a preferred source on Google

Add Drug Discovery News as a preferred Google source to see more of our trusted coverage.

About the Author

Here are some related topics that may interest you:

Loading Next Article...
Loading Next Article...
Subscribe to Newsletter

Subscribe to our eNewsletters

Stay connected with all of the latest from Drug Discovery News.

Subscribe

Sponsored

Researcher using a laptop with a digital DNA helix and molecular biology graphics overlaid, illustrating connected workflows for sequence design, data management, and therapeutic research.
To keep pace with modern drug discovery, researchers need molecular biology approaches that can support complexity without slowing down the science.
Shaping Science graphic featuring the question “How can labs become truly sustainable?” and a photo of James Connelly, Chief Executive Officer of My Green Lab.
Creating more sustainable laboratories depends on practical changes that strengthen scientific performance while reducing environmental impact.
A gloved laboratory technician selects a labeled blood sample tube from a rack containing multiple color-coded collection tubes.
Analytical performance begins long before a sample reaches the instrument, making sample preparation one of the most important determinants of data quality.