Articles

AI toxicity prediction: What machine learning models can (and can't) tell you about drug safety

AI toxicity models catch real safety liabilities, but off-target biology still slips past them.
Written byErika Russell
| 7 min read
Scientist reviewing cardiac toxicity risk data on a monitor in a safety pharmacology laboratory.

AI drug toxicity prediction covers hERG, hepatotoxicity, and genotoxicity risk assessment. Discover what these models reliably flag and where they fall short.

GEMINI (2026)

Register for free to listen to this article
Listen with Speechify
0:00
7:00

AI drug toxicity prediction has become a standard early checkpoint for flagging cardiac, hepatic, and genotoxic liabilities before a candidate advances toward costly in vivo studies. These models can prevent late-stage safety failures that are far more expensive to catch after significant development investment, but toxicity frequently emerges from off-target biology that a model trained on existing assay data was never positioned to capture, which makes knowing a model's coverage boundaries as important as knowing its accuracy score.

Key takeaways

  • Cardiotoxicity models targeting the human ether-a-go-go-related gene (hERG) channel report area under the curve (AUC) scores as high as 0.96 under favorable validation splits.
  • Drug-induced liver injury (DILI) models report more modest and variable performance, with published two-class prediction accuracy often in the low-to-mid 70% range.
  • Genotoxicity assessment under the ICH M7 guideline requires two complementary quantitative structure-activity relationship (QSAR) models, one rule-based and one statistical.
  • AI toxicity models are trained on existing assay data, which means they struggle with toxicity mechanisms that are rare, idiosyncratic, or driven by off-target biology absent from training sets.
  • Regulators, including the FDA, accept in silico toxicity data as supporting evidence within defined frameworks but have not replaced experimental toxicology requirements with model outputs.

Main toxicity endpoints AI models cover

AI toxicity prediction models today concentrate on a defined set of endpoints where public assay data is deep enough to support reliable modeling: cardiotoxicity through hERG channel blockade; hepatotoxicity, including DILI; genotoxicity and mutagenicity; and a smaller set of models addressing nephrotoxicity and phototoxicity. Each endpoint has its own dominant modeling architecture, driven largely by how much curated public data exists and how mechanistically well-characterized the underlying biology is.

Toxicity-related attrition has always been costly precisely because it tends to surface late, after a compound has already cleared efficacy hurdles and consumed substantial development resources. Historical analysis of drug attrition patterns found that pharmacokinetic-driven failures fell sharply once earlier screening became standard practice, a shift that left safety and efficacy as the harder remaining causes of attrition and underscores why AI drug toxicity prediction has become a priority parallel to the broader effort around AI-powered ADMET prediction. That effort sits within the wider shift toward target identification through clinical translation powered by AI, now underway across drug discovery organizations.

Continue reading below...
Researcher using a laptop with a digital DNA helix and molecular biology graphics overlaid, illustrating connected workflows for sequence design, data management, and therapeutic research.
ExplainersExplained: How can molecular biology teams scale therapeutic design with connected workflows?
To keep pace with modern drug discovery, researchers need molecular biology approaches that can support complexity without slowing down the science.
Read More

hERG and genotoxicity modeling benefit from decades of regulatory-driven data generation, since both endpoints have long been mandatory components of preclinical safety packages. Hepatotoxicity modeling, by contrast, has had to work with sparser and more heterogeneous data, in part because DILI in humans often reflects idiosyncratic reactions that do not reproduce consistently across standard preclinical species, complicating both data collection and model training.

The distinction between mechanistic and idiosyncratic toxicity is central to understanding why these endpoints model so differently. hERG blockade has a single, well-characterized molecular mechanism: a compound binds the channel and disrupts cardiac repolarization, a relationship that structure-based descriptors capture reasonably well. DILI, by contrast, can arise through multiple independent mechanisms, including direct hepatocellular damage, immune-mediated reactions, and reactive metabolite formation, and a single model rarely captures all of them with equal fidelity.

The table below summarizes reported performance across the toxicity endpoints discussed in this article.

Toxicity endpointReported benchmark performanceRegulatory framework
hERG channel blockadeAUC-ROC up to 0.956 random split, 0.922 scaffold splitStandard component of preclinical safety packages
Drug-induced liver injuryTwo-class accuracy around 72.9%The FDA's DILI premarketing guidance
Genotoxicity and mutagenicityDual-model requirement rather than a single accuracy figureICH M7 quantitative structure-activity relationship (QSAR) framework

hERG prediction is the most mature toxicity model

hERG channel blockade prediction is one of the more mature areas of AI drug toxicity prediction, reflecting the endpoint's clinical significance as a leading cause of cardiac safety-related drug withdrawals. A directed message-passing neural network model combined with molecular operating environment descriptors reported its best-performing configuration at an AUC-ROC of 0.956 under random split and 0.922 under scaffold split on a widely used hERG dataset, a gap that illustrates how much harder genuine extrapolation is compared with interpolation within familiar chemical space. That scaffold-split score is a strong result even within the hERG literature specifically, where other published models report figures in the high-0.70s to low-0.80s under similar scaffold-split conditions.

Ensemble learning approaches using molecular fingerprints have reported somewhat more modest but still clinically useful performance, with one ensemble fingerprint study finding 84.9% accuracy and an AUC of 0.887 in cross-validation, dropping to 85.0% accuracy and an AUC of 0.786 on external validation. That external validation figure, lower than the cross-validation score, is the more relevant number for a discovery team assessing how the model will perform on its own chemistry.

Hepatotoxicity and DILI prediction lag behind hERG

Hepatotoxicity prediction, and DILI specifically, remains one of the harder toxicity endpoints to model reliably, largely because the underlying biology is more heterogeneous than a single mechanism like ion channel blockade. The FDA's Liver Toxicity Knowledge Base evaluated more than 1,000 drugs for DILI likelihood, classifying over 700 into most-DILI, less-DILI, and no-DILI categories, and that dataset has become a foundational resource for subsequent modeling work.

Performance on DILI prediction has generally lagged behind hERG and CYP inhibition models. A decision forest DILI model built on the DILIrank classification and trained on structural descriptors from more than 1,000 FDA-approved drugs reported a two-class prediction accuracy of 72.9%, with a sensitivity of 62.8% and a specificity of 79.8%, figures that are clinically useful for prioritization but well short of the near-0.95 area under the curve (AUC) scores reported for some CYP inhibition models. The FDA's own DILI premarketing guidance continues to emphasize clinical laboratory monitoring rather than treating any computational prediction as sufficient on its own, a reflection of how much residual uncertainty remains in this endpoint.

Continue reading below...
A gloved laboratory technician selects a labeled blood sample tube from a rack containing multiple color-coded collection tubes.
Technology GuidesTechnology Guide: Sample preparation for modern analytical workflows
Analytical performance begins long before a sample reaches the instrument, making sample preparation one of the most important determinants of data quality.
Read More

DILI has also carried unusually high regulatory weight for a single toxicity endpoint, having been the most frequent cause of safety-related drug marketing withdrawals for decades. That history is a large part of why the FDA built and maintained the Liver Toxicity Knowledge Base in the first place, and why later machine learning work on DILI prediction has consistently used that dataset as its foundation rather than starting from scratch with smaller, less standardized data sources.

Genotoxicity assessment requires dual QSAR models

Genotoxicity assessment is governed by one of the more formalized regulatory frameworks for in silico toxicology anywhere in drug development. The ICH M7 guideline requires two complementary QSAR methodologies, one expert rule-based system and one statistical model, to assess the mutagenic potential of pharmaceutical synthesis impurities and limit potential carcinogenic risk.

This dual-model requirement exists specifically because a single modeling approach, however well validated, has historically missed distinct classes of mutagenic structural alerts that a complementary method catches. The framework has proven durable, with the guideline reaching successive updates while retaining the same core two-model structure, suggesting regulators view the combination as more robust than any single model regardless of how much any individual QSAR platform improves.

When the two QSAR models disagree, or when a prediction falls outside either model's applicability domain, expert human review is required to resolve the ambiguity rather than defaulting automatically to either model's output. That requirement for expert judgment alongside computational prediction is a template that other toxicity endpoints, including hepatotoxicity, may eventually move toward as their own modeling frameworks mature and gain broader regulatory acceptance.

The coverage gap limits toxicity prediction reliability

Every AI toxicity prediction model inherits the boundaries of its training data, and that constraint is more consequential for toxicity than for most other ADMET endpoints. A model trained on hundreds of documented hERG blockers can reasonably generalize across structurally related chemotypes, but toxicity mechanisms driven by rare off-target receptor binding, idiosyncratic immune-mediated reactions, or metabolite-driven liabilities that only manifest after bioactivation are systematically underrepresented in the datasets models are trained on.

This coverage gap explains why AI toxicity prediction functions best as a prioritization layer rather than a clearance gate. A negative prediction on a well-characterized endpoint like hERG blockade provides real, actionable reassurance. A negative prediction on a poorly characterized or rare toxicity mechanism provides much weaker reassurance, simply because the model had comparatively little relevant data to learn from in the first place. A companion piece on weighing in silico ADMET predictions against in vitro assays lays out a practical framework for making exactly this call.

Continue reading below...
Gloved researcher transferring liquid into a microplate using a multichannel pipette.
EbooksA practical guide to better pipetting
Discover practical strategies to improve pipetting accuracy, reproducibility, ergonomics, and instrument performance across diverse laboratory workflows.
Read More

Clinical case history offers a useful reminder of how real this gap is. Several drugs withdrawn from the market for hepatotoxicity or cardiotoxicity performed adequately on the preclinical toxicity assessments available at the time, precisely because the mechanism responsible for the human safety signal was not well represented in the models or assays used to screen the compound. Modern AI toxicity models inherit the same fundamental constraint: they can only flag what resembles something already seen, and idiosyncratic human toxicity by definition resists that kind of pattern matching.

Discovery and safety teams generally address this gap through a layered approach:

  • Applying AI toxicity predictions broadly across virtual libraries to catch well-characterized liabilities early and cheaply.
  • Running targeted in vitro confirmation for compounds flagged by any endpoint with strong benchmark performance, such as hERG or CYP inhibition.
  • Maintaining full experimental toxicology panels for lead candidates regardless of how favorably they score computationally, specifically to catch mechanisms the models were never positioned to see.
  • Updating internal models continuously as new experimental toxicology data accumulates, narrowing the coverage gap over time for a given chemical space.

Regulatory acceptance of in silico toxicity data

Regulators have built a measured, endpoint-specific framework for accepting in silico toxicity data rather than a blanket policy. Under the ICH M7 guideline, QSAR predictions for mutagenic impurity assessment are routinely accepted in regulatory submissions to the FDA, EMA, and other international bodies, and predictions accepted early in development are generally not required to be rerun later unless a specific concern emerges.

More broadly, the FDA's AI development guidance has outlined a risk-based framework emphasizing model transparency, data quality, and validated performance, rather than establishing a single accuracy threshold above which any AI toxicity model becomes automatically acceptable. That framework signals that regulatory acceptance will likely continue to expand endpoint by endpoint, following the same pattern already established for genotoxicity QSAR modeling, as validation evidence accumulates for other toxicity domains.

AI toxicity prediction earns trust one endpoint at a time

AI drug toxicity prediction has reached genuine reliability for well-characterized endpoints with deep public data, particularly hERG channel blockade and, within a formalized dual-model framework, genotoxicity assessment under ICH M7. Hepatotoxicity prediction remains a harder problem, constrained by the heterogeneous and often idiosyncratic biology behind human DILI.

The coverage gap that separates a model's benchmark score from its real-world reliability is not a temporary limitation that better architecture will simply erase. It reflects how much relevant toxicity data actually exists for a given mechanism, and safety teams that keep that distinction in view get the most reliable use out of AI toxicity prediction as a prioritization tool rather than a substitute for experimental toxicology.

This article was produced under Drug Discovery News' AI Editorial Guidelines.

Frequently Asked Questions (FAQs)

  • How accurate is AI drug toxicity prediction?

    Accuracy varies substantially by endpoint, with hERG cardiotoxicity models reporting AUC scores up to 0.96 under favorable validation conditions, while drug-induced liver injury models typically report lower two-class prediction accuracy in the low-to-mid 70% range.

  • What is hERG prediction?

    hERG prediction estimates whether a compound will block the human ether-a-go-go-related gene potassium channel, a mechanism strongly associated with cardiac arrhythmia risk and a leading cause of drug safety withdrawals.

  • Can AI predict drug-induced liver injury?

    AI models can predict DILI risk with moderate reliability, with published decision forest models reporting roughly 73% two-class prediction accuracy, but the FDA's guidance still emphasizes clinical laboratory monitoring rather than treating computational predictions as sufficient evidence alone.

  • Does the FDA accept in silico toxicity predictions?

    The FDA accepts in silico toxicity predictions within defined frameworks, most notably QSAR modeling for genotoxic impurity assessment under ICH M7, and has outlined a broader risk-based approach to evaluating artificial intelligence tools in drug development.

Add Drug Discovery News as a preferred source on Google

Add Drug Discovery News as a preferred Google source to see more of our trusted coverage.

About the Author

Here are some related topics that may interest you:

Loading Next Article...
Loading Next Article...
Subscribe to Newsletter

Subscribe to our eNewsletters

Stay connected with all of the latest from Drug Discovery News.

Subscribe

Sponsored

Researcher using a laptop with a digital DNA helix and molecular biology graphics overlaid, illustrating connected workflows for sequence design, data management, and therapeutic research.
To keep pace with modern drug discovery, researchers need molecular biology approaches that can support complexity without slowing down the science.
Shaping Science graphic featuring the question “How can labs become truly sustainable?” and a photo of James Connelly, Chief Executive Officer of My Green Lab.
Creating more sustainable laboratories depends on practical changes that strengthen scientific performance while reducing environmental impact.
A gloved laboratory technician selects a labeled blood sample tube from a rack containing multiple color-coded collection tubes.
Analytical performance begins long before a sample reaches the instrument, making sample preparation one of the most important determinants of data quality.