- Main toxicity endpoints AI models cover
- hERG prediction is the most mature toxicity model
- Hepatotoxicity and DILI prediction lag behind hERG
- Genotoxicity assessment requires dual QSAR models
- The coverage gap limits toxicity prediction reliability
- Regulatory acceptance of in silico toxicity data
- AI toxicity prediction earns trust one endpoint at a time
AI drug toxicity prediction has become a standard early checkpoint for flagging cardiac, hepatic, and genotoxic liabilities before a candidate advances toward costly in vivo studies. These models can prevent late-stage safety failures that are far more expensive to catch after significant development investment, but toxicity frequently emerges from off-target biology that a model trained on existing assay data was never positioned to capture, which makes knowing a model's coverage boundaries as important as knowing its accuracy score.
Key takeaways
- Cardiotoxicity models targeting the human ether-a-go-go-related gene (hERG) channel report area under the curve (AUC) scores as high as 0.96 under favorable validation splits.
- Drug-induced liver injury (DILI) models report more modest and variable performance, with published two-class prediction accuracy often in the low-to-mid 70% range.
- Genotoxicity assessment under the ICH M7 guideline requires two complementary quantitative structure-activity relationship (QSAR) models, one rule-based and one statistical.
- AI toxicity models are trained on existing assay data, which means they struggle with toxicity mechanisms that are rare, idiosyncratic, or driven by off-target biology absent from training sets.
- Regulators, including the FDA, accept in silico toxicity data as supporting evidence within defined frameworks but have not replaced experimental toxicology requirements with model outputs.
Main toxicity endpoints AI models cover
AI toxicity prediction models today concentrate on a defined set of endpoints where public assay data is deep enough to support reliable modeling: cardiotoxicity through hERG channel blockade; hepatotoxicity, including DILI; genotoxicity and mutagenicity; and a smaller set of models addressing nephrotoxicity and phototoxicity. Each endpoint has its own dominant modeling architecture, driven largely by how much curated public data exists and how mechanistically well-characterized the underlying biology is.
Toxicity-related attrition has always been costly precisely because it tends to surface late, after a compound has already cleared efficacy hurdles and consumed substantial development resources. Historical analysis of drug attrition patterns found that pharmacokinetic-driven failures fell sharply once earlier screening became standard practice, a shift that left safety and efficacy as the harder remaining causes of attrition and underscores why AI drug toxicity prediction has become a priority parallel to the broader effort around AI-powered ADMET prediction. That effort sits within the wider shift toward target identification through clinical translation powered by AI, now underway across drug discovery organizations.
hERG and genotoxicity modeling benefit from decades of regulatory-driven data generation, since both endpoints have long been mandatory components of preclinical safety packages. Hepatotoxicity modeling, by contrast, has had to work with sparser and more heterogeneous data, in part because DILI in humans often reflects idiosyncratic reactions that do not reproduce consistently across standard preclinical species, complicating both data collection and model training.
The distinction between mechanistic and idiosyncratic toxicity is central to understanding why these endpoints model so differently. hERG blockade has a single, well-characterized molecular mechanism: a compound binds the channel and disrupts cardiac repolarization, a relationship that structure-based descriptors capture reasonably well. DILI, by contrast, can arise through multiple independent mechanisms, including direct hepatocellular damage, immune-mediated reactions, and reactive metabolite formation, and a single model rarely captures all of them with equal fidelity.
The table below summarizes reported performance across the toxicity endpoints discussed in this article.
| Toxicity endpoint | Reported benchmark performance | Regulatory framework |
|---|---|---|
| hERG channel blockade | AUC-ROC up to 0.956 random split, 0.922 scaffold split | Standard component of preclinical safety packages |
| Drug-induced liver injury | Two-class accuracy around 72.9% | The FDA's DILI premarketing guidance |
| Genotoxicity and mutagenicity | Dual-model requirement rather than a single accuracy figure | ICH M7 quantitative structure-activity relationship (QSAR) framework |
hERG prediction is the most mature toxicity model
hERG channel blockade prediction is one of the more mature areas of AI drug toxicity prediction, reflecting the endpoint's clinical significance as a leading cause of cardiac safety-related drug withdrawals. A directed message-passing neural network model combined with molecular operating environment descriptors reported its best-performing configuration at an AUC-ROC of 0.956 under random split and 0.922 under scaffold split on a widely used hERG dataset, a gap that illustrates how much harder genuine extrapolation is compared with interpolation within familiar chemical space. That scaffold-split score is a strong result even within the hERG literature specifically, where other published models report figures in the high-0.70s to low-0.80s under similar scaffold-split conditions.
Ensemble learning approaches using molecular fingerprints have reported somewhat more modest but still clinically useful performance, with one ensemble fingerprint study finding 84.9% accuracy and an AUC of 0.887 in cross-validation, dropping to 85.0% accuracy and an AUC of 0.786 on external validation. That external validation figure, lower than the cross-validation score, is the more relevant number for a discovery team assessing how the model will perform on its own chemistry.
Hepatotoxicity and DILI prediction lag behind hERG
Hepatotoxicity prediction, and DILI specifically, remains one of the harder toxicity endpoints to model reliably, largely because the underlying biology is more heterogeneous than a single mechanism like ion channel blockade. The FDA's Liver Toxicity Knowledge Base evaluated more than 1,000 drugs for DILI likelihood, classifying over 700 into most-DILI, less-DILI, and no-DILI categories, and that dataset has become a foundational resource for subsequent modeling work.
Performance on DILI prediction has generally lagged behind hERG and CYP inhibition models. A decision forest DILI model built on the DILIrank classification and trained on structural descriptors from more than 1,000 FDA-approved drugs reported a two-class prediction accuracy of 72.9%, with a sensitivity of 62.8% and a specificity of 79.8%, figures that are clinically useful for prioritization but well short of the near-0.95 area under the curve (AUC) scores reported for some CYP inhibition models. The FDA's own DILI premarketing guidance continues to emphasize clinical laboratory monitoring rather than treating any computational prediction as sufficient on its own, a reflection of how much residual uncertainty remains in this endpoint.
DILI has also carried unusually high regulatory weight for a single toxicity endpoint, having been the most frequent cause of safety-related drug marketing withdrawals for decades. That history is a large part of why the FDA built and maintained the Liver Toxicity Knowledge Base in the first place, and why later machine learning work on DILI prediction has consistently used that dataset as its foundation rather than starting from scratch with smaller, less standardized data sources.
Genotoxicity assessment requires dual QSAR models
Genotoxicity assessment is governed by one of the more formalized regulatory frameworks for in silico toxicology anywhere in drug development. The ICH M7 guideline requires two complementary QSAR methodologies, one expert rule-based system and one statistical model, to assess the mutagenic potential of pharmaceutical synthesis impurities and limit potential carcinogenic risk.
This dual-model requirement exists specifically because a single modeling approach, however well validated, has historically missed distinct classes of mutagenic structural alerts that a complementary method catches. The framework has proven durable, with the guideline reaching successive updates while retaining the same core two-model structure, suggesting regulators view the combination as more robust than any single model regardless of how much any individual QSAR platform improves.
When the two QSAR models disagree, or when a prediction falls outside either model's applicability domain, expert human review is required to resolve the ambiguity rather than defaulting automatically to either model's output. That requirement for expert judgment alongside computational prediction is a template that other toxicity endpoints, including hepatotoxicity, may eventually move toward as their own modeling frameworks mature and gain broader regulatory acceptance.
The coverage gap limits toxicity prediction reliability
Every AI toxicity prediction model inherits the boundaries of its training data, and that constraint is more consequential for toxicity than for most other ADMET endpoints. A model trained on hundreds of documented hERG blockers can reasonably generalize across structurally related chemotypes, but toxicity mechanisms driven by rare off-target receptor binding, idiosyncratic immune-mediated reactions, or metabolite-driven liabilities that only manifest after bioactivation are systematically underrepresented in the datasets models are trained on.
This coverage gap explains why AI toxicity prediction functions best as a prioritization layer rather than a clearance gate. A negative prediction on a well-characterized endpoint like hERG blockade provides real, actionable reassurance. A negative prediction on a poorly characterized or rare toxicity mechanism provides much weaker reassurance, simply because the model had comparatively little relevant data to learn from in the first place. A companion piece on weighing in silico ADMET predictions against in vitro assays lays out a practical framework for making exactly this call.
Clinical case history offers a useful reminder of how real this gap is. Several drugs withdrawn from the market for hepatotoxicity or cardiotoxicity performed adequately on the preclinical toxicity assessments available at the time, precisely because the mechanism responsible for the human safety signal was not well represented in the models or assays used to screen the compound. Modern AI toxicity models inherit the same fundamental constraint: they can only flag what resembles something already seen, and idiosyncratic human toxicity by definition resists that kind of pattern matching.
Discovery and safety teams generally address this gap through a layered approach:
- Applying AI toxicity predictions broadly across virtual libraries to catch well-characterized liabilities early and cheaply.
- Running targeted in vitro confirmation for compounds flagged by any endpoint with strong benchmark performance, such as hERG or CYP inhibition.
- Maintaining full experimental toxicology panels for lead candidates regardless of how favorably they score computationally, specifically to catch mechanisms the models were never positioned to see.
- Updating internal models continuously as new experimental toxicology data accumulates, narrowing the coverage gap over time for a given chemical space.
Regulatory acceptance of in silico toxicity data
Regulators have built a measured, endpoint-specific framework for accepting in silico toxicity data rather than a blanket policy. Under the ICH M7 guideline, QSAR predictions for mutagenic impurity assessment are routinely accepted in regulatory submissions to the FDA, EMA, and other international bodies, and predictions accepted early in development are generally not required to be rerun later unless a specific concern emerges.
More broadly, the FDA's AI development guidance has outlined a risk-based framework emphasizing model transparency, data quality, and validated performance, rather than establishing a single accuracy threshold above which any AI toxicity model becomes automatically acceptable. That framework signals that regulatory acceptance will likely continue to expand endpoint by endpoint, following the same pattern already established for genotoxicity QSAR modeling, as validation evidence accumulates for other toxicity domains.
AI toxicity prediction earns trust one endpoint at a time
AI drug toxicity prediction has reached genuine reliability for well-characterized endpoints with deep public data, particularly hERG channel blockade and, within a formalized dual-model framework, genotoxicity assessment under ICH M7. Hepatotoxicity prediction remains a harder problem, constrained by the heterogeneous and often idiosyncratic biology behind human DILI.
The coverage gap that separates a model's benchmark score from its real-world reliability is not a temporary limitation that better architecture will simply erase. It reflects how much relevant toxicity data actually exists for a given mechanism, and safety teams that keep that distinction in view get the most reliable use out of AI toxicity prediction as a prioritization tool rather than a substitute for experimental toxicology.
This article was produced under Drug Discovery News' AI Editorial Guidelines.













