- The ADMET prediction reliability spectrum
- Endpoints reliable enough to act on
- Endpoints where predictions prioritize, not replace, assays
- The chemical space novelty problem limits reliability
- Building a decision framework
- How major pharma groups navigate ADMET reliability
- A documented framework beats ad hoc trust in ADMET AI
The question of in silico ADMET vs. in vitro reliability rarely has a single right answer, because reliability varies sharply by endpoint, by how closely a compound resembles a model's training data, and by how the reported benchmark score was actually generated. Every ADMET prediction model comes with a confidence interval, whether or not a discovery team can see it, and very few organizations have a documented framework for deciding when a predicted score is trustworthy enough to act on without confirmatory testing.
Key takeaways
- Reliability in in silico ADMET vs. in vitro comparisons depends heavily on the specific endpoint, not on machine learning as a category.
- Endpoints with large, well-curated public datasets, including cytochrome P450 (CYP) inhibition and blood-brain barrier permeability, support higher-confidence prediction-led triage.
- Endpoints with sparse or heterogeneous data, such as idiosyncratic hepatotoxicity, should be treated as prioritization signals rather than pass or fail gates.
- Scaffold-split validation scores are a more honest predictor of real-world reliability than random-split cross-validation scores.
- A documented decision framework, rather than ad hoc judgment calls, produces more consistent outcomes across a discovery organization.
The ADMET prediction reliability spectrum
ADMET endpoints do not sit at a single point on a reliability spectrum; they range from genuinely actionable to merely suggestive, and that range is driven almost entirely by data depth rather than by which specific algorithm a vendor or research group happens to favor. Large pharmaceutical organizations learned this lesson the expensive way: pharmacokinetic and safety-related failures once accounted for a substantial share of clinical-stage attrition, and systematic early ADME screening only reduced that share once companies built disciplined processes around when to trust early data rather than adopting screening in name only. AI ADMET prediction raises the same question earlier in the pipeline, and it sits within the broader move toward target identification through clinical translation powered by AI, which has reshaped how discovery teams allocate experimental resources. CYP inhibition, hERG channel blockade, and blood-brain barrier permeability sit toward the reliable end, benefiting from decades of accumulated, standardized public assay data.
Hepatotoxicity, transporter-mediated drug-drug interactions, and rare or idiosyncratic toxicity mechanisms sit toward the less reliable end, constrained by sparser and more heterogeneous training data rather than by any inherent limitation of machine learning as an approach. Positioning a given endpoint correctly on this spectrum, rather than treating all ADMET AI outputs as equally trustworthy, is the foundation of any workable decision framework.
A few signals reliably indicate where a given endpoint sits on that spectrum:
- The size and structural diversity of the public dataset underlying the model, since larger and more diverse training sets generally support more reliable extrapolation.
- Whether reported performance comes from random-split or scaffold-split validation, with scaffold split offering the more honest estimate.
- Whether the endpoint has a single, well-characterized biological mechanism or arises through multiple independent pathways.
- How many independent research groups have reproduced similar performance on the same endpoint using different modeling approaches.
The table below places several common ADMET endpoints along that reliability spectrum.
| Endpoint | Reliability tier | Typical reported performance |
|---|---|---|
| CYP inhibition | Reliable enough to act on | AUC of 0.92 to 0.95 |
| hERG channel blockade | Reliable enough to act on | AUC-ROC up to 0.922 under scaffold split |
| Blood-brain barrier permeability | Reliable enough to act on | Accuracy above 98% in top-performing models |
| Drug-induced liver injury | Prioritize, do not replace assays | Two-class accuracy in the low-to-mid 70% range |
| Transporter-mediated drug-drug interactions | Prioritize, do not replace assays | No standardized benchmark; earlier-stage models |
Endpoints reliable enough to act on
Some endpoints have matured to the point where a prediction alone can reasonably justify deprioritizing a compound, at least as an early triage step ahead of confirmatory work. CYP inhibition models built on tree-based ensembles have reported AUC values as high as 0.92 to 0.95 for major isoforms, a level of performance that machine learning drug metabolism prediction has sustained across multiple independent studies rather than a single favorable dataset.
Blood-brain barrier permeability prediction has reached similarly strong performance, with deep learning models reporting accuracy above 98% on validation sets, and hERG channel blockade prediction has reported AUC-ROC scores above 0.92 under the more demanding scaffold-split validation. For these endpoints, a clearly unfavorable prediction is generally strong enough evidence to justify redesigning a chemical series before committing to synthesis, provided the compound in question resembles the chemical space the model was trained on.
What unites these higher-confidence endpoints is not architecture but data lineage. Each has been the subject of mandatory or near-mandatory experimental screening for years, generating exactly the large, standardized public datasets that machine learning depends on. A discovery team evaluating whether to trust a new model on an unfamiliar endpoint can reasonably ask whether that same data lineage exists before assuming similar reliability will follow.
Endpoints where predictions prioritize, not replace, assays
Other endpoints have not reached that bar, and treating their outputs with the same confidence as CYP inhibition or hERG scores is a common and costly mistake. Hepatotoxicity prediction, including drug-induced liver injury (DILI) modeling, has generally reported two-class prediction accuracy in the low-to-mid 70% range, figures that are useful for prioritization but well short of a level that should override a decision to run confirmatory experimental screening.
A closer look at what AI toxicity models can and cannot tell a safety team about drug risk is worth reading in full, since the coverage gaps behind these lower scores differ meaningfully by mechanism. Transporter-mediated drug-drug interactions fall into the same category, constrained by data scarcity rather than by any fundamental modeling barrier, and predictions there should reorder which assays run first rather than substitute for those assays entirely.
Treating a data-limited endpoint as though it carried the same confidence as a mature one is one of the more common and costly misapplications of ADMET AI. A team that deprioritizes a compound based solely on an unfavorable hepatotoxicity prediction, without confirmatory testing, risks discarding viable chemistry on the strength of a signal that has not earned that level of trust, since the model's own reported performance suggests meaningful uncertainty remains.
The chemical space novelty problem limits reliability
Every benchmark score describes performance on chemistry structurally similar to a model's training set, and that caveat should govern how much weight a team places on any single prediction. Scaffold-split validation, which deliberately separates structurally distinct compounds between training and test data, consistently produces lower reported performance than random-split validation on the same dataset, and that gap is the most honest available signal of how a model will behave on genuinely novel chemistry.
A hERG model trained with directed message-passing networks illustrates the point clearly, reporting its best-performing descriptor configuration at an AUC-ROC of 0.956 under random split but only 0.922 under scaffold split, a meaningful decline that reflects how much harder true extrapolation is than interpolation within already-familiar chemical territory. That scaffold-split figure represents a strong result within the broader hERG literature, where other published models report scaffold-split performance in the high-0.70s to low-0.80s, and a discovery team pushing into a new scaffold series, a novel target class, or unusual physicochemical property space should discount any single benchmark score accordingly, regardless of how strong that score looked on the original validation set.
The practical takeaway is not that scaffold-split scores make prediction useless for novel chemistry, but that the appropriate response to novelty is more experimental confirmation, not less. Teams sometimes make the opposite mistake, leaning more heavily on computational prediction precisely when a program moves into unfamiliar territory because in vitro capacity is limited, which is exactly the situation where prediction reliability is weakest.
Building a decision framework
A workable decision framework does not require abandoning AI ADMET prediction where it adds real value; it requires matching the level of trust placed in a prediction to the evidence supporting that specific endpoint and that specific chemical series. A practical framework follows a consistent sequence:
- Identify which specific endpoint is being predicted and locate its reported scaffold-split, not random-split, benchmark performance.
- Assess how closely the compound in question resembles the model's training distribution, flagging genuinely novel scaffolds for additional scrutiny.
- For high-confidence endpoints and well-represented chemistry, allow the prediction to drive triage decisions with minimal confirmatory testing.
- For lower-confidence endpoints or novel chemical space, treat the prediction as a prioritization signal and route the compound through targeted experimental confirmation before making a final call.
- Feed confirmed experimental results back into the model's training set, narrowing the gap between predicted and observed performance for that chemical series over time.
This sequence converts an implicit, individual judgment call into a documented, repeatable process, which matters because informal trust calibration tends to vary widely between chemists on the same team, let alone across an entire organization.
How major pharma groups navigate ADMET reliability
Large pharmaceutical discovery organizations have generally converged on a tiered approach that mirrors the framework above, formalized into internal decision trees rather than left to individual judgment. Compounds are typically scored against the full battery of available ADMET predictions early, with clear internal thresholds distinguishing endpoints where a prediction can independently justify deprioritization from endpoints requiring mandatory experimental confirmation regardless of the predicted score.
Many organizations also revisit these internal thresholds periodically rather than treating them as fixed, since a model's reliability for a given endpoint can genuinely improve as more experimental data accumulates and gets fed back into training. A threshold set conservatively when a model was new may reasonably shift once several years of confirmatory assay results have validated its performance on the organization's own chemical space.
This tiered structure echoes how regulators themselves treat computational toxicology. Under the ICH M7 guideline for mutagenic impurity assessment, two complementary quantitative structure-activity relationship (QSAR) models are required specifically because no single computational approach has proven reliable enough to stand alone, and the FDA's AI development guidance similarly emphasizes validated performance evidence over blanket acceptance of any model's output. Discovery organizations that build internal frameworks along the same lines tend to get more consistent, defensible outcomes than those relying on case-by-case judgment calls.
A documented framework beats ad hoc trust in ADMET AI
Comparing in silico ADMET vs. in vitro evidence endpoint by endpoint, rather than treating machine learning outputs as uniformly reliable or uniformly suspect, is the single most useful habit a discovery team can build around AI ADMET prediction. The endpoints with the deepest public data support genuine prediction-led triage, while endpoints with sparser or more heterogeneous data still require the experimental caution that has always governed early ADMET screening.
A documented decision framework, applied consistently across chemists and programs, converts that endpoint-by-endpoint judgment into a repeatable process rather than individual intuition. That consistency is what ultimately determines whether AI ADMET prediction accelerates a discovery program or quietly introduces new, undetected risk into candidate selection.
This article was produced under Drug Discovery News' AI Editorial Guidelines.














