Articles

In silico ADMET vs. in vitro assays: When can you trust the prediction?

Every ADMET model has a confidence interval; few teams have a rule for when to act on it.
Written byErika Russell
| 6 min read
Two scientists reviewing ADMET prediction confidence data during a discovery team meeting.

In silico ADMET vs. in vitro reliability varies sharply by endpoint. Discover a practical decision framework for knowing when to trust a prediction fully.

GEMINI (2026)

Register for free to listen to this article
Listen with Speechify
0:00
6:00

The question of in silico ADMET vs. in vitro reliability rarely has a single right answer, because reliability varies sharply by endpoint, by how closely a compound resembles a model's training data, and by how the reported benchmark score was actually generated. Every ADMET prediction model comes with a confidence interval, whether or not a discovery team can see it, and very few organizations have a documented framework for deciding when a predicted score is trustworthy enough to act on without confirmatory testing.

Key takeaways

  • Reliability in in silico ADMET vs. in vitro comparisons depends heavily on the specific endpoint, not on machine learning as a category.
  • Endpoints with large, well-curated public datasets, including cytochrome P450 (CYP) inhibition and blood-brain barrier permeability, support higher-confidence prediction-led triage.
  • Endpoints with sparse or heterogeneous data, such as idiosyncratic hepatotoxicity, should be treated as prioritization signals rather than pass or fail gates.
  • Scaffold-split validation scores are a more honest predictor of real-world reliability than random-split cross-validation scores.
  • A documented decision framework, rather than ad hoc judgment calls, produces more consistent outcomes across a discovery organization.

The ADMET prediction reliability spectrum

ADMET endpoints do not sit at a single point on a reliability spectrum; they range from genuinely actionable to merely suggestive, and that range is driven almost entirely by data depth rather than by which specific algorithm a vendor or research group happens to favor. Large pharmaceutical organizations learned this lesson the expensive way: pharmacokinetic and safety-related failures once accounted for a substantial share of clinical-stage attrition, and systematic early ADME screening only reduced that share once companies built disciplined processes around when to trust early data rather than adopting screening in name only. AI ADMET prediction raises the same question earlier in the pipeline, and it sits within the broader move toward target identification through clinical translation powered by AI, which has reshaped how discovery teams allocate experimental resources. CYP inhibition, hERG channel blockade, and blood-brain barrier permeability sit toward the reliable end, benefiting from decades of accumulated, standardized public assay data.

Continue reading below...
A scientist in a white lab coat looking into a microscope in a brightly lit modern laboratory.
WebinarsChoosing the right study for developmental and reproductive safety testing
Learn how developmental and reproductive toxicology study selection supports regulatory decision-making and generates meaningful nonclinical safety data.
Read More

Hepatotoxicity, transporter-mediated drug-drug interactions, and rare or idiosyncratic toxicity mechanisms sit toward the less reliable end, constrained by sparser and more heterogeneous training data rather than by any inherent limitation of machine learning as an approach. Positioning a given endpoint correctly on this spectrum, rather than treating all ADMET AI outputs as equally trustworthy, is the foundation of any workable decision framework.

A few signals reliably indicate where a given endpoint sits on that spectrum:

  • The size and structural diversity of the public dataset underlying the model, since larger and more diverse training sets generally support more reliable extrapolation.
  • Whether reported performance comes from random-split or scaffold-split validation, with scaffold split offering the more honest estimate.
  • Whether the endpoint has a single, well-characterized biological mechanism or arises through multiple independent pathways.
  • How many independent research groups have reproduced similar performance on the same endpoint using different modeling approaches.

The table below places several common ADMET endpoints along that reliability spectrum.

EndpointReliability tierTypical reported performance
CYP inhibitionReliable enough to act onAUC of 0.92 to 0.95
hERG channel blockadeReliable enough to act onAUC-ROC up to 0.922 under scaffold split
Blood-brain barrier permeabilityReliable enough to act onAccuracy above 98% in top-performing models
Drug-induced liver injuryPrioritize, do not replace assaysTwo-class accuracy in the low-to-mid 70% range
Transporter-mediated drug-drug interactionsPrioritize, do not replace assaysNo standardized benchmark; earlier-stage models

Endpoints reliable enough to act on

Some endpoints have matured to the point where a prediction alone can reasonably justify deprioritizing a compound, at least as an early triage step ahead of confirmatory work. CYP inhibition models built on tree-based ensembles have reported AUC values as high as 0.92 to 0.95 for major isoforms, a level of performance that machine learning drug metabolism prediction has sustained across multiple independent studies rather than a single favorable dataset.

Blood-brain barrier permeability prediction has reached similarly strong performance, with deep learning models reporting accuracy above 98% on validation sets, and hERG channel blockade prediction has reported AUC-ROC scores above 0.92 under the more demanding scaffold-split validation. For these endpoints, a clearly unfavorable prediction is generally strong enough evidence to justify redesigning a chemical series before committing to synthesis, provided the compound in question resembles the chemical space the model was trained on.

What unites these higher-confidence endpoints is not architecture but data lineage. Each has been the subject of mandatory or near-mandatory experimental screening for years, generating exactly the large, standardized public datasets that machine learning depends on. A discovery team evaluating whether to trust a new model on an unfamiliar endpoint can reasonably ask whether that same data lineage exists before assuming similar reliability will follow.

Endpoints where predictions prioritize, not replace, assays

Other endpoints have not reached that bar, and treating their outputs with the same confidence as CYP inhibition or hERG scores is a common and costly mistake. Hepatotoxicity prediction, including drug-induced liver injury (DILI) modeling, has generally reported two-class prediction accuracy in the low-to-mid 70% range, figures that are useful for prioritization but well short of a level that should override a decision to run confirmatory experimental screening.

A closer look at what AI toxicity models can and cannot tell a safety team about drug risk is worth reading in full, since the coverage gaps behind these lower scores differ meaningfully by mechanism. Transporter-mediated drug-drug interactions fall into the same category, constrained by data scarcity rather than by any fundamental modeling barrier, and predictions there should reorder which assays run first rather than substitute for those assays entirely.

Continue reading below...
3D illustration of a membrane protein embedded within a lipid nanodisc, representing a native-like environment used for membrane protein stabilization and characterization.
Application NoteCharacterizing nanodisc-embedded membrane proteins
Mass photometry supports membrane protein characterization by providing rapid insights into sample composition, purity, and molecular assembly.
Read More

Treating a data-limited endpoint as though it carried the same confidence as a mature one is one of the more common and costly misapplications of ADMET AI. A team that deprioritizes a compound based solely on an unfavorable hepatotoxicity prediction, without confirmatory testing, risks discarding viable chemistry on the strength of a signal that has not earned that level of trust, since the model's own reported performance suggests meaningful uncertainty remains.

The chemical space novelty problem limits reliability

Every benchmark score describes performance on chemistry structurally similar to a model's training set, and that caveat should govern how much weight a team places on any single prediction. Scaffold-split validation, which deliberately separates structurally distinct compounds between training and test data, consistently produces lower reported performance than random-split validation on the same dataset, and that gap is the most honest available signal of how a model will behave on genuinely novel chemistry.

A hERG model trained with directed message-passing networks illustrates the point clearly, reporting its best-performing descriptor configuration at an AUC-ROC of 0.956 under random split but only 0.922 under scaffold split, a meaningful decline that reflects how much harder true extrapolation is than interpolation within already-familiar chemical territory. That scaffold-split figure represents a strong result within the broader hERG literature, where other published models report scaffold-split performance in the high-0.70s to low-0.80s, and a discovery team pushing into a new scaffold series, a novel target class, or unusual physicochemical property space should discount any single benchmark score accordingly, regardless of how strong that score looked on the original validation set.

The practical takeaway is not that scaffold-split scores make prediction useless for novel chemistry, but that the appropriate response to novelty is more experimental confirmation, not less. Teams sometimes make the opposite mistake, leaning more heavily on computational prediction precisely when a program moves into unfamiliar territory because in vitro capacity is limited, which is exactly the situation where prediction reliability is weakest.

Building a decision framework

A workable decision framework does not require abandoning AI ADMET prediction where it adds real value; it requires matching the level of trust placed in a prediction to the evidence supporting that specific endpoint and that specific chemical series. A practical framework follows a consistent sequence:

  1. Identify which specific endpoint is being predicted and locate its reported scaffold-split, not random-split, benchmark performance.
  2. Assess how closely the compound in question resembles the model's training distribution, flagging genuinely novel scaffolds for additional scrutiny.
  3. For high-confidence endpoints and well-represented chemistry, allow the prediction to drive triage decisions with minimal confirmatory testing.
  4. For lower-confidence endpoints or novel chemical space, treat the prediction as a prioritization signal and route the compound through targeted experimental confirmation before making a final call.
  5. Feed confirmed experimental results back into the model's training set, narrowing the gap between predicted and observed performance for that chemical series over time.

This sequence converts an implicit, individual judgment call into a documented, repeatable process, which matters because informal trust calibration tends to vary widely between chemists on the same team, let alone across an entire organization.

How major pharma groups navigate ADMET reliability

Large pharmaceutical discovery organizations have generally converged on a tiered approach that mirrors the framework above, formalized into internal decision trees rather than left to individual judgment. Compounds are typically scored against the full battery of available ADMET predictions early, with clear internal thresholds distinguishing endpoints where a prediction can independently justify deprioritization from endpoints requiring mandatory experimental confirmation regardless of the predicted score.

Continue reading below...
3D illustration of a protein complex composed of clustered spherical subunits arranged in a ring-like oligomeric structure, shown in shades of blue, cyan, and purple against a blue gradient background.
Application NoteUnderstanding protein oligomerization with mass photometry
Automated mass photometry helps reveal the complex dynamics of protein oligomerization and the factors that govern protein assembly.
Read More

Many organizations also revisit these internal thresholds periodically rather than treating them as fixed, since a model's reliability for a given endpoint can genuinely improve as more experimental data accumulates and gets fed back into training. A threshold set conservatively when a model was new may reasonably shift once several years of confirmatory assay results have validated its performance on the organization's own chemical space.

This tiered structure echoes how regulators themselves treat computational toxicology. Under the ICH M7 guideline for mutagenic impurity assessment, two complementary quantitative structure-activity relationship (QSAR) models are required specifically because no single computational approach has proven reliable enough to stand alone, and the FDA's AI development guidance similarly emphasizes validated performance evidence over blanket acceptance of any model's output. Discovery organizations that build internal frameworks along the same lines tend to get more consistent, defensible outcomes than those relying on case-by-case judgment calls.

A documented framework beats ad hoc trust in ADMET AI

Comparing in silico ADMET vs. in vitro evidence endpoint by endpoint, rather than treating machine learning outputs as uniformly reliable or uniformly suspect, is the single most useful habit a discovery team can build around AI ADMET prediction. The endpoints with the deepest public data support genuine prediction-led triage, while endpoints with sparser or more heterogeneous data still require the experimental caution that has always governed early ADMET screening.

A documented decision framework, applied consistently across chemists and programs, converts that endpoint-by-endpoint judgment into a repeatable process rather than individual intuition. That consistency is what ultimately determines whether AI ADMET prediction accelerates a discovery program or quietly introduces new, undetected risk into candidate selection.

This article was produced under Drug Discovery News' AI Editorial Guidelines.

Frequently Asked Questions (FAQs)

  • When can you replace in vitro assays with AI?

    In vitro assays can reasonably be deprioritized, though rarely eliminated entirely, for endpoints with strong scaffold-split benchmark performance and well-represented chemical space, such as CYP inhibition, hERG blockade, and blood-brain barrier permeability.

  • How reliable are in silico ADMET models?

    Reliability varies by endpoint, ranging from area under the curve or accuracy scores above 0.90 for mature endpoints like CYP inhibition to two-class prediction accuracy in the low-to-mid 70% range for harder endpoints like drug-induced liver injury.

  • When should you trust an AI ADMET prediction?

    An AI ADMET prediction warrants higher trust when the endpoint has strong scaffold-split validation performance and the compound closely resembles the model's training data, and warrants lower trust for novel chemical scaffolds or data-sparse endpoints.

Add Drug Discovery News as a preferred source on Google

Add Drug Discovery News as a preferred Google source to see more of our trusted coverage.

About the Author

Here are some related topics that may interest you:

Loading Next Article...
Loading Next Article...
Subscribe to Newsletter

Subscribe to our eNewsletters

Stay connected with all of the latest from Drug Discovery News.

Subscribe

Sponsored

A scientist in a white lab coat looking into a microscope in a brightly lit modern laboratory.
Learn how developmental and reproductive toxicology study selection supports regulatory decision-making and generates meaningful nonclinical safety data.
Illustration of multiple three-dimensional patient-derived organoids suspended against a dark blue background, representing tumor models used in precision oncology research.
By combining organoid biology with precision automation, researchers developed a miniaturized organoid screening platform that could help speed personalized cancer treatment testing.
Illustration of multiple three-dimensional patient-derived organoids suspended against a dark blue background, representing tumor models used in precision oncology research.
By combining organoid biology with precision automation, researchers developed a miniaturized organoid screening platform that could help speed personalized cancer treatment testing.
Drug Discovery News December 2025 Issue
Latest IssueVolume 21 • Issue 4 • December 2025

December 2025

December 2025 Issue

Explore this issue