News

Can better training data fix AI antibody design?

The field has invested heavily in building better models for antibody discovery. The structural interaction data those models are trained on has not kept pace — and that shortfall is now a defining constraint on what AI can reliably do.
Written byAndrea Corona
| 5 min read
Visual representation of antibody-antigen interaction data with AI technology

Whether a model works in the real world of discovery isn't decided at the modeling stage — it's decided by the quality of the data infrastructure behind it. That makes the data factory central to the science, not incidental to it.

GEMINI 

Register for free to listen to this article
Listen with Speechify
0:00
5:00

The last several years of progress in protein artificial intelligence (AI) have been undeniably impactful. AlphaFold's demonstration that protein folding could be predicted with near-experimental accuracy reset expectations across structural biology, and the models that followed, for protein design, interaction prediction, and sequence generation, have moved antibody discovery into a new computational era.

And yet a practical problem persists, one that has become increasingly difficult to ignore. The antibody-antigen structural data those tools depend on is not keeping up with the demands being placed on it. The models built on that foundation perform well on the problems it covers. They struggle, often silently, on the problems it does not.

Dan Benjamin, co-founder and chief technology officer of Immuto Scientific, told DDN that the numbers illustrate the scale of what is missing: There are more than 240,000 structures in the Protein Data Bank, but only approximately 1,800 are antibody-antigen pairs. "This results in models that are very strong at predicting monomeric chain protein folding, but often miss when predicting those same proteins bound to an antibody," he said.

What the training data actually contains

The scarcity of antibody-antigen structures in public databases is compounded by the way those structures were generated. Most were produced to answer specific biological or therapeutic questions, which means they cluster around targets that were experimentally tractable, structurally stable, and scientifically compelling enough for a group to invest in solving. That selectivity, while understandable, produces a training corpus with significant blind spots.

Continue reading below...
3D illustration of a membrane protein embedded within a lipid nanodisc, representing a native-like environment used for membrane protein stabilization and characterization.
Application NoteCharacterizing nanodisc-embedded membrane proteins
Mass photometry supports membrane protein characterization by providing rapid insights into sample composition, purity, and molecular assembly.
Read More

Benjamin described the consequence as a dataset that is "scientifically useful, but uneven." Well-behaved proteins, common interaction geometries, and stabilized experimental systems are overrepresented. Difficult targets, conformationally flexible antigens, non-binding antibodies, and non-functional binders are largely absent. The field often refers to this as the missing "negative" data problem, a training set that shows models what productive interactions look like but gives them limited exposure to the cases where binding fails and why.

"If the training set is biased toward relatively well-behaved proteins, common interaction geometries, or stabilized experimental systems, the model can appear strong in benchmark settings but become less reliable when applied to new targets or more complex discovery problems," Benjamin said. He added that the field took time to recognize this because the first wave of progress in protein AI was so impressive — it was natural to focus on architectures, scale, and compute. "We are now reaching the point where the data layer is becoming a more visible constraint," he said.

This matters because AI models generalize from what they have seen. A 2024 review in Frontiers Discovery on AI applications in antibody discovery noted that a core limitation of models trained on specific antibody-antigen pairs is that they are not target-agnostic — the training set governs the domain in which a model can be expected to perform, and that domain has boundaries that become apparent in real discovery settings.

Those settings are exactly where generalization is required. The targets a discovery team cares about are not necessarily the targets that were crystallized and deposited in 2005. The binding modes relevant to a therapeutic program may not resemble the interaction geometries that dominate the training corpus. A model that performs well on benchmarks constructed from public data may generate plausible-looking structures for new targets while being systematically wrong about the actual binding mode — a failure that is difficult to detect computationally and expensive to discover experimentally.

Why structural diversity is the operative constraint

Capturing structural diversity in antibody-antigen interaction data requires a genuinely broad range of antigens, epitopes, paratope geometries, complementarity-determining region loop conformations, scaffold types, binding orientations, affinities, and dynamic interaction states — the full combinatorial space in which antibodies encounter their targets in biological systems.

The dynamic dimension is particularly important and particularly underrepresented. Conventional structural biology methods, X-ray crystallography and cryo-electron microscopy, (Cryo-EM) provide extraordinary resolution but require conditions that can move proteins away from their biological context. Crystallization demands a stable, homogeneous structure. Cryo-EM samples are flash-frozen. The result is a snapshot of a complex that may have been stabilized, engineered, or concentrated in ways that alter what is being observed. For targets that are flexible, membrane-associated, or conformationally heterogeneous, those conditions can systematically exclude the interaction states most relevant to therapeutic function.

Continue reading below...
3D illustration of a protein complex composed of clustered spherical subunits arranged in a ring-like oligomeric structure, shown in shades of blue, cyan, and purple against a blue gradient background.
Application NoteUnderstanding protein oligomerization with mass photometry
Automated mass photometry helps reveal the complex dynamics of protein oligomerization and the factors that govern protein assembly.
Read More

"When structural diversity is limited, models can still perform well within familiar territory," Benjamin said. "The problem emerges when they are asked to generalize." The model may generate a large set of plausible structures, with the biologically correct answer present somewhere in the set — but ranked incorrectly. "In practice, that means a team may spend time optimizing around an incorrect structural hypothesis," he said. "Even a modest amount of experimentally anchored interaction data, sometimes as few as 20 antibody-antigen pairs, can help distinguish which structure is biologically plausible."

Research published in Molecules in 2024 described the disparity between the abundance of antibody sequence data and the scarcity of antibody structural information as a defining challenge for the field, noting that tools designed to bridge that divide — by predicting structure from sequence — still face fundamental constraints when the training data does not represent the interaction geometries being asked about.

Where current AI models are most likely to fail

The failure modes that emerge from structural data limitations are distributed unevenly across target classes. Certain categories of antibody discovery problems are consistently more exposed to these limitations than others.

Conformational epitopes — where the antibody's binding site is defined by the three-dimensional shape of a folded protein surface rather than a contiguous sequence — are among the most difficult. Models trained primarily on linear epitope structures may not have enough relevant examples to infer how an antibody engages a conformational target correctly. Membrane proteins, intrinsically disordered regions, multimeric complexes, and targets with disease-associated conformational changes present similar challenges: The antigen may exist in multiple structural states, and the model may not have sufficient examples to know which state the relevant epitope inhabits.

Epitope-specific antibody design is where these limitations have the most immediate practical consequence. Generating an antibody that binds a target at all is a solved problem for many target classes. Generating an antibody that binds a defined epitope, from the right angle, with appropriate selectivity and developability properties, is substantially harder — and substantially more dependent on the quality and diversity of the training data underlying the prediction. As a 2025 review in npj Precision Oncology on AI in antibody-drug conjugate development noted, data scarcity limits the robustness of predictive models precisely in the cases where specificity and interaction geometry matter most.

"In these settings, models often do not fail in an obvious way," Benjamin said. "They may generate structures that look reasonable computationally, but are wrong biologically." The error only becomes apparent when the predicted structure is used as the basis for an optimization campaign and the experimental results do not match expectations. By that point, meaningful resources have been committed to the wrong hypothesis. "The future is not likely to be purely computational or purely experimental," he said. "It will be an integrated loop where empirical data help models make better structural decisions, and models help guide the next set of experiments."

Continue reading below...
A 3D rendering illustrates a sandwich ELISA technique, where antigen detection is achieved between two layers of antibodies: a capture antibody and a detection antibody
EbooksThe four essentials of immunoassay quality
Learn the key characteristics that determine whether an immunoassay generates accurate and reproducible data.
Read More

Data as infrastructure

"The next major gains may come from improving the training environment, not only the model architecture," Benjamin said. "It is not enough to generate more examples if those examples reproduce the same biases in existing datasets." The infrastructural dimension of this is substantial. Many AI systems built for structural biology work with established formats — coordinate files that encode three-dimensional atomic positions. High-throughput experimental methods may generate different kinds of structural readouts: solvent accessibility measurements, epitope-level interaction data, hydrogen-deuterium exchange profiles, chemical cross-linking patterns. Converting those empirical readouts into formats that AI models can use directly requires a translation layer that is itself a significant engineering challenge, and one that most structural biology workflows were not designed to address.

AI applications in antibody discovery has been described the most defensible applications of these tools as pragmatic and tightly coupled to experimental workflows — suggesting that the integration of empirical structural constraints with computational prediction is where the field is heading. The logic is that even a modest amount of experimentally anchored interaction data can help a model distinguish which of its generated structures is biologically plausible — a disproportionate return on experimental investment compared to what the same resources would yield spent on additional model training alone.

"AI drug discovery is becoming an infrastructure problem as much as a modeling problem," Benjamin said. "The teams that can generate, standardize, and continuously expand high-quality biological data will have a different kind of advantage. Models will still matter, but the most durable progress may come from building the systems that allow those models to learn from better biology."

In that framing, the data factory is a core capability that determines whether models work in real discovery settings — not a support function operating at the margins of the science.

Add Drug Discovery News as a preferred source on Google

Add Drug Discovery News as a preferred Google source to see more of our trusted coverage.

About the Author

  • Drug Discovery News Placeholder Image

    Andrea Corona is the senior editor at Drug Discovery News, where she leads daily editorial planning and produces original reporting on breakthroughs in drug discovery and development. With a background in health and pharma journalism, she specializes in translating breakthrough science into engaging stories that resonate with researchers, industry professionals, and decision-makers across biotech and pharma.

    Prior to joining DDN, Andrea served as senior editor at Pharma Manufacturing, where she led feature coverage on pharmaceutical R&D, manufacturing innovation, and regulatory policy. Her work blends investigative reporting with a deep understanding of the drug development pipeline, and she is particularly interested in stories at the intersection of science, innovation and technology.

    View Full Profile

Here are some related topics that may interest you:

Loading Next Article...
Loading Next Article...
Subscribe to Newsletter

Subscribe to our eNewsletters

Stay connected with all of the latest from Drug Discovery News.

Subscribe

Sponsored

3D illustration of a membrane protein embedded within a lipid nanodisc, representing a native-like environment used for membrane protein stabilization and characterization.
Mass photometry supports membrane protein characterization by providing rapid insights into sample composition, purity, and molecular assembly.
3D illustration of a protein complex composed of clustered spherical subunits arranged in a ring-like oligomeric structure, shown in shades of blue, cyan, and purple against a blue gradient background.
Automated mass photometry helps reveal the complex dynamics of protein oligomerization and the factors that govern protein assembly.
Illustration of translucent Y-shaped antibodies floating in a soft blue and green background, representing antibody research, development, and biomedical science.
Explore how antibody accessibility and custom development strategies can influence the pace and success of translational research.
Drug Discovery News December 2025 Issue
Latest IssueVolume 21 • Issue 4 • December 2025

December 2025

December 2025 Issue

Explore this issue