The last several years of progress in protein artificial intelligence (AI) have been undeniably impactful. AlphaFold's demonstration that protein folding could be predicted with near-experimental accuracy reset expectations across structural biology, and the models that followed, for protein design, interaction prediction, and sequence generation, have moved antibody discovery into a new computational era.
And yet a practical problem persists, one that has become increasingly difficult to ignore. The antibody-antigen structural data those tools depend on is not keeping up with the demands being placed on it. The models built on that foundation perform well on the problems it covers. They struggle, often silently, on the problems it does not.
Dan Benjamin, co-founder and chief technology officer of Immuto Scientific, told DDN that the numbers illustrate the scale of what is missing: There are more than 240,000 structures in the Protein Data Bank, but only approximately 1,800 are antibody-antigen pairs. "This results in models that are very strong at predicting monomeric chain protein folding, but often miss when predicting those same proteins bound to an antibody," he said.
What the training data actually contains
The scarcity of antibody-antigen structures in public databases is compounded by the way those structures were generated. Most were produced to answer specific biological or therapeutic questions, which means they cluster around targets that were experimentally tractable, structurally stable, and scientifically compelling enough for a group to invest in solving. That selectivity, while understandable, produces a training corpus with significant blind spots.
Benjamin described the consequence as a dataset that is "scientifically useful, but uneven." Well-behaved proteins, common interaction geometries, and stabilized experimental systems are overrepresented. Difficult targets, conformationally flexible antigens, non-binding antibodies, and non-functional binders are largely absent. The field often refers to this as the missing "negative" data problem, a training set that shows models what productive interactions look like but gives them limited exposure to the cases where binding fails and why.
"If the training set is biased toward relatively well-behaved proteins, common interaction geometries, or stabilized experimental systems, the model can appear strong in benchmark settings but become less reliable when applied to new targets or more complex discovery problems," Benjamin said. He added that the field took time to recognize this because the first wave of progress in protein AI was so impressive — it was natural to focus on architectures, scale, and compute. "We are now reaching the point where the data layer is becoming a more visible constraint," he said.
This matters because AI models generalize from what they have seen. A 2024 review in Frontiers Discovery on AI applications in antibody discovery noted that a core limitation of models trained on specific antibody-antigen pairs is that they are not target-agnostic — the training set governs the domain in which a model can be expected to perform, and that domain has boundaries that become apparent in real discovery settings.
Those settings are exactly where generalization is required. The targets a discovery team cares about are not necessarily the targets that were crystallized and deposited in 2005. The binding modes relevant to a therapeutic program may not resemble the interaction geometries that dominate the training corpus. A model that performs well on benchmarks constructed from public data may generate plausible-looking structures for new targets while being systematically wrong about the actual binding mode — a failure that is difficult to detect computationally and expensive to discover experimentally.
Why structural diversity is the operative constraint
Capturing structural diversity in antibody-antigen interaction data requires a genuinely broad range of antigens, epitopes, paratope geometries, complementarity-determining region loop conformations, scaffold types, binding orientations, affinities, and dynamic interaction states — the full combinatorial space in which antibodies encounter their targets in biological systems.
The dynamic dimension is particularly important and particularly underrepresented. Conventional structural biology methods, X-ray crystallography and cryo-electron microscopy, (Cryo-EM) provide extraordinary resolution but require conditions that can move proteins away from their biological context. Crystallization demands a stable, homogeneous structure. Cryo-EM samples are flash-frozen. The result is a snapshot of a complex that may have been stabilized, engineered, or concentrated in ways that alter what is being observed. For targets that are flexible, membrane-associated, or conformationally heterogeneous, those conditions can systematically exclude the interaction states most relevant to therapeutic function.
"When structural diversity is limited, models can still perform well within familiar territory," Benjamin said. "The problem emerges when they are asked to generalize." The model may generate a large set of plausible structures, with the biologically correct answer present somewhere in the set — but ranked incorrectly. "In practice, that means a team may spend time optimizing around an incorrect structural hypothesis," he said. "Even a modest amount of experimentally anchored interaction data, sometimes as few as 20 antibody-antigen pairs, can help distinguish which structure is biologically plausible."
Research published in Molecules in 2024 described the disparity between the abundance of antibody sequence data and the scarcity of antibody structural information as a defining challenge for the field, noting that tools designed to bridge that divide — by predicting structure from sequence — still face fundamental constraints when the training data does not represent the interaction geometries being asked about.
Where current AI models are most likely to fail
The failure modes that emerge from structural data limitations are distributed unevenly across target classes. Certain categories of antibody discovery problems are consistently more exposed to these limitations than others.
Conformational epitopes — where the antibody's binding site is defined by the three-dimensional shape of a folded protein surface rather than a contiguous sequence — are among the most difficult. Models trained primarily on linear epitope structures may not have enough relevant examples to infer how an antibody engages a conformational target correctly. Membrane proteins, intrinsically disordered regions, multimeric complexes, and targets with disease-associated conformational changes present similar challenges: The antigen may exist in multiple structural states, and the model may not have sufficient examples to know which state the relevant epitope inhabits.
Epitope-specific antibody design is where these limitations have the most immediate practical consequence. Generating an antibody that binds a target at all is a solved problem for many target classes. Generating an antibody that binds a defined epitope, from the right angle, with appropriate selectivity and developability properties, is substantially harder — and substantially more dependent on the quality and diversity of the training data underlying the prediction. As a 2025 review in npj Precision Oncology on AI in antibody-drug conjugate development noted, data scarcity limits the robustness of predictive models precisely in the cases where specificity and interaction geometry matter most.
"In these settings, models often do not fail in an obvious way," Benjamin said. "They may generate structures that look reasonable computationally, but are wrong biologically." The error only becomes apparent when the predicted structure is used as the basis for an optimization campaign and the experimental results do not match expectations. By that point, meaningful resources have been committed to the wrong hypothesis. "The future is not likely to be purely computational or purely experimental," he said. "It will be an integrated loop where empirical data help models make better structural decisions, and models help guide the next set of experiments."
Data as infrastructure
"The next major gains may come from improving the training environment, not only the model architecture," Benjamin said. "It is not enough to generate more examples if those examples reproduce the same biases in existing datasets." The infrastructural dimension of this is substantial. Many AI systems built for structural biology work with established formats — coordinate files that encode three-dimensional atomic positions. High-throughput experimental methods may generate different kinds of structural readouts: solvent accessibility measurements, epitope-level interaction data, hydrogen-deuterium exchange profiles, chemical cross-linking patterns. Converting those empirical readouts into formats that AI models can use directly requires a translation layer that is itself a significant engineering challenge, and one that most structural biology workflows were not designed to address.
AI applications in antibody discovery has been described the most defensible applications of these tools as pragmatic and tightly coupled to experimental workflows — suggesting that the integration of empirical structural constraints with computational prediction is where the field is heading. The logic is that even a modest amount of experimentally anchored interaction data can help a model distinguish which of its generated structures is biologically plausible — a disproportionate return on experimental investment compared to what the same resources would yield spent on additional model training alone.
"AI drug discovery is becoming an infrastructure problem as much as a modeling problem," Benjamin said. "The teams that can generate, standardize, and continuously expand high-quality biological data will have a different kind of advantage. Models will still matter, but the most durable progress may come from building the systems that allow those models to learn from better biology."
In that framing, the data factory is a core capability that determines whether models work in real discovery settings — not a support function operating at the margins of the science.













