An antibody can bind its target with high affinity and still stall in development if it aggregates, loses stability, or resists formulation at scale. Predicting those liabilities early could spare drug developers months of work and significant investment, but the computational models built for the job have largely been trained on small, fragmented datasets. A new industry consortium aims to close that gap by characterizing 10,000 antibodies under one set of assays.
On Sept. 29, Ginkgo Datapoints and Apheris announced AbbVie, argenx, Lundbeck, and Takeda as founding members of the Antibody Developability Consortium, which remains open to additional pharmaceutical and biotech companies. Developability covers the biophysical properties that determine whether a candidate antibody can be manufactured, formulated, and advanced into a clinical product.
Why standardized data matters
"Training a useful model requires antibody sequences paired with reliable measurements of their developability properties," said Peter Tessier, an antibody engineer at the University of Michigan who is providing independent scientific oversight for the consortium, speaking with DDN. "Generating those measurements is expensive, and much of the available data is proprietary or was collected using different methods."
Even large internal datasets at individual companies tend to be limited in sequence diversity. Characterizing 10,000 deliberately diverse antibodies under a common set of assays, Tessier told DDN, "should provide a much stronger basis for training and testing models, particularly for assessing how well they perform on antibodies unlike those in their training data."
Erwin Pannecoucke, a principal scientist in discovery at argenx, described the same gap from inside a founding member. Each organization explores a different part of antibody sequence space, he told DDN, so any single dataset captures only part of the picture. "Diversity is critical because predictive models can only reliably evaluate new antibody sequences if they have been trained on similar sequence characteristics," he said.
Athena Hadjixenofontos, who leads artificial intelligence (AI) work in biotherapeutics and genetic medicine at AbbVie, said in a statement that datasets purpose-built for machine learning help address "limitations associated with convenience datasets."
Where late failures cost the most
In a typical discovery process, argenx starts with hundreds of potential antibody sequences and progressively narrows them to a small number of candidates. Pannecoucke told DDN that the steepest price comes when a developability issue surfaces after a program has already committed to a lead molecule. At that point, fixing it can require extensive re-engineering, additional testing, and confirmation that the changes have preserved the molecule's biological activity. "These efforts can take months, add significant cost, and in some cases may not fully solve the problem," he said.
Better predictive models could move those decisions earlier. During discovery, Pannecoucke said, they can help teams prioritize which antibodies to advance, identify candidates that may need engineering, and focus resources on the most promising molecules. At candidate selection, "they enable us to evaluate developability alongside biological function while alternatives are still available."
How the consortium works
Each founding member contributes proprietary antibody sequences, and Ginkgo fills the remaining capacity with publicly available sequences to reach 10,000. Ginkgo designs the sequence selection approach, produces the antibodies, and runs high-throughput wet-lab characterization across core developability endpoints. It then trains a foundation developability model on the resulting dataset.
That training happens inside a federated computing environment run by Apheris. Federated learning lets companies train and refine shared models while each member's raw data stays in its own environment. Members can fine-tune the foundation model on their own molecules, build new models using consortium data, and keep ownership of the sequences and assay data they contribute.
"Pooling standardized developability data across the industry can create stronger predictive models than any one company could build alone," said Yves Fomekong Nanfack, head of AI and machine learning research at Takeda, in a statement.
For argenx, the shared experimental standards were as important as the added sequences. Because all data are generated under common protocols, Pannecoucke told DDN, the consortium provides "the consistency needed to develop robust predictive models."
Tessier and Charlotte Deane, a structural bioinformatician at the University of Oxford, serve as independent scientific advisors. "I have participated in technical discussions with the consortium's computational and experimental scientists about sequence selection, assay design, and modeling, and I expect to continue advising as the data and models are developed," Tessier said.
For academic researchers, the resource will stay within the consortium for now. "The dataset and models are being developed for consortium members; my academic group does not have direct access to them," Tessier said. "Academic access would require a separate research collaboration or a decision by the consortium to make resources available more broadly."
What success would look like
The consortium expects to deliver its initial dataset to members by early 2027 and plans to explore more complex antibody formats over time.
"Success would mean showing that the models improve predictions for new antibodies and help companies identify developability risks earlier, so they can focus experimental work on the most promising candidates," Tessier said. "I would also like to see what we learn about the amount and diversity of data needed for models to generalize beyond the sequences used to train them."
Pannecoucke framed success in two to three years as evidence from new molecules in real discovery programs that the models consistently improve decision-making, with fewer unexpected developability issues and fewer rounds of re-engineering. Ultimately, he told DDN, argenx hopes to become confident enough in certain predictions "to reduce selected early-stage screening experiments because computational assessment has already filtered out molecules unlikely to meet developability criteria." That would free more lab time for understanding biology and disease mechanisms.
Tessier added a caution for teams hoping to lean on the predictions. "These models will still need experimental validation, especially for antibodies, formats, or properties that are poorly represented in the dataset," he said.
Pannecoucke pointed to a similar boundary. The initial focus on conventional immunoglobulin G antibodies is a foundation, he said, but the models are unlikely to predict how multispecifics and other complex formats behave. argenx is already looking toward a follow-on developability consortium dedicated to those multispecific modalities.












