News

Pharma companies pool antibody data to predict developability

A new consortium will characterize 10,000 antibodies to a single standard, aiming to train AI models that flag manufacturing risks before candidates advance.
Written byAndrea Corona
| 4 min read
Illustration of antibodies related to the Antibody Developability Consortium

Antibodies that bind their targets well can still fail in development if they aggregate or prove hard to manufacture. A new consortium is building shared models to flag those risks earlier.

GEMINI (2026)

Register for free to listen to this article
Listen with Speechify
0:00
4:00

An antibody can bind its target with high affinity and still stall in development if it aggregates, loses stability, or resists formulation at scale. Predicting those liabilities early could spare drug developers months of work and significant investment, but the computational models built for the job have largely been trained on small, fragmented datasets. A new industry consortium aims to close that gap by characterizing 10,000 antibodies under one set of assays.

On Sept. 29, Ginkgo Datapoints and Apheris announced AbbVie, argenx, Lundbeck, and Takeda as founding members of the Antibody Developability Consortium, which remains open to additional pharmaceutical and biotech companies. Developability covers the biophysical properties that determine whether a candidate antibody can be manufactured, formulated, and advanced into a clinical product.

Why standardized data matters

"Training a useful model requires antibody sequences paired with reliable measurements of their developability properties," said Peter Tessier, an antibody engineer at the University of Michigan who is providing independent scientific oversight for the consortium, speaking with DDN. "Generating those measurements is expensive, and much of the available data is proprietary or was collected using different methods."

Even large internal datasets at individual companies tend to be limited in sequence diversity. Characterizing 10,000 deliberately diverse antibodies under a common set of assays, Tessier told DDN, "should provide a much stronger basis for training and testing models, particularly for assessing how well they perform on antibodies unlike those in their training data."

Continue reading below...
Abstract 3D-rendered background with glowing blue and purple lines and violet light rays radiating through a dark space.
Technology GuidesTechnology Guide: The fundamentals behind high-parameter flow cytometry
Understand the principles and practices that support high-quality data in increasingly complex flow cytometry experiments.
Read More

Erwin Pannecoucke, a principal scientist in discovery at argenx, described the same gap from inside a founding member. Each organization explores a different part of antibody sequence space, he told DDN, so any single dataset captures only part of the picture. "Diversity is critical because predictive models can only reliably evaluate new antibody sequences if they have been trained on similar sequence characteristics," he said.

Athena Hadjixenofontos, who leads artificial intelligence (AI) work in biotherapeutics and genetic medicine at AbbVie, said in a statement that datasets purpose-built for machine learning help address "limitations associated with convenience datasets."

Where late failures cost the most

In a typical discovery process, argenx starts with hundreds of potential antibody sequences and progressively narrows them to a small number of candidates. Pannecoucke told DDN that the steepest price comes when a developability issue surfaces after a program has already committed to a lead molecule. At that point, fixing it can require extensive re-engineering, additional testing, and confirmation that the changes have preserved the molecule's biological activity. "These efforts can take months, add significant cost, and in some cases may not fully solve the problem," he said.

Better predictive models could move those decisions earlier. During discovery, Pannecoucke said, they can help teams prioritize which antibodies to advance, identify candidates that may need engineering, and focus resources on the most promising molecules. At candidate selection, "they enable us to evaluate developability alongside biological function while alternatives are still available."

How the consortium works

Each founding member contributes proprietary antibody sequences, and Ginkgo fills the remaining capacity with publicly available sequences to reach 10,000. Ginkgo designs the sequence selection approach, produces the antibodies, and runs high-throughput wet-lab characterization across core developability endpoints. It then trains a foundation developability model on the resulting dataset.

That training happens inside a federated computing environment run by Apheris. Federated learning lets companies train and refine shared models while each member's raw data stays in its own environment. Members can fine-tune the foundation model on their own molecules, build new models using consortium data, and keep ownership of the sequences and assay data they contribute.

"Pooling standardized developability data across the industry can create stronger predictive models than any one company could build alone," said Yves Fomekong Nanfack, head of AI and machine learning research at Takeda, in a statement.

For argenx, the shared experimental standards were as important as the added sequences. Because all data are generated under common protocols, Pannecoucke told DDN, the consortium provides "the consistency needed to develop robust predictive models."

Continue reading below...
Puzzle pieces spelling “DATA” alongside connected icons representing data analysis, storage, and reporting.
CompendiumTurning analytical data into proactive method oversight
Earlier insight into analytical data can help laboratories recognize emerging trends and maintain method performance over time.
Read More

Tessier and Charlotte Deane, a structural bioinformatician at the University of Oxford, serve as independent scientific advisors. "I have participated in technical discussions with the consortium's computational and experimental scientists about sequence selection, assay design, and modeling, and I expect to continue advising as the data and models are developed," Tessier said.

For academic researchers, the resource will stay within the consortium for now. "The dataset and models are being developed for consortium members; my academic group does not have direct access to them," Tessier said. "Academic access would require a separate research collaboration or a decision by the consortium to make resources available more broadly."

What success would look like

The consortium expects to deliver its initial dataset to members by early 2027 and plans to explore more complex antibody formats over time.

"Success would mean showing that the models improve predictions for new antibodies and help companies identify developability risks earlier, so they can focus experimental work on the most promising candidates," Tessier said. "I would also like to see what we learn about the amount and diversity of data needed for models to generalize beyond the sequences used to train them."

Pannecoucke framed success in two to three years as evidence from new molecules in real discovery programs that the models consistently improve decision-making, with fewer unexpected developability issues and fewer rounds of re-engineering. Ultimately, he told DDN, argenx hopes to become confident enough in certain predictions "to reduce selected early-stage screening experiments because computational assessment has already filtered out molecules unlikely to meet developability criteria." That would free more lab time for understanding biology and disease mechanisms.

Tessier added a caution for teams hoping to lean on the predictions. "These models will still need experimental validation, especially for antibodies, formats, or properties that are poorly represented in the dataset," he said.

Pannecoucke pointed to a similar boundary. The initial focus on conventional immunoglobulin G antibodies is a foundation, he said, but the models are unlikely to predict how multispecifics and other complex formats behave. argenx is already looking toward a follow-on developability consortium dedicated to those multispecific modalities.

Add Drug Discovery News as a preferred source on Google

Add Drug Discovery News as a preferred Google source to see more of our trusted coverage.

About the Author

  • Headshot

    Andrea Corona is the senior editor at Drug Discovery News, where she leads daily editorial planning and produces original reporting on breakthroughs in drug discovery and development. With a background in health and pharma journalism, she specializes in translating breakthrough science into engaging stories that resonate with researchers, industry professionals, and decision-makers across biotech and pharma. Her work blends investigative reporting with a deep understanding of the drug development pipeline, and she is particularly interested in stories at the intersection of science, innovation and technology.

    View Full Profile

Here are some related topics that may interest you:

Loading Next Article...
Loading Next Article...
Subscribe to Newsletter

Subscribe to our eNewsletters

Stay connected with all of the latest from Drug Discovery News.

Subscribe

Sponsored

Crystal Girod, Senior Product Manager at Beckman Coulter Life Sciences, featured in a “Tell Us What You Know” graphic titled “Making Lab Automation More Accessible,” alongside the Beckman Coulter Life Sciences logo.
Explore how advances in liquid handling are making automation more accessible, flexible, and practical for modern labs.
Abstract 3D-rendered background with glowing blue and purple lines and violet light rays radiating through a dark space.
Understand the principles and practices that support high-quality data in increasingly complex flow cytometry experiments.
Puzzle pieces spelling “DATA” alongside connected icons representing data analysis, storage, and reporting.
Earlier insight into analytical data can help laboratories recognize emerging trends and maintain method performance over time.