Articles

Data sharing and IP in AI drug discovery: Who owns the model, the data, and the drug?

Patent law was not written with generative models in mind. As AI becomes central to drug discovery, questions about ownership, inventorship, and data sharing are being answered in real time, not always consistently.
Written byTrevor J Henderson
| 6 min read
Two gloved hands exchanging a small vial across a lab bench, a patent-style document nearby

When a model, a dataset, and a molecule all touch the same drug candidate, ownership stops being a simple question.

Flow (2026)

A generative model proposes a molecule. A chemist selects it, modifies it, and eventually synthesizes it. Somewhere in that sequence, a question that used to have an obvious answer stops being obvious: who actually owns the result? IP disputes in AI drug discovery exist precisely because artificial intelligence (AI) has inserted a genuinely new kind of contributor into a legal framework, patent law, that was built entirely around human inventors. The commercial question, who owns the compound, is often contractually straightforward. The legal question, who counts as its inventor, is not, and the two do not always point in the same direction.

This connects to the regulatory guide to AI in drug discovery for how FDA, EMA and ICH are approaching AI validation more broadly, and to the guide to AI across drug discovery for the wider technical picture. DDN has previously noted that intellectual property disputes remain one of the field's more unresolved questions even as adoption accelerates; this works through why, and what is actually being done about it.

Who owns an AI-generated drug molecule?

In practice, ownership of the molecule itself is usually the easier half of this problem. Most AI-driven drug discovery (AIDD) companies operate under well-established commercial models, service agreements, licensing deals, or in-house development, in which intellectual property for whatever gets discovered is contractually assigned, typically to the pharmaceutical partner funding the work. That clean default exists precisely because it makes deals possible: a licensee or partner needs confidence that the IP they are paying for is not encumbered by a separate claim from the AI vendor whose tools helped find it.

The complication is that contractual ownership and patent inventorship are legally separate questions. A company can own a patent outright while still being required, under patent law, to correctly name the individual humans who made a significant intellectual contribution to conceiving the invention. Getting that second question wrong, even when ownership is not in dispute, can expose a patent to a later validity challenge. That is the narrower, thornier problem the rest of this piece is really about.

The inventorship problem

The foundational rule, settled after years of test litigation over an AI system called DABUS, is that only a natural person can be named as a patent inventor; no jurisdiction that has substantively considered the question has concluded otherwise. That still leaves the harder question open: when a human and an AI model collaborate closely on a discovery, which human, if any, actually qualifies as the inventor?

The USPTO's February 2024 Inventorship Guidance is the most detailed attempt yet to answer that for AI-assisted inventions generally. It applies the long-standing Pannu factors, requiring a significant, non-trivial intellectual contribution to an invention's conception, and translates them into a handful of guiding principles specific to AI. Paraphrased, the core ideas are: using an AI system does not by itself disqualify a person from being an inventor; simply posing a research goal or running a pre-built model does not qualify as a significant contribution; but a person who builds an essential piece of the system that produces the invention, or who identifies a specific problem and designs an AI tool to solve it, may qualify even without being present for every step that followed.

Insilico Medicine's own General Counsel has written that these principles, while a genuine step forward, were illustrated with simplified, linear examples that do not fully capture how AIDD actually works: a real discovery program typically layers multiple AI platforms, target identification, generative chemistry, property prediction, across iterative rounds involving chemists, biologists and data scientists, none of whom individually resembles the guidance's tidy hypothetical inventor. The practical response many companies have adopted is procedural rather than legal: documenting, at each stage of a discovery program, who made which decision and why, so that inventorship can be reconstructed and defended later if a patent is challenged.

Proprietary data and training data IP

Inventorship is only one of three separable ownership questions running through an AI drug discovery program simultaneously: who owns the trained model, who owns the data that trained it, and who owns whatever molecule the model helped identify. These do not automatically travel together. A model trained on one company's proprietary assay data, then licensed for use on a second company's target, can easily end up with disputed ownership over improvements the model makes during that second engagement, even when the resulting drug candidate's ownership is contractually clear.

This is now a design consideration for informatics platforms themselves, not just a legal afterthought. Revvity Signals has described building its AI features specifically to keep a client's underlying data and intellectual property from becoming exposed to or absorbed by the third-party AI models integrated into its software, a direct acknowledgment that the software layer connecting a company's proprietary data to an outside AI tool is itself a point of IP risk that has to be engineered around, not just contracted around.

Data-sharing consortia: MELLODDY, OpenADMET

Two consortia illustrate genuinely different answers to the same underlying puzzle: how can competitors share enough data to build better models without giving away the proprietary information that makes their data valuable in the first place?

MELLODDY (Machine Learning Ledger Orchestration for Drug Discovery), an EU-funded effort launched in 2019 spanning ten pharmaceutical companies alongside technology and academic partners, completed a three-year federated learning experiment in 2022 built around 2.6 billion confidential experimental data points covering more than 21 million molecules. No partner ever saw another partner's raw data; each trained a shared model locally and only encrypted model updates, not underlying compounds or assay results, were combined into a global model. Every one of the ten partners measurably improved their own predictive models as a result, without any partner's proprietary data ever leaving its own servers.

OpenADMET takes the opposite approach: fully open data and fully open models, rather than private data feeding a shared model. Governed by the Open Molecular Software Foundation and funded by ARPA-H, the Gates Foundation and other backers, it focuses specifically on ADMET and toxicity properties, the pharmacokinetic and safety characteristics responsible for a large share of clinical failures, publishing open datasets, open-source models and public benchmark challenges rather than keeping any of it proprietary to a single sponsor.

MELLODDY

OpenADMET

Data model

Private; never leaves each partner's own infrastructure

Fully open and publicly published

What's shared

Encrypted model updates only

Raw datasets and trained models

Focus

Broad small-molecule bioactivity (QSAR)

ADMET and toxicity specifically

Participants

10 pharma companies plus tech/academic partners

Academic and nonprofit research groups

Status

Concluded 2022; technology extended to new domains

Active; first public model released 2026

Competitive advantage with shared models

The apparent paradox, why competitors would cooperate at all, resolves once the specific thing being shared is separated from the thing being protected. MELLODDY's partners never shared their compounds, targets or assay results; they shared only the statistical signal extracted from training on that data, aggregated in a way no partner could reverse-engineer to recover another's underlying molecules. The competitive asset, the proprietary dataset itself, stayed private throughout; only its generalizable predictive value was pooled.

OpenADMET's model rests on a related but distinct idea: pre-competitive data. ADMET liabilities, the properties that cause a compound to fail for reasons unrelated to how well it hits its target, are widely treated across the industry as a shared cost of doing business rather than a source of competitive edge. A company's actual target selection and lead compounds remain closely guarded; the underlying toxicology and metabolism data used to screen out bad actors increasingly is not, on the logic that better shared ADMET models reduce costly late-stage failures for everyone without giving any single competitor an edge over another.

Emerging legal frameworks

The DABUS litigation, filed across roughly a dozen jurisdictions starting in 2018 by an inventor seeking to name his own AI system on patent applications, has by now run its course in most of them, and the results are strikingly consistent:

Jurisdiction

AI named as sole inventor?

Basis

United States

Not permitted

Federal Circuit, 2022; USPTO guidance requires a natural person's significant contribution

United Kingdom

Not permitted

Supreme Court, 2023; further divisional appeals rejected through 2025

European Patent Office

Not permitted

Legal Board of Appeal, 2021; related refusal upheld Feb. 2026

Germany

Not permitted

Federal Patent Court: AI-assisted inventions are patentable, but a natural person must be named

Australia

Not permitted

Full Federal Court overturned an initial 2021 ruling on appeal in 2022

Japan

Not permitted

IP High Court, Jan. 2025; Supreme Court declined further appeal, Mar. 2026

South Africa

Granted (sole outlier)

Non-substantive registration system; inventorship was not contested at filing

What this settles is narrower than it sounds: it closes off naming an AI system itself as inventor, but it does not resolve the much more common, much murkier question of which human collaborator in a genuinely AI-assisted discovery program qualifies as one. That question is being worked out through patent office guidance and internal documentation practice at a time, not by a single sweeping ruling, and is likely to remain unsettled longer than the DABUS question ever was.


Key takeaways

  • Owning an AI-discovered compound and being correctly named as its patent inventor are legally separate questions; getting the second one wrong can expose an otherwise valid patent to challenge.
  • USPTO's 2024 guidance confirms AI cannot be a named inventor but leaves real ambiguity about which human collaborators in a multi-platform AIDD program qualify as one.
  • MELLODDY and OpenADMET solve the same cooperation problem two different ways: sharing only encrypted model signal from private data, versus publishing pre-competitive data and models outright.
  • Every jurisdiction that has substantively ruled on the question has barred AI from being named a patent inventor; South Africa's lone exception reflects a filing quirk, not a considered legal position.

This article was produced in accordance with Drug Discovery News’ AI Editorial Policies.

Frequently Asked Questions (FAQs)

  • Can AI be named as a drug inventor?

    No. Every jurisdiction that has substantively ruled on the question, including the United States, the United Kingdom, the European Patent Office, Germany, Australia, and Japan, has concluded that only a natural person can be named as a patent inventor. South Africa granted a patent naming an AI system, but that resulted from a non-examining registration system rather than a considered legal ruling.

  • Who owns an AI-generated drug molecule?

    Ownership of an AI-generated molecule is typically resolved contractually, usually assigned to the pharmaceutical partner funding the discovery work, regardless of which AI tools were involved. That commercial ownership is legally separate from patent inventorship, which depends on which specific humans made a significant intellectual contribution to conceiving the invention, a question current guidance leaves partly unresolved.

  • What is MELLODDY?

    MELLODDY, Machine Learning Ledger Orchestration for Drug Discovery, was an EU-funded federated learning consortium of ten pharmaceutical companies plus technology and academic partners. Running from 2019 to 2022, it let each partner improve its own predictive models using the collective statistical signal from more than 21 million molecules and 2.6 billion data points, without any partner's proprietary data ever leaving its own infrastructure.

Add Drug Discovery News as a preferred source on Google

Add Drug Discovery News as a preferred Google source to see more of our trusted coverage.

About the Author

  • Drug Discovery News Placeholder Image

    Trevor Henderson is the Creative Services Director for the Laboratory Products Group at LabX Media Group. With over two decades of experience, he specializes in scientific and technical writing, editing, and content creation. His academic background includes training in human biology, physical anthropology, and community health. Since 2013, he has been developing content to engage and inform scientists and laboratorians.

    View Full Profile

Here are some related topics that may interest you:

Related Articles

Subscribe to Newsletter

Subscribe to our eNewsletters

Stay connected with all of the latest from Drug Discovery News.

Subscribe

Sponsored

A scientist in a white lab coat looking into a microscope in a brightly lit modern laboratory.
Learn how developmental and reproductive toxicology study selection supports regulatory decision-making and generates meaningful nonclinical safety data.
Illustration of multiple three-dimensional patient-derived organoids suspended against a dark blue background, representing tumor models used in precision oncology research.
By combining organoid biology with precision automation, researchers developed a miniaturized organoid screening platform that could help speed personalized cancer treatment testing.
Illustration of multiple three-dimensional patient-derived organoids suspended against a dark blue background, representing tumor models used in precision oncology research.
By combining organoid biology with precision automation, researchers developed a miniaturized organoid screening platform that could help speed personalized cancer treatment testing.
Drug Discovery News December 2025 Issue
Latest IssueVolume 21 • Issue 4 • December 2025

December 2025

December 2025 Issue

Explore this issue