Articles

We have more biological data than ever. Are we asking the right questions?

We have never had more biological data, or more ways to connect it. The programs that win will be the ones that ask better questions of it, not bigger ones.
Written byAbhishek Jha
| 4 min read
Abstract image of connections.

More biological connections do not necessarily mean better drug targets.

credit: istock.com/royyimzy

Register for free to listen to this article
Listen with Speechify
0:00
4:00

Drug discovery has spent a decade getting bigger. A single program today can draw on millions of catalogued gene-disease associations, functional screens across hundreds of cancer cell lines, and large language models that surface connections in seconds. By any measure of scale, we have never had more to work with. And yet the odds of success have not improved.

Only about one in ten drugs that reach clinical testing has historically been approved, and recent analyses suggest that the success rate has fallen below seven percent. That gap is not a data-volume problem. It is a question problem. Drug discovery often asks what a target is connected to, when the question that decides a program is harder: can we move a cell out of the disease state, and does the evidence support that move?

The shape of the problem has not changed in years. A program still starts from something like twenty thousand candidate genes and has to narrow that down to a handful that are worth real investment. The information needed to make that call usually exists already, but it is scattered across full-text papers, supplementary tables, and clinical readouts, and most of it never gets structured into anything searchable. So, teams lose weeks manually reconciling evidence that is technically already public. New biomedical research also appears faster than any team can read it, so the evidence that matters is usually buried, not missing. Scaling up the associations does not fix this. It just hands you a longer list to wade through.

Bigger is not the same as better

Most knowledge graphs are built for the easy question. A knowledge graph is a structured map of biological entities, genes, proteins, diseases, and drugs, with the links between them, and it is very good at telling you what associates with what. The catch is that an association has no direction. Many graphs score a gene-disease link by co-occurrence, how often the two are named together, which cannot separate a paper reporting a strong effect from one reporting no effect. Add more papers, and you inherit the same blind spot at greater scale. More links of the same shallow kind do not make a hypothesis any easier to trust.

A better graph does more than connect the dots. It shows how strong the evidence is behind each connection. A cell in disease is often a cell stuck in the wrong state, and the question worth answering is whether an intervention can push it into a healthier one. That takes more than a co-occurrence count. It requires evidence from full-text results, model systems, and clinical data: what kind of relationship the link represents, whether a gene causes a disease, a drug treats it, or a marker simply tracks it, along with which experiment a finding came from, which way the effect ran, and whether any other studies contradict it. In practice, this means distinguishing an established result from a promising one without having to spend a week piecing the evidence together. A graph that carries that evidence, and not just the headline link, is what makes it evidence-rich.

A short story from AML

Acute myeloid leukemia (AML) is a good place to see the difference. It is a cancer in which blood cells stop maturing and pile up in an immature state, driven in about a third of adults by a single recurrent mutation. Oncology is also the hardest arena in drug development, with the lowest approval odds of any major disease area, which makes AML a demanding test rather than an easy one. In a recent exercise, an evidence-rich graph was given nothing but the disease name and asked to reconstruct the biology. It did, tracing a path from the disease to an approved drug, revumenib, that releases this maturation block in genetically defined patients. None of that biology was new. What mattered was that the graph assembled it from evidence rather than recall, with every step traceable to a source.

In that exercise, quality meant the difference between an edge that merely says a gene is associated with the disease and one that specifies how the two are related: whether the gene carries a somatic mutation in the disease (has_somatic_mutation_in), is a drug target (has_drug_target), or is only a biomarker (has_biomarker_gene). That distinction helped the graph prioritize a potentially treatable step ahead of a dead end. It trusted that call because six independent kinds of evidence agreed, not because any one score was high.

A second mark of quality was what the graph would not overclaim. The same question surfaced the next steps along that axis, including a related target still in trials and a regulator that no drug can yet reach, but it marked them for what they are: promising, but unproven. A bigger graph would have listed them beside everything else at equal weight. A better one showed where the evidence was solid and where it ran out.

That is also where this gets interesting. Once the evidence is connected and weighed, a team can start asking the forward-looking question the field actually cares about: not only what is linked to a disease, but what might move a cell out of it, and whether the evidence would support trying. The job of the infrastructure is to let people pose that kind of hypothesis and stress-test it early, while it is still cheap to be wrong.

The point is falsifiability, not volume

The value of an evidence-rich graph is not that it finds more hypotheses. It is that it lets a team challenge them early. It brings contradictory evidence forward before the budget is committed, shows which claims rest on a single experiment and which on ten, and gives researchers a target thesis they can defend when it reaches a skeptical reviewer. Because every claim points back to its source, that review becomes a conversation about the biology, rather than whether the summary can be trusted. Ruling out a bad idea in week one is worth more than ranking a thousand associations, because the cost of a wrong target compounds the longer it goes unnoticed.

None of this is an argument against scale. Bigger maps have value, and the field will keep building them. But a bigger map alone does not tell researchers which connections matter or what to do next. . A better one, built on evidence that can be inspected and challenged, can show where a hypothesis is worth persuing and where it is likely to fail. That is the shift worth making in drug discovery: not more connections, but connections we can trust.

Add Drug Discovery News as a preferred source on Google

Add Drug Discovery News as a preferred Google source to see more of our trusted coverage.

About the Author

  • Headshot of Abhishek Jha, in front of some sailboats.

    Abhishek Jha is Co-founder and CEO of Elucidata, building data-centric infrastructure for biomedical discovery. He has over 20 years of experience spanning life sciences, data science, and machine learning, with training from MIT, the University of Chicago, and IIT Bombay. At Agios Pharmaceuticals, he helped advance four first-in-class therapies to the clinic, all now FDA-approved. He writes about AI, data, and the future of drug discovery and its evolving landscape.

    View Full Profile

Here are some related topics that may interest you:

Loading Next Article...
Loading Next Article...
Subscribe to Newsletter

Subscribe to our eNewsletters

Stay connected with all of the latest from Drug Discovery News.

Subscribe

Sponsored

Crystal Girod, Senior Product Manager at Beckman Coulter Life Sciences, featured in a “Tell Us What You Know” graphic titled “Making Lab Automation More Accessible,” alongside the Beckman Coulter Life Sciences logo.
Explore how advances in liquid handling are making automation more accessible, flexible, and practical for modern labs.
Abstract 3D-rendered background with glowing blue and purple lines and violet light rays radiating through a dark space.
Understand the principles and practices that support high-quality data in increasingly complex flow cytometry experiments.
Puzzle pieces spelling “DATA” alongside connected icons representing data analysis, storage, and reporting.
Earlier insight into analytical data can help laboratories recognize emerging trends and maintain method performance over time.