Advancing biomarker discovery with spatial data means confronting a problem conventional discovery pipelines never had to solve: the datasets are too large and too information-rich to search by eye, and the candidates worth finding are frequently defined by patterns across many tissue positions at once rather than by any single measurement. That combination, more information and less ability to inspect it manually, is reshaping how biomarker discovery is actually done.
Key takeaways
|
The biomarker discovery bottleneck
Conventional biomarker discovery, built on bulk or dissociated single-cell measurements, faces a data volume problem that is real but manageable: a researcher can inspect a table of differentially expressed genes or proteins and reason about which candidates merit follow-up. Spatial biomarker discovery faces the same problem at a different scale entirely, because the candidate space is not a list of molecules but a list of spatial patterns, positions, distances and neighborhood arrangements, multiplied across every position in every tissue section in a study.
A peer-reviewed imaging mass spectrometry study states the resulting bottleneck precisely. Manually examining the spatial mapping of thousands of molecular species across the surface of a single sample is laborious and risks introducing human subjectivity into the process, producing results whose reproducibility cannot necessarily be guaranteed. The volume of data generated by these experiments is large enough that computationally searching for biomarker candidates among a multitude of signals has become more efficient, and in many cases outright necessary, rather than an optional enhancement to manual review.
Our earlier coverage of this challenge from a detection-technology angle is available in Advancing cancer biomarker detection, and the broader strategic question of matching a biomarker approach to an increasingly diverse set of therapeutic modalities is addressed in The Compass and the Map: Navigating Biomarker Strategies for Novel Modalities. This spoke focuses specifically on the computational discovery step those broader strategic questions ultimately depend on.
Spatial signatures of disease
What makes spatial biomarker discovery different from a conventional differential expression analysis is that a genuine signature can exist entirely in the relationship between measurements, not in any single measurement’s value. A specific ratio of two cell types within a fixed distance of each other, a particular texture of tissue organization, or an unusual spatial entropy pattern within a region can each carry predictive information indistinguishable from noise if examined as an isolated data point, but become a coherent signature once the underlying spatial structure is captured explicitly.
That is precisely what a discovery workflow needs to be built to find: patterns that a researcher scanning individual images, or a bulk assay averaging across a whole sample, would have no way to notice, because the pattern’s existence depends on structure that only a spatially aware method preserves in the first place.
Discovery workflows
A 2026 study in Cancer Cell addresses the discovery bottleneck directly with an approach built around interpretability from the start, rather than attempting to explain a black-box model after the fact.
Solving the black-box problem by design, not after the factMost deep learning models in computational pathology remain black boxes, given the volume of visual information embedded in whole-slide images and the complexity of model parameters, offering limited insight into which features actually drive their predictions. Attempts to derive candidate biomarkers from such models using post hoc explainability techniques often produce explanations that are diffuse and ambiguous, since the underlying model relies on raw pixels or latent representations that were never designed to be human-interpretable, making it difficult to define, quantify, and validate the resulting signal as a genuine biomarker. PathPrism takes a different approach. It converts whole-slide histopathology images into a spectrum of 628 interpretable spatial features, including tissue spatial fractions, spatial entropy, and slide-level graph-based topology features that capture intra-tissue and inter-tissue spatial relationships. Because this representation is interpretable by construction rather than extracted from a black box afterward, it supports transparent linear modeling of prognosis, molecular alterations, and treatment response, with each prediction directly attributable to specific, named spatial biomarkers rather than an opaque combination of pixels. The resulting biomarker spectrum further supports hypothesis generation and in silico perturbation testing, letting a researcher explore candidate biomarkers computationally before committing to wet-lab validation. |
From discovery to candidate assay
Once a discovery method has generated a large set of candidate spatial features, the next problem is ranking them, since not every discovered pattern is equally likely to matter, and a discovery workflow that produces hundreds of candidates without prioritizing among them has not actually solved the researcher’s problem.
A separate peer-reviewed method demonstrates one solution to this ranking problem in an entirely different data modality: imaging mass spectrometry rather than histopathology images. The approach translates biomarker candidate discovery into a feature-ranking problem directly. Given a classification model that assigns tissue pixels to different biological classes based on their mass spectra, Shapley additive explanations, a technique from explainable artificial intelligence, quantify each molecular species’ relative predictive importance to that classification. The molecular species the model relies on most are ranked in descending order of importance, producing a shortlist of top candidates rather than requiring a researcher to manually scan every detected ion signal for significance.
That two different discovery approaches, developed for two different spatial data types, both converge on the same underlying logic, replacing black-box or manual pattern-finding with a method that produces a ranked, interpretable shortlist, suggests this is a structural requirement of spatial biomarker discovery generally rather than a solution specific to either histopathology or mass spectrometry imaging.
Validation requirements
A discovered and ranked spatial candidate is not yet a biomarker ready for use, and the distinction matters. Discovery produces a hypothesis, a specific spatial feature statistically associated with an outcome in the dataset it was found in; validation confirms that association holds in independent data and that the feature can be measured reproducibly outside the discovery environment.
The full requirements for that validation step, including analytical and clinical validation standards and how they differ for a spatial signature compared with a single-molecule biomarker, are covered in depth in Spatial Biomarkers and Companion Diagnostics: The Next Frontier. The specific path from a validated discovery candidate toward a regulatory-grade companion diagnostic assay is developed further in From Spatial Biomarker to Companion Diagnostic: The Development Path.
For the foundational taxonomy of spatial biomarker types this discovery process generates candidates within, see What Are Spatial Biomarkers? A Primer for Drug Developers, and for where biomarker discovery sits within the full spatial biology pipeline, Spatial Biology in Drug Discovery: From Target Discovery to Translational Medicine.
This article was produced under Drug Discovery News’s AI editorial policies.


















