Articles

​Why sample preparation determines your data quality 

Before sequencers run, before spectrometers fire, before algorithms analyze — sample preparation has already determined how much of the truth you'll actually see
Written byAndrea Corona
| 7 min read
Laboratory equipment

The quality of the conclusions drawn from modern biomedical datasets is bounded by the quality of the samples that generated them.

GEMINI

Register for free to listen to this article
Listen with Speechify
0:00
7:00

Before sequencers run, before spectrometers fire, before algorithms analyze — sample preparation has already determined how much of the truth you'll actually see.

Drug discovery has never had better tools. Next-generation sequencing can resolve transcriptomes at the single-cell level. Mass spectrometry-based proteomics can detect thousands of proteins in a single run. High-content imaging can capture cellular phenotypes at scale, and AI can interrogate datasets that would have overwhelmed even the largest bioinformatics teams a decade ago.

And yet, across all of these platforms, one stubborn source of variability persists — often unacknowledged, rarely standardized, and frequently the reason that experiments fail to reproduce: sample preparation.

The problem is not new, but its consequences are becoming harder to ignore. A 2015 meta-analysis estimated that $28 billion per year is spent on preclinical research that cannot be reproduced — and methodological inconsistency, including variability in how samples are handled, collected, and processed, is a core contributor. As discovery programs grow more complex and datasets grow larger, the gap between what instruments can measure and what researchers reliably detect is increasingly determined upstream, before any data is acquired.

Why the upstream matters

In biomedical research, "data quality" is often discussed in terms of instrument performance, sequencing depth, or analytical pipeline design. These are important, but they are downstream of a more fundamental question: is the biological material entering the workflow representative of the biology being studied?

Continue reading below...
A 3D illustration showing stylized antibody molecules with blue and purple surfaces interacting in a dark, abstract molecular environment.
EbooksMass photometry in antibody analytics: from discovery to development
As antibody formats grow more complex, mass-based, single-particle measurements are providing clearer and faster insight into antibody structure and behavior.
Read More

Sample preparation encompasses every step between biological collection and analytical measurement — lysis, extraction, enrichment, digestion, library preparation, and more. Each step is an opportunity for signal loss, contamination, degradation, or bias. And because these effects compound across a multi-step workflow, small inconsistencies early in the process can translate into large, systematic errors by the time data is collected.

Researchers at the University of Washington described this challenge directly in a 2024 framework for quality control in quantitative proteomics, noting that sample processing variability stems from differences in collection and storage conditions, digestion parameters, enzyme efficiency, contaminants, and unidentified protocol issues — and that identifying these sources is critical to ensuring observed results reflect biology rather than technical artifacts.

The distinction between biological and technical variation is not just a statistical concern — it is a scientific one. When technical noise is mistaken for biological signal, researchers may pursue targets that do not exist, dismiss effects that do, or draw cross-study comparisons that are fundamentally invalid.

The proteomics problem

Proteomics offers some of the clearest examples of how sample preparation shapes data quality — and how difficult the problem is to solve.

Bottom-up proteomics, the dominant approach for large-scale protein identification and quantification, requires converting proteins into peptides through enzymatic digestion before mass spectrometry analysis. The most widely used enzyme for this is trypsin, but as researchers at Leiden University Medical Center and the University of Southern Denmark noted in a 2024 review, tryptic digestion remains a variable biochemical process that leaves missed cleavage sites at varying degrees. Variability in digestion efficiency directly affects which peptides are detected and in what quantities — without any change to the instrument or the downstream analysis.

The problem compounds. The same review found that mass spectrometry methods contend with high variability in sample preparation, instrumentation, and data analysis workflows, resulting in false positives and poor reproducibility unless robust statistical controls are applied.

Contaminants introduced during preparation are another persistent challenge. Common laboratory materials — pipette tips, skin care products, surfactant-based cell lysis reagents — can introduce polymer contamination at levels sufficient to compromise liquid chromatography–mass spectrometry (LC-MS) data. Working with particularly limited sample amounts makes this worse. Researchers studying paucicellular samples have found that multistep proteomics workflows can lead to substantial sample loss at each stage from cell lysis through peptide recovery — a problem that becomes especially acute in clinical settings where limited material is the norm.

For drug discovery teams relying on proteomics to characterize drug mechanism, identify biomarkers, or profile patient-derived samples, these sources of variability are not merely technical nuisances. They represent potential misreadings of biology that shape which compounds advance and which are deprioritized.

Different workflow, same constraint

The challenges in genomics differ in their specifics but not in their essence. RNA sequencing (RNA-seq), increasingly used for target identification, mechanism of action profiling, and biomarker discovery, is sensitive to a range of preanalytical variables that can undermine the reliability of transcriptomic data.

A 2025 study developing a quality control framework for blood-based RNA-seq biomarker discovery found that preanalytical metrics — encompassing specimen collection, RNA integrity, and genomic DNA (gDNA) contamination — showed the highest failure rates across all quality control checkpoints evaluated. The researchers found that gDNA contamination was significant enough to warrant a secondary DNase treatment step, which substantially reduced intergenic read alignment and improved the reliability of downstream analysis.

The implication is straightforward: sequencing depth and computational pipeline sophistication cannot compensate for RNA that has degraded during collection or DNA that was not adequately removed before library preparation. The study concluded that preparation standards are essential for improving reliability, accelerating biomarker discovery, and translating findings into clinically actionable diagnostics and therapeutics.(5)

This point has direct relevance for drug developers. RNA-seq is now routinely used to assess drug effects via transcriptomic changes, guide target identification, and establish expression-based biomarker signatures that support patient stratification in clinical trials. If the preparation of RNA samples is not consistent — across operators, time points, or sites — the signal being used to make those decisions may not reflect consistent biology.

Where density matters

Cell-based assays, including phenotypic screens, functional genomics studies, and patient-derived models, introduce a different set of preparation variables. Cell seeding density, for instance, is known to affect gene expression at a global level, yet protocols vary widely across laboratories.

Researchers examining reproducibility in chromatin immunoprecipitation sequencing found protocols recommending cell densities ranging across several orders of magnitude — variation that may contribute meaningfully to inconsistent results between studies. Cell culture conditions, passage number, media formulation, and time between plating and assay execution are all sources of variability that are rarely fully controlled or reported, but that can substantively affect what a cell-based assay actually measures.

As discovery teams increasingly turn to more complex biological systems — organoids, co-cultures, patient-derived models — the preparation challenges multiply. These systems are inherently more variable than immortalized cell lines, and the gap between what is biologically interesting and what is technically reproducible narrows.

The meta-problem

Individual experiments are not conducted in isolation. Datasets are pooled across studies, institutions, and time points to build the statistical power needed for robust target validation and biomarker discovery. That aggregation only works if the underlying data is comparable — which requires that the preparation methods generating it are consistent.

The proteomics community has been grappling with this directly. A multiyear longitudinal harmonization study published in 2025 tracked quality control metrics across mass spectrometry-based proteomics core facilities, finding that consistent instrument monitoring and standardized sample preparation were essential to achieving inter-laboratory comparability. Without that consistency, pooled datasets may reflect technical batch effects rather than biological differences — a problem that sophisticated normalization strategies can mitigate but not eliminate.

The same principle applies to genomics. Without standardized collection and processing, RNA-seq datasets from different sites or time points may be incompatible, even when generated on the same sequencing platform using the same analysis pipeline.

Automation as a partial answer

One approach gaining traction across both academic and industry settings is the use of automated sample preparation platforms to reduce operator-driven variability. Liquid handling automation, in particular, has been applied across proteomics, genomics, and cell-based workflows to standardize pipetting steps that are prone to human error.

Researchers reporting on a fully integrated automated proteomics sample preparation platform in a 2025 study in Analytical Chemistry found that the system outperformed established manual and semiautomated workflows in accuracy, precision, and reproducibility, achieving consistent peptide identification and quantification across batches. The platform's output was shown to directly support decision-making in experimental drug discovery approaches.

A parallel effort in physicochemical property assays — often used to assess drug candidates early in development — found that implementing an automated robotic system improved sample throughput six to 10-fold while producing consistent, reproducible assay data that outperformed semiautomated preparation. The work also highlighted that AI-based predictive models for compound properties are only as good as the experimental data they are trained on, making reliable sample preparation a prerequisite for effective in silico decision-making.

At industry forums and in the scientific literature, the conversation around automation has evolved from throughput to quality. Research teams across pharma and biotech have increasingly applied automation not simply to run more experiments, but to enforce experimental consistency — addressing reproducibility, data integrity, and early decision-making as primary goals rather than secondary benefits. The emphasis has shifted toward generating high-fidelity evidence that supports confident decisions — which depends on controlling the variables that automation can address.

That said, automation is not a complete solution. It addresses pipetting precision and reduces operator-to-operator variability, but it does not resolve the upstream choices that define a preparation workflow: which lysis buffer to use, how to handle degradation-sensitive analytes, whether and how to deplete abundant proteins in plasma proteomics, or what quality thresholds to apply before a sample proceeds to analysis. Those decisions require scientific judgment, and they require that the field develop and adopt shared standards.

What reliable data generation now depends on

The field is beginning to formalize what reliable sample preparation looks like. In proteomics, the Clinical Proteomic Tumor Analysis Consortium developed a system suitability protocol to evaluate targeted assays across 11 institutions and 15 instruments, establishing metrics for chromatographic reproducibility and instrument performance that could anchor cross-site comparability. In RNA-seq, the development of end-to-end quality control frameworks now provides a structured approach for monitoring preparation quality at every stage from collection to sequencing.

For drug developers, the practical implications are several. First, sample preparation protocols should be treated as a critical experimental variable rather than a fixed background condition. Changes in operator, reagent lot, storage duration, or processing time should be tracked and assessed for their impact on data quality. Second, quality control samples — internal standards incorporated at the protein or peptide level in proteomics, bulk RNA controls in sequencing workflows — should be used to differentiate preparation failures from biological effects. Third, when comparing datasets across studies, sites, or time points, preparation method comparability should be assessed as a prerequisite to interpretation.

The broader point is that the quality of the conclusions drawn from modern biomedical datasets is bounded by the quality of the samples that generated them. Advances in sequencing depth, instrument sensitivity, and computational analysis raise the ceiling of what is possible — but preparation variability establishes the floor below which reliable results cannot be obtained. For a field where discovery decisions are made on the basis of data that will eventually guide clinical development, that floor matters enormously.

Add Drug Discovery News as a preferred source on Google

Add Drug Discovery News as a preferred Google source to see more of our trusted coverage.

About the Author

  • Headshot

    Andrea Corona is the senior editor at Drug Discovery News, where she leads daily editorial planning and produces original reporting on breakthroughs in drug discovery and development. With a background in health and pharma journalism, she specializes in translating breakthrough science into engaging stories that resonate with researchers, industry professionals, and decision-makers across biotech and pharma. Her work blends investigative reporting with a deep understanding of the drug development pipeline, and she is particularly interested in stories at the intersection of science, innovation and technology.

    View Full Profile

Here are some related topics that may interest you:

Loading Next Article...
Loading Next Article...
Subscribe to Newsletter

Subscribe to our eNewsletters

Stay connected with all of the latest from Drug Discovery News.

Subscribe

Sponsored

Gloved researcher transferring liquid into a microplate using a multichannel pipette.
Discover practical strategies to improve pipetting accuracy, reproducibility, ergonomics, and instrument performance across diverse laboratory workflows.
Multichannel pipette dispensing a serial dilution into a 96-well microplate.
Discover practical strategies for performing reliable serial dilutions with optimized liquid handling and mixing.
Serial dilution series in microcentrifuge tubes showing progressively decreasing concentrations of a purple solution.
Learn best practices for improving the accuracy, precision, and reproducibility of automated serial dilution workflows.