Bioinformatics & omics

Mass-spectrometry proteomics

The dynamic range of plasma, stochastic under-sampling in data-dependent acquisition versus DIA, and why peptide measurements do not cleanly reconstruct protein-level truth.

There is no polymerase for proteins. Every method for detecting a nucleic acid can make more of it first, so a single molecule can become a measurable quantity. Proteomics has no such step: whatever is in the sample is all there will ever be, and the instrument must detect it directly. Sensitivity is therefore a hard physical limit rather than a matter of cycling longer.

Dynamic range in plasma

Human plasma is the extreme case. Protein concentrations span roughly ten orders of magnitude, from albumin at tens of milligrams per millilitre down to tissue-leakage proteins and cytokines at picograms. Albumin alone is around half the total protein mass, and the top handful of species account for the large majority. A mass spectrometer sampling such a mixture spends its acquisition time on the abundant proteins, and the low-abundance proteins that carry most clinical interest are simply never selected. Depletion of abundant species and fractionation both help and both introduce their own problem: depletion antibodies remove bound partners along with their target, and fractionation multiplies instrument time per sample.

Data-dependent acquisition samples stochastically

In the classical workflow the instrument surveys the ions eluting at a moment, picks the most intense ones and fragments them for identification. When more peptides coelute than the duty cycle can address, selection is effectively a race, so a peptide near the threshold may be picked in one run and missed in the next. The result is the missing-value pattern familiar from any large data-dependent study: matrices with substantial gaps that are not random but abundance-dependent, which biases both imputation and differential testing.

Data-independent acquisition attacks exactly this. Instead of choosing precursors, the instrument fragments everything within a series of wide mass windows, cycling across the range on a fixed schedule. Every peptide in the sample is fragmented in every run, so the acquisition is reproducible and missingness falls sharply. The cost moves to the data: the resulting spectra are chimeric mixtures of fragments from many precursors, and extracting individual peptides requires either a spectral library or a library-free deconvolution.

The protein inference problem

Bottom-up proteomics digests proteins to peptides and measures peptides; protein-level results are reconstructed afterwards, and the reconstruction is not clean. Many peptides are shared between members of a protein family or between splice isoforms and cannot be assigned uniquely. A protein is often reported on the evidence of a small number of peptides, which may sit in accessible regions and say nothing about the rest of the molecule. Post-translational modifications change a peptide’s mass and move it out of the search unless explicitly considered, so an unmodified peptide can under-report a protein whose pool is heavily modified. Proteolysis itself is incomplete and sequence-dependent.

Consequently a “protein quantity” from a bottom-up experiment is a model-derived summary of the peptides that happened to be observed. Where the distinction matters — isoform-specific biology, modification stoichiometry, degradation products — targeted methods measuring named peptides against isotope-labelled standards remain the reference, and affinity-based platforms trade the specificity question for reagent specificity rather than removing it.

Last updated: