Bioinformatics & omics

Single-cell omics

Droplet capture and barcoding, why a zero is usually sampling depth rather than biological absence, doublets and ambient contamination, and the stress signature that dissociation itself induces.

Bulk sequencing measures the average of a tissue, which is a number no cell in that tissue necessarily holds. Single-cell methods recover the distribution instead. The dominant approach isolates cells into nanolitre droplets with a barcoded bead, so every transcript captured in that droplet inherits the same cell barcode and a unique molecular identifier that marks it as one original molecule rather than a duplicate from amplification. After that the material is pooled and sequenced conventionally; the barcodes reconstruct which molecule came from which cell.

Dropout is sampling, not absence

Capture efficiency is limited: only a fraction of the messenger RNA in a cell is reverse-transcribed and recovered, and typical experiments end with a few thousand distinct molecules per cell. A gene expressed at a handful of copies therefore has a substantial chance of contributing nothing, purely by the statistics of drawing a small sample from a small pool. This is what a zero in a single-cell matrix usually means. Treating those zeros as biological absence produces spurious bimodality and inflated differences between cell groups; treating them as a known sampling process — modelling counts, aggregating over cells, or comparing at the level of cell populations rather than individual cells — is what makes the data behave.

Doublets and ambient RNA

Droplet loading is Poisson, so a fraction of droplets receive two cells. Their combined profile can look like a genuine intermediate or transitional state, which is the more damaging failure because it is biologically plausible. Loading cells sparsely reduces the rate at the cost of throughput; computational detection and, where feasible, genetic or hashtag multiplexing to identify cross-sample doublets are the standard defences.

Ambient RNA is the mirror problem. Cells that lysed during handling release transcripts into the suspension, and that soup is partitioned into every droplet. A highly expressed marker of one abundant cell type therefore appears at low level across all clusters, and can be mistaken for widespread low expression. Empty-droplet profiles estimate the ambient composition and allow it to be subtracted, which is why retaining those barcodes matters.

Dissociation is an experiment on the cells

Getting cells out of a tissue requires enzymatic and mechanical treatment, usually at elevated temperature, and cells respond transcriptionally within minutes — induction of immediate-early genes and heat-shock proteins is well documented, and it varies by cell type, so the artefact is not uniform and does not cancel. Dissociating cold with protease active at low temperature, or fixing before dissociation, reduces it; scoring the stress module explicitly and reporting it is the minimum honest practice. Cell types differ in how well they survive: neurons and adipocytes are fragile, so single-nucleus protocols are often used instead, at the cost of losing cytoplasmic transcripts.

The upshot is that a single-cell dataset describes the cells that survived dissociation, in proportions set by that survival, measured at a depth that hides low-expressed genes. Every one of those qualifiers is tractable, and each one is a place where a confident biological conclusion can be an artefact of handling.

Last updated: