Monitoring & conservation

Reference genomes as measurement infrastructure

Why reference genomes and pangenomes are measurement infrastructure: alignment interprets raw reads, representation bias propagates into identifications and variant calls, pangenomes restore what a single reference compresses away, and why sovereignty pulls against the sharing the method needs.

Sequencers do not read genomes; they read fragments — and everything downstream, from variant calls to species identifications, is a comparison of those fragments against reference sequences that somebody has already assembled and published. A genome database is therefore less like a library and more like a calibration standard: the instrument is the sequencer, but the scale is the database — the same role certified reference materials play elsewhere in measurement.

Why inference needs a reference

A variant is called where a read disagrees with the reference while the sample agrees with it elsewhere; a metabarcoding read becomes a species because a near-identical barcode already sits in the library. Both operations silently assume the reference spans the truth. Assembly itself is the hard part: long reads bridge repeats that short reads collapse, chromosome-level structure requires scaffolding from proximity ligation, and annotation — gene models, pathways — is transferred from relatives and inherits their errors. A finished reference is an estimate with structure, not a photograph.

Representation bias is measurement bias

Coverage of the world’s biodiversity and of human genetic variation is skewed — model organisms, crops and livestock, and in human genetics populations of European ancestry — and the skew propagates into every act of comparison. A species absent from the barcode library is not merely “unidentified”: its reads are assigned to the nearest sequenced relative. A population absent from variant databases has its variants read as rare and therefore suspect, and risk scores trained on the represented majority lose accuracy outside it. The general law: an estimator built on a reference inherits the reference’s blind spots, and missing data does not stay neutral — it turns into confident misclassification. The consequence for conservation is concrete: when a species enters the library as a single specimen, every other population of that species is measured as a deviation from one individual, and the genetic diversity that lets the species survive is stored as the anomaly.

Pangenomes and the sovereignty tension

The single-reference design compresses all variation absent from the chosen individual into anomalies. A pangenome — a set of assemblies or a graph carrying the variation of a whole population — restores the missing baseline: every individual maps somewhere, and “difference from reference” stops being an artefact of who happened to be sequenced first. Building them is expensive but mechanical; the harder tension is political. Access control over genomic data — for biosecurity, for benefit-sharing, for national advantage — is real and defensible, yet the method’s accuracy is a network good: every database that goes dark, every population sequenced only behind a wall, degrades inference everywhere else. Sovereignty and reference quality pull against each other, and the honest position is to name the trade-off rather than pretend that a fragmented set of closed databases measures the same thing an open one does.

Last updated: