Bioinformatics & omics
Metabolomics
Chemical diversity as the defining constraint of metabolite measurement: extraction and platform coverage, annotation as the real bottleneck and what confidence levels mean, matrix effects, and why quenching at sampling is mandatory.
Genomes, transcriptomes and proteomes are polymers of a small alphabet, so one chemistry extracts them, one instrument reads them and identification reduces to sequence matching. The metabolome has no such backbone. It spans ions, sugars, amino acids, organic acids, nucleotides, lipids and xenobiotics, across a polarity range from highly water-soluble to entirely hydrophobic and a mass range of a few tens to a couple of thousand daltons. Every practical consequence follows from this.
No extraction and no platform sees everything
An extraction solvent is a selection. Cold methanol recovers polar metabolites and precipitates protein; a chloroform or methyl-tert-butyl-ether partition is needed for lipids; volatile compounds require headspace sampling. Chromatography makes the same choice again: reversed-phase columns retain non-polar species and let polar ones elute unresolved in the void, while hydrophilic-interaction chromatography does the opposite, and gas chromatography requires derivatisation to make polar analytes volatile and thermally stable. Nuclear magnetic resonance is quantitative and highly reproducible without a standard curve, but is far less sensitive and sees only abundant species. A study that reports “the metabolome” has in practice reported the window defined by its extraction and its column.
Annotation is the bottleneck
An untargeted run yields thousands of features — a mass-to-charge ratio at a retention time — and most are never identified. Some are adducts, isotopes, in-source fragments or contaminants of the same underlying compound; many are genuine molecules absent from spectral libraries. The community reports confidence explicitly for this reason, following the Metabolomics Standards Initiative levels: an identification confirmed against an authentic standard run on the same system is the highest level, while a match to a spectral library, a putative compound class, or an unknown feature are progressively weaker claims. Reading a metabolomics result begins with asking at what level each named compound was assigned; a formula from accurate mass alone does not distinguish structural isomers, and isomers frequently differ in biology.
Matrix effects break the intensity-to-concentration link
In electrospray ionisation, coeluting compounds compete for charge, so the signal from an analyte depends on what else arrives at the source with it. The same concentration gives different intensity in plasma than in urine, and in a diseased sample than a healthy one. This is why absolute quantification requires isotope-labelled internal standards that coelute with the target and experience the same suppression, and why untargeted data should be treated as relative comparison across samples processed identically, not as concentration.
Turnover forces quenching at sampling
Metabolite pools turn over on timescales of seconds; glycolytic intermediates in particular. Between taking a sample and stopping its enzymes, the composition changes, and the change is not noise but a systematic drift toward whatever the remaining enzymes catalyse. Rapid quenching — cold organic solvent, liquid nitrogen — and controlled handling times are therefore part of the measurement, not sample logistics. The same discipline governs blood: tube type, time to centrifugation and freeze-thaw history are documented preanalytical variables with effects large enough to dominate a comparison between groups if they differ systematically.