Diagnostics & medtech
Epigenetic clocks
How bisulfite conversion makes methylation readable at scale, why a few hundred CpG sites suffice where eighty-five thousand wouldn't be believed, how first-, second- and pace-of-aging clocks differ in their training targets, and where prediction silently fails to become cause.
Cells keep a diary in chemistry. Along the genome sit millions of cytosine-guanine doublets whose attached methyl groups switch genes’ readability without touching sequence, and their pattern shifts so reliably across decades that statistical models recover someone’s age from a speck of tissue within a few years. Epigenetic clocks are those models — and the deepest thing about them is definitional: a clock means whatever it was trained on.
The chemistry that remembers
Reading methylation at scale relies on conversion chemistry: treating DNA so that every unmethylated cytosine transforms into a form sequenced as thymine while methylated ones resist, turning a chemical modification into four-letter text difference easily measured genome-wide on commercial arrays covering hundreds of thousands of sites simultaneously. The age signal itself splits by direction: developmental gene regions tend toward additional methylation through life, structural heterochromatin loses it, and both drift quantitatively enough to serve as coordinates. A wrinkle worth admiring: these patterns prove largely shared across tissues, so a blood sample’s lymphocytes report on the body’s trajectory rather than merely their own, even though each cell type carries its own baseline accent.
Training targets define meaning
The measurable matrix vastly exceeds any honest sample size, so machine learning performs radical compression: regularised regression searches the site inventory for the handful carrying most predictive load, yielding working clocks from fewer positions than a shopping list. What varies between clock generations is not chemistry but ambition. First-generation models were taught to reproduce calendar age — technically impressive, scientifically circular if taken alone, since matching a passport proves nothing about health. Second-generation versions instead learned targets like hospital blood-panel values or, better, mortality itself, scoring people against observable consequences. Third-generation pace measures abandon accumulation entirely: built from cohorts followed for decades, they estimate slope — how fast this person currently burns through biological years, whether 0.8 or 1.2 per calendar annum. Pace answers the question aging medicine actually asks, which is why recent trial designs prefer it.
Prediction is not power
Here honesty must bite hard. That accelerated methylation age accompanies diseases says nothing about steering them: it may be an engine, a sensor, or exhaust — causality demands intervention experiments. Several exist now, treating repeat clock readings as trial endpoints around exercise regimens, dietary protocols or candidate geroprotectors. The arrangement harbours a specific logical trap: a therapy could move the molecular needle — the chosen yardstick — while leaving survival untouched, which is why regulators and epidemiologists distinguish surrogate-endpoint wins from life-extension proof and treat the former skeptically until mortality data arrive. Within-pulse reliability poses quieter problems too: replicate measurements scatter by a year or two from instrumentation plus genuine short-term fluctuation, so changes smaller than that are noise wearing a lab coat.
Commercially the field nests inside this uncertainty comfortably: modest panel costs, multi-algorithm stacking to fill reports, and mechanistic language ahead of interventional evidence. The measurement itself deserves admiration — decades of physiological history reconstructed from modified letters — but its consumer-facing cousins share the direct-testing pattern: precision about a quantity whose medical relevance still waits on trials designed to establish whether moving it moves anything that matters.