Diagnostics & medtech

AI-driven biomarker discovery: when the pattern finds the marker

How cohort-scale measurement and supervised models discover biomarkers, why assay locking and regulatory review are the actual bottleneck, and what the NHS-Galleri readout and Caris's approvals establish.

A classical biomarker begins with a hypothesis: a protein is measured because biology says it should move. AI-driven biomarker discovery inverts the direction of inference. Population-scale molecular measurement — methylation arrays across hundreds of thousands of plasma samples, whole-exome and transcriptome profiles over comparable cohorts — is placed under supervised learning against outcomes, and the model returns patterns nobody named in advance: a cancer signal and its predicted tissue of origin carried in cell-free DNA methylation, a therapy-response signature distributed across thousands of genes. The discovery step, once the rate-limiting imagination of the field, has become the cheap step. Everything downstream is not.

The reason is the two gates. A discovered classifier must first be locked into an assay — a fixed analytical object whose inputs, thresholds and reporting a regulator can inspect — and proven reproducible under CLIA and CAP accreditation. Then it must survive clinical validation and coverage review: FDA PMA examination, Medicare coverage through MolDX, IVDR certification in Europe. The field’s recent history is the story of these gates operating on AI-born markers at scale. GRAIL’s Galleri, a methylation-pattern classifier for multi-cancer early detection, reported full NHS-Galleri trial results at the 2026 ASCO Annual Meeting — a four-fold higher cancer detection rate against standard screening and a substantial reduction in stage IV diagnoses, with Pathfinder 2 results across more than 35,000 participants — and has an FDA PMA application submitted, an advisory committee anticipated. Caris Life Sciences carries the conversion proof on the therapy-selection side: FDA approval for MI Cancer Seek in November 2024, MolDX approval for Caris ChromoSeq, Q2 2026 revenue of $263.7 million with full-year guidance of $1.03–1.04 billion, and AI models predicting immunotherapy-relevant molecular features and brain-metastases risk in breast and lung cancer.

The economics deserve as much attention as the algorithms, because the business is laboratory medicine at population scale. GRAIL operates a CLIA-certified laboratory in Research Triangle Park sized for up to one million tests per year, backed by $110 million of equity financing completed with Samsung C&T and Samsung Electronics in June 2026 and Q2 2026 revenue of $44.7 million. Caris runs over 275,000 square feet of laboratory space across four laboratories. Capacity of this kind is built ahead of covered demand — the wager being that a regulatory win converts infrastructure into margin.

What the method cannot do is the honest boundary of the field. A model trained on a cohort inherits its cohort’s blind spots: populations absent from training do not get validated markers. Correlation at discovery scale is cheap, so the differentiating evidence is interventional — whether finding cancers earlier changes stage distribution and survival, which is precisely what the NHS-Galleri program tested inside a national health system rather than a registry. And Europe still holds its gate: any European deployment certifies under IVDR 2017/746, a bar the cfDNA field has passed and multi-cancer tests have not yet faced. The field’s near-term competition is not who trains the largest model but who converts — who walks a discovered pattern through assay lock, review and coverage, and keeps a million-test laboratory running while doing it.

Last updated: