Diagnostics & medtech
Next-gen bio-forensics
The science behind sequencing-based forensics: why STR matching could only confirm, what allele-frequency structure lets SNP profiles predict, how identity-by-descent search finds relatives no database contains, and where the statistical and ethical limits sit.
Classic DNA forensics was an answering machine. It could only confirm: amplify the suspect’s STR profile, compare it to evidence, and report a match probability against an unrelated population. The method’s power rested on identity already supplied by investigation. Sequencing changed the epistemic position of the trace itself, and it did so through three mechanisms.
The first is population structure made readable. Human allele frequencies are correlated across the genome within populations — the residue of drift and shared ancestry — and hundreds of thousands of SNPs record that structure faithfully. Principal-component analysis of a crime-scene profile therefore places its owner on the map of continental and regional ancestries without any reference sample from the person. This is a statement about statistics, not destiny: the prediction is a probability contour, sharper where the person’s ancestors came from a well-sampled region, blurrier where they did not — and the reference panels’ uneven coverage is a known calibration bias, strongest for populations outside Europe.
The second is phenotype prediction, bounded by heritability. Models like HIrisPlex-type SNP sets predict eye and hair colour because those traits have few loci of large effect; skin tone and height, being broader and more environmental, resolve much more coarsely. The bounds are honest and measurable — a predicted category carries a likelihood, and a trace that excludes the witness’s assumption is exactly as informative as one that confirms it. What sequencing added is not clairvoyance but a shift from exclusionary evidence to an investigative hypothesis with stated uncertainty.
The third is the genealogy mechanism, the genuinely new instrument. Relatives share DNA in segments inherited intact from common ancestors; the fraction shared falls by roughly half each meiosis removed, so third and fourth cousins still carry long stretches of identity-by-descent. A crime-scene profile uploaded to a consumer-genealogy dataset does not need the perpetrator in the database — it needs someone in it who shares such segments, after which conventional genealogical work walks the family tree inward. This is how a decades-old serial case was closed in days from a public dataset in 2018, and it works because human mating structure makes even distant kinship measurable. The same sequencing-paradigm logic drives pathogen forensics, where genomic pathogen surveillance applies it to outbreaks instead of suspects.
The limits define the ethics. Mixtures of contributors and degraded samples, trivial in STR practice, still tax SNP pipelines; kinship search reaches people who never consented to be findable, which is why jurisdictions regulate forensic phenotyping unevenly — permitted, restricted or banned — while the databases feeding it are built from consumer genetic testing opted in by relatives, not by suspects.