Genome engineering
DNA synthesis biosecurity screening
Homology-based screening of synthesis orders, the 200-nucleotide threshold in US guidance, and the defeats: split orders, functional-equivalence redesign, and benchtop synthesisers outside the screened supply chain.
Commercial DNA synthesis is the point at which a pathogen sequence stored as text becomes physical material. Screening is the control placed there: before a provider manufactures an order, its sequence is compared against curated databases of agents and toxins of concern, and the customer is checked as well.
How the comparison works, and what fixes the threshold
Screening is homology search. The submitted sequence is aligned against reference sequences from pathogens and toxins, and a hit above a defined similarity and length is flagged for human review. The length matters because short matches are meaningless: sequence fragments shared with harmless organisms are common, and a screen that flags them produces more false alarms than a review team can absorb.
The threshold in United States guidance is the well-known number. The 2010 Screening Framework Guidance for Providers of Synthetic Double-Stranded DNA defines a sequence of concern as one with a best match, over a window of 200 nucleotides or more, to a select agent or toxin. That figure is a judgement about the false-positive rate, not a biological boundary — it is the length at which similarity begins to mean something operationally, and it is also, by construction, a published specification of what will not be examined.
The three ways the control is defeated
Splitting. A sequence long enough to trigger review can be ordered as fragments, each below threshold or each individually unremarkable, from several providers, and assembled by the customer. Screening is per order and per provider; the assembled construct is nobody’s order.
Functional-equivalence redesign. Homology detects resemblance to known sequences, so a sequence that has been recoded to be far in sequence space while encoding the same protein evades detection while retaining function. Codon reassignment is trivial to perform. This is the structural limit of the whole approach: the control is defined over sequences, but the hazard lives in function, and the two are not the same object. Screening on predicted structure or function rather than sequence identity is the research direction that follows, and it is not yet a deployed control.
Synthesis leaving the screened channel. Benchtop synthesisers put oligonucleotide production in the customer’s own laboratory, where an order never reaches a provider at all. The control point then has to move — into the instrument’s firmware, its consumable supply, or its remote authorisation — and that relocation is not equivalent, because it depends on the device manufacturer rather than on a small number of providers.
Structural properties of the control
Screening is voluntary in most jurisdictions and coordinated through industry association membership rather than statute, so its coverage is a function of who joins. It also has an unavoidable tension with commercial confidentiality: a customer’s sequence is proprietary, and screening requires disclosing it, which is why cryptographic screening schemes that compare against a hazard database without revealing either side have been developed.
The information-technology attack surface of these systems — instrument networks, order pipelines, laboratory data integrity — is a distinct subject, treated on the bio-cybersecurity page rather than here.