Bioinformatics & omics

Genomics and DNA design

Designing sequence to be built and to work — the synthesis constraints of repeats, GC extremes and secondary structure, why codon optimisation is a real but overstated lever, and how regulatory elements set expression.

Writing DNA is now routine enough that design, not construction, is usually the limiting step. But the design space is bounded twice: by what the chemistry and assembly can physically produce, and by what the host cell will do with the sequence once it has it.

Synthesis constraints decide what can be ordered

Vendors reject or surcharge sequences for reasons that are entirely physical. Long repeats are the worst offender: identical stretches cause mispriming during amplification and misassembly during homology-based joining, so a construct with tandem repeats or many copies of the same promoter may be unbuildable even though every base is legal. Extreme GC content is the second constraint — GC-rich regions form stable secondary structure that stalls polymerases and coupling, while AT-rich regions destabilise and can be lost. Hairpins and G-quadruplexes interfere with both synthesis and sequencing verification. Practical design therefore breaks repeats deliberately, using synonymous recoding to make functionally identical elements sequence-distinct, and smooths GC across the construct. This is why a design tool’s first job is buildability scoring, not optimisation.

Codon optimisation is real and routinely oversold

Codon usage varies between organisms and correlates with the abundance of the matching transfer RNAs, so recoding a gene toward the host’s preferred codons can raise protein yield substantially — this is well established for heterologous expression. The overstatement lies in treating the effect as universal and monotonic. Translation initiation is frequently rate-limiting, and there the dominant variable is not codon usage but the folding energy of the messenger RNA around the start codon: strong local structure occludes ribosome binding, and a change of a few bases near the 5’ end can outweigh whole-gene recoding. Rare codons also serve functions — slowing elongation at domain boundaries appears to assist cotranslational folding — so maximally optimised genes sometimes yield more insoluble protein, not more active protein. Recoding additionally changes messenger stability, can create or destroy internal ribosome entry sites and splice signals in eukaryotes, and may introduce restriction sites or repeats that break assembly.

Regulatory elements set the range

Expression level is mostly a property of the parts around the coding sequence: promoter strength, ribosome binding site or Kozak context, terminator efficiency, and copy number. These behave as approximately composable parts within a well-characterised chassis, which is what libraries of measured promoters are for, but context dependence is persistent — the same promoter gives different output depending on adjacent sequence, plasmid backbone, growth phase and host strain. Reported part strengths are therefore conditional measurements, not constants.

Whole-genome design

Building at chromosome scale, as the synthetic yeast genome effort has done, changes the problem again: the design decisions are deletions and recodings across megabases, and the check is whether the organism remains viable. That work has shown both that genomes tolerate extensive redesign — removal of introns and transfer-RNA genes, wholesale codon substitution — and that a small number of individually reasonable changes can prove lethal in combination, requiring debugging by mapping the defect back to a specific edit.

The general lesson is that sequence design is an engineering discipline with real constraints, and its predictions get weaker the further the construct sits from the characterised context in which its parts were measured.

Last updated: