Crop biotech

Genetic engineering and breeding

Additive variance, the breeder's equation and genomic selection: why marker-assisted selection failed for yield, what genomic prediction actually does, and where single-gene engineering fits.

This is the page the rest of the cluster leans on. Crop improvement has two toolkits — recombination and selection on one side, direct sequence change on the other — and which one applies is decided by the genetic architecture of the trait, not by how modern the method is.

Why architecture decides the method

A trait controlled by one or two loci with large effects behaves like a Mendelian character. It segregates visibly, it can be tracked with a marker, and it can be introduced by backcrossing or, now, written directly. Disease resistance genes, herbicide tolerance, waxy starch and dwarfing alleles are of this kind, and they are where both marker-assisted backcrossing and gene editing work cleanly.

Yield is not of this kind, and neither are most things correlated with it. Yield is the summed output of hundreds to thousands of loci of small effect, interacting with each other and with the environment; no individual allele accounts for enough of the variance to be worth selecting on its own. Marker-assisted selection, which succeeded for the first class, largely failed for the second — the QTL detected in one population were small, environment-dependent, and often did not reappear in the next.

What genomic selection changed

Genomic selection answered that failure by giving up on identifying genes. A training population is genotyped densely and phenotyped carefully; a statistical model is fitted that predicts breeding value from all markers at once, each contributing a small estimated effect; and thereafter candidates are ranked from genotype alone. Nothing in the model is a discovery about biology, and that is deliberate — it captures the additive variance without needing to attribute it.

The gain shows up in the breeder’s arithmetic. Genetic gain per unit time rises with selection intensity, with the accuracy of selection, and with the amount of additive genetic variance, and falls with the length of the breeding cycle. Genomic prediction usually improves accuracy modestly for a complex trait; its larger contribution is that it lets a breeder select before the plant is grown to maturity, which attacks cycle length — the denominator, and the term with the most room in it.

Its limits are structural. Prediction accuracy depends on relatedness between the training set and the candidates and decays as that relatedness weakens, so the model must be retrained as the population moves. It captures additive effects well and epistasis poorly. And it converts the bottleneck into phenotyping: a model is only as good as the trait data behind it, which is why high-throughput phenotyping is a prerequisite for this method rather than an accessory to it.

Where engineering fits

Direct sequence change is not a competitor to this machinery but a way to introduce variation it can then select on. Editing supplies alleles that do not exist in the germplasm or that would arrive with linkage drag if introgressed; transgenes supply functions no plant relative has. Both then enter the same recurrent selection process, in locally adapted backgrounds, and are evaluated the same way — because a trait in the wrong genetic background is not a variety.

Last updated: