Foundries & design
Generative protein design
From co-evolution signals to diffusion over backbones, the design-test cycle that produces real binders, and an honest account of what still fails: specificity, membrane proteins, function.
A protein folds because its chain, jostled by thermal motion, settles into the shape with the lowest free energy. That single fact is what makes design possible in principle: a protein can be built by writing a sequence whose energy minimum is the shape you want. Structure prediction and generative design are the two directions of the same mapping — sequence to structure and back.
Where the models get their physics
The models are not given the laws of molecular mechanics; they learn the statistics of what evolution has already tried. Genomes preserve a geometric record: residues that touch in the folded structure co-evolve, because a mutation in one is usually compensated by a mutation in its partner, so a deep alignment of related sequences carries the contact map inside its correlation patterns — the signal that made structure prediction work. Design inverts the flow. Diffusion models, trained on the public archive of solved structures, learn the distribution of plausible protein backbones and generate new ones by starting from a random cloud of atoms and repeatedly removing noise, steered toward a chosen symmetry, motif or binding geometry around a target. Then inverse folding does the simpler half of the original problem: with the backbone fixed, the packing constraints narrow which amino-acid sequences can hold it.
The cycle does the real work
A designed sequence is a hypothesis, and only measurement converts it into a protein. The genes are synthesised, the proteins expressed, and the library screened — display on the surface of yeast or phage, sorting for binding, then direct affinity measurement on the survivors. Folding now succeeds far more often than it did in the physics-simulation era, which is the real change; binding usefully succeeds far less often than folding, which has not changed. What makes the field move is that every round of failures is data: measured outcomes feed the next generation of models, and the design-build-test loop — same discipline as in metabolic engineering — is the actual product, not any single binder.
What still fails
Specificity is the standing problem. A small binder touches its target through a handful of residues, and a handful of contacts cannot encode exclusivity — a molecule that binds one member of a protein family tends to bind its relatives, and off-target binding is discovered in the assay, not in the model. Membrane proteins lag far behind soluble ones: they are underrepresented in structural archives, and the physics of the lipid bilayer — hydrophobic embedding, no ordinary water, strong electrostatics — is poorly captured by models trained mostly on soluble folds. Function beyond sticking remains harder still: catalysis needs a precisely tuned reaction geometry, not just a well-packed pocket, and conformational switching and immunogenicity are largely screened for rather than designed. The honest summary: generative models have made geometry cheap, but geometry was never the whole problem.