Strain engineering (synbio CRO)
01Overview and value chain
Markers: [EC: EU industrial-biotechnology programmes / TRL 1–6 scale-up | OECD: 2.8 Industrial biotechnology | Regulator: EPA (US)]
A strain-engineering contract research organisation sells the biological catalyst rather than the product: given a target reaction or molecule, it delivers an enzyme variant or a production strain that makes the process economic. The work is organised as a design-build-test-learn cycle, and the modern differentiator is what trains the design step — proprietary experimental datasets accumulated over fifteen to twenty years rather than public sequence databases alone. The gains are measurable at process level: in one documented pharmaceutical cascade the engineered enzymes raised substrate loading from 5 g/L to 67 g/L, cut the enzyme load of one enzyme from 100 weight percent to 8 and of another from 100 to 3, shortened reaction time from 21 hours to 12, and held stereochemical purity above 99% enantiomeric and diastereomeric excess. That is the value proposition in a sentence: the same chemistry, run at an order of magnitude higher concentration with a fraction of the catalyst. The sector splits into commercial platform companies selling engineered enzymes, public-private biofoundries carrying projects from research to pre-industrial scale, and large gene-synthesis houses that grew from oligonucleotides into full molecular-biology service stacks.
The key directions of strain engineering services are:
- Directed Evolution Platforms: integrated enzyme-engineering platforms combining machine learning, bioinformatics and laboratory workflows in a closed design-to-validation loop, trained on two decades of in-house experimental data rather than public data alone.
- Computational Protein Design: physics-based de novo design driven by AI agents over a proprietary structured dataset, aiming at hit rates ten to a hundred times higher than mutagenesis-led screening, and carried through to manufacturability and formulation rather than stopping at a sequence.
- Genome Editing of Production Strains: CRISPR nuclease families developed specifically to edit industrial hosts — E. coli, Bacillus, Pichia, Aspergillus — with activity across prokaryotic and eukaryotic organisms, used to optimise the chassis rather than to treat disease.
- Public-Private Biofoundries: consortium-funded facilities that carry microbial cell-factory projects across TRL 1 to 6, covering synthetic biology, microbial resource management, bioprocess development and scale-up in one place.
Sectoral value chain
[Target Definition] ──> [Design] ──> [Build & Screen] ──> [Strain/Enzyme Delivery]
│
(test data feedback)
│
▼
[Industrial Process] <─── [Bioprocess Scale-Up] <─────────┘Value chain levels
| Level | Description | Key inputs/outputs |
|---|---|---|
| Target Definition | The client’s reaction, molecule or process constraint is translated into a protein or strain specification, including the conditions the catalyst must survive | In: Target reaction, process conditions, commercial constraints. Out: Design specification for enzyme or strain. |
| Design | Computational design proposes variants — physics-based de novo design or ML models trained on proprietary experimental data — rather than random mutagenesis | In: Specification, proprietary design dataset, models. Out: Ranked variant library for construction. |
| Build and Screen | Variants are synthesised and expressed in the chassis, then screened under application-relevant conditions; results feed back into the design model | In: Variant designs, gene synthesis, expression hosts. Out: Measured variant performance, updated model. |
| Strain and Enzyme Delivery | The selected variant or production strain is handed over with its performance data, having been validated under real process conditions rather than assay conditions | In: Best-performing variants, validation data. Out: Delivered enzyme or production strain. |
| Bioprocess Scale-Up | The biology is carried from bench to pre-industrial scale, where titre, substrate loading and catalyst load determine whether the process is economic | In: Delivered strain, fermentation and downstream development. Out: Scalable process at pre-industrial TRL. |
| Industrial Process | The client runs the engineered biology in production, with the gains realised as higher substrate loading, lower enzyme load and shorter cycle time | In: Scaled process package. Out: Commercial production at improved process economics. |
Cross-cutting technologies of the sector:
- Proprietary Design Datasets: the structured record of what has actually been built and measured in-house over fifteen to twenty years, which is what separates a design platform from a public-database screen.
- Closed-Loop Design-Build-Test-Learn: the engineering cycle that turns each screening round into training data for the next design round, increasingly driven by AI agents rather than manual iteration.
- Industrial-Host Genome Editing: nuclease families selected for activity in production organisms rather than in mammalian cell lines, enabling targeted edits in the chassis that actually makes the product.
02US
The United States holds the commercial platform layer: privately held companies that sell engineered proteins as a product, backed by long-accumulated proprietary datasets and validated through named pharmaceutical partnerships.
Codexis CodeEvolver, Arzeda computational design, documented process gains
- Codexis (CodeEvolver): an integrated, proprietary enzyme-engineering platform combining machine learning, bioinformatics and laboratory workflows in a fully connected design-to-validation workflow, with models trained on more than twenty years of Codexis-specific experimental data rather than public data alone, applied across pharmaceuticals, RNA therapeutics and diagnostics.
- Documented process gains: in a collaboration with Bristol Myers Squibb on BMS-986278, an LPA1 antagonist for pulmonary fibrosis, Codexis engineered a three-enzyme biocatalytic cascade, evolving an ERED and a KRED to hold above 99% enantiomeric and diastereomeric excess while raising substrate loading from 5 g/L to 67 g/L, cutting the ERED load from 100 weight percent to 8 and the KRED from 100 to 3, and shortening reaction time from 21 hours to 12.
- Arzeda (Intelligent Protein Design): AI agents drive a closed-loop design-build-test-learn cycle over a proprietary multi-year dataset with physics-based protein design rather than mutagenesis or public sequences; the company reports more than fifteen years of structured design data and targets hit rates ten to a hundred times higher, carrying designs through productisation, formulation and manufacturability rather than delivering a sequence.
03CN
China’s strain-engineering service base grew out of gene synthesis rather than enzyme evolution: the anchor company scaled from oligonucleotide and gene-synthesis work into a full molecular-biology and biologics service stack, inside a policy environment that now names bio-manufacturing as a priority future industry.
GenScript gene-synthesis origin, full service stack, bio-manufacturing policy priority
- GenScript (金斯瑞): founded in 2002 in New Jersey and built on gene synthesis, it expanded from gene-cloning efficiency, cost and quality work into antibodies, reagents, peptide synthesis, protein expression, antibody and protein engineering, and further into immunotherapy and biologics CDMO operations, serving scientists in more than one hundred countries.
- Domestic synbio CRO sector (qualitative): Chinese industry analysis describes the field organised around the design-build-test-learn cycle applied to chassis cells such as E. coli and Saccharomyces cerevisiae, with the value chain split into upstream read-write-edit technologies, midstream platform companies and downstream applications — but no second domestic strain-engineering CRO was independently confirmed in 2026 sourcing, so the sector is described here qualitatively rather than tabled.
- Policy pull: bio-manufacturing was named first among the future industries in the 2025 Chinese government work report’s commitment to build an investment-growth mechanism for future industries, and the field sustains its own industrial conference series, with the fourth China synthetic-biology and bio-manufacturing conference held in Shenzhen in January 2026.
04EU
Europe’s distinctive contribution is institutional rather than corporate: a public-private biofoundry model that carries projects from research to pre-industrial scale, alongside a listed industrial-biotech company whose genome-editing IP targets production strains directly.
TWB public-private biofoundry, BRAIN Biotech BMC nucleases, TRL 1–6 coverage
- Toulouse White Biotechnology (TWB): a public-private biofoundry in France jointly managed by INRAE, INSA and CNRS, created in 2012 and operating as a joint service unit with a consortium of roughly 45 to 50 members. It covers synthetic biology, microbial resource management, bioprocess development and scale-up across TRL 1 to 6, running around 50 R&D projects a year and having completed more than 400 since inception.
- BRAIN Biotech (BMC nuclease family): its BRAINBiocatalysts segment supplies enzymes, microorganisms and ingredients for industrial use alongside CRO activity and production-strain development, and it holds a granted European patent (EP4301852 B1) for the CRISPR-BMC nuclease. The nuclease is active across prokaryotic and eukaryotic organisms — bacteria, yeasts, fungi, plants and mammalian cells — and is aimed explicitly at optimising microbial production strains such as E. coli, Bacillus, Pichia and Aspergillus.
- The pre-industrial gap: the European model is built around the step that neither a university nor a commercial client wants to own — carrying a validated strain from laboratory result to a bioprocess that can be handed to industry — which is why the biofoundry sits as a shared consortium asset rather than a private service company.
05Leading companies and research institutes
| Company / Institute | Country | Key products / platforms | Tech features | Status 2026 |
|---|---|---|---|---|
| Codexis | 🇺🇸 USA | CodeEvolver enzyme-engineering platform | ML plus bioinformatics in a closed design-to-validation loop; models trained on 20+ years of in-house data; BMS cascade raised loading 5 to 67 g/L, cut enzyme load to 3-8 wt% | commercial; pharma, RNA therapeutics and diagnostics |
| Arzeda | 🇺🇸 USA | Intelligent Protein Design Technology | AI agents driving closed-loop DBTL; physics-based de novo design; 15+ years structured design data; targets 10-100x hit rates; design through to formulation | commercial; full-stack design to product |
| Toulouse White Biotechnology | 🇫🇷 France | Public-private biofoundry, TRL 1-6 | Joint service unit of INRAE, INSA and CNRS; created 2012; ~45-50 consortium members; ~50 R&D projects/year, 400+ completed | operating; pre-industrial scale-up focus |
| BRAIN Biotech | 🇩🇪 Germany | BMC nuclease family, BRAINBiocatalysts | European patent EP4301852 B1 for CRISPR-BMC; active in bacteria, yeasts, fungi, plants and mammalian cells; targets E. coli, Bacillus, Pichia, Aspergillus strains | commercial; CRO plus own product initiatives |
| GenScript | 🇨🇳 China | Gene synthesis and molecular-biology services | Founded 2002; gene synthesis, peptide synthesis, antibody customisation, protein expression and engineering; expanded into immunotherapy and biologics CDMO | commercial; serves 100+ countries |
06Tech stack and innovations
The stack is organised around one loop — design, build, test, learn — with the competitive difference lying in what trains the design step and how far the provider carries the result toward manufacture.
- Directed Evolution with Machine Learning:
- An integrated platform connects computational design to experimental validation in one workflow, with proprietary models trained on twenty-plus years of the provider’s own experimental results rather than only on public sequence data.
- The output is measured in process terms: a three-enzyme cascade delivering above 99% enantiomeric and diastereomeric excess while substrate loading rises from 5 g/L to 67 g/L and enzyme loads fall from 100 weight percent to 8 and to 3, with reaction time down from 21 hours to 12.
- Physics-Based Computational Protein Design:
- De novo design guided by physics and a proprietary structured dataset, with AI agents running the design-build-test-learn cycle, is positioned against mutagenesis and public-sequence approaches and targets hit rates ten to a hundred times higher.
- The scope extends past the sequence: designs are validated under application-relevant conditions and taken through formulation and manufacturability, because a protein that cannot be produced or formulated is not a product.
- Genome Editing Aimed at Industrial Hosts:
- The BMC nuclease family, protected by a granted European patent, provides high activity across prokaryotic and eukaryotic organisms and is aimed at defined edits in the production chassis — E. coli, Bacillus, Pichia, Aspergillus — rather than at therapeutic editing.
- This matters because the limiting factor in an industrial fermentation is usually the host’s own metabolism, not the pathway enzyme in isolation.
- Biofoundry Infrastructure and Scale-Up:
- A public-private biofoundry spans synthetic biology, microbial resource management, bioprocess development and scale-up in one unit across TRL 1 to 6, sustaining roughly 50 projects a year and more than 400 completed.
- The consortium structure exists because the pre-industrial step is expensive, shared across roughly 45 to 50 members, and is the stage at which most laboratory results fail to become processes.
07Value chains and production pipelines
Industrial pipeline of a strain-engineering engagement (TRL 1–6 progression, EPA oversight of engineered organisms)
┌───────────────────────────┐ ┌───────────────────────────┐
│ 1. Target Specification │ ───> │ 2. Computational Design │
└───────────────────────────┘ └───────────────────────────┘
│
▼
┌───────────────────────────┐ ┌───────────────────────────┐
│ 4. Screening Under │ <─── │ 3. Build & Expression │
│ Process Conditions │ │ │
└───────────────────────────┘ └───────────────────────────┘
│
▼
┌───────────────────────────┐ ┌───────────────────────────┐
│ 5. Bioprocess Scale-Up │ ───> │ 6. Handover & Production │
└───────────────────────────┘ └───────────────────────────┘Stage 1: Target Specification
The client’s reaction or molecule is translated into a protein or strain specification that includes the conditions the catalyst must survive — substrate concentration, temperature, solvent, cycle time. Specifying against real process conditions rather than assay conditions is what determines whether the eventual variant survives contact with the plant.
Stage 2: Computational Design
Variants are proposed by physics-based de novo design or by machine-learning models trained on the provider’s proprietary dataset — fifteen to twenty years of structured in-house results — rather than by random mutagenesis. This is where the ten- to hundred-fold hit-rate claims are made, and where public-data-only approaches are explicitly said to fall short.
Stage 3: Build and Expression
Designed variants are synthesised and expressed in the chosen chassis. Where the host itself limits performance, genome editing with nucleases selected for activity in E. coli, Bacillus, Pichia or Aspergillus alters the production strain rather than only the pathway enzyme.
Stage 4: Screening Under Process Conditions
Variants are measured under application-relevant conditions, and the results are fed back as training data for the next design round — the closed loop that distinguishes a platform from a one-off screening campaign. The metrics that matter are the industrial ones: substrate loading, catalyst load, stereochemical purity and reaction time.
Stage 5: Bioprocess Scale-Up
The validated biology is carried from bench toward pre-industrial scale, the TRL 1 to 6 span a biofoundry is built to cover. This is the stage the consortium model exists to fund, because it is expensive, shared across roughly 45 to 50 members, and the point at which most laboratory results fail to become processes.
Stage 6: Handover and Production
The engineered enzyme or strain is delivered with its validation data, and the gains appear in the client’s plant as higher substrate loading, lower enzyme load and shorter cycle times — in the documented case, an order-of-magnitude concentration increase with catalyst use cut to a few weight percent, which is what makes a biocatalytic route competitive with a chemical one.
| Supplier | Price | Lead time | Certificates | Risk | Confidence |
|---|---|---|---|---|---|
| Codexis | custom | 12 wk | Low | HIGH | |
| GenScript | custom | 8 wk | Low | HIGH | |
| Arzeda | custom | 16 wk | Medium | HIGH | |
| Toulouse White Biotechnology | custom | 20 wk | Low | HIGH | |
| BRAIN Biotech | custom | custom | Commercial BMC Nuclease (EP4301852 B1) | Low | HIGH |
AI note: strain-engineering-synbio-cro (EN)
Key directions:
- Directed evolution platforms — Codexis CodeEvolver integrates ML, bioinformatics and lab workflows in a closed design-to-validation loop, with models trained on 20+ years of Codexis-specific experimental data rather than public data alone.
- Computational protein design — Arzeda’s Intelligent Protein Design Technology runs AI agents over a proprietary multi-year dataset with physics-based design (not mutagenesis or public sequences); 15+ years of structured design data, targeting 10-100x hit rates, carried through to productisation, formulation and manufacturability.
- Genome editing of production strains — BRAIN Biotech’s BMC nuclease family (granted European patent EP4301852 B1) is active across bacteria, yeasts, fungi, plants and mammalian cells, aimed explicitly at optimising E. coli, Bacillus, Pichia and Aspergillus production strains rather than at therapeutic editing.
- Public-private biofoundries — TWB spans synthetic biology, microbial resource management, bioprocess development and scale-up across TRL 1-6, with ~50 R&D projects a year and 400+ completed.
The headline evidence is the Codexis/Bristol Myers Squibb cascade for BMS-986278 (LPA1 antagonist, pulmonary fibrosis): a three-enzyme cascade in which an evolved ERED and KRED held >99% ee and >99% de while substrate loading rose from 5 g/L to 67 g/L, ERED load fell from 100 wt% to 8 wt%, KRED from 100 wt% to 3 wt%, and reaction time dropped from 21 h to 12 h.
Regulatory:
- US: EPA oversight of engineered organisms; the commercial validation route is named pharmaceutical partnerships rather than a regulatory approval of the service itself.
- EU: the distinctive structure is institutional — TWB is a joint service unit of INRAE, INSA and CNRS (created 2012, ~45-50 consortium members) covering the pre-industrial TRL 1-6 gap that neither universities nor clients want to fund. BRAIN Biotech’s position rests on granted European patent IP.
- China: no service-specific regulator surfaced; the driver is policy — bio-manufacturing named FIRST among future industries in the 2025 government work report’s future-industry investment-growth commitment, with the 4th China synthetic-biology and bio-manufacturing conference held in Shenzhen in January 2026.
Companies not in table: Selexis (probed, 0 of 5 retrieved sources mentioned it — dropped by the relevance gate and REPLACED by BRAIN Biotech as the one permitted EU alternate). Synbio Technologies was also dropped, and the reason is worth recording: its alias list included the generic term “合成生物” (= “synthetic biology”), which scored 12/14 name_hits against Chinese industry-report sources that discuss the FIELD, not the company. That is a FALSE CONFIRMATION mode of the relevance gate — a generic alias inflates the score. Aliases must be company-specific tokens (“金斯瑞” is; “合成生物” is not).
Processing note: target specification against real process conditions (substrate concentration, temperature, solvent, cycle time) -> computational design from a proprietary dataset -> build and expression in the chassis, with host-level genome editing where the host itself limits performance -> screening under application-relevant conditions, feeding results back as training data -> bioprocess scale-up across TRL 1-6 (the consortium-funded step where most lab results fail) -> handover, with gains realised as higher substrate loading, lower catalyst load and shorter cycle time.
Relevance: SVC-001 is placed by the catalog under section 1.1 Crop biotech, so clusters: [“crop-biotech”] is correct as given even though the subject is an industrial-biotech service — cluster is given by the catalog, not chosen. Honest scope note: the CN region is written qualitatively beyond GenScript because no second domestic strain-engineering CRO was independently confirmed.