Biocomputing and DNA computing
01Overview and value chain
Markers: [EC: research-stage, no dedicated conformity route | OECD: 1.6 Genomics and bioinformatics | Regulator: FDA (US), EMA (EU)]
Biocomputing covers two distinct propositions that share a substrate. The first is storage: DNA holds data at densities no magnetic or optical medium approaches and remains readable for centuries, so the engineering problem is not capacity but access — writing cheaply, finding one file among billions, and recovering the data despite a synthesis and sequencing process that corrupts it. Published benchmarking of six representative codecs shows the coding layer can tolerate error rates up to roughly 14% and sequence loss up to roughly 65% in isolation, which is the margin that makes an inherently lossy medium usable at all. The second proposition is computation: genetic circuits inside living cells performing logic, with 2026 work demonstrating three-input combinatorial gates, half adders, full adders and a dynamic multiplexer built from orthogonal trans-splicing AND gates and synthetic microRNAs — arithmetic executed by a cell rather than a transistor. Around both sits a smaller unconventional-computing tradition that treats slime mould and fungal mycelium as computing substrates. The honest framing for all of it in 2026 is that this is research and patent activity, not a product market: the named participants are corporate research labs, universities and platform companies filing patents, and no regulatory conformity route exists for a biological computer.
The key directions of biocomputing are:
- DNA Data Storage and Random Access: encoding data in synthesised polynucleotides and retrieving specific files from a pool, with patent work on whole-pool amplification combined with in-sequencer random access so that access and read collapse into a single workflow step.
- Error-Correction Coding for DNA Storage: the codec layer that makes a lossy medium reliable, benchmarked across six representative schemes tolerating error rates to roughly 14% and sequence loss to roughly 65% in isolation, with realistic deployment densities lower than the isolated figures suggest.
- Genetic Circuit Computing: logic implemented inside mammalian cells using orthogonal trans-splicing AND gates, native-synthetic hybrid promoters and synthetic microRNAs for inhibitory logic, demonstrated in 2026 as three-input gates, half and full adders and a dynamic three-to-one multiplexer within few computational layers.
- Unconventional and Biohybrid Computing: living systems as computing and sensing substrates — Physarum polycephalum cellular-automata models driving robot navigation, and fungal mycelial networks proposed as distributed sensing, self-healing and anomaly-detection layers.
Sectoral value chain
[Encoding & Design] ──> [Molecular Write] ──> [Storage/Execution] ──> [Molecular Read]
│
(error correction)
│
▼
[Recovered Data / Output] <─── [Decoding] <─────────────┘Value chain levels
| Level | Description | Key inputs/outputs |
|---|---|---|
| Encoding and Design | Data or logic is mapped onto nucleotide sequence — codeword design, constraint handling and error-correction scheme chosen before anything is synthesised | In: Source data or logic specification, codec choice. Out: Constrained sequence design with redundancy. |
| Molecular Write | Sequences are synthesised, the cost and fidelity bottleneck of the whole field; high-throughput, low-cost, high-fidelity synthesis is the explicit research target | In: Sequence design, synthesis platform. Out: Physical DNA pool encoding the payload. |
| Storage or Execution | The pool is archived, or in the computing case the circuit is expressed in living cells where logic is evaluated by the cell’s own machinery | In: DNA pool or engineered cells, storage or culture conditions. Out: Stable archive, or executed biological computation. |
| Molecular Read | Retrieval by sequencing, with random access to locate the addressed subset rather than reading the whole pool; access and read can be combined in-sequencer | In: Archived pool, access primers or in-sequencer method. Out: Raw reads of the addressed data. |
| Error Correction and Decoding | The codec reconstructs the payload despite substitution, insertion, deletion and outright sequence loss — the layer that determines whether the medium is usable | In: Raw reads with errors and dropout. Out: Corrected codewords, recovered payload. |
| Recovered Data or Output | The original file is returned, or the cellular circuit’s logical output is read as a biological signal | In: Decoded data or circuit output. Out: Verified data, or an actuated cellular response. |
Cross-cutting technologies of the sector:
- Molecular Electronics Sensing: single-molecule electronic sensors using DNA reporter tags for multiplex genetic analysis, providing an electrical rather than optical readout of molecular events.
- High-Fidelity DNA Synthesis: the write step common to storage and circuit construction, where throughput, cost and fidelity jointly determine what is economically possible.
- Bio-Inspired and Cellular-Automata Algorithms: computational models drawn from organisms — notably Physarum polycephalum — used both to run biohybrid hardware and as algorithms in conventional robotics.
02US
The US contribution is split between a corporate research programme treating DNA as an archival storage tier and a platform company building molecular-electronic readout, with both visible primarily through patent filings rather than shipping products.
Microsoft DNA storage random access, Roswell molecular electronics, patent-stage maturity
- Microsoft (DNA data storage): a patent application by Microsoft Technology Licensing describes copying all polynucleotide-encoded data in a storage container while preserving random-access read capability, combining whole-pool amplification with in-sequencer random access so that access and read occur in a single workflow step — addressing the retrieval problem that separates a DNA archive from a DNA sample.
- Roswell Biotechnology (molecular electronics): the San Diego company’s filings cover molecular electronic sensors for multiplex genetic analysis using DNA reporter tags, and N-terminal multifunctional peptide conjugation with alpha-helical peptides for biosensing. The documents describe a technology platform rather than a finished product, with roots in provisional filings from 2020.
- Maturity, stated plainly: what is visible in 2026 sourcing for both is intellectual property and platform description, not revenue, deployment scale or funding disclosure — which is the accurate status of molecular computing and storage in the US commercial sector.
03CN
China’s activity is concentrated in university programmes that treat DNA information storage as an engineering pipeline to be optimised end to end, with recruitment aimed squarely at the synthesis and error-correction bottlenecks.
SJTU Bio-X DNA information storage, synthesis platform optimisation, encoding and error correction
- Shanghai Jiao Tong University (Bio-X Institute): the Shi Yongyong group recruited postdoctoral researchers in 2026 explicitly for DNA information storage and DNA synthesis platform optimisation, covering DNA encoding design, data writing, sequencing readout, error correction and information recovery as one connected workflow.
- Synthesis as the stated bottleneck: the same programme frames its objective as high-throughput, low-cost, high-fidelity DNA synthesis, with experimental design, process optimisation and performance evaluation named as the work — an accurate reflection that the write step, not the read step, currently limits the field.
- Disciplinary breadth of recruitment: candidates are sought from molecular biology, synthetic biology, chemical biology, biochemistry, nucleic-acid chemistry, bioengineering, bioinformatics, computational biology, information storage and sequencing technology, with DNA information storage and coding/error-correction algorithms named as preferred backgrounds — the interdisciplinary profile the field requires.
04EU
Europe supplies the theoretical layer that makes DNA storage trustworthy and the longest-running unconventional-computing tradition, both academic rather than commercial.
ETH Zurich codec benchmarking, UWE Bristol unconventional computing, biohybrid fungal systems
- ETH Zurich (error-correction benchmarking): a March 2026 Nature Communications study systematically benchmarked six representative codecs for sequence-based DNA data storage under both in-silico and in-vitro conditions, finding that in isolation codecs tolerate error rates up to roughly 14% and sequence loss up to roughly 65%, while realistic deployment implies lower usable storage densities than the isolated figures suggest.
- UWE Bristol (Unconventional Computing Laboratory): researchers including Genaro J. Martínez and Andrew Adamatzky built a vigilance robot, McIntosh I, whose navigation is driven by a bio-inspired algorithm based on a Physarum polycephalum complex cellular-automata model, with collaborators from Mexico’s Artificial Life and Robotics Lab and the Czech Unconventional Algorithms and Computation Lab.
- Fungal biohybrid systems: work from the same tradition proposes living fungal mycelial networks as biohybrid substrates for security and resilience — distributed sensing, self-healing materials and low-observability anomaly detection, leveraging decentralised control, embodied memory and autonomous repair for infrastructure protection and environmental monitoring.
05Leading companies and research institutes
| Company / Institute | Country | Key products / platforms | Tech features | Status 2026 |
|---|---|---|---|---|
| Shanghai Jiao Tong University | 🇨🇳 China | Bio-X DNA information storage programme | End-to-end pipeline: encoding design, data writing, sequencing readout, error correction and recovery; targets high-throughput, low-cost, high-fidelity synthesis | research; active 2026 recruitment |
| ETH Zurich | 🇨🇭 Switzerland | DNA storage error-correction benchmarking | Systematic comparison of six codecs in silico and in vitro; ~14% error tolerance, ~65% sequence loss in isolation; realistic densities lower | research; Nature Communications, March 2026 |
| UWE Bristol | 🇬🇧 United Kingdom | Unconventional Computing Laboratory, McIntosh I | Physarum polycephalum cellular-automata navigation; fungal mycelial networks as distributed sensing and self-healing substrates | research; international collaborations |
| Roswell Biotechnology | 🇺🇸 USA | Molecular electronic sensors | Single-molecule electronic sensing with DNA reporter tags for multiplex genetic analysis; N-terminal peptide conjugation biosensing | platform stage; patent filings, no product launch stated |
| Microsoft | 🇺🇸 USA | DNA data storage random access | Whole-pool amplification combined with in-sequencer random access, collapsing access and read into one step | research; patent application |
06Tech stack and innovations
The stack is best read as two columns — storage and computation — resting on one shared physical layer, molecular synthesis and sequencing, with an unconventional-computing tradition running alongside.
- DNA Storage: Write, Address, Read:
- Data is encoded into constrained nucleotide sequences and synthesised; the write step is the cost and fidelity bottleneck, which is why the explicit research objective is high-throughput, low-cost, high-fidelity synthesis rather than higher density.
- Retrieval is the differentiator between an archive and a sample: whole-pool amplification combined with in-sequencer random access preserves the ability to read a specific addressed subset while copying the pool, collapsing two operations into one.
- Error-Correction Coding:
- Benchmarking of six representative codecs under in-silico and in-vitro conditions shows tolerance of error rates up to roughly 14% and sequence loss up to roughly 65% when each is assessed in isolation.
- The same work is explicit that realistic deployment implies lower usable storage densities than isolated codec performance suggests — the gap between a benchmark and a working archive.
- Genetic Circuit Computing:
- Mammalian gene circuits built from orthogonal trans-splicing AND gates, native-synthetic hybrid promoters and synthetic microRNAs implementing inhibitory logic operate within few computational layers, which is what makes them scalable.
- Demonstrated circuits in 2026 include a three-input combinatorial logic gate, a half adder, a full adder and a dynamic three-to-one multiplexer — the same primitives as digital electronics, evaluated by cellular machinery.
- Molecular Electronics and Biohybrid Substrates:
- Molecular electronic sensors read single-molecule events electrically using DNA reporter tags for multiplex genetic analysis, bypassing optical detection entirely.
- Biohybrid computing takes the organism as the machine: a Physarum polycephalum cellular-automata model drives robot navigation, and fungal mycelial networks are proposed as distributed sensing and self-healing layers with embodied memory and autonomous repair.
07Value chains and production pipelines
Industrial pipeline of a DNA information-storage campaign (research stage; no dedicated conformity route, FDA/EMA relevant only where biological outputs are clinical)
┌───────────────────────────┐ ┌───────────────────────────┐
│ 1. Codec & Sequence │ ───> │ 2. DNA Synthesis (Write) │
│ Design │ │ │
└───────────────────────────┘ └───────────────────────────┘
│
▼
┌───────────────────────────┐ ┌───────────────────────────┐
│ 4. Random Access │ <─── │ 3. Pool Storage │
│ Retrieval │ │ │
└───────────────────────────┘ └───────────────────────────┘
│
▼
┌───────────────────────────┐ ┌───────────────────────────┐
│ 5. Sequencing Readout │ ───> │ 6. Error Correction & │
│ │ │ Recovery │
└───────────────────────────┘ └───────────────────────────┘Stage 1: Codec and Sequence Design
The payload is mapped onto nucleotide sequences under biochemical constraints, and the error-correction scheme is selected before synthesis, because redundancy has to be designed into the codewords rather than added afterwards. The choice is consequential: benchmarked codecs differ in how much substitution error and outright sequence loss they survive.
Stage 2: DNA Synthesis (Write)
Designed sequences are synthesised into a physical pool. This is the field’s binding constraint — throughput, cost and fidelity together — and it is the stated target of dedicated research programmes rather than a solved commodity step.
Stage 3: Pool Storage
The synthesised pool is archived, where DNA’s density and longevity are the advantages that motivate the whole approach. In the computing branch, the equivalent stage is expression of the circuit in living cells, where the logic is evaluated by cellular machinery instead of being stored.
Stage 4: Random Access Retrieval
Locating one file among many without reading everything is what makes an archive practical. Patent work combines whole-pool amplification with in-sequencer random access, preserving addressed retrieval while the pool is copied and merging access with the read step.
Stage 5: Sequencing Readout
The addressed subset is sequenced, producing reads carrying substitutions, insertions, deletions and dropout. The raw readout is expected to be corrupted; the design assumption of the whole pipeline is that the medium is lossy and the coding layer compensates.
Stage 6: Error Correction and Recovery
The codec reconstructs the payload, tolerating error rates to roughly 14% and sequence loss to roughly 65% in isolated benchmarks. Realistic end-to-end deployment yields lower usable densities than those isolated figures imply, which is the honest measure of where DNA storage stands in 2026.
| Supplier | Price | Lead time | Certificates | Risk | Confidence |
|---|---|---|---|---|---|
| Microsoft Research | custom | null | Medium | HIGH | |
| Shanghai Jiao Tong University | custom | null | Medium | HIGH | |
| ETH Zurich | custom | null | Medium | HIGH | |
| UWE Bristol | custom | null | High | HIGH | |
| Roswell Biotechnology | custom | null | High | HIGH |
AI note: biocomputing-dna-computing (EN)
Key directions:
- DNA data storage and random access — Microsoft Technology Licensing patent on whole-pool amplification combined with in-sequencer random access, so copying the pool preserves addressed retrieval and access merges with the read step. The problem in DNA storage is access, not capacity.
- Error-correction coding — a March 2026 Nature Communications study (ETH Zurich) benchmarked six representative codecs in silico and in vitro: in ISOLATION they tolerate error rates up to ~14% and sequence loss up to ~65%, but the same work states realistic deployment implies LOWER usable densities than the isolated figures suggest. Quote both halves; the caveat is the finding.
- Genetic circuit computing — 2026 Nature Communications work on modular mammalian gene circuits operating within few computational layers: orthogonal trans-splicing AND gates, native-synthetic hybrid promoters, synthetic microRNAs for inhibitory logic; demonstrated three-input combinatorial gate, half adder, full adder, dynamic 3-to-1 multiplexer. NOT attributed to a tabled institution — see below.
- Unconventional and biohybrid computing — UWE Bristol’s Unconventional Computing Lab (Genaro J. Martínez, Andrew Adamatzky) built the McIntosh I vigilance robot navigating on a Physarum polycephalum complex cellular-automata model, with Mexican and Czech collaborators; separately, fungal mycelial networks proposed as biohybrid distributed sensing, self-healing and low-observability anomaly detection.
Regulatory: there is NO conformity route for a biological computer. FDA/EMA are carried in front matter and are relevant only where biological outputs become clinical. The accurate 2026 status across all three regions is research and patent activity, not a product market — say so rather than implying commercial maturity.
Companies not in table: MIT was DROPPED. Its original 5/5 relevance score was an artefact of substring matching (“MIT” inside “transmitted”/“limited”/“submitted”); after the word-boundary fix (bf65cbc4) it scored 1/5, and that single match is a Nature Biomedical Engineering paper on drug-responsive self-amplifying RNA — off-topic for biocomputing. The gene-circuit result in direction 3 is therefore cited as a 2026 field finding WITHOUT institutional attribution, because the retrieved sources do not establish whose work it is. Do not re-add MIT without a source that names it on DNA/cellular computing.
Processing note: codec and sequence design under biochemical constraints (redundancy designed into codewords, not added later) -> DNA synthesis, the binding constraint on throughput, cost and fidelity -> pool storage, or expression of the circuit in living cells for the computing branch -> random-access retrieval, what separates an archive from a sample -> sequencing readout, expected to be corrupted by substitution, insertion, deletion and dropout -> error correction and recovery, where the ~14%/~65% tolerances apply and where realistic densities fall below benchmark figures.
Relevance: INT-010 sits in bioinformatics-omics. Honest scope caveats stated in all three languages: Microsoft and Roswell are visible only as patents and platform descriptions with no revenue, deployment scale or funding disclosed, and the article says so; ETH Zurich and UWE Bristol are research groups rather than vendors, and SJTU is tabled on a 2026 recruitment programme rather than a product.