# Biodiversity genomics

Sequencing and archiving the genomes of all eukaryotic life to support ecosystem restoration and bioprospecting.

Source: https://en.bioecon.ru/technology/biodiversity-genomics/
Updated: 2026-08-18



## Overview and value chain

Markers: [EC: Convention on Biological Diversity | OECD: genomics-bioinformatics | Regulator: USDA-APHIS, EFSA, MOA]

Biodiversity genomics aims to decode the genetic blueprints of the planet's entire eukaryotic biodiversity. The Earth BioGenome Project (EBP) — often called biology's Moon program — is the central effort: a 10-year consortium target to sequence, catalogue and openly publish chromosome-level reference genomes for all of the roughly 1.5-1.8 million eukaryotic species (plants, animals, fungi and protists). This archive underpins three application domains: bioprospecting (mining secondary-metabolite pathways to port into synthetic-biology chassis for drugs and green chemistry), conservation genetics (detecting inbreeding depression and resolving disputed taxonomy for genetic-rescue programs), and food security (transferring stress-tolerance genes from crop wild relatives into commercial varieties via genome editing). With sequencing costs dropping below $1,000 per high-quality reference genome, the focus has shifted to automated ultra-long DNA extraction and open, decentralized data sharing.

The key directions of biodiversity genomics are:
1. **Reference Genome Generation:** Producing chromosome-level assemblies for keystone and endangered species.
2. **Environmental DNA (eDNA) Monitoring:** Tracking biodiversity changes by sequencing DNA shed into soil and water.
3. **Bioprospecting:** Mining genomic databases for novel enzymes and secondary metabolites for industrial use.
4. **Conservation Genetics:** Assessing genetic diversity within isolated populations to guide breeding programs.

### Sectoral value chain

```
[Sample Collection] ──> [DNA Extraction] ──> [Sequencing] ──> [Bioinformatics Assembly]
                                  │
                          (Annotation)
                                  │
                                  ▼
[Ecosystem Monitoring] <─── [Open Data Archiving] <─────┘
```

### Value chain levels

| Level | Description | Key inputs/outputs |
|:---|:---|:---|
| **Sample Collection** | field collection of tissues under the Nagoya Protocol, with morphological voucher archiving. | **In:** flora, fauna, collection permits.<br>**Out:** flash-frozen voucher tissues. |
| **DNA Extraction** | isolating ultra-high-molecular-weight DNA without shearing. | **In:** frozen tissue, Nanobind magnetic disks.<br>**Out:** UHMW DNA above 100 kb. |
| **Sequencing** | long-read sequencing on PacBio HiFi and Oxford Nanopore platforms. | **In:** UHMW DNA.<br>**Out:** long reads above Q30. |
| **Bioinformatics Assembly** | reconstructing entire chromosomes de novo. | **In:** long reads, Hi-C contacts, Hifiasm/YaHS.<br>**Out:** chromosome-level assemblies. |
| **Open Data Archiving** | storing sequences in public databases. | **In:** assemblies.<br>**Out:** curated reference genomes (NCBI, ENA). |
| **Ecosystem Monitoring** | applying data to conservation policies. | **In:** genomes, eDNA.<br>**Out:** species recovery plans. |

Cross-cutting technologies of the sector:
- **Long-Read Sequencing:** Technologies like Nanopore and PacBio enabling highly continuous genome assemblies.
- **Spatial Transcriptomics:** Mapping gene expression directly within tissue sections to understand complex biology.
- **Cloud Computing:** Massive scalable infrastructure required to process petabytes of genomic data.

---

## US

The United States acts as the central administrative hub for global biodiversity initiatives, coordinating standards and data protocols for the Earth BioGenome Project.

### EBP coordination, High-throughput sequencing, Cloud infrastructure
- **EBP coordination hub:** UC Davis (Harris Lewin's working group) coordinates the global EBP network and its assembly-quality metrics (contig N50 above 1 Mb, base accuracy above Q40).
- **Vertebrate Genomes Project:** Erich Jarvis's group at Rockefeller University drives the VGP, targeting reference genomes for all roughly 70,000 vertebrate species with PacBio HiFi reads.
- **Smithsonian collections and NCBI:** the Smithsonian Institution anchors specimen and voucher curation for US ecosystem-sequencing initiatives, while NCBI archives the open genomic data.

---

## CN

China is heavily investing in massive sequencing infrastructure, leveraging automation to rapidly sequence thousands of Asian plant and animal species.

### Stereo-seq, 10KP Project, Mass automation
- **Spatial genomics:** BGI-Research pairs its proprietary Stereo-seq spatial-transcriptomics platform with whole-genome assembly of rare Asian plants.
- **Large-scale projects:** the 10,000 Plant Genomes (10KP) and companion eukaryote programmes sequence mosses, ferns, algae and endemic animals on BGI's DNBSEQ platforms; the Kunming Institute of Zoology holds the leading databases of Chinese wild-mammal evolutionary zoogeography.
- **National databases:** the China National GeneBank (CNGB) securely hosts the country's immense biodiversity datasets.

---

## EU

The European Union focuses on deep ecological genomics, sequencing entire ecosystems to understand the impacts of climate change and discover novel bioproducts.

### Darwin Tree of Life, Marine genomics, Data sovereignty
- **Darwin Tree of Life:** the Wellcome Sanger Institute-led DtL project aims to sequence all roughly 70,000 eukaryotic species of Britain and Ireland, running one of the world's highest-throughput de novo assembly pipelines.
- **Marine bioprospecting:** France's Genoscope sequences ocean microbiomes (notably the Tara Oceans expeditions), screening for extremophile enzymes useful to green chemistry.
- **Assembly algorithms and open data:** the Max Planck Institute (Myers lab) develops the overlap-layout graph mathematics behind modern assemblers, and EMBL-EBI hosts open European bioinformatics infrastructure.

---

## Leading companies and research institutes

| Company / Institute | Country | Key products / platforms | Tech features | Status 2026 |
|:---|:---|:---|:---|:---|
| **Wellcome Sanger** | 🇬🇧 UK | *Darwin Tree of Life* | High-volume de novo assembly | commercial |
| **BGI** | 🇨🇳 China | *Stereo-seq platforms* | Spatial transcriptomics | commercial |
| **UC Davis** | 🇺🇸 USA | *EBP Hub coordination* | Standardizing >Q40 metrics | commercial |
| **Max Planck Institute** | 🇩🇪 Germany | *Bioinformatics tools* | Complex polyploid assemblers | commercial |
| **Genoscope** | 🇫🇷 France | *Marine metagenomics* | Environmental screening | commercial |
| **Oxford Nanopore** | 🇬🇧 UK | *MinION/PromethION* | Ultra-long read sequencing | commercial |

---

## Tech stack and innovations

Modern biodiversity genomics depends entirely on breakthroughs in extracting ultra-long DNA molecules and algorithmic assembly of highly repetitive regions.

1. **UHMW DNA Extraction:**
   - Nanobind magnetic disks (a nanosilica-coated surface) gently wind ultra-long DNA under mild rocking, keeping fragments above 100 kb that long-read sequencing needs.
   - gentle, low-shear lysis preserves megabase-scale chromosomes that spin-column kits would fragment.
2. **Sequencing Technologies:**
   - PacBio HiFi and Oxford Nanopore generate multi-kilobase reads above Q30 that resolve repetitive regions short reads cannot.
   - Hi-C chromatin conformation capture fixes chromatin with formaldehyde and proximity-ligates physically close DNA, producing the contact map that scaffolds contigs into chromosomes.
3. **Bioinformatics Algorithms:**
   - Hifiasm builds accurate contigs from HiFi reads, Purge_dups removes haplotype duplicates, and YaHS orders contigs into chromosome-scale scaffolds using Hi-C contacts.
   - AI annotation pipelines such as Braker3 rapidly label genes, regulatory elements and biosynthetic gene clusters across novel species.

---

## Value chains and production pipelines

### Industrial pipeline of Reference Genome Generation (EBP Standards)

```
┌───────────────────────────┐      ┌───────────────────────────┐
│ 1. Tissue Sampling        │ ───> │ 2. UHMW DNA Extraction    │
└───────────────────────────┘      └───────────────────────────┘
                                                 │
                                                 ▼
┌───────────────────────────┐      ┌───────────────────────────┐
│ 4. Bioinformatics Assembly│ <─── │ 3. Long-Read Sequencing   │
└───────────────────────────┘      └───────────────────────────┘
              │
              ▼
┌───────────────────────────┐      ┌───────────────────────────┐
│ 5. Gene Annotation        │ ───> │ 6. Public Release         │
└───────────────────────────┘      └───────────────────────────┘
```

#### Stage 1: Tissue Sampling
Field biologists collect fresh tissue samples, instantly freezing them in liquid nitrogen to perfectly preserve cellular DNA integrity.

#### Stage 2: UHMW DNA Extraction
Laboratories utilize specialized magnetic nanodisks to gently isolate ultra-high molecular weight DNA without mechanical shearing.

#### Stage 3: Long-Read Sequencing
Extracted DNA is processed through Nanopore or PacBio systems, generating massive datasets of continuous, multi-kilobase reads.

#### Stage 4: Bioinformatics Assembly
Supercomputers run complex graph algorithms to piece the reads together, utilizing Hi-C data to scaffold the contigs into full chromosomes.

#### Stage 5: Gene Annotation
Automated AI pipelines scan the assembled chromosomes to identify genes, regulatory elements, and structural variants.

#### Stage 6: Public Release
The fully curated, high-quality reference genome is uploaded to central databases like GenBank or ENA for open scientific access.


