Biodiversity genomics
01Overview and value chain
Markers: [EC: Convention on Biological Diversity | OECD: genomics-bioinformatics | Regulator: USDA-APHIS, EFSA, MOA]
Biodiversity genomics aims to decode the genetic blueprints of the planet’s entire eukaryotic biodiversity. The Earth BioGenome Project (EBP) — often called biology’s Moon program — is the central effort: a 10-year consortium target to sequence, catalogue and openly publish chromosome-level reference genomes for all of the roughly 1.5-1.8 million eukaryotic species (plants, animals, fungi and protists). This archive underpins three application domains: bioprospecting (mining secondary-metabolite pathways to port into synthetic-biology chassis for drugs and green chemistry), conservation genetics (detecting inbreeding depression and resolving disputed taxonomy for genetic-rescue programs), and food security (transferring stress-tolerance genes from crop wild relatives into commercial varieties via genome editing). With sequencing costs dropping below $1,000 per high-quality reference genome, the focus has shifted to automated ultra-long DNA extraction and open, decentralized data sharing.
The key directions of biodiversity genomics are:
- Reference Genome Generation: Producing chromosome-level assemblies for keystone and endangered species.
- Environmental DNA (eDNA) Monitoring: Tracking biodiversity changes by sequencing DNA shed into soil and water.
- Bioprospecting: Mining genomic databases for novel enzymes and secondary metabolites for industrial use.
- Conservation Genetics: Assessing genetic diversity within isolated populations to guide breeding programs.
Sectoral value chain
[Sample Collection] ──> [DNA Extraction] ──> [Sequencing] ──> [Bioinformatics Assembly]
│
(Annotation)
│
▼
[Ecosystem Monitoring] <─── [Open Data Archiving] <─────┘Value chain levels
| Level | Description | Key inputs/outputs |
|---|---|---|
| Sample Collection | field collection of tissues under the Nagoya Protocol, with morphological voucher archiving. | In: flora, fauna, collection permits. Out: flash-frozen voucher tissues. |
| DNA Extraction | isolating ultra-high-molecular-weight DNA without shearing. | In: frozen tissue, Nanobind magnetic disks. Out: UHMW DNA above 100 kb. |
| Sequencing | long-read sequencing on PacBio HiFi and Oxford Nanopore platforms. | In: UHMW DNA. Out: long reads above Q30. |
| Bioinformatics Assembly | reconstructing entire chromosomes de novo. | In: long reads, Hi-C contacts, Hifiasm/YaHS. Out: chromosome-level assemblies. |
| Open Data Archiving | storing sequences in public databases. | In: assemblies. Out: curated reference genomes (NCBI, ENA). |
| Ecosystem Monitoring | applying data to conservation policies. | In: genomes, eDNA. Out: species recovery plans. |
Cross-cutting technologies of the sector:
- Long-Read Sequencing: Technologies like Nanopore and PacBio enabling highly continuous genome assemblies.
- Spatial Transcriptomics: Mapping gene expression directly within tissue sections to understand complex biology.
- Cloud Computing: Massive scalable infrastructure required to process petabytes of genomic data.
02US
The United States acts as the central administrative hub for global biodiversity initiatives, coordinating standards and data protocols for the Earth BioGenome Project.
EBP coordination, High-throughput sequencing, Cloud infrastructure
- EBP coordination hub: UC Davis (Harris Lewin’s working group) coordinates the global EBP network and its assembly-quality metrics (contig N50 above 1 Mb, base accuracy above Q40).
- Vertebrate Genomes Project: Erich Jarvis’s group at Rockefeller University drives the VGP, targeting reference genomes for all roughly 70,000 vertebrate species with PacBio HiFi reads.
- Smithsonian collections and NCBI: the Smithsonian Institution anchors specimen and voucher curation for US ecosystem-sequencing initiatives, while NCBI archives the open genomic data.
03CN
China is heavily investing in massive sequencing infrastructure, leveraging automation to rapidly sequence thousands of Asian plant and animal species.
Stereo-seq, 10KP Project, Mass automation
- Spatial genomics: BGI-Research pairs its proprietary Stereo-seq spatial-transcriptomics platform with whole-genome assembly of rare Asian plants.
- Large-scale projects: the 10,000 Plant Genomes (10KP) and companion eukaryote programmes sequence mosses, ferns, algae and endemic animals on BGI’s DNBSEQ platforms; the Kunming Institute of Zoology holds the leading databases of Chinese wild-mammal evolutionary zoogeography.
- National databases: the China National GeneBank (CNGB) securely hosts the country’s immense biodiversity datasets.
04EU
The European Union focuses on deep ecological genomics, sequencing entire ecosystems to understand the impacts of climate change and discover novel bioproducts.
Darwin Tree of Life, Marine genomics, Data sovereignty
- Darwin Tree of Life: the Wellcome Sanger Institute-led DtL project aims to sequence all roughly 70,000 eukaryotic species of Britain and Ireland, running one of the world’s highest-throughput de novo assembly pipelines.
- Marine bioprospecting: France’s Genoscope sequences ocean microbiomes (notably the Tara Oceans expeditions), screening for extremophile enzymes useful to green chemistry.
- Assembly algorithms and open data: the Max Planck Institute (Myers lab) develops the overlap-layout graph mathematics behind modern assemblers, and EMBL-EBI hosts open European bioinformatics infrastructure.
05Leading companies and research institutes
| Company / Institute | Country | Key products / platforms | Tech features | Status 2026 |
|---|---|---|---|---|
| Wellcome Sanger | 🇬🇧 UK | Darwin Tree of Life | High-volume de novo assembly | commercial |
| BGI | 🇨🇳 China | Stereo-seq platforms | Spatial transcriptomics | commercial |
| UC Davis | 🇺🇸 USA | EBP Hub coordination | Standardizing >Q40 metrics | commercial |
| Max Planck Institute | 🇩🇪 Germany | Bioinformatics tools | Complex polyploid assemblers | commercial |
| Genoscope | 🇫🇷 France | Marine metagenomics | Environmental screening | commercial |
| Oxford Nanopore | 🇬🇧 UK | MinION/PromethION | Ultra-long read sequencing | commercial |
06Tech stack and innovations
Modern biodiversity genomics depends entirely on breakthroughs in extracting ultra-long DNA molecules and algorithmic assembly of highly repetitive regions.
- UHMW DNA Extraction:
- Nanobind magnetic disks (a nanosilica-coated surface) gently wind ultra-long DNA under mild rocking, keeping fragments above 100 kb that long-read sequencing needs.
- gentle, low-shear lysis preserves megabase-scale chromosomes that spin-column kits would fragment.
- Sequencing Technologies:
- PacBio HiFi and Oxford Nanopore generate multi-kilobase reads above Q30 that resolve repetitive regions short reads cannot.
- Hi-C chromatin conformation capture fixes chromatin with formaldehyde and proximity-ligates physically close DNA, producing the contact map that scaffolds contigs into chromosomes.
- Bioinformatics Algorithms:
- Hifiasm builds accurate contigs from HiFi reads, Purge_dups removes haplotype duplicates, and YaHS orders contigs into chromosome-scale scaffolds using Hi-C contacts.
- AI annotation pipelines such as Braker3 rapidly label genes, regulatory elements and biosynthetic gene clusters across novel species.
07Value chains and production pipelines
Industrial pipeline of Reference Genome Generation (EBP Standards)
┌───────────────────────────┐ ┌───────────────────────────┐
│ 1. Tissue Sampling │ ───> │ 2. UHMW DNA Extraction │
└───────────────────────────┘ └───────────────────────────┘
│
▼
┌───────────────────────────┐ ┌───────────────────────────┐
│ 4. Bioinformatics Assembly│ <─── │ 3. Long-Read Sequencing │
└───────────────────────────┘ └───────────────────────────┘
│
▼
┌───────────────────────────┐ ┌───────────────────────────┐
│ 5. Gene Annotation │ ───> │ 6. Public Release │
└───────────────────────────┘ └───────────────────────────┘Stage 1: Tissue Sampling
Field biologists collect fresh tissue samples, instantly freezing them in liquid nitrogen to perfectly preserve cellular DNA integrity.
Stage 2: UHMW DNA Extraction
Laboratories utilize specialized magnetic nanodisks to gently isolate ultra-high molecular weight DNA without mechanical shearing.
Stage 3: Long-Read Sequencing
Extracted DNA is processed through Nanopore or PacBio systems, generating massive datasets of continuous, multi-kilobase reads.
Stage 4: Bioinformatics Assembly
Supercomputers run complex graph algorithms to piece the reads together, utilizing Hi-C data to scaffold the contigs into full chromosomes.
Stage 5: Gene Annotation
Automated AI pipelines scan the assembled chromosomes to identify genes, regulatory elements, and structural variants.
Stage 6: Public Release
The fully curated, high-quality reference genome is uploaded to central databases like GenBank or ENA for open scientific access.
| Supplier | Price | Lead time | Certificates | Risk | Confidence |
|---|---|---|---|---|---|
| Wellcome Sanger | custom | custom | eu | Low | HIGH |
| BGI | custom | custom | cn | Low | HIGH |