# Biological AI & bioinformation services

The AI, machine-learning and bioinformatics software-as-a-service layer — cloud platforms, generative biological foundation models and managed analysis pipelines — sold to pharma, biotech and industrial-biology teams rather than run as in-house drug pipelines.

Source: https://en.bioecon.ru/technology/biological-ai-bioinformation-services/
Updated: 2026-08-18



## Overview and value chain

Markers: [EC: FDA AI/ML in Drug Development + EMA AI reflection paper | OECD: Bioinformatics & genomics / Bio-pharmaceuticals | Regulator: FDA (USA), EMA (EU), NMPA (China)]

Biological AI and bioinformation services is the compute-and-software layer
that turns the torrents of sequence, structure, image and omics data generated
by modern biology into discoverable, computable assets — delivered as cloud
platforms, foundation models and managed analysis pipelines rather than as
in-house drug programmes. Where the sibling Industry of AI drug design covers
firms running their own clinical pipelines (Isomorphic Labs, Recursion,
Insilico Medicine), this projection covers the platform and infrastructure
providers that sell those capabilities as a service to pharma, biotech,
academic and increasingly industrial-biology teams. NVIDIA's BioNeMo Agent
Toolkit, announced on 23 June 2026, is being adopted by Dassault Systèmes,
Eli Lilly, OpenAI, Schrödinger and Snowflake, among dozens of others; the
17-year multiomics veteran DNAnexus reports 90% customer satisfaction as it
reframes its managed-data cloud as "AI fuel"; and AWS HealthOmics can now
submit up to 100,000 bioinformatics workflow runs in a single API call. The
segment is shaped by four forces: hyperscaler bio-clouds, generative
biological foundation models, federated multi-modal data networks, and
protein- and molecular-design-as-a-service.

The key directions of biological AI and bioinformation services are:
1. **Hyperscaler bio-clouds (Cloud bioinformatics):** fully-managed,
   GPU-accelerated analysis services — AWS HealthOmics runs up to 100,000
   workflows per API call and is HIPAA-eligible; NVIDIA BioNeMo serves
   biological models through NIM microservices.
2. **Generative biological foundation models (Bio foundation models):**
   large models trained on molecules and proteins and served as callable
   services — NVIDIA GenMol v2.0 is an 89M-parameter masked diffusion model
   for fragment-based molecule generation, commercial-ready within the
   BioNeMo NIM family.
3. **Managed omics-data platforms (Omics data platforms):** curated,
   harmonized, ML-ready multi-omics — DNAnexus, QIAGEN Digital Insights and
   Elucidata's Polly make fragmented biomedical data analysis-ready.
4. **Design-as-a-service (Protein & molecular design):** generative protein
   and small-molecule design sold to others — Cradle raised USD 73 million in
   November 2024 to scale its protein-engineering platform; Dassault BIOVIA's
   MARIE companion is built on the BioNeMo Agent Toolkit.

### Sectoral value chain

```
[bio/omics data] ──> [data platform / FAIR curation] ──> [foundation models / analysis]
                                    │                                │
                          (federated, governed)          (GPU-accelerated inference)
                                    │                                │
                                    ▼                                ▼
                       [cloud / HPC infrastructure] <─── [agentic workflow orchestration]
                                                                    │
                                                                    ▼
                                                    [insights / candidates delivered to R&D]
```

### Value chain levels

| Level | Description | Key inputs/outputs |
|:---|:---|:---|
| **Data generation** | sequencers, mass spectrometers and high-content imagers produce raw biological signal | **In:** biological samples, instruments. **Out:** raw reads, spectra, images. |
| **Data platforms & curation** | ingest, harmonize and make data FAIR and ML-ready | **In:** raw reads, metadata. **Out:** structured, queryable datasets. |
| **Foundation models & algorithms** | generative and discriminative biological models deliver predictions | **In:** structured data, model weights. **Out:** structure, function and hit predictions. |
| **Cloud & HPC infrastructure** | elastic, GPU-accelerated compute hosts the models and data | **In:** models, datasets. **Out:** scalable compute, accelerated inference. |
| **Agentic orchestration** | LLM agents chain tools into multi-step scientific computations | **In:** scientist query, tool catalog. **Out:** executed pipeline, traceable report. |
| **Decision support & delivery** | deliver validated insights and candidates to R&D under governance | **In:** computed results, wet-lab feedback. **Out:** targets, candidates, designs. |

Cross-cutting technologies of the sector:
- **GPU-accelerated inference (NIM):** biological models packaged as callable, containerized microservices (BioNeMo NIM) for hosted or local deployment.
- **Federated data networks:** secure, distributed analysis that moves algorithms to governed data rather than pooling it (Velsera Global Data Network).
- **Foundation models for biology:** masked-diffusion and language models over molecular and protein representations (GenMol, protein language models).

---

## US

The United States anchors the service layer, led by the GPU and hyperscaler platforms that the rest of the field builds on, plus a mature managed-data cloud segment.

### BioNeMo & GPU model services, managed omics cloud, hyperscaler bio-cloud
- **NVIDIA BioNeMo:** the open developer platform for AI-driven life-science research packages structure prediction, molecular generation, docking, sequence analysis and genomics as callable BioNeMo Skills; its GenMol v2.0 (NV-GenMol-89M-v2) masked-diffusion model is commercial-ready, and the BioNeMo Agent Toolkit announced on 23 June 2026 is being adopted by Dassault Systèmes, Eli Lilly, OpenAI, Schrödinger and Snowflake, with Anthropic and OpenAI integrating it.
- **DNAnexus:** a 17-year multiomics pioneer positioning its managed-data cloud as "AI fuel" for precision health; under CEO Thomas Laur it pushes a federated "move algorithms to data" model and reported 90% customer satisfaction in July 2026 amid global expansion.
- **AWS HealthOmics (Amazon):** a fully-managed, HIPAA-eligible bioinformatics service that now submits up to 100,000 workflow runs in a single API call; VPC-connected workflows (March 2026) and ephemeral scratch storage (June 2026) cut cost and speed sequence alignment, BAM sorting and variant calling.
- **Velsera:** formed in 2023 from the Seven Bridges, PierianDx and UgenTec combination, its federated Global Data Network spans 175 million-plus patients and powers the knowledge base behind Illumina's TruSight Oncology Comprehensive assay (FDA-cleared 2024, Japan MHLW-approved 2025).

---

## CN

China's presence in this Industry is felt less through standalone service vendors than through state-coordinated AI-for-science platforms and a fast-maturing regulatory frame; no dedicated commercial service vendor cleared live source confirmation for this projection, so the market is described qualitatively.

### AI-for-science platforms, sovereign bio-clouds, AI+drug regulation roadmap
- **AI-for-science platforms:** domestic "AI for science" players such as DP Technology (深势科技) and the BGI cloud stack extend molecular and materials foundation models (Uni-Mol-class) to academic and industrial users, though company-specific 2026 commercial facts were not independently confirmed.
- **Sovereign bio-clouds:** national genomics and bioinformatics clouds built on the BGI and CAS ecosystems prioritise data sovereignty and feed domestic drug-discovery and breeding programmes.
- **AI+drug regulation roadmap:** the Ministry of Science and Technology's 2026 "AI+Drug" regulation roadmap, reported in April 2026, sets out a convergence path for AI-assisted discovery under NMPA oversight, paralleling the FDA and EMA AI/ML frameworks.

---

## EU

Europe's contribution centres on established bioinformatics software houses and a distinctive protein-design-as-a-service cluster, increasingly built on top of US GPU platforms.

### Bioinformatics software SaaS, protein-design-as-a-service, agentic modeling
- **QIAGEN Digital Insights (Germany):** the bioinformatics software portfolio spans QIAGEN CLC Genomics Workbench, QCI cloud-based secondary analysis, Ingenuity Pathway Analysis (IPA) and OmicSoft, covering secondary analysis, single-cell and microbial/metagenomic workflows.
- **Cradle (Netherlands):** the Amsterdam protein-engineering company raised USD 73 million in November 2024 to build out its generative AI platform and wet lab, with a February 2026 follow-on funding round; its CRADLE-1 model demonstrates automated lead optimization of proteins and it serves both drug-discovery and industrial-biotechnology customers.
- **Dassault Systèmes BIOVIA (France):** the 3DEXPERIENCE life-sciences brand launched the MARIE agentic AI companion for drug discovery on the NVIDIA BioNeMo Agent Toolkit (23 June 2026), building on a 30-plus-year Discovery Studio modeling lineage that now combines physics-based methods with generative AI.

---

## Leading companies and research institutes

| Company / Institute | Country | Key products / platforms | Tech features | Status 2026 |
|:---|:---|:---|:---|:---|
| **NVIDIA** | 🇺🇸 USA | *BioNeMo / NIM model services* | GenMol v2.0 89M-param diffusion model; Agent Toolkit (Jun 2026); Lilly, Schrödinger, OpenAI adopting | commercial |
| **DNAnexus** | 🇺🇸 USA | *Managed multiomics data cloud* | 17-year pioneer; 90% customer satisfaction (Jul 2026); federated data model | commercial |
| **Amazon** | 🇺🇸 USA | *AWS HealthOmics managed cloud* | Up to 100,000 workflow runs per API call; VPC workflows; HIPAA-eligible | commercial |
| **QIAGEN** | 🇩🇪 Germany | *Digital Insights: CLC / IPA / QCI cloud* | Secondary analysis, IPA, OmicSoft; single-cell & metagenomic modules | commercial |
| **Cradle** | 🇳🇱 Netherlands | *Protein-design-as-a-service* | USD 73M Series B (Nov 2024); CRADLE-1 automated optimization; industrial biotech | operating |
| **Elucidata** | 🇮🇳 India | *Polly MLOps / BioAgent* | Founded 2015; BioAgent compresses a 6-hour omics workflow to 20 minutes | operating |

---

## Tech stack and innovations

The stack rests on three layers: GPU-served foundation models that turn biological representations into predictions, managed omics-data clouds that make fragmented data computable, and design services that generate new molecules and proteins on demand.

1. **Foundation models & NIM serving:**
   - NVIDIA GenMol v2.0 (NV-GenMol-89M-v2) is an 89M-parameter masked diffusion model trained on SAFE fragment representations for fragment-based molecule generation, hit generation and lead optimization, served as a commercial BioNeMo NIM.
   - The BioNeMo Agent Toolkit (23 June 2026) wraps structure prediction, molecular generation, docking, sequence analysis and genomics as agentic Skills accessible through hosted and local NIM, with Anthropic and OpenAI integrating.
2. **Managed omics-data clouds:**
   - AWS HealthOmics runs up to 100,000 bioinformatics workflow runs from a single API call, with VPC-connected workflows (March 2026) and ephemeral scratch storage (June 2026) for faster alignment, BAM sorting and variant calling.
   - DNAnexus applies a federated "move algorithms to data" model across multiomics, and Elucidata's BioAgent compresses a 6-hour omics workflow into roughly 20 minutes on its Polly platform.
3. **Protein & molecular design-as-a-service:**
   - Cradle's generative platform (USD 73 million Series B, November 2024) delivers automated protein lead optimization (CRADLE-1) to both pharma and industrial-biotechnology teams.
   - Dassault BIOVIA's MARIE agentic companion, built on the BioNeMo Agent Toolkit, shortens the path from a scientist's question to a correctly executed computation on the 3DEXPERIENCE platform.

---

## Value chains and production pipelines

### Industrial pipeline of an AI-augmented bioinformatics service engagement (FAIR data · ISO 9001 · 21 CFR Part 11)

```
┌───────────────────────────┐      ┌───────────────────────────┐
│ 1. Data ingestion          │ ───> │ 2. Harmonization &        │
└───────────────────────────┘      │    ML-readiness            │
                                   └───────────────────────────┘
                                                  │
                                                  ▼
┌───────────────────────────┐      ┌───────────────────────────┐
│ 4. Agentic orchestration  │ <─── │ 3. Foundation-model        │
└───────────────────────────┘      │    inference               │
                                   └───────────────────────────┘
               │
               ▼
┌───────────────────────────┐      ┌───────────────────────────┐
│ 5. Human-in-the-loop      │ ───> │ 6. Delivery & regulatory   │
│    validation             │      │    traceability            │
└───────────────────────────┘      └───────────────────────────┘
```

#### Stage 1: Data ingestion and governance
Raw sequencing reads, mass-spectrometry spectra and high-content images are ingested into a governed store under FAIR and 21 CFR Part 11 data-integrity controls; a managed cloud such as AWS HealthOmics or DNAnexus provides the HIPAA-eligible landing zone.

#### Stage 2: Harmonization and ML-readiness
The platform curates metadata, harmonizes diverse assays and converts fragmented biomedical data into analysis-ready, ML-ready tensors; Elucidata's Polly applies a tech-enabled curation engine to multi-omics and assay data at this stage.

#### Stage 3: Foundation-model inference
Generative and discriminative biological models deliver predictions — NVIDIA GenMol for fragment-based molecule generation, protein language models for function, and AlphaFold-class structure services — served as BioNeMo NIM microservices for hosted or local use.

#### Stage 4: Agentic orchestration
LLM agents chain the model services and data tools into multi-step scientific computations; the BioNeMo Agent Toolkit and BIOVIA MARIE turn a scientist's natural-language question into a correctly executed, traceable pipeline.

#### Stage 5: Human-in-the-loop validation
Computed candidates and designs are reviewed against wet-lab feedback and domain constraints before release, with Velsera's clinicogenomic knowledge base and Illumina TSO interpretation supplying the real-world evidence layer.

#### Stage 6: Delivery and regulatory traceability
Validated targets, candidates and protein designs are delivered to the sponsor's R&D pipeline with full provenance for FDA, EMA or NMPA submission, closing the loop from raw biological data to a decision-ready asset.

