Analytics & PAT

Bioprocess data historians and analytics

Time-series compression, batch alignment and latent-variable statistics — why multivariate models see deviations that no single alarm limit can, and why they still cannot tell you the cause.

A stirred-tank fermentation is instrumented with dozens of channels — temperature, pH, dissolved oxygen, agitation, off-gas carbon dioxide, feed rate, pressure, weight — but it does not have dozens of independent behaviours. Nearly all of them are driven by the same few underlying things: how much biomass is present, how fast it is respiring, how the controller is responding. The measurements are therefore heavily collinear, and that redundancy is what makes historian data analysable at all.

The archive is lossy by design

A historian does not store every sample. It applies a deadband or swinging-door compression: a new point is written only when the signal departs from a linear interpolation of the stored points by more than a set tolerance. On a slow, well-controlled loop this discards enormous volumes of redundant data at no cost. But the tolerance is chosen when the tag is configured, often years before anyone needs the data, and a deadband wide enough to look tidy will erase exactly the small transient a later investigation is trying to reconstruct. Because the archive is also a GMP record, compression settings, timestamp source and clock synchronisation are not IT housekeeping — they determine whether the record satisfies data-integrity expectations under 21 CFR Part 11 and the ALCOA+ principles.

Why batch data need aligning before they can be compared

Batch records form a three-way array: batch by variable by time. Multivariate statistical process control unfolds that array and fits a latent-variable model — principal component or partial least squares regression on the batch trajectories, in the form set out by Nomikos and MacGregor in the mid-1990s.

The awkward step is the time axis. Batches do not last the same number of hours, so aligning on wall-clock time compares a batch in mid-exponential growth against another already in stationary phase. Alignment is done instead against a monotonic maturity indicator — cumulative oxygen uptake, integrated feed, or dynamic time warping onto a reference trajectory — and the choice of that indicator quietly determines what the model afterwards regards as normal.

Two statistics, and what each catches

A fitted model yields two monitoring quantities. Hotelling’s T² measures how far a batch has moved within the correlation structure the model learned — unusual, but of a familiar kind. The squared prediction error, the residual outside the model, measures departure from that structure: variables that have stopped moving together as they always did. The second is the one that earns the method its keep, because a fault can break the relationship between pH and oxygen uptake while both remain comfortably inside their individual alarm limits.

The limit that does not go away

These models are correlational and are trained on historical production, where variation is small precisely because control is good. A well-run process contains little excitation, so the model interpolates within a narrow envelope and extrapolates badly, and a flagged deviation identifies which variables contributed to the residual, not why. Contribution plots are a starting point for investigation, not a cause. Making causal claims requires designed variation that a commercial campaign is not permitted to introduce — which is why process understanding still comes from scale-down models, and the historian’s role is detection, not explanation.

Last updated: