Monitoring & conservation
Passive acoustic monitoring
Why passive acoustic monitoring works: species-specific call structure and spectrogram classification, how spreading loss and masking set the detection range, and why abundance estimation from sound alone is genuinely hard.
Many animals announce themselves. Their calls are loud, repeatable and species-specific, and they travel through air, water and darkness that defeat visual surveys. A weatherproof recorder left in a canopy or a wetland collects that advertisement for months without an observer present, and a classifier turns the archive into detections. The physics that makes this work, however, is also what limits it.
Why a call identifies its maker
Vocal signals evolve under selection for recognition by the right audience at the right distance, so each species transmits in a characteristic combination of frequency band, bandwidth, modulation and rhythm; bats echolocate in ultrasonic bands shaped by prey size, and songbirds occupy bands that usually sit clear of the loudest background noise. The measurable object is the spectrogram — energy distributed over frequency and time. Species identification is pattern recognition on that object: template matching and hand-built features first, neural networks trained on large labelled archives now. The classifiers work where the discriminative structure is real and stable within a species, and fail where it is not: dialects that shift between regions, individual variation, calls degraded by distance, and species that were never in the training set.
Propagation and noise set the range
Sound loses energy with distance through spherical spreading — intensity falls with the square of range — and, especially at high frequency, through absorption by air or water. Wind, rain, surf, traffic and the choruses of other animals add noise in overlapping bands, and a receiver hears only what survives both losses: the signal-to-noise ratio, not the animal, decides whether a call is recordable. High-frequency signals pay twice, which is why bat detectors work at short range and why the effective survey radius of a recorder is a property of the soundscape, not of the microphone. Duty cycle, storage and power add a temporal sampling problem: a recorder that listens for seconds every few minutes misses species whose calling bouts are brief or tied to particular weather.
Why counting animals from audio is hard
A detection answers “this species called here” — occupancy, not a census. Three gaps separate the two. Call rate varies with season, time of night, breeding state and motivation, so detections per hour cannot be read as animals per hectare. Distance is unknown to a single microphone: the same call at ten and at a hundred metres differs mainly in amplitude, which the channel degrades unpredictably, so without localisation or a calibrated attenuation model the sampled area is undefined. And individuals are indistinguishable — twenty birds singing together are one mass of overlapping signal. The standard answer is the same as in eDNA work: repeated deployments, explicit detection probability, occupancy models. Where true abundance is required, the honest routes are synchronised microphone arrays that localise individual callers, or call-rate calibration on the ground — both orders of magnitude more expensive than the recorder, which is precisely why “we have the audio” overstates what the audio knows.