A missed failure means unplanned downtime, secondary damage, a blown schedule, and at the limit a safety incident. A false alarm is a wasted inspection, tolerable once, but repeated it teaches the crew to ignore the system, and then a real warning lands on deaf ears. The second failure mode is the one that kills these programs, because it is gradual and nobody records the day the system stopped being believed.
Accuracy is the wrong target by construction. Failures are rare, so a model that always reports healthy scores over ninety-nine percent and catches nothing. Underneath that: the same healthy machine produces different vibration and current signatures at different loads and speeds, so a reading judged against a global baseline flags normal operation as anomalous every time the regime changes, which is the dominant false-alarm source. And sensor data from a real plant is not clean, because a dropout or a slow calibration drift mimics a fault convincingly if nothing upstream conditions the signal.
The reflex, and the fix.
Tune for accuracy, alert on anomalies
Standard, quick, and demonstrates a high score on historical data.
Misses rare failures or floods the crew with false alarms, and either way the system gets switched off within a month.
Cost-weighted, with actionable lead time
Operating-context normalization, a cost-tuned threshold, a lead-time gate, and calibration against real run-to-failure history. More to build and more to tune.
Warnings the crew believes, early enough to act on.
Optimize the cost of the plant's errors and the timeliness of its warnings, not a metric on a slide.
Condition at the signal layer first
Denoise, impute and resample gaps, remove drift, and align sensors to a clean baseline before anything else runs, because conditioning removes false signals that no amount of downstream modeling can recover from. Half of what looks like a detection problem is a signal processing problem nobody addressed.
Normalize for operating context
The signal is judged against the baseline for the current load and speed rather than against a global one. This single decision removes the largest source of false alarms, and it requires understanding the equipment rather than the dataset.
Predict only what leaves a trail
The system tracks condition indicators that trend as a machine degrades and predicts the crossing before it happens. Part of its honesty is knowing which failure modes degrade gradually and can be seen coming, and which fail suddenly and cannot.
A decision layer, not a score
An alert fires only when a cost-weighted threshold and a lead-time gate both agree, calibrated on real run-to-failure history and held just short of the false-alarm rate that would cost the crew its trust. Everything else stays quiet, deliberately.
A single headline number hides where a system fails. This work was scored on the dimensions that actually decide whether it holds in production, measured on real, held-out cases rather than the demo path.
A high score on a slide and, in the plant, either silence on the failures that matter or a stream of false alarms that gets the system muted within a month.
False alarms fell 41% across the 87-unit fleet, and the alerts that remained were believed and acted on, early enough to prevent the failures that mattered. The operating point was tied to the real cost of each error rather than to a generic accuracy figure.
What it owns, and what it hands to a person.
It predicts the degradation modes it has signal and history for. A truly sudden or random failure leaves nothing to warn on, and no system will change that. Conditions outside its training degrade it, the cost threshold has to be maintained as the economics and the equipment change, and it informs a maintenance decision rather than making one.
Three of the five deciding decisions on this program are below the model: signal conditioning, operating-context normalization, and health-indicator selection. A specialist engaged to build the model would have built a correct model on an uncorrected signal, judged against the wrong baseline, and would have been entirely within scope while the system failed. The discipline carries to any detection problem with rare, expensive events: weight the errors by what they cost, normalize for the conditions, demand actionable lead time, calibrate for trust, and be honest about what has no signal at all.