Industrial AutomationField engagement

The two ways to be wrong on a plant floor are nowhere near the same size, and a model tuned for accuracy is solving a problem the plant does not have.

Fractional CTO and lead engineer.
41%
fewer false alarms
87
unit fleet
cost-weighted
operating point
Missed failuredowntime, damage, safety
False alarma wasted inspection
The two ways to be wrong are not the same size, so the threshold is set by cost, not by accuracy.
What was at stake

A missed failure means unplanned downtime, secondary damage, a blown schedule, and at the limit a safety incident. A false alarm is a wasted inspection, tolerable once, but repeated it teaches the crew to ignore the system, and then a real warning lands on deaf ears. The second failure mode is the one that kills these programs, because it is gradual and nobody records the day the system stopped being believed.

The constraint

Accuracy is the wrong target by construction. Failures are rare, so a model that always reports healthy scores over ninety-nine percent and catches nothing. Underneath that: the same healthy machine produces different vibration and current signatures at different loads and speeds, so a reading judged against a global baseline flags normal operation as anomalous every time the regime changes, which is the dominant false-alarm source. And sensor data from a real plant is not clean, because a dropout or a slow calibration drift mimics a fault convincingly if nothing upstream conditions the signal.

alertgradualsuddenpredicted before it failstimecondition indicator
What degrades gradually can be predicted before the crossing. What fails suddenly leaves nothing to warn on.
The fork

The reflex, and the fix.

Road not taken

Tune for accuracy, alert on anomalies

Pull

Standard, quick, and demonstrates a high score on historical data.

Why not

Misses rare failures or floods the crew with false alarms, and either way the system gets switched off within a month.

Road taken

Cost-weighted, with actionable lead time

Accepted

Operating-context normalization, a cost-tuned threshold, a lead-time gate, and calibration against real run-to-failure history. More to build and more to tune.

Bought

Warnings the crew believes, early enough to act on.

Decision

Optimize the cost of the plant's errors and the timeliness of its warnings, not a metric on a slide.

How it was built
01Conditioning
02Context normalize
03Health indicators
04Decision
05Evaluate
measured on real, held-out cases, then tuned
01

Condition at the signal layer first

Denoise, impute and resample gaps, remove drift, and align sensors to a clean baseline before anything else runs, because conditioning removes false signals that no amount of downstream modeling can recover from. Half of what looks like a detection problem is a signal processing problem nobody addressed.

02

Normalize for operating context

The signal is judged against the baseline for the current load and speed rather than against a global one. This single decision removes the largest source of false alarms, and it requires understanding the equipment rather than the dataset.

03

Predict only what leaves a trail

The system tracks condition indicators that trend as a machine degrades and predicts the crossing before it happens. Part of its honesty is knowing which failure modes degrade gradually and can be seen coming, and which fail suddenly and cannot.

04

A decision layer, not a score

An alert fires only when a cost-weighted threshold and a lead-time gate both agree, calibrated on real run-to-failure history and held just short of the false-alarm rate that would cost the crew its trust. Everything else stays quiet, deliberately.

How it was measured

A single headline number hides where a system fails. This work was scored on the dimensions that actually decide whether it holds in production, measured on real, held-out cases rather than the demo path.

Cost-weighted errorRecall on real failuresFalse-alarm rateLead-time distributionCalibration
too early
actionable window
too late
failure
vague, ignored
order the part, schedule the stoppage
no time to act
A correct warning at the wrong moment is useless. The target is the window where there is still time to act.
figures

What it produces
Without this discipline

A high score on a slide and, in the plant, either silence on the failures that matter or a stream of false alarms that gets the system muted within a month.

This system

False alarms fell 41% across the 87-unit fleet, and the alerts that remained were believed and acted on, early enough to prevent the failures that mattered. The operating point was tied to the real cost of each error rather than to a generic accuracy figure.

cost-weightedactionable lead timecontext-awarecalibrated for trust
The operating envelope

What it owns, and what it hands to a person.

Handled with confidence
Gradual degradation modes with a signal
Within trained operating conditions
With actionable lead time
Flagged for review
Low confidence
An operating regime not seen before
Out of scope by design
Sudden or random failures
Modes with no signal or no failure history
The honest limit

It predicts the degradation modes it has signal and history for. A truly sudden or random failure leaves nothing to warn on, and no system will change that. Conditions outside its training degrade it, the cost threshold has to be maintained as the economics and the equipment change, and it informs a maintenance decision rather than making one.

What it generalizes to

Three of the five deciding decisions on this program are below the model: signal conditioning, operating-context normalization, and health-indicator selection. A specialist engaged to build the model would have built a correct model on an uncorrected signal, judged against the wrong baseline, and would have been entirely within scope while the system failed. The discipline carries to any detection problem with rare, expensive events: weight the errors by what they cost, normalize for the conditions, demand actionable lead time, calibrate for trust, and be honest about what has no signal at all.

How we engage

You have a system like this one.
Tell us where it stands.

Whether it is failing, not yet built, or about to meet a scale it has never seen, we can tell you what we see.

Start a conversation
mostafa@opulion.dev · Response within 24 hours · By inquiry