Service · 01 / 05

Diagnose.
Know what you are facing before you commit to it.

A paid, independent investigation into what a system is actually doing and what it will actually take. Scoped in days, and structured so the finding is defensible to whoever has to approve the next decision.

One continuous lifecycle. Enter at any stage. None of them requires the others.
When

When a program needs a straight answer before it commits, or when a system is failing and the people closest to it have run out of explanations.

What it is

Most of the value here goes to buyers who have not failed at anything yet. Before committing a program budget to a build, you can commit a much smaller one to establishing what you are actually facing: whether the architecture you are about to fund will hold at the scale you will reach, whether the hardware already selected supports what the software assumes, whether the system needs the model everyone has been planning around, and what the real critical path is.

It works in the other direction too. A production model once tested at 95% in staging and sat at 44% a week after launch. Two experienced contractors reviewed it independently and both recommended a multi-month rebuild of the model. The cause was in the data: the training set had been recorded almost entirely by one person, so the model had learned that person rather than the task. The fix took five days. The expensive decision was never the rebuild. It was the quarter that would have gone into it, on the wrong thing.

Data and distribution mismatch60%
Hardware and inference behavior25%
Pipeline and sensor10%
Model architecture5%
Where production failures actually live, in our experience across dozens of systems. Most teams spend their time on the smallest slice of the chart.
How it works
01Confirm the ground
02Compare across the boundary
03Replay the pipeline
04Name the cause, sequence the fix
01

Confirm the ground

Start from the system as it actually runs, not the architecture diagram or the team's theory of the problem. Establish what is true before forming any hypothesis.

02

Compare across the boundary

Run the comparison nobody ran: the production path against the reference path on identical inputs, staging against the field, a hundred nodes against twenty-five thousand.

03

Replay the pipeline

Reproduce the failure under controlled conditions so the cause can be observed and measured rather than guessed at.

04

Name the cause, sequence the fix

Name the cause, demonstrate it, rank the failure modes by impact, and lay out the fix in the order that actually holds.

What this looks like in practice

Anonymized, all real delivered work. Described by domain, architecture, and outcome, never by client.

01

A document system stuck at 44 percent

Accuracy was low enough that a person re-checked nearly everything, which defeated the automation entirely. The reflex was to make the model more accurate. The real question was whether the system knew when it was unsure.

rebuilt to 95 percent in five days
02

A wearable days from a locked bill of materials

The firmware work was scoped as firmware work. The real risk sat one layer below, in an undocumented controller routing issue and a charger-detect front end that, left as drawn, would have become a fabricated-in defect the moment the bill of materials froze.

caught on the bench, before the freeze
03

A fleet that found its failures late

Health was split across three tools and failures were found after impact by a human. The stated ask was a unified dashboard. The real problem was detecting the absence of a signal, because a node that goes quiet is the one that matters and naive polling never notices silence.

the ask was not the problem

The pattern under all three: it is almost never the thing everyone is blaming. The visible component gets the attention, and the invisible one, in the data, a seam, a schematic, or a missing heartbeat, gets ignored.

The comparison nobody ran. What the production path computes against the reference path, on identical inputs. The moment they diverge is the cause.
What you walk away with
The cause, named and demonstrated, not theorized
The failure modes ranked by impact
A fix sequence in the order that actually works
A straight answer: repair or rebuild, and what it takes
Where it connects
How we engage

Tell us the system,
the situation, and the stakes.

We will tell you whether the diagnose stage is what the work actually needs, and whether we are the firm to do it.

Start a conversation
mostafa@opulion.dev · Response within 24 hours · By inquiry