Methodology

How a Diagnose gets to a straight answer in days.

This is the investigation behind one of the five ways to engage us, published in full because the value was never in keeping it private. Most of what we do is architecture and delivery on programs that have not failed. This is the method for the cases where something already has, and where somebody needs a defensible answer before the next decision gets made.

The principle

Most investigations start from the team's theory of the problem and the architecture diagram. We start from neither. We start from the system as it actually runs and work outward from what we can measure, on the other side of the boundary the team is standing inside.

The comparison nobody ran is almost always the one that finds the cause. What the production path computes against what the reference path computes, on identical inputs. The behavior in staging against the behavior in the field. The assumption that held at a hundred nodes against the reality at twenty-five thousand. The steps below are how we set up that comparison, in the order that keeps the signal clean.

The method, demonstrated

44% to 95%, in five days.

A production model tested at 95% in staging and sat at 44% a week after launch. Two experienced contractors reviewed it independently and both recommended a multi-month rebuild of the model. Step one of the sequence below, confirm the system is what you think it is, found the cause instead.

The training set had been recorded almost entirely by one person. The model had learned that person rather than the task, and it collapsed the first time it met real users. The model was never the problem, which is why reviewing the model twice found nothing. The fix took five days against the months that had been scoped.

Four steps. The order is not optional.

01 / 04

Confirm the ground

Before anything else, we establish that the environments you believe are identical actually are. The same artifacts, the same versions, the same configuration, the same hardware behavior. A surprising share of production failures resolve here, in the layer everyone assumes is fine and no one checked, and it is the fastest check we run.

02 / 04

Compare across the boundary

We take representative real inputs and run them through both sides, the reference path and the production path, and compare the outputs directly, before any post-processing. Agreement tells us the problem is downstream. Divergence with identical inputs tells us it is in the runtime or the environment. Either way, the comparison narrows the search to a specific layer.

03 / 04

Replay the pipeline

The check that is almost never run, and where the cause most often lives. We take what production actually computed and compare it against what the reference pipeline would have computed on the identical raw input. Any divergence is a confirmed skew between the two pipelines, the silent kind that degrades a system for months without throwing a single error.

04 / 04

Name the cause, sequence the fix

We resolve the investigation into specific, demonstrated failure modes, each with its root cause and a fix. Then we order them, because order is part of the answer. Fix the pipeline first. Get a clean signal. Then measure what is actually there. Then address it. The sequence is the difference between a fix that holds and a month back at the start.

What the comparison looks like

One line, until it is two.

Training and serving are meant to compute the same thing. On a chart they sit as a single line, until they do not. The moment they diverge on identical inputs is the skew: the silent kind that degrades a system for months without throwing a single error. Finding it is a matter of running the replay and reading the gap.

training vs serving, identical input
Why we publish it

A method you can read is not a method that stops working once it is known. The value was never secrecy. It is the discipline to run every step, in order, when the pressure is on and the temptation is to skip to the part you already suspect.

If you want to run this yourself, run it. If you would rather have the people who built it run it on a system that matters, bring it to us.

How we engage

Better to own the program
than to investigate it later.

If you have a system that has already defeated everyone who has touched it, bring it and we will run this. If you have one that has not been built yet, there is a cheaper conversation to have first.

Start a conversation
mostafa@opulion.dev · Response within 24 hours · By inquiry