Ground Truth

You Cannot Hire a Gap: Why Cross-Layer Failures Have No Owner

Mostafa DhouibMostafa Dhouib··9 min read
The short answer

The analog engineer designed a correct reference loop. The actuation engineer designed a correct current source with a safety limit. The mechanical engineer placed the electrodes correctly. Every one of them did excellent work inside their lane, and the device failed, because the failure lived in the loop connecting all three and that loop was on nobody's schematic. A seam is not a gap in anyone's competence. It is the region between all of your competent people.

You Cannot Hire a Gap: Why Cross-Layer Failures Have No Owner

The short answer. Three engineers each did excellent work. The analog engineer built a correct reference loop, the actuation engineer built a correct current source with a proper safety limit, and the mechanical engineer placed the electrodes where they would fit and stay in contact. The device failed anyway, because the failure lived in the loop that connected all three, and that loop was not on anyone's schematic, in anyone's job description, or in anyone's review. It got decided by default, by whatever path the current happened to find. A seam is not a gap in anyone's competence. It is the region between all of your competent people, and it is the one place none of their maps include.

I want to be precise about a failure mode that gets misdiagnosed as a people problem, because the misdiagnosis leads to hiring decisions that make it worse.

The device that failed with three correct designs

A closed-loop biosignal device. It reads a small signal off the body and acts back on it.

The analog engineer designed the front end, including a driven reference electrode holding the body's common mode where the amplifier can see the signal cleanly. In isolation, correct. Genuinely good work, and the rejection figures were something to be proud of.

The actuation engineer designed the path that injects current, with a large series resistor limiting the current to a safe value. In isolation, correct. The safety reasoning is exactly right.

The mechanical engineer placed the electrodes so they fit the enclosure and maintained skin contact. In isolation, correct.

The front end went completely flat the instant the actuator fired.

The analog engineera correct reference loop, with rejection figures worth being proud of
The actuation engineera correct current source with a proper safety limit
The mechanical engineerelectrodes placed to fit the enclosure and stay in contact
The loop connecting all threeon nobody schematic, in nobody job description, in no review
So it was decided by default, by whatever path the current found on the board. Anything decided by default is almost always decided wrong.
FigureThree engineers, three correct designs, one dead device. Each could see their half. No one could see the loop, and the loop was not assigned to anyone, because assignments are made in terms of components and a loop is not a component.

The injected current is a loop. It has to come home, and it comes home by the lowest-impedance path available, which is the driven reference electrode, precisely because that electrode is a deliberately low-impedance connection to the body. The reference amplifier acquired a second job it was never specified for, could not do both, and railed. The common mode wandered out of range, rejection collapsed, and the signal read as gone.

Nobody made a mistake. The mistake was in the loop between them, and nobody had drawn it.

Why nobody drew it

This is the part worth sitting with, because it is not carelessness and treating it as carelessness produces exactly the wrong response.

The analog engineer could not see the whole loop, because half of it lived in the actuation schematic. The actuation engineer could not, because half of it lived in the front end. To the mechanical engineer it was two electrode positions, not a circuit at all.

Each person could see their half. No one could see the loop.

And the loop was not assigned to anyone, because assignments are made in terms of components and subsystems, and the loop is not a component. It does not appear on a schematic sheet, it does not have a part number, it does not belong to a discipline, and there is no review that takes it as its subject.

So it was decided by default, by whatever path the current found on the board.

Anything decided by default is almost always decided wrong. Not because defaults are malicious, but because a default is what happens when no one is optimising, and there is no reason for an unoptimised outcome to be a good one.

The shape, everywhere else

If this were only a hardware story it would be a good war story and not a thesis. It is worth stating as a thesis because the identical structure produces the identical failure in layers that have nothing to do with each other.

Between agent stepseach step reasons flawlesslycorrect reasoning is the transmission mechanism
Between model and toolthe call is correct, and droppedevery component reports success
Between planner and executortwo valid termination proofscomposed, it never terminates. Both panels green
Between retrieval and generationidentical outputs, opposite fixesthe blended score hides which half broke
Between capped runsevery run boundednothing bounds the depth
Between training and deploymentboth pipelines finediffering by a constant nobody compared
FigureSix layers with no shared vocabulary and one structure. In every case the components are correct, the failure is in the relationship, and no component owner is responsible for the relationship.

Between agent steps. Every individual step reasons flawlessly. One early wrong assumption propagates, and correct reasoning at each stage is the transmission mechanism. Audit every step and each passes; the reliability was lost in the connections.

Between the model and the tool. The model emits a correct tool call. The tool is ready to run it. The call is dropped in the streaming reassembly, the index collision, the proxy translation, or the finish flag, and every component reports success. No faulty part exists to find.

Between a planner and an executor. Each has a genuine, valid termination proof. Composed, the planner's replanning resets the executor's measure, and the system never terminates. Both dashboard panels are correct and green.

Between retrieval and generation. Search failing and generation ignoring a good page produce identical outputs. A blended quality score averages the two and hides which half is broken, and the boundary between them is what nobody instrumented.

Between capped runs. Every agent run is provably bounded. Nothing bounds the depth of runs calling runs, so a system of bounded loops is unbounded.

Between training and deployment. The model is fine and the preprocessing is fine, and they differ from each other by a normalisation constant nobody compared.

Six layers, no shared vocabulary, one structure. In every case the components are correct, the failure is in the relationship, and no component owner is responsible for the relationship.

Why hiring deeper does not fix it

Here is the mistake this article exists to prevent.

The natural response to a hard failure is to hire more expertise in the area where it appeared. Bring in a better analog engineer. Get someone who really knows retrieval. Find an agent framework specialist.

That response is aimed at depth, and depth is not the missing quantity.

A deeper specialist has a taller lane. They will do better work inside the same boundaries, and the failure is not inside the boundaries. Hire the best analog engineer available and they will still not see the actuation schematic, because it is not theirs, and the loop still runs through both.

You cannot hire a seam, because the seam is by definition the place between all of your specialists. Adding another specialist adds another lane and therefore another two boundaries. Past a certain point, more specialists is more seams.

This is also why the failure survives excellent process. Every discipline reviews its own work competently. Nobody is assigned the region between reviews, so the region between reviews is where the defect lives, and it lives there specifically because that is the only place nothing is looking.

What actually fixes it: staffing for span

The fix, when it came, required the analog person, the actuation person, and the mechanical person in one room looking at one loop at the same time.

Not because any of them had done poor work. Because the loop only becomes visible when you look at it as a loop, and no single seat has that view.

Generalised, that is the structural answer: somebody's explicit job has to be tracing the paths that cross lanes.

Hire deeper
A deeper specialist has a taller lane
Better work inside the same boundaries
The failure is not inside the boundaries
Each addition creates two more boundaries
Staff for span
Someone whose explicit job is tracing paths across lanes
Enough real depth in each adjacent layer to follow a signal across the boundary
Not a coordinating manager, not a box-drawing architect
Assign each seam to a named person, in writing
When something goes wrong, ask whether the answer to whose is it is a name. If it is between two people, you have found a seam the expensive way.
FigureThe natural response to a hard failure is to hire more depth where it appeared. Depth is not the missing quantity, and past a point more specialists means more seams.

That role is not a manager, who coordinates people rather than tracing loops, and it is not an architect who draws boxes without going into any of them. It is someone with enough real depth in each adjacent layer to follow a signal, a current, a request, or a piece of state across the boundary and understand what happens to it on both sides.

Range, in that sense, is not a flex. It is a risk-reduction mechanism, and it is worth being blunt about why it is rare: it is expensive to acquire, it looks unfocused on a resume, and every incentive in a career pushes toward depth in one lane, because that is what gets hired and promoted.

Which is precisely why the seams stay unowned.

The practical version

You do not need to restructure your organisation to get most of the benefit. Three things.

Make the crossings explicit artifacts. For each boundary in your system, write down what crosses it, in what format, with what assumptions on both sides. A current and its return path. A piece of state and its preconditions. A tool call and what happens to it in transit. This document is nobody's deliverable today, which is the problem, and it is usually a page.

Run at least one review whose subject is a path, not a component. Take one thing that crosses three lanes and trace it end to end with all three owners in the room. This is the single highest-value hour available in most programs, and its output is usually one finding that nobody could have produced alone.

Assign the seams by name. Not implicitly, not to a committee. If the return path, the model-to-tool boundary, or the retrieval-to-generation handoff is not in someone's written responsibilities, it is being decided by default.

And a diagnostic worth applying to your own program: when something goes wrong, ask whether the answer to "whose is it" is a name. If the honest answer is that it is between two people, you have found a seam, and you have found it the expensive way.

The glossary defines the seam and the other terms this body of work relies on.

FAQ

What is a seam in engineering, and why do cross-layer failures have no owner? A seam is the region between competent specialists that none of their maps include. Failures live there because responsibilities are assigned in terms of components and subsystems, and a relationship between components is not a component: it has no schematic sheet, no part number, and no discipline, so it gets decided by default.

Why doesn't hiring a better specialist fix an integration failure? Because a deeper specialist has a taller lane, and the failure is not inside any lane. The missing quantity is span, not depth. Adding another specialist also adds another two boundaries, so past a point more specialists means more seams.

How do I find seam failures before they happen? Write down, for each boundary, what crosses it and with what assumptions on both sides. Run at least one review whose subject is a path rather than a component, with every lane owner present. And assign each seam to a named person, since anything not explicitly assigned is decided by default.

What does staffing for span mean? Having someone whose explicit job is tracing the paths that cross lanes, with enough real depth in each adjacent layer to follow a current, a request, or a piece of state across a boundary and understand both sides. It is not a coordinating manager and not a box-drawing architect.

Is this specific to hardware? No. The identical structure appears between agent steps, between a model and its tools, between a planner and an executor, between retrieval and generation, between capped runs in a call tree, and between training and deployment pipelines. Different vocabularies, same failure: correct components, broken relationship, no owner.

Carrying a program like this one?

Tell us the system, the stakes, and the date that matters. You get a straight technical reply from the person who would lead the work, within 24 hours.

Bring us the program