Firmware & OTAField engagement

A staged rollout limits how many devices a bad build reaches. It cannot make a bad build safe.

Architecture and release-ownership lead.
2%
where the release was held
no-downgrade
platform constraint
0
models in it
ObserveDetect driftTuneDeploythe live system, over time
A standing loop around the live system. Watch, catch drift while it is small, tune, and deploy, before a customer ever sees a problem.
What was at stake

A connected-hardware fleet needs firmware delivered over the air across multiple SKUs and factories, on a platform with a strict no-downgrade rule and modules whose recovery behavior depends on the exact silicon. A power cut mid-update can brick a unit to a bench with no path back over the air, in somebody's home, as a return and a review. A clean staged rollout of a bad build still bricks its percentage of the fleet.

The constraint

The stated ask is run the over-the-air updates and a staged rollout: allowlist, then one percent, then wider. The real gap sits one layer below the rollout ladder. The ladder limits how many devices a bad build reaches, and that is all it does. It cannot make the build safe, and nothing in the ladder proves the build survives an interrupted update before it reaches the first device.

BMC session limit: 4 to 8 concurrent
naive fan-out14 sessions
opens far more sessions than the controller supports, and crashes it
scheduled polling6 sessions
capped per controller, with backoff, so monitoring does not degrade what it monitors
The constraint that separates people who have run a real fleet from people who have read the specification.
The fork

The reflex, and the fix.

Road not taken

Treat the staged rollout as the safety mechanism

Pull

It is the industry-standard answer, and it is genuinely useful for limiting blast radius.

Why not

A one percent rollout of a build that bricks on power loss bricks one percent of the fleet. Limiting reach is not the same as being safe.

Road taken

Prove survivability below the ladder first

Accepted

A module-specific failure-mode assessment, an executable failure-injection matrix, and a go or no-go gate before device one.

Bought

A build that has been shown to survive a botched update, and an honest boundary between what an update can recover and what needs a bench.

Decision

Pin the module and its recovery boundary before touching the rollout, because the ladder assumes a safe build and cannot produce one.

How it was built
01Survivability gate
02Identity binding
03Staged rollout
04Field evidence
05Recovery
01

Survivability, the gate before device one

The single-bank unpack-over-application-partition brick, no fallback, bench-only recovery, stated at the part-family level, with an executable failure-injection matrix and a go or no-go gate.

02

Diagnosis behind a misleading progress bar

A unit sitting at 98 percent has almost always finished downloading, so the stuck layer is install, reboot, or the version report, not the bar. Locating the failing layer is the first diagnostic move.

03

Identity, so the right image reaches the right device

Firmware-key and product-ID binding, which is where the multi-SKU key-mismatch trap lives.

04

Control, a ladder that advances on evidence

Allowlist to a small proportion to wider, advancing on positive evidence rather than on a calendar. Missing evidence counts as a fail, so unknown never gets promoted to go.

05

Recovery, forward-versioned rollback

On a platform that only accepts higher versions, take the known-good older binary, relabel it to a higher version, and the device installs it as a normal upgrade and runs the old safe code. Then draw the recovery boundary honestly between what an update can fix and what needs a bench.

How it was measured

A single headline number hides where a system fails. This work was scored on the dimensions that actually decide whether it holds in production, measured on real, held-out cases rather than the demo path.

Survival under injected power lossCorrect image to correct SKUField offline rate against baselineRecovery coverage without a bench
figures

The release held at 2 percent is a measured, delivered outcome.

What it produces
Without this discipline

A textbook staged rollout that does everything right procedurally and still bricks its percentage of the fleet, because nothing upstream ever asked whether the build survives an interruption.

This system

A survivability gate with a module-specific failure-mode assessment, an executable failure-injection matrix, a recovery-boundary determination with a no-downgrade forward-version rollback, and a go or no-go gate, plus real release ownership on top.

survivability before rolloutevidence, not calendarforward-versioned rollbackhonest recovery boundary
The operating envelope

What it owns, and what it hands to a person.

Handled with confidence
Over-the-air delivery across SKUs and factories
Interrupted-update survival within the assessed part family
Forward-versioned rollback
Flagged for review
Modules outside the assessed part family
Field signals diverging from bench behavior
Out of scope by design
Recovery of a bricked single-bank unit without a bench
The honest limit

The recovery boundary is drawn honestly, which means some failure modes are named as needing a bench rather than being papered over with an optimistic recovery claim. A release was held at 2 percent when field data showed updated units dropping offline below baseline, a signal the bench never produced, which is the argument for field evidence over bench confidence.

What it generalizes to

A process that limits blast radius is not a substitute for a property that makes the payload safe. Ask what the procedure assumes, then go prove the assumption one layer down before relying on the procedure.

How we engage

You have a system like this one.
Tell us where it stands.

Whether it is failing, not yet built, or about to meet a scale it has never seen, we can tell you what we see.

Start a conversation
mostafa@opulion.dev · Response within 24 hours · By inquiry