What a Regulator Asks About Your Biosignal Front End That Your Bench Never Did
Your bench asks whether the signal looks right. A reviewer asks what the system does when it is wrong, across every unit you will ship, in conditions you did not choose. Those are different questions, and the second one is answered in the analog front end months before there is software to test.
What a Regulator Asks About Your Biosignal Front End That Your Bench Never Did
The short answer. A bench test asks whether the signal looks right under conditions you chose, on the unit you built. A reviewer asks what happens when it is wrong: across manufacturing variation, at the limits of the input range, when an electrode degrades, when the patient is not the one you recruited, and how the system knows. Those questions are answered by the analog front end and the signal chain, months before there is software, and no amount of downstream processing recovers what the front end did not acquire.
The bench is a friendly environment. You built the board, you know its quirks, you place the electrodes properly, and the subject is usually someone on the team.
A review is an unfriendly environment by design, and the questions it asks are ones a bench setup structurally cannot answer.
One: what is your input range, and what happens outside it?
The bench answers: the signal sits comfortably in range.
The question is: what happens when it does not? Motion artifact, electrode pop, defibrillation recovery, an electrosurgery unit in the room, a patient whose physiology sits outside the range you assumed.
The three behaviours that matter, and they need to be distinguishable in the record:
Saturation. The front end clips. What does the downstream see, and can it tell clipping from a genuine large signal?
Recovery. After saturation, how long until the signal is trustworthy again? A front end with a long recovery time constant produces minutes of plausible-looking, meaningless data.
Detection. Does anything mark the interval as invalid, or does it flow into the algorithm as though it were signal? This is the one that decides whether a hardware limitation becomes a clinical error.
The most useful design property here is that out-of-range should be detectable rather than merely survivable. A system that saturates and flags is fine. A system that saturates and recovers silently produces a wrong answer that nobody can trace.
Two: how does the reference path behave, and what is your margin?
Common-mode rejection is what makes a biosignal measurement possible at all, and it depends on the reference path being able to do its job across the whole range of conditions.
The question is about margin: how much return current must the reference path sink, what voltage must the driving amplifier reach to do it, and does that clear the supply rail with room left, including under fault conditions.
- How much return current must the reference path sink?
- What voltage must the driving amplifier reach to do it?
- Does that clear the supply rail, with room left?
- And does it still clear under fault conditions?
This is checkable before a board exists, from a series limit, a set point, and the supply rails. It is also the kind of defect that is fabricated in: a marginal reference path yields a device that works on the bench, works on most units, and degrades on the units where component tolerances stack the wrong way.
That is the worst possible failure profile for a fleet, because it is intermittent, unit-specific, and impossible to reproduce on the unit you have in front of you.
Three: what is the variance across units, and how do you know?
The bench has one board, or three. A submission covers a manufacturing process.
Component tolerances stack. Electrode impedance varies between units and between placements. Gain, offset, filter corner frequency and reference behaviour all move within tolerance, and the combination moves further than any single one.
The question a reviewer asks is not whether you thought about this. It is what evidence you have across units, and whether the algorithm was validated on the spread rather than on one exemplar.
Two practices answer it well:
Characterise the spread, then validate at the edges. Not the nominal unit. The corner cases of the tolerance stack, built deliberately.
Split your evaluation by unit as well as by subject. If the same device appears in training and test, your number includes how well the algorithm has learned that device.
Four: is your performance number subject-disjoint?
This one is not about hardware but it is asked about the same evidence, and it invalidates more claims than any other single item.
If samples from one subject appear on both sides of the train and test split, the reported number includes how well the system recognises that subject rather than the condition. Biosignals carry strong individual signatures: anatomy, electrode placement habits, skin properties, baseline physiology. A model has every incentive to learn the person.
A subject-disjoint split, holding out whole subjects, reports what happens with someone new. That is the situation the device meets on day one in the field, so it is the only number that describes the product.
The uncomfortable property of this: the honest number is lower, and it is lower for a system that is exactly as good as it was before you fixed the split. Teams discover this late, usually near a submission, which is the worst time to find out that the headline was a fiction.
Five: what is the physical ceiling, and where is your claim relative to it?
The question underneath all the others, and the one that arrives too late most often.
What the sensing can acquire is bounded the moment the sensing approach is chosen: the electrode type and placement, the bandwidth, the noise floor, the sampling. That happens months before there is a dataset or a prototype, and no downstream processing recovers information that was never acquired.
- Sensing approach chosenelectrode type, placement, bandwidth, noise floor, sampling
- The physical ceiling is now fixedno downstream processing recovers what was never acquired
- Months pass
- A dataset exists
- A claim is madeand if it sits above the ceiling, only changing the sensing rescues it, which at that point means changing the product
So a claim made above that ceiling cannot be rescued by engineering, by a better model, or by more data. It can only be rescued by changing the sensing, which at that point in a program means changing the product.
The check is arithmetic and it is available very early: what is the amplitude of the phenomenon you are measuring, what is the noise floor of your front end, what bandwidth does the feature occupy, and does the acquisition preserve it. If that calculation has never been written down, it is the highest-value hour available in the program.
What the bench should be asked to do instead
The change is small and it happens early.
Test the failures, not the successes. Deliberately saturate, deliberately degrade an electrode, deliberately place it wrong. Record what the system does and whether it knows.
Build the corner units. Not the good board. The board at the edges of the tolerance stack.
Instrument the front end's own health. Electrode impedance, saturation events, reference path headroom, invalid-interval count. These are the four numbers that let a fielded device tell you it is degrading rather than tell you a wrong answer confidently.
Do the ceiling arithmetic before the sensing is locked.
The pattern across all four is the same one that runs through medical device work generally: the questions that decide the program are answered at the seam between the physics and the software, and that seam is usually owned by nobody.
FAQ
What does a regulator ask about a biosignal front end? What happens outside the input range, how the reference path behaves and with what margin, what the variance is across manufactured units, whether performance numbers are subject-disjoint, and where the claim sits relative to the physical ceiling of the sensing approach.
Why is a subject-disjoint split necessary for biosignal machine learning? Because biosignals carry strong individual signatures, so a model trained and tested on overlapping subjects reports how well it recognises those people rather than the condition. Holding out whole subjects reports what happens with someone new, which is the situation the device meets on day one.
What is the physical ceiling of a sensing approach? The upper bound on what can be acquired, fixed when the electrode type, placement, bandwidth, noise floor and sampling are chosen. No downstream processing recovers information never acquired, so a claim above that ceiling can only be rescued by changing the sensing.
Why does reference path margin matter so much? Because a marginal reference path yields a device that works on the bench and on most units, then degrades on units where component tolerances stack unfavourably. That is intermittent, unit-specific and unreproducible on the board in front of you, which is the worst failure profile for a fleet.
What should a bench test do differently? Test failures rather than successes: deliberately saturate, degrade an electrode, misplace it, and record whether the system knows. Build units at the edges of the tolerance stack rather than a good one, and instrument electrode impedance, saturation events, reference headroom and invalid-interval counts.
Tell us the system, the stakes, and the date that matters. You get a straight technical reply from the person who would lead the work, within 24 hours.
Bring us the program