Document IntelligenceField engagement

The fix was not making the model more accurate. It was making the system know when it was unsure.

Diagnosis and rebuild lead.
44 to 95%
extraction accuracy
5 days
elapsed
the 5%
routed to a person
before44%
after, in five days95%
Extraction accuracy before and after the rebuild. The gain is real, and the gate on the doubtful tail is what made it usable.
What was at stake

An extraction system pulling structured data out of messy real-world documents, sitting at 44% accuracy. Low enough that a person re-checked nearly everything, which defeats the automation entirely. A wrong extraction is expensive, and a model that is confidently wrong on the hard cases is more dangerous than one that flags them.

The constraint

The stated ask was to make the model more accurate. The real problem is whether the system knows when it is unsure, so a person can be pointed at the doubtful cases instead of re-checking everything. Raw accuracy alone does not fix an automation that nobody trusts, because at any accuracy short of perfect the reviewer still has to look at all of it unless the system can say which ones to look at.

training distributiondeployment distributionunseen, where it fails
The fork

The reflex, and the fix.

Road not taken

Push raw accuracy

Pull

It is the stated ask, it is measurable, and it feels like the obvious engineering response.

Why not

Even a large accuracy gain leaves a reviewer checking everything, because without calibrated uncertainty there is no way to know which outputs to trust.

Road taken

Make the system honest about its own doubt

Accepted

Rebuilding the extraction path and adding calibrated uncertainty and a human review gate, rather than only tuning for a higher number.

Bought

A 95% system where the 5% it is unsure about is routed to a person, which is what makes the automation usable.

Decision

Fix the doubt, not just the accuracy, because the reviewer's time is the thing the automation was supposed to buy back.

How it was built
01Varied documents
02Extraction
03Confidence scoring
04Review gate
05Downstream
01

The extraction path rebuilt

From 44% to 95% accuracy on messy real-world documents, in five days.

02

Output that carries calibrated uncertainty

The model returns a confidence that means something, rather than a bare answer that a reader has to take on faith.

03

A human review gate on the doubtful tail

Low-confidence cases are caught before they flow downstream. The gate is what turns the model is 95% into the 5% it is unsure about is routed to a person.

How it was measured

A single headline number hides where a system fails. This work was scored on the dimensions that actually decide whether it holds in production, measured on real, held-out cases rather than the demo path.

Extraction accuracy on real documentsCalibration of the confidence signalReview volume against the old baselineErrors reaching downstream
figures

44 to 95% in five days is a measured, delivered result.

What it produces
Without this discipline

A more accurate model that a reviewer still checks in full, because nothing tells them which outputs to distrust.

This system

44 to 95% extraction accuracy in five days, with calibrated uncertainty and a human review gate on the low-confidence tail, so review effort lands where the doubt is.

44 to 95% in five dayscalibrated uncertaintyhuman gate on the tailthe model is a component
The operating envelope

What it owns, and what it hands to a person.

Handled with confidence
Structured extraction from the supported document mix
Confidence-scored output
Routing of the low-confidence tail
Flagged for review
Low-confidence extractions
Document types outside the supported mix
Out of scope by design
Documents with no recoverable structure
Fully unattended operation
The honest limit

Accuracy is measured on the document mix this system actually receives. Calibration has to be maintained as the document population shifts, and the review gate assumes there is a person available to take the flagged cases.

What it generalizes to

This is the second clean proof of the thesis that it is almost never the model. The stated problem was accuracy, the real problem was self-knowledge, and the fix took five days against the multi-month rebuild that a model-first reading would have scoped. The discipline carries anywhere a human reviews machine output: make the system say what it does not know, and the review effort collapses onto the part that needs it.

How we engage

You have a system like this one.
Tell us where it stands.

Whether it is failing, not yet built, or about to meet a scale it has never seen, we can tell you what we see.

Start a conversation
mostafa@opulion.dev · Response within 24 hours · By inquiry