Regulated DocumentsField engagement

In a regulated domain a wrong answer is not a lower score, it is a liability, so the system has to know when it does not know.

Fractional CTO and lead engineer.
closed-domain
grounded
human
in the loop
abstention
as a correct output
Fluent summary

Either party may terminate on material breach with thirty days notice.

nothing to check it against
Cited extraction

Termination on material breach, thirty days notice. section 9.3

traceable to the source
The same fact, two outputs. One asks to be believed. The other can be checked in seconds against the source.
What was at stake

The output is not the end of the process. It is the start of a decision a professional is accountable for, and a confident wrong answer is worse than no answer, because nobody downstream can see it is wrong. Verifiability is the product.

The constraint

Fluency is the wrong target, because a fluent answer is not a verifiable one. The domain-specific failures are particular: an unsupported claim, a misattributed citation, and most dangerously the material provision the system never mentions, which is invisible because there is nothing on screen to check against.

ARBITRARY CHUNK
Party B shall indemnify Party A|except where loss arises from gross negligence
carve-out split from the obligation, meaning lost
CLAUSE SEGMENT
section 7.1Party B shall indemnify Party A, except where loss arises from the gross negligence of Party A.
obligation and carve-out, retrieved as one unit
An obligation and its exception are a single legal thought. Segment on a token count and the model retrieves the duty without the carve-out that guts it.
The fork

The reflex, and the fix.

Road not taken

Summarize the document with a language model

Pull

A polished summary in seconds, and it demonstrates beautifully on a clean document.

Why not

No grounding, no recall guarantee, and no way to verify. It produces exactly the output a professional cannot rely on.

Road taken

Grounded extraction with review built in

Accepted

Closed-domain scoping, grounding on retrieved spans, hallucination control, and a human-in-the-loop review surface.

Bought

An answer a professional can confirm quickly, and an explicit signal when the system is not confident.

Decision

Build a tool that supports professional judgment with checkable evidence, rather than one that asks for it to be outsourced.

How it was built
01Parse
02Ground
03Extract
04Verify
05Review
01

Closed-domain scoping

The system answers within a defined document domain rather than open-endedly, which is what makes grounding and completeness checks possible at all.

02

Grounding and hallucination control

Every claim is bound to retrieved source content, and an unsupported claim is withheld or flagged rather than presented, which is precisely the failure a plain pipeline lets through.

03

Human in the loop by design

The review surface is part of the system rather than an afterthought, so the tool does the finding and the professional does the judging.

How it was measured

A single headline number hides where a system fails. This work was scored on the dimensions that actually decide whether it holds in production, measured on real, held-out cases rather than the demo path.

Support rate for extracted claimsCitation accuracyFlagged-for-review rateRecall on material provisions
figures

What it produces
Without this discipline

A clean paragraph the professional still has to verify by reading the whole document, with no signal about what was missed.

This system

Grounded, checkable output with uncertainty surfaced rather than hidden, and a review surface where a professional confirms in seconds instead of re-reading.

closed-domaingroundedhallucination controlhuman-in-the-loop
The operating envelope

What it owns, and what it hands to a person.

Handled with confidence
Defined provisions on supported document types
Claims bound to retrieved source content
Flagged for review
Low confidence
Unusual or conflicting language
Out of scope by design
Professional judgment and strategy
Wholly novel document types
The honest limit

It accelerates review and makes output checkable. It does not exercise professional judgment, the professional remains accountable for what they sign, and flagged items exist to be verified by a person. It is not advice.

What it generalizes to

Where correctness is legally load-bearing, an answer you cannot trace to its source is an answer nobody can use. Bind every claim to something checkable, let the system abstain, and put the output where it can be verified quickly.

How we engage

You have a system like this one.
Tell us where it stands.

Whether it is failing, not yet built, or about to meet a scale it has never seen, we can tell you what we see.

Start a conversation
mostafa@opulion.dev · Response within 24 hours · By inquiry