// What we build

A summary you cannot verify is one you redo.

Medical records, attending physician statements, demand packages, forensic reports. Four hundred scanned pages of narrative and clinical text read down to diagnoses, dates, amounts and liability indicators, every extracted fact linked to the page it came from.

The strongest legitimate use of a large language model in insurance, for a specific reason: it is genuinely summarization of unstructured text, with a human making the call at the end.

// The situation

Sometimes the document everyone waits for was not needed.

In a retrospective study of life underwriting files, a decision made from foundational data plus the electronic health record matched the decision made once the attending physician statement was added in 87% of cases.

Milliman IntelliScript retrospective study · vendor research · sample and case mix not disclosed.

We cite it as directional and would replicate it on your own book first, because it is somebody else’s case mix. The point that survives replication is narrower: the chokepoint document is often read at full length to confirm what the file already said, and page-level citation makes that checkable in minutes rather than hours.

// The stack you actually run

The document types, and why each one is hard differently.

The failure modes here are named, documented, and specific.

Claims professionals describe them precisely: faulty claim summaries, omitted details from medical reports leading to inaccurate payouts, and hallucination from smudged or low-quality scans. Extraction quality collapses on exactly the pages a person would have squinted at, and a system that writes in one uniform tone hides which pages those were. That is an interface and confidence-reporting problem rather than a model-quality one.

And the incoming document is getting harder, not easier.

Bodily injury now takes a larger share of the loss dollar than physical damage does, and plaintiff firms are adopting generative AI of their own. The demand package arriving on the desk is itself machine-accelerated, which raises the standard the file has to meet.

// The build

Four builds, and what each one is measured on.

01

Page-level citation on every extracted fact

Diagnoses, dates, medications, severity markers, amounts and liability indicators each link to the source page, so checking a fact costs one click rather than a re-read. This is a design requirement rather than a feature: a summary you have to redo has negative value, because it cost you the reading as well.

Surface: the document store and the claim or underwriting file, with citations written back as part of the record.

02

Low-confidence output labeled rather than smoothed

Smudged scans, poor-quality faxes, handwriting and pages arriving rotated or out of sequence are where hallucination happens. Those pages are scored, flagged and routed for human read instead of being guessed at in the same confident register as the clean pages.

Surface: the extraction pipeline, with a routed exception queue your own supervisors control and can re-tune.

03

A structured case record instead of a block of prose

Timeline, entities, treatment sequence, coverage-relevant facts and contradictions, as fields that can be queried, compared across documents and diffed when a supplemental arrives. A narrative summary cannot be reconciled against a second document. A structured record can.

Surface: the document set into the claim or underwriting record as structured fields with citations retained.

04

The decision record, with a named human on it

No reserve set automatically. No coverage position generated. No adverse decision issued. The system assembles and surfaces; a licensed human decides; the record shows who decided, on what evidence, and what the system recommended if anything.

Surface: the decision record in the claim or policy system, with the human decision-maker named on every write.

What we measure

Correction rate on extracted facts, tracked as a first-class metric rather than absorbed into a support queue, and reported separately for clean and low-confidence pages because the average of the two is meaningless. Verification cost, in clicks and seconds to check a fact against its source page. Share of pages routed as low confidence, and the false-positive rate on that routing, since a flag on everything is the same as no flag. Reading time per file, observed rather than estimated. Definitions and baselines, set in the first week.

“A summary of a medical record is an input to a decision somebody has to defend, which is why the check-path is the first thing built on this capability rather than the last.”

// Where this comes from

// What's hard about this

Two limits, and what we do about each.

A fully automated adverse decision is already unlawful in a growing list of states.

The statutes on the books are health-insurance utilization-review laws, and they are where the drafting pattern is being set. California SB 1120, Texas SB 815, Maryland HB 820, Nebraska LB 77, Washington SB 5395, Arizona HB 2175 and Alabama SB 63 all say some version of the same three things: artificial intelligence may inform, a licensed human must decide, and the insurer must be able to show its work. Maryland adds quarterly reporting on involvement in adverse decisions; Arizona requires a state-licensed medical director to personally review. The pattern is uniform enough to design against.

Any system here that cannot produce an audit trail and a named human decision-maker is unsellable in those states, so the decision record is architecture rather than documentation: built first, naming the person, retaining what the system recommended alongside what the human decided, and producing the statutory report as a byproduct rather than a quarterly project.

Privilege structure is part of the document, not a wrapper around it.

A forensic report is typically retained through breach counsel for a reason, and the same is true of investigative material in a litigated claim. Ingesting privileged documents into a general-purpose store, indexing them beside ordinary claim material, or surfacing facts from them to users outside the privileged group can damage a protection that was deliberately constructed.

Nothing built here changes who holds the privileged document. The design keeps custody where counsel put it: privileged material is processed inside the boundary counsel specifies, access is scoped to the privileged group rather than to the claim, extracted facts inherit their source document’s classification, and the privilege log is generated from the pipeline rather than reconstructed later. Where a document’s status is unclear it routes to counsel rather than to a model.

// What ships with it

The governance file, scoped to a decision a person signs.

A model inventory entry with its risk tier, data lineage from source document and page to extracted field, pre-deployment testing results, an annual re-test procedure, drift thresholds with remediation triggers, a human-override specification, and a named decision-maker on every path touching an adverse outcome. Roughly half the states have adopted the NAIC AI Model Bulletin and the variations are material: Connecticut requires an annual compliance certification attested by a named officer, and Iowa formally defines bias and outcomes testing.

Bring us one file.

We will tell you where extraction quality collapses on your documents, and what the check-path has to catch. How the read works →