// Claims operations automation
A summary you cannot check is a summary you redo.
Police reports, medical records, demand packages, forensic reports. The path to a liability decision runs through documents an adjuster has to be able to verify in one click, inside ClaimCenter, or the tool is abandoned by month three and the file is worse for having had it.
Dearborn Labs is an AI-native software development firm built specifically for insurance. There is no Dearborn Labs product in this — you hire the engineers, and your claims organization owns the file the day we leave.
// The situation
The document on the desk is getting harder.
Two trends are running against each other on this desk. Medical expense per claim and bodily-injury severity keep climbing, and hiring at the junior end keeps thinning. The file gets harder as the bench gets lighter, and week one measures both on your own book rather than on an industry average.
So the first thing we build on a claims file is the check-path, not the summary. If an adjuster cannot confirm a diagnosis, a date or a liability indicator against the page it came from, the summary has to be re-read rather than trusted, and every hour the system saved is spent twice.
It is also worth knowing what your core vendor is about to ship. Guidewire’s Qusar release shipped pre-built claim summarization and FNOL agents, and Duck Creek’s April 2026 platform ships an AI gateway supporting MCP and A2A. Part of an honest read on a claims workflow is telling you which of these to wait for.
Vendor release notes · Guidewire, August 2026 · Duck Creek, April 2026.
// The stack you actually run
The file, and the systems it is scattered across.
Two structural facts about this work that a demo does not surface.
The first is that extraction quality collapses on exactly the pages a person would have squinted at. Smudged scans and low-quality faxes are where hallucination happens, and a system that writes in one uniform tone hides which pages those were. That is an interface and confidence-reporting problem before it is a model-quality problem, and it is the failure mode adjusters name first when they are asked.
The second is that a claims summary is an exhibit. The file gets litigated, the demand package arriving on the desk is itself machine-accelerated, and a coverage position generated by a system nobody can trace is a problem in a deposition rather than a bad user experience. So the summary assembles and surfaces. It does not set a reserve and it does not generate a coverage position. The adjuster decides, and the record shows who decided and on what.
// The build
Four builds, and what each one is measured on.
Structured file with page-level citation
Estimates, police reports, medical records and demand packages turned into a structured, searchable file where every extracted fact — diagnosis, date, amount, liability indicator — links to the source page. Checking costs one click instead of a re-read of 300 pages.
Surface: the document repository into the ClaimCenter file, with the page image retained as the citation target.
Loss-run normalization
Valuation-date alignment so two runs valued three months apart stop being compared as if they were not. ALAE included, excluded or unstated, with unstated labeled as what it is. Open-reserve development surfaced rather than buried in a total. Claim counts reconciled against occurrence counts, with the convention each source used labeled.
Surface: inbound loss runs in one format per carrier into a normalized history with its assumptions stated per field.
Coverage-to-claim check when the file opens
Policy terms compared against the claim facts at intake rather than after the team has worked it, so reserves get set on the right basis the first time. The check cites the form and endorsement it read.
Surface: the policy record and the claim record, joined at FNOL, written back as a flagged coverage note.
Confidence labeling and routing on poor source pages
Low-quality pages are flagged as low confidence and routed for a human read rather than silently guessed at. The confidence distribution is a reported metric, not an internal threshold, so the team can see where the system is weakest before it costs them a payout.
Surface: the extraction pipeline into a named review queue, with the reason for routing attached.
What we measure
Correction rate on system output, tracked as a first-class metric rather than hidden in a support queue. Time from FNOL to coverage position and from FNOL to liability determination, on the same cohort definition on both sides. Reserve development on files where the coverage check fired against files where it did not. Verification cost, meaning the time an adjuster spends confirming a fact against its source page. Share of pages routed as low confidence, and the accuracy on those pages specifically. Baselines are taken before the build, because a baseline reconstructed afterward is an argument rather than a measurement.
// We ran one of these
The check-path gets built before the summary, because the file has to hold up on the day it is examined.
The operating record →// What's hard about this
Two limits, and what we do about each.
Automation buys capacity. It does not buy judgment.
Hiring is being cut at the junior end while senior postings hold, so the pipeline thins at exactly the moment the surviving work gets harder. The assumption that technology compensates for a less experienced adjuster has a limit, and it shows up in indemnity rather than in cycle time.
So the build is scoped against the judgment it cannot replace. Severity thresholds and complexity markers route files to experience rather than to whoever is free, the routing rules are explicit and editable by your team instead of living inside a model, and the review seat — who checks output, who owns the correction loop, who is named on the escalation — is written into the statement of work with a runbook and the training behind it.
A cleaned file reads as authoritative whether or not the values were ever right.
This is the specific failure mode of AI on claims documents. Normalization raises confidence faster than it raises accuracy, and a loss history that reads clean is more dangerous than one that reads messy, because nobody re-checks it in the third year when it is feeding a rate filing.
The refutation is to build the confidence signal alongside the data rather than after it. Every output states which fields were extracted, which were inferred, and which were defaulted. Where a source document is silent, ALAE treatment is the common case, the output says silent rather than choosing. A summary with four caveats attached is worth more than one that reads clean and is quietly wrong.
// What ships with it
The governance file, scoped to a claim that gets litigated.
A model inventory entry with a risk tier for the extraction models and, separately, for anything that scores or ranks a claim. Data lineage from source page to extracted field. Retention of the extraction record and the human decision beside it, because the defensibility of the file depends on being able to show who decided and on what. Drift thresholds with remediation triggers, and a human-override specification naming who can act outside the boundary. Roughly half the states have adopted the NAIC AI Model Bulletin and a claims decision sits squarely inside it.
NAIC and state bulletins · current at September 2026.
What the governance file contains →Bring us one claim type.
We will tell you what the claims file needs before a summary is worth trusting.
