// AI-assisted software delivery
Software written with AI. Judged by your engineers.
This is how we build, not something you buy. Software written with AI and judged by people, shipped through evaluation gates, with a test record and an audit trail a state examiner can read before anything touches PolicyCenter.
Dearborn Labs is an AI-native software development firm built specifically for insurance. There is no Dearborn Labs product in this — you hire the engineers, and the repository, the gates, the tests and the evaluation set are your team's when we go.
// The situation
Carriers know what to build. Delivery is the constraint.
Datos Insights puts it plainly: carriers broadly know what they need to build, and bandwidth constraints, data quality gaps and governance demands are what limit their ability to deliver — data being the top challenge for large insurers and IT operations constraints for midsize ones.
Datos Insights CIO survey · June 2026.
So the first thing we do on a build is not architecture. It is agreeing what would count as evidence that a piece of this system is correct, on your own documents, before any of it is written — because the honest answer to “software projects take years and fail” is not a faster team. It is a shorter distance between writing something and knowing whether it works.
This is also the page that answers the build-or-buy question in the direction most vendors will not. A carrier that can run and extend its own systems has a genuine structural advantage, and the reason that is realistic now is that the cost of writing software has moved. The reason it is still hard is that the cost of judging software has not.
// The stack you actually run
Where the gates run, and what they run against.
An evaluation set is not a test suite, and the difference is the whole discipline.
A test suite asserts that a function returns what it should. An evaluation set is a labeled corpus of your own documents — the smudged fax, the SOV with forty thousand rows, the loss run where ALAE is unstated, the demand package a plaintiff firm assembled with a model of its own — where the right answer has been written down by one of your people. It is the only thing that can tell you whether the system got better between Tuesday and Friday, and it is the artifact a carrier almost never has.
Two things follow from building it first. The work becomes deterministic where it must be and judgment-based only where that is genuinely better: premium arithmetic, effective dating, ACORD field mapping, bordereaux totals and statutory field formats are code with tests, not inference, because being right 99% of the time on a premium total is being wrong. Reading roof shape out of a narrative appraisal, or classifying what a broker meant in the third paragraph of an email, is judgment, and it is measured on the evaluation set rather than asserted.
And the model behind each judgment step sits behind a stable interface — a contract in your repository, with the model as a configuration line and a register recording which version produced which output. Model generations change on somebody else’s schedule. The interface, the tests and the evaluation set do not, which is what lets the system outlive the generation it was built on.
// The build
Four builds, and what each one is measured on.
An evaluation set from your own documents
One document class, one state, a defined number of real files, each with the correct answer recorded by a named person on your side — an underwriter, an adjuster or a reporting lead rather than an engineer. Pass thresholds set per field, because a wrong garaging address and a wrong total insured value do not cost the same.
Surface: your own historical documents, held in your infrastructure, versioned in your repository as the corpus the gate reads.
The gate in your pipeline
No change reaches a production surface without a run against the evaluation set and the deterministic test suite, with the result recorded, attributable and reviewable. A regression on a field class blocks the release rather than generating a ticket. The gate runs in your pipeline under your credentials, so it keeps running after we leave.
Surface: your continuous-integration pipeline, gating deploys to Cloud API, Application Events and the semantic layer.
Deterministic boundaries around the regulated arithmetic
Premium calculation, effective-dated transaction logic, statutory field formats, ACORD mappings and bordereaux totals implemented as code with full test coverage and no model in the path. The boundary between deterministic and judgment-based is documented as an architectural decision with the reason, so a later team can see why the line is where it is.
Surface: the calculation and mapping layers in your repository, with coverage reported per module rather than as a project average.
Stable interfaces and a model register
Each judgment step behind an interface with a fixed contract, multiple model providers configurable behind it, and a register recording model, version, prompt version and date for every production output. A model change is a configuration change that must pass the gate — not a migration project, and not a silent drift.
Surface: the interface definitions and the register in your repository, joined to the audit trail on every written decision.
What we measure
Evaluation pass rate by document class, by field and by state, on a versioned corpus so two runs are comparable. Test coverage on the deterministic modules, reported per module. Share of production releases with a recorded gate run and a named reviewer. Time from a model-version change to a green gate, which is the honest measure of whether the system is portable across generations. Defect escape rate on effective-dated transactions specifically, because that is where a wrong answer reaches finance. And the labeling cost itself, hours of your people’s time per hundred labeled files, reported rather than absorbed, because it is the real budget line under all of this.
// We ran one of these
A bad release in regulated insurance operations is a regulatory event rather than a rollback, and the people judging the software are the underwriters and adjusters using it.
The operating record →// What's hard about this
Two limits, and what we do about each.
The gate is only as good as the labels, and labeling is your underwriters’ time.
An evaluation set built by engineers measures whether the software does what engineers assumed. Only an underwriter can say that the total insured value on that SOV should have been the schedule’s figure and not the ACORD 140’s, and only an adjuster can say which of two dates in a medical record is the date of service. That is real time from people whose calendars are already full, and a project that pretends otherwise produces a gate that passes while the output is wrong.
So the corpus starts narrow — one document class, one state, a hundred files labeled by one senior person — and widens on evidence that the narrow version is holding. The labeling hours are scoped in the statement of work as a client task with a named owner and a number attached, rather than assumed into a blended rate. And the labels stay in your repository, because they are the most expensive thing produced by the engagement and they are the reason the second build costs less than the first.
A green gate here can still break on your core vendor’s release train.
The pipeline cannot see the vendor’s upgrade calendar. A system that passes every gate against a snapshot of Cloud API can fail the week the platform upgrades, and the failure lands on your team at whatever moment the vendor chose.
So gates run against the versioned contract and the published event schemas rather than against a captured response, upgrade windows in your change-control calendar are treated as gate events with a rehearsal run before them, and general-availability status is verified against the vendor’s own documentation before anything is architected on a surface. Where a build would depend on an early-access surface, the dependency is named in the proposal with the breakage it implies, and you decide.
// What ships with it
The delivery record, written for your compliance review.
The evaluation corpus with its labels, its version history and its pass thresholds. The record of every gate run tied to the release it authorized and the reviewer who accepted it. Test coverage on the deterministic modules, reported as a number rather than a claim. The model register — model, version, prompt version, date — joined to the audit trail on every written decision, which is what makes a model inventory entry and a human-override specification testable instead of declarative. Runbook, monitoring, alerting, retraining procedure, escalation path, and an owner named on your side before go-live, with exit criteria written at kickoff. Roughly half the states have adopted the NAIC AI Model Bulletin, and this is the machinery that makes its requirements auditable.
NAIC and state bulletins · current at September 2026.
What the governance file contains →Bring us one workflow and one week of its documents.
We will tell you what an evaluation set on your own documents would have to contain.
