// AI governance and model risk

Virginia says eliminate the risk, not mitigate it.

Roughly half the states have adopted the NAIC AI Model Bulletin and the variations are material. Connecticut wants an annual certification attested by a named officer. Colorado's Regulation 10-1-1 outcomes testing is enforceable. New York DFS says an insurer cannot rely solely on a third party's claim of non-discrimination.

Dearborn Labs is an AI-native software development firm built specifically for insurance. There is no Dearborn Labs product in this — you hire the engineers, and the file belongs to your compliance function.

// The situation

Examination is standardizing while you build.

The NAIC’s AI Systems Evaluation Tool is piloting in 12 states, with adoption expected at the 2026 Fall National Meeting, and Colorado’s expanded Regulation 10-1-1 compliance deadline passed on 1 July 2026.

NAIC and state bulletins · current at September 2026.

So the practical question on any build is not whether governance is required. It is whether the file exists in a shape a market-conduct examiner can read, for the states you actually write in. That is a gap read we run in week one against your state footprint and your model inventory as it stands, and the common finding is not an absent policy — it is a policy with no evidence underneath it.

Responsibility here is non-delegable, which is the whole reason this is engineering work rather than a procurement question. Buying a product does not transfer the obligation. Somebody still has to build the file.

// The stack you actually run

The regulatory surfaces, and where they disagree.

An inventory entry is only useful if it names the surface.

The chips above are not a category list, they are where the regulated decisions actually happen on a carrier’s estate, and each one tiers differently. An extraction model reading an ACORD 140 is a data-quality risk. A model that scores appetite at intake can produce a declination, which is an adverse action in several states. A servicing model writing an effective-dated change into PolicyCenter can create a notice failure. A conversion between CEDE and EDM can move a modeled loss estimate without anybody choosing to. One inventory with one tier across all of that is a document, not a control.

The variations are the work.

A single governance policy written to the NAIC template satisfies the bulletin’s shape and none of the state-specific obligations underneath it. Virginia’s substitution of eliminate for mitigate changes what a pre-deployment test has to demonstrate. Connecticut’s certification means a named officer signs, which means the evidence has to be legible to someone who did not build the model. Iowa’s formal definition of bias and outcomes testing changes what counts as a test at all. Colorado’s outcomes testing is enforceable rather than advisory.

And New York DFS states the position plainly: an insurer cannot rely solely on a third party’s claim of non-discrimination. That sentence is why a vendor-validation record is a build artifact. A vendor’s marketing claim is not evidence, a vendor’s questionnaire response is not a test, and the carrier is the one being examined.

// The build

Four builds, and what each one is measured on.

01

Model inventory with risk tier

Every model in production, its purpose, its decision path, its owner, its risk tier, and the states it operates in. Extraction models and scoring models are separate entries because they carry different tiers, and a model that can stop a submission or decline a claim is tiered on that consequence rather than on its architecture.

Surface: the inventory as a maintained artifact in your own systems, not a document produced for an audit and abandoned.

02

Data lineage to source page and field

Every value a model consumed, traced to the document and page or the system, table and field it came from, including the provenance of every external source. Lineage is the part of the file an examiner can actually test, and it is the part most often asserted rather than built.

Surface: the extraction and integration pipelines, instrumented so lineage is a by-product of running rather than a reconstruction.

03

Pre-deployment testing and the annual re-test procedure

Bias and outcomes testing before deployment, written to the strictest standard among the states in your footprint, with the procedure documented so the annual re-test is repeatable by your team rather than by whoever ran it first. The vendor-validation record sits here too: what was tested, by whom, on what data, with what result.

Surface: a test suite and an evidence pack, versioned alongside the model it tests.

04

Drift thresholds, remediation triggers and the override specification

Thresholds set with a named owner and a defined action, so a drift alert has a person attached rather than a dashboard. And a human-override specification naming who can act outside the model on each decision path, which is both what the bulletins require and what makes the system defensible in a file that gets examined or litigated.

Surface: monitoring and alerting in your infrastructure, with the escalation path and the named authority written into the runbook.

What we measure

Share of production models with a current inventory entry and a tier. Lineage completeness, meaning the proportion of model-consumed fields with a traceable source. Days from a drift trigger to acknowledgment by the named owner. Override rate by decision path, reported rather than minimized, because an override rate of zero on a path that can decline a person is a finding rather than a success. Time to assemble the examination pack, measured by running the assembly rather than estimating it. Gap count against each state in your footprint, closed and open. All of it defined before the model ships.

// We ran one of these

The governance file gets built with the model rather than reconstructed after somebody asks for it, because the request arrives with an examination attached.

The operating record →

// What's hard about this

Two limits, and what we do about each.

One file does not satisfy every state you write in.

The adopted bulletins are not copies of each other and the differences are not cosmetic. Written to the mildest standard, the file fails in Virginia and Connecticut. Written state by state, it becomes an unmaintainable set of documents that drift apart within two quarters, and the drift is invisible until an examination finds it.

So the file is built once to the strictest adopted standard in your footprint, and the state-by-state delta is held as a reviewable table with an owner and a date rather than as institutional memory. When a state moves, and they have been moving every quarter, the change lands in one table and the affected artifacts are named by the table rather than found by a search.

A vendor’s attestation is not a defense, and most of your models have a vendor in them.

New York DFS is explicit that an insurer cannot rely solely on a third party’s claim of non-discrimination, and the practical problem is that the carrier often cannot see inside the vendor’s model to test it. Refusing every vendor model is not a real option; neither is accepting a questionnaire response as evidence.

The workable answer is outcome-side validation. Where the model cannot be opened, it is tested on its outputs against your own book — held-out data, protected-class proxies where the state defines them, and the disparity measures the state’s own definition uses — and the vendor-validation record states plainly what was tested, what could not be tested, and what the carrier is therefore relying on. A file that names its own untested surface is stronger under examination than one that implies there is none.

// What ships with it

Eight artifacts, on every build in every discipline.

Model inventory with risk tierData lineagePre-deployment bias testingAnnual re-test procedureDrift thresholdsRemediation triggersHuman-override specificationVendor-validation record

These do not ship as a separate governance engagement. They are produced by the build that creates the model, they live in your repository and your systems, and they are legible to the officer who has to sign. If a deliverable does not include them, it cannot deploy in Connecticut — which makes the governance file a delivery requirement rather than a compliance afterthought.

What ships for your compliance file →

Bring us the model you have already deployed.

We will tell you where the gap sits against the states you write in and the bulletins each has adopted.