The Process You Wrote Down Is Not the Process You Run

Kyle Nakatsuji·July 1, 2026·9 min read

The next pilot that fails won't fail because the model was wrong. It'll fail because it was aimed at a process that exists in your documentation but not in your operation.


The short version (by humans, for busy humans)

We know how reading works now: you skim the top, delegate the rest, move on. Fair. This part is for you. Under 400 words. The full version is below if you want the depth — or let your AI read it.

The actual point:

Most AI pilots fail because they're aimed at the wrong target. Every carrier has a documented process — a SOP, a workflow diagram, a wiki entry that describes how claims get handled or how a submission gets underwritten. Almost every AI project starts there. The vendor reads the documentation. The internal team describes the steps. The pilot gets built to match.

The documentation is wrong. Not because anyone lied about it, but because real operations drift away from their paperwork the moment they meet reality. A new state gets added and the workaround never gets written down. A regulation changes and the adjuster adapts. A system integration breaks and someone builds a manual fix that becomes the de facto process for the next three years. The gap between the documented process and the lived process compounds quietly until it's the normal operating state.

Build automation against the documented process and you automate the share of volume that follows the documentation. You break on the share that doesn't. The cases that don't follow the documentation are not random — they're the complex claims, the unusual risks, the multi-state policies. The large losses, the litigation exposure, and the most senior judgment were already concentrated there.

The fix happens before the model, not after it. Map the real workflow first. Sit with the people who do the work. Watch what they actually do, not what the SOP says they do. This is process archaeology, not technology selection. It takes weeks, not years. Almost every failed pilot skipped it.

That's the post. If you're approaching an AI deployment — or diagnosing a pilot that didn't hold — that mapping work is how DL engages with new carrier partners. If it sounds relevant, reach out.


The full version (human on the loop — for depth, yours or your AI's)

Opening scenario

The email came in on a Tuesday at 6am. Subject: "AI Pilot — Update Request." The VP of Claims at a regional carrier had sent it to herself six months earlier with a calendar reminder: check in on production performance.

She already knew what the update would say. The demo had been exceptional. Sixty percent of incoming claims triaged automatically in under two minutes, accuracy well above the manual baseline, the vendor's project team in the room applauding. She greenlit the production deployment. The business case penciled out. Headcount from the first-notice-of-loss team was reallocated. The pilot was declared a success.

What happened after the cameras left was harder to explain.

The automated triage handled the clean volume beautifully. But the claims that landed outside the clean paths — complex injuries, multi-state policies, coverage questions that required judgment on the facts — went somewhere the automation hadn't been designed to send them. The queue for those claims grew. The senior adjusters who used to absorb the exceptions were no longer in those roles. Cycle times on the hardest files deteriorated. Two claims that would have been closed in 30 days were now approaching 90.

The pilot was still running. The vendor still called it a success. On aggregate metrics, it probably was. But the VP of Claims was sitting with a problem nobody had anticipated: her operation was now worse on the cases that mattered most.


The problem

What happened at that carrier happens regularly. The pattern has a name — the conformance gap — and it is the most underestimated risk in any carrier AI program.

Every operation has a documented process. Carriers are particularly diligent about documentation: SOPs, workflow diagrams, state-specific procedures, compliance-approved decision trees. The documentation serves real purposes — training, auditing, regulatory filings. It represents how work is supposed to happen.

Real work looks different. Operations drift from their documentation the moment they encounter the full range of reality. A new state gets added to the book and the handling protocol changes, but the workaround never makes it into the formal SOP. A regulation shifts and the adjuster adapts on the fly. A system integration breaks and someone builds a manual bridge — a spreadsheet, an email thread, a standing agreement with a specific vendor — that becomes the real process for years. Over time, the documented process and the lived process diverge until the gap is the normal operating state.

One widely-shared industry analysis estimated that a meaningful share of cases — roughly a third in typical operations, higher in exception-heavy workflows like complex claims — don't follow the documented process. The number is almost certainly understated for most carriers, because the people doing the estimation are often the same people who wrote the documentation.

Insurance feels this harder than most industries. That requires an honest explanation, because the conformance gap is not an insurance disease — banks have it, hospitals have it, any operation that runs on documented procedures and human exception-handling has it. Insurance feels it harder because of exception density.

Insurance work is judgment under uncertainty applied to highly variable inputs. A claims operation spans first notice, scene reports, adjuster notes, medical records, injury reports, reserves, and reinsurance attachment. An underwriting operation spans loss runs, statements of value, bordereaux, agent submissions, appetite documents, and jurisdiction filings. None of those inputs has a compiler that tells you in milliseconds whether the work was done correctly. The "right answer" on a complex claim takes a senior adjuster hours to confirm. The share of work that lives in exceptions — the share the documentation fails to fully describe — is higher in insurance than in most industries.

Which means that when you build automation against the documented process, you're often building against a description that accurately characterizes the routine majority and almost completely ignores the exception-dense minority. The minority is where the dollar exposure lives.


Here's what we learned

We've built and run production AI inside a live P&C carrier for nearly a decade — in auto, the most structured and documented line in insurance. Even there, the conformance gap was real. We had to map the actual workflow before any of the automation held. Here's what that experience taught us.

1. The document describes the process you wish you ran

The SOP isn't a lie. It's a snapshot that stopped being updated the moment operations met reality.

Documentation gets written at the moment a process is designed. At that moment, it's accurate. But processes encounter reality in motion: new states, new regulations, new system behaviors, new edge cases. The people doing the work are fast and adaptive. They find workarounds, develop judgment shortcuts, build informal protocols that handle the exceptions the SOP didn't anticipate. The documentation stays static while the operation evolves.

Operations move faster than documentation can follow. That's physics, not negligence. In any large claims or underwriting operation, the actual process exists as a hybrid: the written procedure handles the clean paths, and the institutional knowledge in the people doing the work handles everything else.

One widely-shared industry analysis found that in exception-heavy workflows, the gap between documented and actual process could run considerably higher than in routine operations — the more complex the work, the wider the conformance gap.

What this means for your Tuesday morning: When someone tells you your documentation is a reliable starting point for an AI deployment, ask when it was last validated against actual observed behavior. Not against policy changes — against what the people doing the work actually do. The answer is usually illuminating.

2. The cases outside the documented flow are where your losses live

Build against the documented process and you automate the wrong work.

The automation handles the documented cases well. That's not an accident — those are the cases the documentation describes. The clean majority goes through the automated path, metrics look strong, the pilot gets declared a success. Then the exceptions arrive.

The cases that fall outside the documented process are not random. They're the ones that required judgment. Complex injuries. Coverage questions with no clear answer. Claims with unusual policy forms. Multi-state fact patterns. These are exactly the cases where the largest loss dollars sit, where the regulatory exposure is highest, and where the most senior human judgment was already being applied. The documented process didn't describe them not because they're rare — they're not — but because they're hard. Writing down how to handle a case that requires judgment is difficult because the answer is: it depends.

When automation handles the routine volume and routes the complex cases into a path designed for the documented workflow, one of two things happens. Either the complex cases fail through the automation and create manual exceptions nobody planned for, or they route into a queue that the humans who used to handle them are no longer positioned to absorb. The pilot that looked successful in aggregate has made the hardest cases harder.

There is a second-order problem worth naming, because insurance executives think in second-order effects. When the pilot's business case books headcount savings and those savings come from the team that used to handle exceptions, the capacity quietly holding the hard cases together is gone. The backlog grows. Cycle times on the hardest files deteriorate. The pilot didn't just fail to improve the operation. It removed the safety valve.

What this means for your Tuesday morning: Before you finalize the headcount assumptions in a pilot's business case, map where your exception cases currently go and who handles them. That's the capacity you can't cut until the automation can actually absorb them — and it won't absorb them until you've mapped the gap.

3. The fix happens before the model, not after it

Carriers that get AI to production do one thing the others skip: they map the real workflow before anyone touches a model.

This means sitting with the people who do the work and watching what they actually do. Not reading the SOP. Not interviewing managers about the process. Watching adjusters work through a morning's claims queue. Sitting with underwriters as they review a stack of submissions. Building a picture of the real process — including the undocumented exception paths, the informal protocols, the judgment calls that happen below the documented surface.

This is process archaeology. It produces a map of where the documented process and the real process diverge. That map is the most valuable input to any AI deployment, because it tells you three things the SOP cannot: where automation pays off, where human judgment has to stay, and where deterministic rules beat a model.

Two reasons this step gets skipped, and both are worth naming because any carrier executive will recognize them.

First, most AI vendors won't do it. Their competency is the model and the platform. Process archaeology means sitting in claims rooms, talking to adjusters, mapping exception paths — slow, unglamorous work that doesn't produce a demo. So vendors take the SOP at face value and build to it. If you want this work done, you have to either do it yourself or find a partner who does it as a first step.

Second, most internal teams skip it because they believe they already know the workflow. They wrote the SOP. They've run the operation for years. But the people who own the documentation are usually not the people living the exceptions. The people living the exceptions are the adjusters and underwriters doing the work — and their knowledge of how the work actually runs often never makes it back into any document.

"We already know our process" is the most expensive assumption in any carrier AI program.

A note on scope: this is a bounded engagement, not a two-year data-readiness project. Mapping the real workflow in a specific claims or underwriting function takes weeks. It produces a clear picture of where the conformance gap is widest and where automation is safe to deploy. That picture changes everything downstream — model selection, integration design, headcount planning, and the business case.

What this means for your Tuesday morning: If you're evaluating a vendor and they haven't asked to sit with your adjusters or underwriting team before proposing a solution, that's your signal about whether they understand your actual workflow. The SOPs they've read describe the process you wish you ran.

So what

The next pilot that fails won't fail because the model was wrong. Carrier AI models have gotten very good. The next failure will come from the same source as the last one: the pilot was aimed at a process that exists in documentation but not in practice, and the exceptions — the hard, expensive, judgment-laden cases — went somewhere the automation couldn't handle them.

Map the gap before you build. That's the sentence that separates the carrier AI deployments that hold from the ones that don't. It's not about the model. It's not about the vendor. It's about knowing what you're actually automating before you automate it.

The carriers that get to production did this. They treated workflow mapping as a funded phase, not an assumption. They found out where their documented process and their real process diverged before they discovered it in production.

The ones that didn't are the ones with the demo that worked and the backlog that didn't.


One more thing

If this matches a decision you're working through — or a pilot you're trying to diagnose — the mapping work described above is the first phase of how DL engages with new carrier partners. We call it Discovery. It's a bounded engagement, roughly six weeks, at the end of which you have a clear picture of where your documented processes diverge from your real ones, and a concrete view of where AI deployment is low-risk and where it isn't.

It's a small bet against a large risk. If it's the right moment to have this conversation, reach out.


Kyle Nakatsuji is CEO of Clearcover and founder of Dearborn Labs. Dearborn Labs partners with P&C carriers navigating the AI shift — not as a vendor, but as an embedded operator.

// Key Questions

What is the conformance gap in insurance AI?

The conformance gap is the distance between a carrier's documented process — its SOPs, workflow diagrams, and compliance-approved decision trees �� and the process its people actually run day to day. Operations drift from their documentation the moment they meet reality: new states, changed regulations, broken integrations, and manual workarounds that never get written down. The gap is widest in exception-heavy work like complex claims, which is exactly where the largest losses and the most senior judgment sit, so building automation against the documentation quietly ignores the most expensive cases.

Why do carrier AI pilots break on the cases that matter most?

Carrier AI pilots break on the hardest cases because they are built against the documented process, which describes the routine majority and barely describes the exceptions. The complex claims, unusual risks, and multi-state fact patterns fall outside the documented flow, so the automation either fails through them and creates unplanned manual work, or routes them into a queue the humans who used to handle them can no longer absorb — especially after a business case reallocated that headcount. The aggregate metrics still look good while cycle times on the highest-exposure files deteriorate.

How should a carrier prepare its workflow before deploying AI?

A carrier should map the real workflow before touching a model by sitting with the people who do the work and observing what they actually do, rather than reading the SOP or interviewing managers. This process archaeology surfaces the undocumented exception paths, informal protocols, and judgment calls that live below the documented surface, and it produces a map of where automation pays off, where human judgment must stay, and where deterministic rules beat a model. It should be treated as a funded phase, not an assumption.

Why is the documented process usually inaccurate?

The documented process is usually inaccurate because it is a snapshot taken when a process was designed and then largely frozen, while the operation keeps evolving. People doing the work adapt to new states, new regulations, and broken systems faster than documentation can follow, building workarounds and judgment shortcuts that never make it back into the SOP. The result is a hybrid where the written procedure handles the clean paths and institutional knowledge in the workforce handles everything else — which is why 'we already know our process' is such an expensive assumption.

What is process archaeology in an AI deployment?

Process archaeology is the practice of reconstructing how work actually happens by observing practitioners directly instead of relying on documentation. In an insurance AI deployment it means watching adjusters work a claims queue or underwriters review submissions, mapping the exception paths and informal protocols, and identifying where the documented process and the real process diverge. That divergence map is the most valuable input to the deployment because it drives model selection, integration design, headcount planning, and the business case.

How long does workflow mapping take before an AI deployment?

Workflow mapping is a bounded engagement measured in weeks, not a multi-year data-readiness project. Dearborn Labs runs this as a Discovery phase of roughly six weeks, at the end of which a carrier has a clear picture of where its documented processes diverge from its real ones and a concrete view of where AI deployment is low-risk and where it isn't. It is a small, time-boxed investment made before model work begins, precisely because that picture changes everything downstream.

Share
← Back to Insights