// What we build
You have the data. You do not have a model.
Cloud Data Access streams every InsuranceSuite change into your own bucket. What lands is a replica of the operational schema: wide, deeply normalized, effective-dated, and built for transactions rather than analysis. A replica is not a model.
This is the actual first project on most Guidewire Cloud AI engagements, budgeted or not.
// The situation
Data is the barrier nobody scopes.
Only 22% of insurers have artificial intelligence in production. The named barriers are skills and resources at 52%, data at 40% and technology not delivering at 38%, while 90% of stakeholders describe themselves as pro-AI.
Roots · State of AI Adoption 2025 · vendor research.
Enthusiasm is not the constraint and neither is the platform. The constraint is that the first model has to be trained on a transactional schema never designed to answer a portfolio question, and the person who discovers this is three weeks into a use case sold as eight. So we scope the layer first, then the use case, with the layer specified as a deliverable you own.
// The stack you actually run
The replica, the retained stack, and the gap between them.
Two structural facts about the replica that decide the whole design.
First, it is effective-dated and transactionally versioned, because the source system has to answer what the policy said on a given date as well as what it says now. A model that flattens them produces a premium figure that is neither.
Second, nearly every carrier above roughly $500M of direct written premium runs a hybrid: a modern policy administration system on some lines, and a retained legacy stack holding runoff books, older lines and the fifteen years of history a model needs to be worth training. The problem is rarely the new system. It is that the history lives elsewhere.
// The build
Four builds, and what each one is measured on.
A conformed model over the replica, effective dating intact
Policy, location, coverage, exposure, transaction, claim and party modeled as dimensions and facts a person can query without knowing the operational schema, with version semantics preserved rather than collapsed.
Surface: the Cloud Data Access replica plus EventBridge lifecycle events, out to your warehouse, in your repository.
Version-semantics tests that run on every load
Endorsement, audit, mid-term cancellation and reinstatement each get an explicit test asserting what the model should say before and after — the four transactions that silently break a premium extract.
Surface: tests, transformations and documentation in your repository, running on your infrastructure.
The retained stack, joined rather than referenced
Extraction from AS/400 and DB2 into the same conformed model, with the coding conventions of the older lines documented as reference data rather than tribal knowledge. Scoped as its own project with its own baseline.
Surface: batch extracts from the retained platform into the conformed model, on a documented and owned schedule.
Reconciliation to the numbers finance already publishes
Written, earned and unearned premium tied to the general ledger and the statutory exhibits, with the variance explained line by line before anything downstream reads it.
Surface: model output against the general ledger and the statutory exhibits, on the close calendar.
What we measure
Freshness lag from the operational transaction to the model, per entity rather than as a platform average. Share of premium and exposure reconciling to finance without manual adjustment, and the named cause of every residual. Effective-dated entities covered by an explicit version-semantics test, as a share of those in use. Downstream reports still sourced from a direct read of the replica or from a spreadsheet — the honest measure of whether the layer is used.
“The same premium number has to satisfy an underwriter, a reserving actuary and a statutory filing, which is why reconciliation to finance is in the first scope rather than a later phase.”
// Where this comes from// What's hard about this
Two limits, and what we do about each.
Effective dating will make a naive extract wrong.
Endorsements change limits and add locations mid-term. Audits restate exposure after the fact. Any model touching premium or exposure has to understand your transaction and version semantics, or it produces a confident figure nobody in finance recognizes.
The answer is structural rather than careful: transaction and version semantics modeled explicitly, the four breaking transactions each carrying a test that runs on every load, and reconciliation to the ledger before a single downstream consumer is pointed at it.
If the data behind the workflow is on a retained legacy stack, that is the first project.
Fifteen years of claims and premium data sitting on AS/400, DB2 or COBOL is the training set, and getting at it is neither glamorous nor fast. Every proposal that skips this step is either scoped on the new system only or wrong about its own timeline. So it gets scoped as a project with a name, a baseline and a deliverable, ahead of the use case that depends on it.
// What ships with it
The governance file, scoped to data lineage.
Here the load-bearing artifact is lineage: every model field traced to the source transaction and platform it derives from, with the transformation documented and versioned. That plus a model inventory entry for anything inferential in the layer, pre-deployment testing results, drift thresholds, and a named human decision-maker specification for fields feeding a rating or reserving decision. Lineage is the first thing an examiner asks for when a modeled number is challenged.
Bring us one portfolio question.
We will tell you what it takes to make one premium number satisfy underwriting, actuarial and the filing. How the read works →
