// What we build
The version you were sent is not the one you model.
CEDE, EDM and OED, across the versions actually in circulation rather than the current ones, with a validation report on every conversion naming what was lost, coerced or defaulted. Three schema families, multiple live versions each, and everything downstream inherits whatever came in.
// The situation
Schema versioning is a business blocker, not a nuisance.
Moody’s documents a real case: a reinsurer received Florida exposure in CEDE 8.0, needed EDM to model it, and its conversion tool supported only CEDE 9.0 and 10.0. The options were a months-long tool update with revalidation, declining the business, or surrendering analytical control to somebody else’s conversion.
Moody’s RMS, documented case · publication date not stated in our source.
All three are bad, and the third is the one people take. So the first thing built is the translator itself, held in your repository with its test suite, rather than a one-time conversion you buy again next quarter when the next cedent sends the next version.
// The stack you actually run
Coexisting versions are normal, not a failure.
Multiple format versions coexist in production simultaneously, at the same reinsurer, in the same portfolio, in the same treaty year. That is the ordinary state of the world rather than evidence that a migration went wrong, and it stays that way because a cedent has no reason to upgrade its export on your schedule. Which is why schema version translation is a recurring, budgeted line item at every reinsurer and cat-exposed carrier.
That implies something about how it should be built.
A translator inside a vendor tool inherits the vendor’s supported-version list, which is exactly the constraint that produced the case above. A translator in your repository, with mappings as configuration and behavior pinned by tests, can be extended by your own team the week a new version arrives. The difference matters most when you are deciding whether to write a piece of business.
// The build
Four builds, and what each one is measured on.
Translation across the versions actually in circulation
CEDE to EDM, EDM to OED, OED to CEDE, and the version pairs inside each family, driven by declarative mappings rather than conversion code per pair. Coverage is stated explicitly: which pairs are supported, which are partial, and which fields are unsupported in each direction. An honest coverage matrix beats a claim of universality.
Surface: cedent and program exposure files in, the model exposure set out, in your repository on your infrastructure.
A validation report on every single conversion
Field counts in and out. Values that would not fit the target schema, and what happened to them. Enumerations mapped to a nearest neighbor, both codes shown. Fields defaulted because the target requires them and the source did not carry them. Records dropped, with the reason.
Surface: generated with every run, attached to the exposure set and retained for the treaty year.
Secondary modifiers recovered from where they actually live
Roof shape, deck attachment, covering material, opening protection and year of last roof replacement usually are not schedule columns at all. They sit in narrative appraisals, engineering surveys and inspection reports, in prose. Those documents get read, the modifiers extracted with a citation to the page, and what could not be found is marked absent rather than defaulted quietly.
Surface: the appraisal and survey document set into the exposure record as tagged fields with page-level provenance.
The translator as code you own
Mappings, tests, coverage matrix and documentation in your repository, on your infrastructure, with a round-trip suite that converts a reference portfolio out and back and asserts on the difference. When a cedent sends a version nobody has seen, your team extends the mapping rather than opening a vendor ticket.
Surface: your repository and build pipeline. Build, operate, transfer — measured on your team running it without us.
What we measure
Hours per submission from cedent file to model-ready exposure, baselined against your own range rather than an industry figure. Share of inbound files convertible without manual intervention, by source family and version. Fields lost, coerced and defaulted per conversion, as a rate rather than a log nobody reads. Secondary-modifier completeness before and after narrative extraction. Days from receipt to filing-ready aggregate. Round-trip fidelity on the reference portfolio, per release. Definitions and baselines, set in the first week.
“The validation report is a required output of every conversion here rather than an optional check somebody can switch off, because the aggregate has to reconcile to a statutory filing.”
// Where this comes from// What's hard about this
Two limits, and what we do about each.
The modifiers that move modeled loss are usually not schedule columns.
Roof shape, deck attachment, covering material, opening protection and first-floor elevation are among the strongest drivers of modeled wind and water loss, and are frequently absent from every structured file in the chain. They live in narrative appraisals and engineering surveys, one building at a time. A translation that only moves columns faithfully carries forward the fact that these fields are empty, and the model then substitutes conservative regional defaults on the exposures where precision mattered most.
So the build reaches into the narrative sources where they exist, extracts the modifiers with a citation to the page, and records confidence per field. Where the document does not contain the answer, the field is marked absent and reported in the completeness summary rather than filled with a plausible value. The output names which exposures are modeled on defaults, ranked by share of total insured value.
Translation cannot improve the quality of the underlying data.
A CEDE file with missing construction classes converts cleanly into an EDM file with missing construction classes. The conversion is not the problem and cannot be the solution. Worse, a successful conversion produces a file that looks native to the target platform, removing the last visible clue that its contents arrived from somewhere else with gaps in them. That is the failure mode here: format fidelity reading as data quality.
So the validation report distinguishes three things usually collapsed into one: what was extracted from the source, what was inferred and at what confidence, and what was defaulted because the target schema demanded a value the source never had. Those three categories travel with the exposure set into the model run and into the aggregate, so a modeled number can be read alongside the share of its inputs nobody ever observed.
// What ships with it
The governance file, scoped to inferred exposure.
Data lineage from the source file and version through every mapping step to the model field. A model inventory entry, with its risk tier, for anything that infers a modifier from narrative text. Pre-deployment testing results, drift thresholds with remediation triggers, and a named human decision-maker specification for the defaults a rate or a ceded recovery rests on. The retained validation report is what answers how a modeled number was arrived at, the question asked when an aggregate is challenged. Roughly half the states have adopted the NAIC AI Model Bulletin, and the variations are material.
Bring us one cedent file.
We will tell you what your conversions are inferring, and what the validation report has to report. How the read works →
