// What we build
Every broker's schedule is a different schema.
Merged cells. Embedded subtotals sitting in data rows. One row describing four buildings. Footnotes inside value fields. Column headers where BLDG_VAL, BldgValue and RCV Building all mean the same thing, and sometimes do not. Arbitrary Excel in, one canonical location model out.
The schedule is not an input to property underwriting. On most accounts it is most of the underwriting.
// The situation
The correction gave back rate, not attachment.
Property catastrophe reinsurance has corrected on price, but attachment points stayed where the 2023 restructuring put them. Your own treaty tells you this faster than any market commentary: the rate moved, the retention did not.
Which means the primary carrier still absorbs secondary-peril frequency, and what decides how much is a set of fields on a spreadsheet somebody typed. So the first output here is not a cleaner file. It is a list of the fields in your own schedules that are blank, contradicted or unresolvable, ranked by how much modeled loss each one moves.
// The stack you actually run
The artifacts, the models, and the unpriced concentration.
Value is radically concentrated. Attention is not.
Concentration on a schedule is usually far sharper than the page count suggests. On a 150-property schedule it is unremarkable for 2 locations to carry 10% of total insured value, 6 to carry 30%, and 16 to carry half of it. Underwriters spend attention evenly across a schedule where value is not, because nothing in the file says where the value went. A normalized schedule says so in the first paragraph of its output, and that ranking is worth more than most of the fields it came from.
And geocode resolution is the single input that most changes a coastal loss estimate.
A ZIP-centroid geocode on a coastal risk produces a fundamentally different loss estimate than a rooftop geocode, and nothing in a standard exposure file records which one you got. Resolution is a data-quality fact that belongs in a field, not an implementation detail of whichever geocoder ran that day.
// The build
Four builds, and what each one is measured on.
Arbitrary schema to canonical location model
Merged cells unmerged with the intended row scope recorded. Embedded subtotals identified as subtotals rather than counted as locations — the most common way a total insured value gets overstated. Multi-building rows split, with the split recorded. Footnotes stripped out of value fields and kept as notes on the field they contaminated.
Surface: the broker email and document store, out to the canonical location model in your policy administration system.
Geocode resolution recorded as a field
Addresses corrected and standardized, resolution stored per location as rooftop, parcel, street or ZIP centroid, and multi-building campuses reconciled to policy records. The exposure set carries its own resolution profile, so a modeled number can be read alongside the quality of the geocoding underneath it.
Surface: your address and location service, written back to policy location records and the model exposure set.
COPE gap filling with provenance on every filled field
Aerial imagery and property intelligence fill roof condition, geometry, material and degradation where they genuinely can, and every filled field is tagged with its source, vintage and confidence. Blank fields make the model substitute conservative regional defaults, and the insured then pays a data-quality penalty inside the premium that neither side can quantify or argue about.
Surface: imagery and property-intelligence APIs, written back as tagged fields, versioned per refresh.
The concentration and quality report that ships with the file
Share of total insured value in the top 2, 6 and 16 locations. Share of fields extracted, inferred and defaulted. Geocode resolution profile. Locations where the schedule total and the ACORD 140 figure disagree, both values shown. One page, generated on every run, and the page an underwriter reads first.
Surface: generated with each schedule and attached to the submission record, versioned per revision received.
What we measure
Hours per submission from broker file to model-ready schedule, baselined against your own range rather than an industry figure. Percentage of locations at rooftop resolution versus ZIP centroid, before and after. Share of fields extracted, inferred and defaulted, per schedule rather than in aggregate. Count of subtotal rows previously counted as locations, because that error has a direction and the direction is upward. Share of total insured value in the top 2, 6 and 16 locations. Definitions and baselines, set in the first week.
“What decides a rate here is a pipeline. A blank field becomes a conservative default, the default becomes a price, and somebody has to explain to a regulator how the field got filled.”
// Where this comes from// What's hard about this
Two limits, and what we do about each.
Imagery only sees what is visible from above.
Sprinkler type and coverage, alarm grade, central-station monitoring, interior occupancy, equipment bracing, and business-interruption dependencies on a single supplier or utility feed are all invisible to aerial data, and several move loss more than anything a camera can see. The highest-value judgment on a property risk is the part computer vision cannot reach, and a gap-filling system that does not say so will produce a full schedule made partly of guesses.
So the build fills what imagery can genuinely fill and stops there. Unreachable fields are marked unreachable rather than inferred, and each is routed to the artifact that can answer it: the engineering survey, the appraisal, the loss-control visit, or a question back to the broker. The output reports what share of the schedule is beyond imagery, which is the number to read when spending survey budget.
A cleaned schedule looks authoritative regardless of whether the values were ever right.
This is the specific failure mode of artificial intelligence in exposure work: normalization raises confidence faster than accuracy. Insurance-to-value drift is a human and appraisal problem, and no model can tell you an insured’s reported building value is five years stale. A margin clause or an Occurrence Limit of Liability Endorsement then caps recovery at a multiple of that stale figure, which is how a clean-looking file becomes a coverage argument at partial loss.
So the confidence signal is built alongside the data rather than added afterward. Every value states whether it was extracted, inferred or defaulted. Every filled field carries its source, vintage and confidence. Schedules not refreshed against replacement cost are flagged as drifted, with the coinsurance exposure quantified at partial loss so the gap is a number rather than a worry.
// What ships with it
The governance file, scoped to inferred exposure values.
A model inventory entry for every gap-filling model, data lineage from the imagery or property-intelligence vendor to the specific schedule field, pre-deployment testing results, drift thresholds with remediation triggers, and a named human decision-maker specification. Here the load-bearing artifact is the provenance record on inferred values, because that is what gets asked about when a modeled loss estimate or a rate is challenged. Roughly half the states have adopted the NAIC AI Model Bulletin, and New York’s Department of Financial Services states the position plainly: an insurer cannot rely solely on a third party’s claim of non-discrimination.
Bring us one schedule.
We will tell you which fields on the schedule are blank, contradicted or unresolvable, ranked by what they move. How the read works →
