
Undisclosed Specialty Insurer
Insurance / AI Discovery / Opportunity Assessment / Data Readiness / Regulated Industries
Executive brief
A specialty commercial insurer had board approval for an AI program and a shortlist assembled from vendor demos, but no evidence that the shortlisted workflows were the expensive ones or that the underlying data could support them. Over six weeks we baselined the real work, sized the opportunities, tested the hardest technical unknowns against production-shaped data, and reviewed each candidate with compliance in the room. The engagement ended with one funded build, one governance workstream, seventeen documented deferrals, and success criteria the sponsors had agreed to in advance.
Candidate use cases assessed
19 → 1 funded
6 weeks
Build budget released before evidence
$0
across the engagement
Executive top-three assumptions disproved
2 of 3
weeks 3–5
Mapped 19 candidate AI use cases against real workflow cost, data availability, and regulatory exposure — then scored and ranked them in the open with the sponsoring executives.
Ran three two-week feasibility spikes against production-shaped data, which disproved two of the executive team’s top-three assumptions before any build budget was committed.
Found that the highest-value candidate (automated claims triage) was blocked by an unresolved data-retention question, and rerouted it to a governance workstream instead of a build.
Produced a funded build brief for submission intake triage with agreed success criteria, a measurement plan, and an explicit stop condition.
Engagement shape
Duration
6 wks
Phases
4
Deliverables
7
Candidate use cases assessed
19 → 1 funded
How it ran
Three overlapping evidence workstreams converging on a single scored decision session
A six-week discovery engagement scored 19 candidate AI use cases against evidence — data, workflow cost, and regulatory exposure — and sent exactly one of them to a funded build.
Starting position
The carrier had a board-approved AI budget and a shortlist of initiatives assembled largely from vendor demonstrations. Nobody could say which workflows actually cost the most, whether the data behind the shortlisted ideas existed in usable form, or what a successful outcome would look like twelve months later. Two prior AI pilots had been declared successful and then quietly abandoned, and the executive team wanted to understand why before spending again.
Constraints we worked inside
5 in playEvery one of these ruled an easier option out. They are the reason the approach looks the way it does.
A shortlist that came from vendor conversations rather than from any measurement of internal workflow cost.
Underwriting and claims data spread across a policy administration system, a document management system, and roughly a decade of email and shared drives.
Regulatory obligations around explainability, record retention, and third-party data processing that had never been mapped to an AI context.
Two abandoned pilots that had left the business skeptical, so the engagement had to produce evidence rather than enthusiasm.
Subject-matter experts available for interviews only in short blocks during renewal season.
What had to be true to finish
5 goalsAgreed up front, so the engagement could be called finished on evidence instead of opinion.
Establish what the highest-cost, highest-friction workflows actually are, measured rather than asserted.
Determine whether the data required by each candidate use case exists, is accessible, and is of usable quality.
Test the hardest technical unknown behind the leading candidates before committing build budget.
Map each candidate against regulatory and model-governance obligations with compliance participating, not reviewing afterwards.
Deliver a ranked roadmap with agreed success criteria, sizing, and an explicit stop condition for each funded item.
Industry context
Insurance
AI Discovery
Opportunity Assessment
Data Readiness
Regulated Industries
Specialty commercial insurance runs on judgement-heavy, document-heavy workflows — submissions, endorsements, claims, and broker correspondence — with regulators, reinsurers, and auditors all entitled to ask how a decision was reached. That makes it fertile ground for AI and unusually unforgiving of AI deployed without a defensible record of why it was deployed.
Approach
We ran discovery as three overlapping workstreams — workflow evidence, data evidence, and technical evidence — converging on a single scored decision session with the sponsoring executives. Nothing advanced on the strength of an interview alone; every candidate that reached the shortlist had to survive a data check and, for the top candidates, a working spike.
Delivery track
Segment width reflects the number of workstreams inside each phase — where the engagement actually spent its effort.
Phase 1 — Baseline the work (Weeks 1–2)
4 workstreams
Interviewed 22 practitioners across underwriting, claims, broker services, and operations, structured around observed work rather than opinions about AI.
Shadowed submission intake and first-notice-of-loss handling to timebox the real steps, including the rework and chasing that never appears in process documentation.
Pulled 18 months of workflow telemetry from the policy administration system to establish volumes, cycle times, and touch counts.
Built a cost-per-workflow baseline in fully loaded hours, which reordered the executive shortlist before any AI question was asked.
Phase 2 — Test the data (Weeks 2–4)
4 workstreams
Profiled the source systems behind each candidate: field coverage, document formats, label availability, retention rules, and lineage.
Sampled 900 historical submissions and 400 claims files to measure extraction difficulty against real artefacts rather than clean examples.
Documented where ground truth existed for evaluation and where it would have to be created — a cost most candidates had not accounted for.
Flagged two candidates as data-blocked and one as blocked pending a retention and reprocessing decision with legal.
Phase 3 — Spike the unknowns (Weeks 3–5)
4 workstreams
Ran three timeboxed two-week feasibility spikes against production-shaped data in an isolated environment.
Spike A confirmed that submission documents could be extracted and normalized to a reviewable structure at usable accuracy.
Spike B showed that broker-email intent classification was accurate enough in aggregate but unstable on the long tail that generated the actual cost.
Spike C established that the claims triage candidate would require historical labels the carrier could not lawfully reprocess without a retention decision.
Phase 4 — Govern and rank (Weeks 4–6)
4 workstreams
Reviewed every surviving candidate with compliance, legal, and the model risk function against explainability, retention, and third-party processing obligations.
Scored all 19 candidates on value, data readiness, technical confidence, and regulatory exposure using criteria agreed in week one.
Ran a live decision session where sponsors saw the scores, the evidence behind each one, and the dissenting views.
Wrote success criteria, a measurement plan, and a stop condition for the single funded candidate before the engagement closed.
What the team kept
Artefacts handed over at the end of the engagement — owned and operable by Undisclosed Specialty Insurer without us.
Workflow cost baseline covering 9 processes in fully loaded hours, touch counts, and cycle times
Data readiness assessment per candidate use case, with named blockers and owners
Three feasibility spike reports, including the two that returned negative results
Opportunity scorecard for all 19 candidates with the scoring rubric and evidence trail
Regulatory and model-governance review mapped to each candidate
Funded build brief for submission intake triage: scope, success criteria, measurement plan, and stop condition
Twelve-month sequenced roadmap that any delivery partner could execute
Results
The carrier left discovery with less on its roadmap than it started with, and considerably more confidence in what remained. One candidate was funded with agreed success criteria, one was rerouted into a governance workstream, and the rest were documented and deferred with reasoning the CFO could defend.
Outcome ledger
Every figure we measured on this engagement, including the ones that are ranges rather than headlines.
Candidate use cases assessed
One candidate funded, one rerouted to governance, 17 documented and deferred.
19
1 funded
Build budget released before evidence
Discovery cost roughly 4% of the approved program budget.
$0
across the engagementExecutive top-three assumptions disproved
Both failed on data availability rather than on model capability.
2 of 3
weeks 3–5Submission intake cycle time baselined
No baseline existed before discovery, so no prior pilot could prove impact.
11.4 days
median, 18-month sampleTime to a funded, evidence-backed decision
The prior two pilots ran 7 and 9 months without producing a decision.
6 weeks
end to endData blockers surfaced before build
Three had named owners and dates by the close of the engagement.
6
weeks 2–4What changed day to day
The part of the result that never shows up in a dashboard, but is the reason the numbers held.
The executive team could explain, with evidence, why the AI roadmap got shorter — which ended the credibility problem left by the two abandoned pilots.
Compliance moved from a late-stage approval gate to a participant in scoring, removing the review cycle that had stalled earlier work.
The funded build started with an agreed baseline and a stop condition, so its success or failure would be legible either way.
The workflow cost baseline outlived the engagement and became the reference point for later automation decisions unrelated to AI.
Your turn
A six-week discovery engagement scored 19 candidate AI use cases against evidence — data, workflow cost, and regulatory exposure — and sent exactly one of them to a funded build. If that shape looks familiar, the first conversation is a working session, not a pitch.
Workflow Mapping & Process Baselining
Opportunity Sizing & Prioritization
Data Readiness Assessment
Technical Feasibility Spikes
Risk, Compliance & Model Governance Review
Success Criteria & Measurement Design
What Undisclosed Specialty Insurer got
Candidate use cases assessed
19 → 1 funded
6 weeks
Build budget released before evidence
$0
across the engagement
Executive top-three assumptions disproved
2 of 3
weeks 3–5