AI prototyping for insurers: testing an AI feature inside a regulated product

Articles › AI prototyping

AI prototyping for insurers means building a rough, working version of an AI feature and watching real people use it, using dummy data and a human check on every output, before the business case is written or the vendor is chosen.

This is for product, claims and operations people at insurers who have an AI idea sitting in scoping. Underwriting triage, claims first notice of loss, policy servicing replies, document checking. The idea has been discussed. Nothing has been tested.

Here is the short way past that: build a rough working version, put it in front of the people who do the job, and watch what happens. Days, not a quarter. No production data, no vendor contract, no committee.

At the end you have a working version people have used, a list of where it broke, and a number you can compare against later. That is usually more convincing to a risk committee than an estimate in a document.

The meeting that keeps repeating

The pattern looks the same in most insurers. Someone writes up an AI idea in a PowerPoint. It goes to a meeting. Compliance asks a fair question about data. IT asks which model. Someone asks for costings. The slides go back for another pass, and the same meeting happens in three weeks with more detail and no more certainty.

Nobody in the room has seen the feature work or fail. So every question is answered with an opinion.

Meanwhile the underwriter is still copying details out of a PDF by hand, and the claims assessor is still waiting on the same document she chased last week.

The cheapest way out is to build a rough version and let it be wrong in front of people. We wrote about this more fully in Build a rough working version of an AI feature instead of writing the business case first. The short of it: building the thing shows you what the slides left out.

Regulated does not mean you cannot test

The usual objection is data. The answer is simple. You do not need real policyholder data to find out whether an AI feature helps.

Make up twenty policies and twenty claims that look like yours. Same fields, same messy handwriting on the scanned forms, same missing documents, same Tagalog and English mixed in the customer notes. Feed those to the prototype. If the feature cannot handle your made-up mess, it will not handle the real thing.

Two rules keep this safe. Nothing leaves the prototype without a person approving it, and no live customer record goes in during the build. That is enough for most internal pilots.

Put compliance or risk in the room while you build, not in a review afterwards. When they watch it run, they usually tell you the one control that matters, and you build it in then instead of rebuilding it much later.

Prototype the task, not the product

People try to prototype the whole claims system. Then nothing gets built.

Pick one task with a clear start and finish. "Read the submitted documents and tell the assessor what is missing." "Draft the reply to a policy change request for the agent to check." "Sort overnight submissions into simple and complex."

One task, one screen, one output. Give it to someone who does that task every day and say nothing while they use it. Do not explain it. If it needs explaining, that is your first finding.

Expect the first version to be wrong somewhere you did not plan for. Usually it handles the clean cases and falls over on the exceptions, which is where all the time goes anyway. That is useful. You now know the feature is worth building only if it handles exceptions, and you know that before signing anything.

Get a before number, or you will be shrugging later

Most teams we work with have no baseline. The pilot runs, the boss asks whether it worked, and the honest answer is a shrug.

Take the number before you build. Turnaround time on that one task. How many cases go back for rework. How many need a second person. Twenty or thirty recent cases is enough. An hour of someone's afternoon, not a measurement project.

Then you can say something specific afterwards, and small honest numbers hold up better than big vague ones.

For what measurement looks like when the thing actually ships: a large insurer ran a series of five-day design sprints on customer-facing web journeys. Journeys that used to take six months or more were designed, tested with customers in two weeks, and live in four. Completion rates rose 80 per cent and the old drop-off points disappeared. The client measured it. They could only say that because someone wrote down the old numbers first.

Who needs to be in the room

Four kinds of people. Someone who does the task daily, usually a claims assessor, an underwriter or branch staff. Someone from compliance or risk. Someone who can build. Someone who can decide.

Leave out anyone who is only there to report back.

A mixed room finds more than a room of specialists. The assessor spots the exception nobody modelled. The compliance officer spots the log that has to exist. The developer says which bit is two hours and which bit is two weeks. We run this shape in design thinking training for cross-functional bank teams and the same mix holds for insurers.

When the prototype is worth keeping, the build carries on with your own developers rather than being handed over. That is described in software development for banks and insurers in the Philippines.

What to do this week

  • Pick one task, not a system. Write it in a single sentence that names who does it.
  • Measure it now. Turnaround time or rework rate on the last twenty or thirty cases.
  • Make twenty fake policies or claims that include your worst real-world mess.
  • Book a session with one person who does the task, one from compliance, and one who can build.
  • Build a rough version, watch someone use it in silence, and write down the first three things that broke.

More on AI prototyping

Questions people ask

Can we prototype an AI feature without touching policyholder data?
Yes. Use made-up policies and claims that look like the real thing. A prototype only has to be real enough for someone to try the task, so live data adds risk without adding learning.

What do we measure to prove it worked?
Take a before number on the same task first, usually turnaround time or how many cases need rework, measured on twenty or thirty recent cases. Then run the same measure after the pilot.

Who should be in the room?
The people who do the work daily, one person from compliance or risk, and someone who can build. Compliance in the room early stops the rebuild later.

Is a prototype enough to get budget approved?
It is usually stronger than a document, because the finance or risk sponsor can watch a claims assessor use the thing and see where it fails, rather than reading an estimate.

On-Off Group trains teams, tests products with real customers, finds where a transformation has stalled and builds what gets it moving, for banks, insurers and enterprises in the Philippines, since 2015. How we help with ai prototyping.