How to build a prototype with AI in one evening
Pick one task, write the smallest version of it you can build, use an AI tool to generate a working screen and some fake data, then watch two or three people try it the same evening. Write down what broke before you change anything.
This is for anyone who wants one AI feature in front of real users this month and keeps ending up in another meeting about scope instead.
You can build a working prototype with AI in one evening. Something a person can open and type into, and get stuck on while you watch.
Below is the order On-Off Group uses on these builds. It works because the evening is short enough that you have to cut the idea down to one task.
Scope meetings cannot answer the questions you have
A normal week looks like this. Someone writes the idea up in a PowerPoint. The next meeting is about what it should include. Then a meeting about which team owns it. Weeks in, nobody has seen the thing working, and the arguments are still about the slides.
Most of those arguments cannot be settled by talking. They are questions about what happens when a customer types something odd, or what the AI answer actually looks like when the data is messy. Building the thing answers them.
There is a longer piece on building first instead of writing the business case first: Build a rough working version of an AI feature instead of writing the business case first. The short version: an idea can fall apart in one evening, or sit safely in a deck for months.
Cut it down to one task before you open a tool
Write the task in one sentence, in the customer's words or the staff member's words. "A branch officer checks whether a submitted document is complete." "A customer asks why their claim is still pending."
One task. One user. One outcome you can watch someone reach or not reach.
Then strip out everything that is not that task. No login. No settings. No admin view. No integration with the core system. Fake data typed by hand is fine, and it is usually better, because you can make the messy cases on purpose.
If you cannot get the task into a sentence, the evening goes on setup and logins, and the feature never gets built.
Write the prompt as instructions the tool can follow
When people ask an AI tool to "build a claims assistant", they get something generic and then spend an hour arguing with it.
Give it the shape instead. Say what the screen has on it. Say what the user types. Say what happens when they press the button. Say what the data looks like, and paste an example row.
Then build in small steps. One screen working. Then the input. Then the AI response. Test each step before asking for the next one. When something breaks, tell the tool exactly what you saw, not "it doesn't work".
For the AI part itself, write the instruction the way you would brief a new hire: what it should do, what it must never say, what to do when it does not know. That written instruction is what the feature actually does. Most of your evening should go there, not on the visual design.
Watch two or three people try it the same evening
This is the step people skip. Skip it and you only have your own opinion about whether the thing works.
Find two or three people who do the real task. Give them the task in one line and then be quiet. Do not demo it. Do not explain the screen. Note where they stop, what they re-read, and what they type that you did not expect.
A short session like this beats internal review, because people type things you would never have put in a test case.
Write down what broke before you change anything. That list is the honest output of the evening.
Get one number before you go home
The reason a project ends in a shrug is that nobody wrote down a before number. Six months later the boss asks whether it worked, and there is nothing to compare.
So take one measure. How long the current process takes, end to end. How many of the people who start it finish it. How many come back with a missing document. Measure it on the old way this week, even roughly, even by hand from ten cases.
It does not need to be sophisticated, and the payoff is that you can state the change, like this. A large insurer used to take six months or more to launch a customer journey on the web. With five-day design sprints, the new journeys were designed and tested with customers in two weeks and live in four. Completion rates on the new journeys rose 80 per cent and the old drop-off points disappeared. None of that could be claimed without the earlier numbers to compare against.
A small honest number beats a big vague claim. If the gap turns out to be small, say so, and don't fund the full build.
Decide what the prototype is for next
There are two ways a rough version is worth the evening. Either it settles an argument and you throw it away, or people can finish the task in it and you commission the real thing.
Keep those separate in your head. The rough version is not a first release. It has no error handling, no security review, no load testing, and the data is made up. Push it into production and someone is stuck maintaining an evening's work.
If you get to the point where the questions are about running it properly rather than what it should do, that is the moment to price a proper build. Prototype or production build: which to commission next sets out how the two differ in scope and cost.
What to do this week
- Write the one task in a single sentence, naming the person who does it.
- Pull ten real cases and time the current process by hand. That is your before number.
- Set aside an evening, build the smallest working version, and stop adding features when the time is up.
- Get two or three people who do the task to try it while you stay quiet and take notes.
- Write the list of what broke, and decide whether the gap is big enough to fund a real build.
More on AI prototyping
- Build a rough working version of an AI feature instead of writing the business case first
- What AI prototyping is and what you get at the end
- AI prototyping sessions for insurance product teams
- Prototype or production build: which to commission next
- What a custom workshop is and how we build one with your team
- Everything on AI prototyping
Questions people ask
What counts as a prototype here?
Something a person can click through and attempt the task in, with fake data. Not a PowerPoint, not a clickable mockup with no logic behind it.
How many users do we need to test with?
Two or three people who do the real task is enough to find the obvious breaks. Five is better if they are easy to reach.
When should we stop prototyping and commission a real build?
When people can finish the task in the rough version and the arguments have moved from what it should do to how to run it properly. That is the point to price a production build.
On-Off Group trains teams, tests products with real customers, finds where a transformation has stalled and builds what gets it moving, for banks, insurers and enterprises in the Philippines, since 2015. How we help with ai prototyping.

