Why test with 5 users, and when AI features need more
Five users usually turn up about 85 percent of usability problems, because after that the same problems keep repeating. AI features need more when you have separate user groups, or when the AI answers differently each time.
Why test with 5 users? Because five people trying a rough version of your product will show you most of what is wrong with it. That holds up well for AI features too, with two exceptions you should know about before you book any sessions.
This is for anyone with an AI feature somewhere between a working prompt and a pilot. Maybe you run operations and have a side product you want in front of real users. Maybe you lead digital at a bank and need to show your boss something real from the customer journey work, not another vendor deck.
Writing the prompts is usually the easy part. The harder calls are how many people to test with, when five falls short, and what to do once the fifth person has gone home. Each of those gets a section below.
Why five users find most of the problems
On a normal working day, testing stalls on the number. Someone asks for thirty users so the results feel solid. Recruiting takes weeks, the budget conversation drags on, and in the end nobody tests at all. Or it swings the other way and one colleague tries the feature at their desk, says it looks fine, and that counts as testing.
Usability researcher Jakob Nielsen found that five users turn up about 85 percent of the problems. The reason is simple. People get stuck in the same places. After the third or fourth session you can usually predict where the next person will pause, and the sixth person mostly confirms what you already wrote down.
So recruit five people who would actually use the feature. Give each of them the same real task, in the words a customer would use. Then watch, and do not help. A session like this measures where people get stuck, and what they tell you they like matters much less. Five is enough for that because the problems repeat.
When five users is not enough to test an AI feature
There are two cases where five falls short, and AI features tend to hit both.
The first is separate user groups. Take an onboarding flow used by branch staff and by customers. Branch staff know the product names and the order of the steps. Customers do not. Put five mixed people through it and you might only watch two customers, which is not enough to see their problems repeat. When the groups really differ, Nielsen suggests at least three people per group.
The second is specific to AI. A normal screen looks the same every time. An AI feature can give a different answer to the same question each time it is asked. Five sessions show you five of its answers, and you have no idea how the rest look.
The fix for this one does not need more people. Keep a list of every question users typed or asked during the sessions. Afterwards, run each of those questions through the feature several times yourself and read the answers side by side. The sessions tell you what people ask, and the reruns tell you how much the answers wander.
What to do after the fifth session
Most testing ends with a write-up. The findings go into a PowerPoint, the PowerPoint goes round by email, and the feature stays exactly as it was. The other common mistake is booking ten more users on the same version, then watching them hit the same wall as the first five.
Fix what you saw, then test with five more. Nielsen's advice is three rounds of five rather than one round of fifteen. The second group shows you whether your fixes worked. It also finds the problems that were hidden behind the ones you just fixed. If nobody gets past the first screen, you never learn what is wrong with the third.
Short rounds with fixes in between are what make things move. A large insurer ran a series of five-day design sprints on its customer-facing web journeys. Journeys that used to take six months or more were designed and tested with customers in two weeks and live in four. Completion rates on the new journeys rose 80 percent, and the old drop-off points disappeared.
Before you fix anything, write down how many of the five finished the task. That becomes your before number, so when someone asks whether the changes worked, you have an answer.
What to do this week
- Pick one task your AI feature has to handle, and write it in the words a customer would use.
- Recruit five people who would really use the feature. If branch staff and customers both use it, get at least three of each.
- Watch each session without helping. Note where people stop and what the AI said back to them.
- Count how many of the five finished the task and keep that number somewhere safe.
- Fix the biggest problem you saw, then book the next five.
Questions people ask
Where does the five user rule come from?
From usability researcher Jakob Nielsen, who found that five users turn up about 85 percent of the problems in a design. After that, each new person mostly gets stuck in places you have already seen.
How many users do I need if branch staff and customers both use the feature?
Nielsen suggests at least three people per group when the groups really differ. Five people split across both groups leaves you with too few from one of them.
Why is testing an AI feature different from testing a normal screen?
A normal screen looks the same every time, but an AI feature can give a different answer to the same question. Five sessions only show you some of its answers, so rerun the questions from the sessions yourself and read the answers side by side.
Should I test with fifteen people at once?
Nielsen's advice is three rounds of five rather than one round of fifteen. Fixing between rounds shows you whether the fixes worked and finds problems the first ones were hiding.
On-Off Group trains teams, tests products with real customers, finds where a transformation has stalled and builds what gets it moving, for banks, insurers and enterprises in the Philippines, since 2015. Who we are.

