Why test with 5 users, and when a usability test needs more

Articles

Five users usually find most usability problems, because the same problems keep repeating after the first few people. Test with five, fix, then test again. Use more when you have very different user groups or need a reliable number.

If you have searched for why test with 5 users, you have probably found the Nielsen Norman Group article that everyone quotes, and then sat in a meeting where someone said five people cannot possibly be enough.

This page is for the people who run usability tests at work or have to defend them. That might be a head of digital who needs to show her boss something real from the customer journey work, or an operations lead who wants one feature in front of real users instead of another scoping meeting.

You will get the reason five works, four beliefs that keep teams from testing small, and the cases where five is not enough.

“Five people is too few to trust”

On a normal working day this sounds sensible. Someone says five is not statistically significant, the test gets folded into a bigger research project, and that project waits for budget sign-off. Months go by and nobody watches a customer use the form.

The maths behind five is simple. Jakob Nielsen's 2000 article, built on research with Tom Landauer, found that a single user shows you about 31 percent of the usability problems in a design. By the fifth user you have seen about 85 percent. After that you mostly watch the same problems again.

A usability test is looking for problems. It is not measuring how common they are. If three out of five people cannot find the button to continue on an account opening form, you do not need a hundred more to know the button is in the wrong place. You need to move it.

“One big test is better than several small ones”

The usual pattern is one large round of testing near the end. Fifteen people, a long report, a few fixes, and nobody checks whether the fixes worked.

Nielsen's own advice is the opposite. If you have budget for fifteen people, run three tests of five. Test, fix, test again. The second round tells you whether the changes helped, and it often finds problems the first set of issues was hiding, because people now get further into the flow.

This is how our design sprints work. With a large insurer, a series of five-day sprints on customer-facing web journeys meant journeys that used to take six months or more were designed and tested with customers in two weeks, and live in four. Completion rates on the new journeys rose 80 percent, and the old drop-off points disappeared. Short rounds of testing and fixing did most of that work.

“Only the designers need to watch”

In most organisations a researcher runs the sessions, writes up a PowerPoint and presents it to managers. The managers argue with the slides. The developer who built the form never sees a customer get stuck on it.

In our workshops the whole team watches customers try a prototype. Developers, marketers, product owners and senior people sit in the same room. Watching one person give up on a screen settles arguments that a report never will.

A simple way to do it:

  • Put the session on a big screen in a meeting room.
  • Everyone writes each problem they see on a sticky note, one problem per note.
  • After the last session, put the notes on a wall and group the ones that match.

With five sessions the groups show up quickly. The problem that four people hit is obvious to everyone, including whoever has to approve the fix.

“We need a finished product before we can test”

Teams often book testing after the build is done. By then the launch date is fixed, the budget is spent, and only small changes are allowed.

Test earlier, on something rough. Clickable screens work. Paper sketches work. If you are building an AI feature, a rough version wired up from a prompt is enough to learn whether people understand what it does and trust what it gives back.

A useful test for when a rough version is ready: someone who has never seen it can attempt the main task without you explaining how. It does not need to look finished. It needs to let a person try the thing you are unsure about.

Five people on a rough prototype will tell you more than a roadmap that nobody has tried.

When five users are not enough

Five works for finding problems in one flow used by one kind of person. There are clear cases where you need more.

Different user groups. If customers and branch staff both use the same onboarding system, they will hit different problems. Nielsen's advice is three or four people from each group.

You need a number. Five people show you what is broken, not how often it happens. If your boss will ask whether turnaround time or completion rate improved, you need a proper sample. Nielsen Norman Group's guidance for quantitative studies is around 20 users. You also need a before number, taken before you change anything, or the answer to “did it work?” is a shrug.

Everyone looks like your team. Five colleagues from head office will miss what an older customer on a cheap phone runs into. Recruit people who match your actual customers, including the ones at the edges.

What to do this week

  • Pick one task in one customer journey, such as opening an account on mobile.
  • Book five people who match your real customers, not colleagues.
  • Show the sessions on a screen and invite the developer and one manager to watch.
  • Fix the top problems, then test again with five different people.
  • If you will need to prove the change worked, record completion rate or turnaround time before you touch anything.

Questions people ask

Where does the five users rule come from?
It comes from a Jakob Nielsen article published by Nielsen Norman Group in 2000, based on research with Tom Landauer. It found that one user shows about 31 percent of the problems and five users show about 85 percent.

How many users do you need if you have different types of customer?
Nielsen's advice is three or four people from each group that uses the product very differently, for example customers and branch staff using the same onboarding system.

Can five users give you a completion rate or a before and after number?
No. Five people show you what is broken, not how often it happens. Nielsen Norman Group's guidance for quantitative studies is around 20 users, and you need a before measurement to compare against.

On-Off Group trains teams, tests products with real customers, finds where a transformation has stalled and builds what gets it moving, for banks, insurers and enterprises in the Philippines, since 2015. Who we are.