How to run a usability test (step by step, with a worked example)
A usability test puts your product or prototype in front of real users doing real tasks, so you can see where they succeed, hesitate, and fail — before you ship. It's the highest-signal research method in UX, and you don't need an enterprise budget or a lab to run one. This guide covers the full process, with a worked example and the numbers you'll use to judge results.
Moderated vs unmoderated: pick your method first
| Moderated | Unmoderated | |
|---|---|---|
| Format | Live session, you facilitate | Self-guided, asynchronous |
| Best for | Depth, prototypes, "why" | Speed, scale, "how many" |
| Participants per round | ~5 | 20–50+ |
| Turnaround | Days (scheduling) | Hours |
| Cost per session | Higher (your time) | Lower |
| Follow-up questions | ✅ Live | ❌ |
- Moderated — you run a live session (video + screen share), watch in real time, and ask follow-up questions. Best for depth, early prototypes, and understanding why something happened. ResearchRocket includes moderated sessions with recording and highlight clipping.
- Unmoderated — participants complete tasks on their own; you review the results after. Best for speed and scale and quantified metrics.
Many teams do both: a few moderated sessions for the why, then an unmoderated round for the how many.
Step 1: Define the goal and tasks
Start with a research question ("Can new users set up their first project without help?"), then write 3–5 realistic tasks tied to real user goals. Phrase tasks as outcomes, not instructions, and know what success looks like for each.
Worked example — a project-management app onboarding. Research question: Can a first-time user get to their first created project and invite a teammate? Tasks:
| # | Task (what the participant reads) | Success = |
|---|---|---|
| 1 | "You've just signed up. Create your first project called 'Website Redesign'." | Project created with that name |
| 2 | "Add a teammate, alex@example.com, to that project." | Invite sent to that address |
| 3 | "Find where you'd change the project's due date." | Lands on the date/settings control |
Notice these are goals ("add a teammate"), not click-by-click instructions ("click the Members tab, then…"). Leading tasks quietly inflate success and hide the real friction.
Step 2: Choose moderated or unmoderated
Prototype needing rich feedback → moderated. Validating a known flow at volume → unmoderated. For our onboarding example, we'll run 5 moderated sessions first (to hear the confusion out loud), then an unmoderated round of 30 to quantify task success.
Step 3: Set your sample size
- Qualitative / moderated: ~5 participants per audience catches the majority of major usability issues — the classic Nielsen "5 users" heuristic. Run more rounds rather than more people per round; you learn more from 3 rounds of 5 than one round of 15.
- Quantitative / unmoderated: aim for 20–50+ if you're reporting success rates and times as numbers. Size it with the sample-size calculator.
Step 4: Recruit participants
Use your own users where possible — cheapest and most representative — via a shareable link, email list, or QR code. Screen for the right audience if the task depends on it (e.g., "new users only"). Incentives for recruited testers typically run $20–$100+ per session depending on audience; your own users often participate for a small thank-you. (See how much usability testing costs.)
Step 5: Run the sessions
Moderated facilitation do's and don'ts:
- Do stay quiet while they work. Silence is uncomfortable, but it's where the real behavior shows.
- Do ask "What are you thinking right now?" and "What did you expect to happen?" — open, non-leading.
- Don't rescue them the moment they struggle — the struggle is the data.
- Don't ask "Was that easy?" (leading). Ask "How did that go?" (neutral).
- Do mark timestamped highlights as key moments happen, so you can find them in a 45-minute recording later.
Unmoderated: let the tool capture task success, time on task, and misclicks automatically. Add a one-line open-ended question after each task ("What, if anything, was confusing?") to get a little qualitative signal at scale.
Step 6: Track the right metrics
- Task success rate — did they complete it? (Binary, or partial/assisted/failed.)
- Time on task — how long did it take? Compare against a realistic target.
- Error / misclick rate — where and how often did they go wrong?
- Qualitative friction — hesitation, confusion, and verbatim quotes.
Worked example — reading the results. Your unmoderated round of 30 comes back:
- Task 1 (create project): 93% success, avg 41s — healthy.
- Task 2 (invite teammate): 57% success, avg 2m 18s, high misclicks on the "Members" area — a problem. In the moderated sessions, 4 of 5 users looked for "Invite" on the project page, but it was buried in Settings.
- Task 3 (change due date): 70% success, but low directness — people found it eventually via trial and error.
The combined story: invite flow is broken (users expect "Invite" on the project, not in Settings), and due-date discoverability is weak. Task 1 is fine. That's a prioritized, evidence-backed fix list — not opinions.
Step 7: Analyze and decide
Cluster observations into themes (AI thematic coding speeds this up), cut highlight clips of the clearest moments for stakeholders (a 20-second clip of four users failing the invite flow is more persuasive than any bar chart), and translate findings into prioritized fixes. In ResearchRocket you can carry these straight into a Decision Brief with risk and a documented go/no-go.
Common mistakes
- Leading tasks that hint at the answer ("Use the Members tab to…"). Phrase as goals.
- Wrong sample size for the goal — 5 for qualitative discovery, 20–50+ for quantitative claims. Don't report "80% success" from 5 people.
- No highlight capture — you'll never re-find the key moment in an hour of video, and stakeholders won't watch the whole thing.
- Rescuing struggling participants — you lose the exact data you came for.
- Stopping at observations instead of prioritized, decided fixes. A usability test that doesn't change a roadmap was a waste.
Run usability tests free
ResearchRocket includes moderated sessions (with recording + highlight clipping) and unmoderated prototype tests, plus the repository, thematic coding, statistics, and decision tools to act on results — self-serve, free to start, with no per-seat minimum.
FAQ
How do you run a usability test? Define 3–5 realistic, goal-based tasks, choose moderated (depth) or unmoderated (scale), recruit participants (about 5 for qualitative, 20–50+ for quantitative), run the sessions capturing success/time/errors, then analyze into themes and prioritized fixes.
How many participants do I need for a usability test? About 5 per audience for qualitative discovery (Nielsen's "5 users" rule — run more rounds rather than more people per round); 20–50+ when you're reporting success rates and times as numbers.
What's the difference between moderated and unmoderated usability testing? Moderated is a live, facilitated session — best for depth, prototypes, and "why." Unmoderated is self-guided and asynchronous — best for speed, scale, and quantified metrics.
What metrics should I track in a usability test? Task success rate, time on task, error/misclick rate, and qualitative friction (hesitation and quotes). Read success alongside directness — completing a task slowly and by trial-and-error is still a warning sign.
How do I write good usability test tasks? Phrase them as goals, not instructions; make them realistic; and don't name the UI elements the user is supposed to find. "Add a teammate to this project," not "Click Members, then Invite."
How long should a usability test take? Moderated sessions typically run 30–60 minutes; unmoderated task sets should take a participant 5–15 minutes to keep completion rates high.
Is there a free way to run usability tests? Yes — ResearchRocket is self-serve with a free trial and free public tools, and includes both moderated and unmoderated testing without an enterprise contract.
Start usability testing free on ResearchRocket →
Related: How much does usability testing cost? · First-click testing · UserTesting alternative · Free sample-size calculator