How to run a tree test (step by step, with a worked example)
A tree test checks whether people can find things in your site or app's structure — before you spend design and engineering time building navigation around it. You give participants a stripped-down, text-only version of your menu hierarchy (the "tree") and ask them to complete findability tasks. It's the fastest, cheapest way to catch an information architecture that makes sense to you but not to your users.
This guide walks through the whole process end to end, with a real worked example and the benchmark numbers you'll use to judge your results.
When to use tree testing
Reach for a tree test when:
- You're redesigning navigation or an IA and want to validate it before visuals or code.
- Users can't find features or content and you suspect the structure, not just the wording.
- You want to compare two IA options objectively — run a tree test on each and let success rates decide.
- You've just run a card sort and want to confirm the structure it produced actually works. (Card sorting generates an IA; tree testing validates it. See card sorting vs tree testing.)
Tree testing is evaluative and quantitative — it produces percentages you can defend to stakeholders, not just anecdotes.
Step 1: Build your tree
Recreate your proposed hierarchy as nested text labels — top-level categories, their subcategories, and the leaf pages underneath. Keep it text-only; the entire point is to test structure and labels in isolation, with no visual design, search bar, or branding to lean on.
Worked example — an outdoor-gear e-commerce site. Say your proposed top level is:
Home
├─ Shop
│ ├─ Men's
│ │ ├─ Jackets
│ │ ├─ Footwear
│ │ └─ Accessories
│ ├─ Women's
│ │ ├─ Jackets
│ │ ├─ Footwear
│ │ └─ Accessories
│ └─ Camping
│ ├─ Tents
│ ├─ Sleeping bags
│ └─ Cookware
├─ Deals
├─ Help
│ ├─ Shipping & Returns
│ ├─ Order Tracking
│ └─ Contact Us
└─ About
Include enough depth that each task has a realistic "correct" destination. A tree that's too shallow makes every task trivially easy and tells you nothing.
Step 2: Write findability tasks
Write 5–10 tasks phrased as goals, not instructions, and crucially, avoid using the exact category words — that hands participants the answer and inflates your success rate.
Using the example tree:
| # | Task (what the participant reads) | Correct destination |
|---|---|---|
| 1 | "You bought hiking boots last week and they don't fit. You want to send them back. Where would you go?" | Help → Shipping & Returns |
| 2 | "You're planning a weekend camping trip and need something to sleep in outdoors." | Shop → Camping → Sleeping bags |
| 3 | "You want to see if there are any current discounts." | Deals |
| 4 | "You're buying a rain jacket for your wife." | Shop → Women's → Jackets |
| 5 | "You want to check where your order is right now." | Help → Order Tracking |
Notice task 1 says "send them back," not "returns," and task 2 says "sleep in outdoors," not "sleeping bags." That's deliberate — you're testing whether the structure leads people to the right place, not whether they can pattern-match your labels.
Step 3: Decide your sample size
Tree testing is quantitative, so more participants sharpen your success and directness metrics. Rough guidance:
- 30–50 participants per audience for reliable directional results.
- 50+ if you're comparing two trees or need tighter confidence intervals.
- Segment if you have distinct audiences (e.g., new vs returning customers) — 30+ per segment.
Use a sample-size calculator to set the exact number for your confidence level and margin of error.
Step 4: Recruit and distribute
Share the test link with your panel, an email list, or a QR code. Unmoderated tree tests run asynchronously, so participants complete them on their own time, on any device. Tree tests are quick for participants (usually a few minutes), so completion rates tend to be high — you can collect 50 responses faster than almost any other method.
Step 5: Read the results
Three metrics carry the analysis:
1. Success rate — the percentage who ended at a correct destination.
- 80%+ = the structure works for that task.
- 60–79% = a warning sign; investigate.
- Below 60% = a real structural problem worth fixing.
2. Directness — the percentage who got to the answer without backtracking up the tree. Low directness with high success means people eventually find it but the path is confusing — often a sign the right answer lives under a non-obvious parent.
3. First click / path analysis — where participants go first is the single strongest predictor of success. If most first clicks for task 1 land on "Shop" instead of "Help," you've learned that users expect returns under shopping, not support.
Worked example — reading a result. Suppose task 1 (returns) comes back like this:
- Success: 52% · Directness: 38% · Most common first click: Shop (61%)
The story writes itself: users overwhelmingly expect "returns" to live under Shop, not Help. Only about half eventually find it, and most who do backtrack to get there. Fix: either surface "Shipping & Returns" under Shop, add it to both locations, or rename the top-level "Help" to something users associate with orders. Then re-test.
Contrast with task 3 (deals): Success 96%, Directness 94%, first click Deals (95%). That branch is healthy — leave it alone.
A simple prioritization framework
After the test, plot each task on two axes — success and directness:
| High directness | Low directness | |
|---|---|---|
| High success | ✅ Healthy — ship it | ⚠️ Works but confusing — minor label/placement tweak |
| Low success | ⚠️ Split decision — right answer unclear | 🚨 Broken — top priority IA fix |
Spend your redesign effort on the bottom-right quadrant first.
Common mistakes to avoid
- Leaking the answer by using category names in the task wording. This is the #1 tree-test error — it quietly inflates every success rate and makes a bad IA look fine.
- Too few participants. A tree test with 8 people is an anecdote wearing a lab coat. Get to 30+.
- Testing labels and structure at once. If a task fails, you won't know whether the label or the hierarchy broke it. Keep labels realistic and change one variable per round.
- Only testing the happy path. Include a couple of tasks you expect to be hard — those reveal the most.
- Ignoring first-click data and looking only at final success. First click tells you why a task failed.
From tree test to redesign, in one loop
- Run the tree test on your candidate IA.
- Rank tasks by the prioritization grid above.
- Fix the broken branches (move, rename, or duplicate).
- Re-test the same tasks on the revised tree.
- Compare success/directness before and after — that delta is your evidence the change worked.
Run a tree test free
ResearchRocket includes tree testing alongside card sorting, usability testing, surveys, and real statistics — so you can build an IA, validate it, and analyze the results in one place. Many tools (Maze, for example) reserve tree testing for an Enterprise plan; ResearchRocket includes it on core tiers. Pair it with the free card sorting tool to build the structure in the first place.
FAQ
What is a tree test? A tree test evaluates how easily people find items in a proposed site or app hierarchy by giving them a text-only version of the menu and asking them to complete findability tasks. It measures the structure in isolation, before any visual design.
How many participants do I need for a tree test? 30–50 per audience gives reliable directional results; use 50+ when comparing two trees or when you need tighter confidence. It's a quantitative method, so larger samples tighten the numbers.
What is a good tree test success rate? 80%+ indicates the structure works for that task; 60–79% is a warning; below 60% signals a structural problem. Always read success alongside directness and first-click data.
What's the difference between success and directness? Success is whether participants reached a correct destination at all. Directness is whether they got there without backtracking. High success but low directness means the path is findable but confusing.
What's the difference between card sorting and tree testing? Card sorting generates an information architecture from how users group items; tree testing validates a proposed architecture by measuring whether users can navigate it. Use them together — card sort first, tree test second.
How is a tree test different from a first-click test? A first-click test measures the first click on a single screen or design; a tree test measures first click plus the full navigation path across an entire hierarchy.
Is there a free tree testing tool? Yes — ResearchRocket offers tree testing on its trial, plus a free sample-size calculator and free card sort with no account required.
Run a tree test free on ResearchRocket →
Related: Free tree testing tool · Card sort analysis guide · Card sorting vs tree testing · ResearchRocket vs Maze