Talk to potential customers from Dubai, Kuwait City, Doha, Riyadh, Jeddah, Abu Dhabi, and other Arab cities to gain insights
UsersArabia
All posts
Research Methods

Card sorting vs tree testing: which one do you need?

August 26, 2026

Card sorting and tree testing get mentioned in the same breath so often that many teams assume they are variations of the same thing. They are close to opposites.

Card sorting is generative. You have content and no structure, and you want users to help you build one.

Tree testing is evaluative. You have a structure and you want to know whether people can find things in it.

Run them in the wrong order and you will either validate a structure nobody helped shape, or generate a structure you never checked.

Card sorting: how people group your content

You give participants a set of items, typically page names, product categories or features, and they sort them into groups that make sense to them.

Open sort. Participants create and name the groups. Use this when you have no structure yet and want to discover the users' mental model, including the words they use for things.

Closed sort. You supply the categories and they place items into them. Use this when the top level is settled and you want to know where individual items belong.

Hybrid. You supply categories and let people add their own. A good default when you have a draft structure but suspect it is incomplete.

What the results tell you

The number that matters is agreement per item. For each card, how many participants put it in the same group?

  • High agreement, above roughly 70 percent. The item is well named and obviously belongs somewhere. Leave it alone.
  • Middling agreement. People are split between two plausible homes. Often the item genuinely belongs in both, and cross-linking is the answer.
  • Low agreement. The item scattered across many groups. This is almost always a naming problem rather than a placement problem. Nobody knows what it is, so everyone guesses differently.

The group names participants invent in an open sort are as valuable as the groupings themselves. If eight of fifteen people create a group called "Account" and your team calls it "My Profile", you have just been handed your navigation label.

When to run it

  • Designing a new site or product structure.
  • Restructuring content that has grown organically.
  • Checking whether internal vocabulary matches user vocabulary.

Participants: 15 to 30. Beyond thirty you rarely learn anything new.

Tree testing: can people find things

Tree testing strips away everything except structure. Participants see a text-only hierarchy, no visual design, no search, no colours, and they are given a task such as "where would you go to change your delivery address?" They click down through the levels until they decide they have arrived.

Removing the visual design is the point. A beautiful page can rescue a bad structure temporarily. Tree testing tells you whether the structure itself works.

What the results tell you

Three numbers, and the relationship between them is where the insight lives.

  • Success rate. Did they end up in the right place?
  • Directness. Did they get there without backtracking?
  • Time. How long did it take?

The combinations are diagnostic:

  • High success, high directness. The structure works. Move on.
  • High success, low directness. People eventually find it but wander first. The structure is roughly right and the labels are ambiguous. Fix the wording, not the hierarchy.
  • Low success. The item is in the wrong place, or it is not where anyone expects to look. This is a structural problem.

Where people went instead is often more useful than the failure rate itself. If a third of participants looked for delivery settings under "Account" when you filed it under "Orders", they have told you where it should live.

When to run it

  • Validating a proposed navigation before you build it.
  • Diagnosing why support keeps getting asked where something is.
  • Comparing two candidate structures head to head.

Participants: 30 or more, because you are reading percentages rather than observing behaviour.

The order that works

  1. Card sort to generate the structure and learn the vocabulary.
  2. Build the draft hierarchy from what you learned.
  3. Tree test to check people can navigate it.
  4. Fix the labels or the placement based on where they went wrong.
  5. Tree test again if the changes were significant.

That second tree test is the step teams skip. It is also the cheapest possible way to find out whether your fix actually fixed anything.

Common mistakes

Testing too many cards. Above roughly fifty items, participants fatigue and the quality of the sort collapses. Sample representative items rather than including everything.

Using page titles instead of user language. If your card says "Post-Purchase Fulfilment Enquiry", you are testing your jargon, not your structure.

Treating a card sort as validation. It is not. People grouping items is not the same as people finding items.

Ignoring the near misses in a tree test. A participant who lands one level away from the target has almost succeeded, and the reason they stopped short usually points straight at the fix.

Running both

UsersArabia supports open, closed and hybrid card sorting, and tree testing with success, directness and destination analysis, alongside participant recruitment across MENA.

Set up a study or read the full method guide.

Research MethodsUser Research