How to do thematic analysis (UX research, with a worked example)
Thematic analysis turns messy qualitative data — interview transcripts, open-ended survey responses, support tickets, usability notes — into a structured set of themes you can act on. It's the backbone of rigorous qualitative research, and it's where most insight is won or lost. This guide covers the full process, shows you a worked coding example, and explains how to use AI to speed it up without gutting the rigor.
What thematic analysis is (and isn't)
Thematic analysis is the systematic process of coding qualitative data (tagging meaningful chunks with short labels) and grouping codes into themes (recurring patterns of meaning). It is not "reading through and noting what stood out" — that's summarizing, and it quietly smuggles in your biases. The discipline of coding every relevant excerpt is exactly what makes the findings defensible when a stakeholder pushes back.
The six-phase process (Braun & Clarke)
The most widely used framework in UX research has six phases:
- Familiarize — read through all the data end to end before coding. Get a feel for the whole before you dissect the parts.
- Generate codes — tag each meaningful excerpt with a short, descriptive label. Code openly at first; don't force a framework.
- Search for themes — cluster related codes into candidate themes. A theme is a pattern that says something significant about your research question, not just a frequent word.
- Review themes — check each candidate theme against the coded data and the full dataset. Split, merge, or drop. Good themes are internally coherent and distinct from one another.
- Define and name themes — write a one-sentence essence for each and give it a clear, memorable name (ideally in participants' own language).
- Report — tie each theme to evidence (representative quotes) and translate it into implications and decisions.
Worked example: coding onboarding feedback
Say you asked 6 new users one open-ended question: "What was frustrating about getting started?" Here are three responses and how you'd code them:
P1: "I couldn't tell which plan I was on, and I got charged before I even finished setup." Codes:
plan visibility,unexpected charge,setup incomplete
P2: "The invite flow was buried. I looked all over the project page for it." Codes:
invite discoverability,wrong mental model (expected on project)
P3: "I didn't know if my data was saved. No confirmation anywhere." Codes:
missing confirmation,uncertainty / trust
Now search for themes by clustering codes across all 6 participants:
plan visibility+unexpected charge+ (from P4)didn't understand trial→ Theme: "Billing anxiety — users don't know what they'll be charged."invite discoverability+wrong mental model+ (from P5)couldn't find settings→ Theme: "Discoverability — key actions aren't where users look."missing confirmation+uncertainty / trust+ (from P6)unsure it worked→ Theme: "Lack of feedback — the system doesn't confirm actions."
Three themes, each grounded in multiple participants and quotable evidence. That's thematic analysis — not "users found onboarding confusing," but three specific, evidenced, addressable patterns.
Inductive vs deductive (open vs closed) coding
- Inductive (open): codes emerge from the data. Best for exploration, and what the example above uses.
- Deductive (closed): you start from a predefined framework (e.g., Nielsen's heuristics, or a set of hypotheses) and code to it. Best when you're testing specific questions.
Many studies blend both — code openly first, then organize the codes under a light framework.
How to keep it rigorous
- Code the whole dataset, not just the vivid quotes. Frequency and spread matter — a theme raised by 5 of 6 participants is stronger than one dramatic outlier.
- Track how many participants each theme touches (prevalence), not just how many mentions — one person saying something ten times isn't a pattern.
- Keep a codebook — a living list of your codes with definitions — so coding stays consistent, especially across multiple researchers.
- Look for disconfirming evidence. Actively check whether any data contradicts a theme before you commit to it.
Where AI helps (and where it doesn't)
AI auto-coding can propose an initial set of codes and themes from a transcript in seconds — a genuine time-saver on the first pass (phases 2–3). But you should always review and refine: AI reliably catches surface patterns and keyword clusters; human judgment catches nuance, sarcasm, mixed sentiment, and the crucial why behind a comment.
The rigor-preserving way to use AI:
- Let AI generate a first-pass coding and candidate themes.
- Review every theme against the actual quotes — merge duplicates, split overloaded themes, kill anything the evidence doesn't support.
- Apply human definition and naming (phases 4–5) — this is where expertise adds the most value.
- Keep AI out of the final interpretation call. It accelerates; it doesn't decide.
Used this way, AI cuts the most time-consuming phase without sacrificing the credibility of the output.
Turning themes into decisions
Themes are the output, but not the point — the point is a decision. Attribute each theme to how many participants raised it and how often, then carry the high-impact themes into a prioritized recommendation. From the example: "Billing anxiety" (5/6 participants) outranks a one-off gripe and becomes a roadmap item. In ResearchRocket, coded insights flow into a repository, a bubble chart of mentions vs insights, and Decision Brief / Decision Studio.
Common mistakes
- Coding too narrowly (one code per participant) or too broadly (everything is one giant theme).
- Cherry-picking quotes that fit a preferred story instead of representing the data.
- Counting mentions, not participants — inflates themes driven by one talkative person.
- Stopping at themes without translating them into implications and decisions.
- Trusting AI output uncritically — always review against the raw quotes.
Do thematic analysis free
ResearchRocket includes an Insights Hub for thematic coding — search across all insights, color-coded two-level theme groups, a word cloud, and a bubble chart of mentions vs insights — plus AI auto-coding to jump-start the first pass. It's part of an all-in-one platform, so the data you collect (usability tests, surveys, interviews) is the data you code, with no import step.
FAQ
What is thematic analysis? A systematic method for coding qualitative data and grouping those codes into recurring themes, producing defensible findings from interviews, open-ended survey responses, tickets, and notes.
What are the six steps of thematic analysis? Familiarize, generate codes, search for themes, review themes, define and name themes, and report — the widely used Braun & Clarke framework.
What's the difference between a code and a theme? A code is a short label on a specific excerpt ("unexpected charge"); a theme is a pattern that groups related codes into something meaningful ("Billing anxiety"). Codes are the raw tags; themes are the findings.
Can AI do thematic analysis? AI can propose an initial coding and candidate themes quickly, but human review is essential for nuance, accuracy, and interpretation. Use AI to accelerate the first pass (coding), then apply expertise to define, name, and decide.
How many participants do you need for thematic analysis? It depends on the method feeding it — for interview-based analysis, patterns often stabilize around 6–12 participants per audience; for open-ended survey responses, more is better since each response is shorter.
Should I count mentions or participants? Participants (prevalence) is the stronger signal — a theme raised by most participants matters more than one raised many times by a single person. Track both.
Is there a free thematic analysis tool? Yes — ResearchRocket's Insights Hub offers thematic coding with AI auto-coding, free to start, and it also collects the research you're analyzing.
Start thematic analysis free on ResearchRocket →
Related: How to run a usability test · HeyMarvin alternative · ResearchRocket vs Marvin