| Dimension | Score |
|---|---|
| Output reliability | 3/5 |
| Workflow fit | 4/5 |
| Context handling | 3/5 |
| Learning curve vs. payoff | 4/5 |
| Failure transparency | 2/5 |
| Overall | 3.2/5 |
What AI does well
Session transcription and tagging. Looppanel and Dovetail AI both auto-tag interview transcripts against a code frame in minutes, not hours. I ran a 12-session study through Looppanel and it caught most theme clusters a human coder would find. The caveat: it over-tags neutral small talk as “friction,” so you still have to prune the tag list by hand.
First-pass synthesis drafts. Claude and ChatGPT can turn raw notes into a first-draft insight summary with themes grouped and rough severity ranked. This saves a full afternoon on a typical 8-participant study. The output reads well, but it flattens contradictions between participants into one tidy narrative, which is the opposite of what a researcher needs.
Discussion guide generation. ChatGPT and Claude write solid starting drafts for interview scripts and screener questions when you give them the research goal and stakeholder questions. Editing a draft is faster than writing from scratch. It still needs a researcher to cut leading questions, because the models default to closed, confirmatory phrasing.
Recruiting screener logic. Dovetail AI and ChatGPT both handle branching logic for screener surveys reasonably well, catching contradictory quota rules a tired researcher might miss at 11pm before a study launch.
Where it fails
Quote attribution breaks under pressure. I asked Dovetail AI to pull supporting quotes for a “trust in onboarding” theme across a 6-participant study. Two of the five quotes it returned were said by a different participant than the one it cited. In a stakeholder readout, a wrong attribution undermines the whole finding, because someone in the room usually remembers who said what.
Sarcasm and hedging get read as literal agreement. ChatGPT and Claude both tend to score a hedged “I guess that’s fine” the same as an enthusiastic “yes.” In usability studies that measure satisfaction, this inflates positive sentiment and can mask real friction that the researcher would have flagged by watching the recording.
Cross-study pattern memory does not exist. None of these tools retain what they learned about your product’s recurring pain points from study to study. Every new project starts from zero, so a researcher ends up re-explaining the same product context every time, which erodes the time saved on synthesis.
My take
My take (August 2026, Bernat Sampera)
The verdict holds up under real use. AI genuinely speeds up the grunt work: transcription, tagging, first-draft synthesis. A researcher running back-to-back studies gets real hours back. But the reliability gap is not cosmetic. Wrong quote attribution and flattened contradictions are exactly the failures that erode stakeholder trust in research, and that trust is the actual product a UX researcher sells. I would not let any of these tools present in a readout unsupervised. I compared ChatGPT and Claude directly on synthesis tasks, and Claude held context across a longer transcript more reliably, but neither caught the sarcasm problem. Use AI to draft, never to finalize. The tool that gets this right will need to say “unsure” instead of guessing, and none of them do that today.
What I would build
The real gap is memory across studies, not synthesis speed. A UX research team runs the same product through five, ten, twenty studies a year, and every tool treats each one as an isolated file. What is missing is a persistent research memory: a system that carries forward known pain points, past severity ratings, and prior contradictions between participant segments, so a new study’s AI-drafted synthesis gets checked against what the team already knows instead of restating it as new. That system would need to flag when a new finding contradicts a past one, not just summarize the new session in isolation. Nobody has built this well yet, because it requires holding structured research history, not just chat history.
Verdict (August 2026, Bernat Sampera): AI tools cut UX research synthesis time roughly in half, but they still misattribute a quote to the wrong participant in more than 1 of every 5 sessions I checked. Overall: 3.2/5.