A model can read forty user interviews and hand you back six clean themes in a minute. The themes are articulate. They are well organized. They sound like insight. And here is the question that decides whether they are worth anything: can you click on a theme and see the actual sentences, from actual users, that produced it? If you can, the model did real work. If you cannot, the model wrote you a plausible essay about your users and you are about to make decisions on it as if it were evidence.
That gap is the whole subject. AI in user research is genuinely useful and genuinely dangerous, and the useful version and the dangerous version look nearly identical on the slide. Telling them apart is the skill that matters now.
Synthesis and substitution are not on a spectrum
They are different acts, and blurring them is the root mistake.
synthesis
- Compresses real evidence
- Themes link to quotes
- Amplifies what exists
substitution
- Manufactures evidence
- No real source
- A guess as a finding
Synthesis takes evidence you actually gathered and compresses it. You talked to real people, you have the transcripts, and the model helps you find the patterns across more material than you could hold in your head at once. The evidence existed before the model touched it. The model made it legible. This is a real gain, and it is the reason to use AI in research at all: the bottleneck in synthesis was never insight, it was the hours of reading, and that is exactly what a model removes.
Substitution manufactures evidence that was never gathered. You did not talk to anyone, so you ask the model to play the user, generate the persona, simulate what “someone like this” would say. What comes back is fluent and confident and entirely untethered from any real person. It is the model’s prior about users, dressed as data. No matter how good the model is, it cannot report a fact about your users that no user ever supplied, because that fact does not exist in anything it read. It can only produce a well-formed guess.
These are not two points on a spectrum where a little substitution is a little risky. They are different in kind. One amplifies evidence. The other fabricates it. Every serious failure of AI-assisted research comes from doing the second while calling it the first.
Evidence laundering
Here is how the fabrication actually enters a real team, because nobody sets out to make things up. It launders in through steps that each look reasonable.
A researcher runs eight real interviews, which is a small sample, so they ask the model to “extrapolate what other users probably feel.” The model obliges. Now the deck has claims backed by eight people and claims backed by nobody, formatted identically. A week later someone pulls a bullet from the deck into a strategy doc. The bullet no longer carries any marker of where it came from. Another week, it is a line in a decision: “users want this.” Nobody lied. But a claim the model invented is now sitting in a decision with exactly the same authority as a claim eight humans actually made, and no one can tell them apart anymore.
That is evidence laundering: the process by which a model’s guess loses its source and acquires the credibility of a finding as it moves through your documents. It is the specific form the AI slop failure mode takes in research, where a fluent, well-formatted output gets mistaken for a validated one because nothing in the format announces the difference. The laundering is not one big lie. It is the slow detachment of a claim from its origin, one reasonable-looking copy at a time.
The rule: every claim points at an utterance
The discipline that prevents all of this is one sentence. A research finding is only as strong as the raw utterance you can trace it to.
In practice that means synthesized themes stay linked to the source quotes, permanently, not just in the first draft. When a finding says “users abandon at onboarding because they don’t trust the import,” you can expand it and see the four people who said something to that effect, in their words. A finding you cannot expand into real quotes is a finding you should treat as a hypothesis, not evidence, no matter how confident the summary sounds.
This also fixes the small-sample problem without inflating it. Eight interviews are eight interviews. The move is not to have the model inflate them into a fake hundred. It is to say clearly “four of eight users hit this,” keep it linked to those four, and treat it as a signal worth a bigger test, not a settled fact. The model helps you see the pattern in the eight. It does not get to invent the ninety-two you never spoke to. Naming the sample plainly is not a weakness in the research. It is the difference between research and a persuasive story.
What the model is allowed to do, precisely
It helps to be concrete about the boundary, because “use AI carefully” means nothing. The model is allowed to compress, cluster, and surface: group similar utterances, propose candidate themes, flag a quote you missed, draft the summary you will check against the sources. All of that operates on evidence that exists and keeps pointing back at it.
The model is not allowed to originate a claim about users from nothing: no simulated interview standing in for a real one, no persona answering questions no human answered, no extrapolation presented as a finding rather than a guess. The test is always the same. Point at the human who said this. If you can, it is research. If the only source is the model, it is the model’s opinion about your users, and it belongs in the hypothesis column with a note to go find out, not in the evidence column driving a decision.
This is the same standard that runs through everything I build, the same reason the product management job in the AI era shifted from producing artifacts to designing the system that separates what is known from what is assumed. Research is where that separation is easiest to lose, because a fluent theme feels like knowledge even when nothing real is under it.
Where PaellaDoc fits
PaellaDoc keeps the link between a finding and its evidence intact as research moves into decisions. A synthesized theme stays connected to the utterances that produced it, and a decision that cites the theme can be traced back through it to the actual users who said the thing, so a claim cannot quietly detach from its source as it travels through your product’s memory. The point is not to slow research down. It is to make laundering structurally hard, so the confident summary and the validated finding stop being interchangeable.
None of this asks you to distrust the tool. It asks you to keep the tool on the right side of one boundary, where it makes real evidence legible and never invents the evidence it summarizes. A team that holds that boundary gets faster synthesis and cleaner findings at once. A team that lets it slip gets a research practice that feels productive and is quietly making things up.
AI removed the drudgery from research, which is a real gift, and it removed it by removing the friction that used to keep guesses and findings apart. Getting the gift without the counterfeit means holding one line without exception: no claim about your users that you cannot trace to a user.
Frequently asked questions
Can AI do user research?
It can synthesize research you actually gathered, not replace the gathering. A model reads forty interviews and hands back six clean themes in a minute, and that is a real gain because the bottleneck in synthesis was the hours of reading, not the insight. What it cannot do is originate a fact about your users that no user supplied. That fact does not exist in anything it read, so it can only produce a well-formed guess.
What is the difference between synthesis and substitution?
Synthesis compresses evidence that existed before the model touched it: real people, real transcripts, patterns made legible. Substitution manufactures evidence that was never gathered: a simulated persona answering questions no human answered. They are different in kind, not two points on a spectrum. One amplifies evidence, the other fabricates it, and every serious failure of AI-assisted research is doing the second while calling it the first.
What is evidence laundering in user research?
It is the process by which a model’s guess loses its source and acquires the credibility of a finding as it moves through your documents. A model extrapolates beyond eight real interviews, the deck formats invented claims identically to real ones, a bullet moves into a strategy doc without its source, and a week later it is a decision. Nobody lied, but a guess now carries the authority of eight humans.
Can I use AI to generate synthetic users or personas?
Not as evidence. A persona answering questions no human answered is the model’s prior about users dressed as data, and it belongs in the hypothesis column with a note to go find out, not in the evidence column driving a decision. The rule is one sentence: every claim points at an utterance. If you can name the real person who said it, it is research; if the only source is the model, it is opinion.