Skip to content
Back to all field notes pm · 8 min

Product discovery with AI: partner for thinking, not for deciding

AI is an excellent thinking partner in discovery and a terrible deciding partner. The line between the two is where good product work now lives, and it is easy to cross without noticing.

field note 8 min

Discovery has two halves that feel similar and are not. One is thinking: generating options, questioning assumptions, framing the problem, widening the space of what you might build. The other is deciding: figuring out what is actually true about real users and betting on it. AI is genuinely excellent at the first half. It is quietly dangerous in the second. And the reason it is dangerous is precisely that it is so good at the first that the boundary between them blurs.

You ask a model to help you think through a problem and it does, well. Then, in the same conversation, in the same confident register, it tells you what users want. Nothing signals that you crossed a line from reasoning to claiming. The output looks identical. That seam, where thinking-help slides into deciding-for-you without any change in tone, is where AI-assisted discovery goes wrong.

thinking partner

  • Widens the options
  • Stress-tests arguments
  • Reframes the problem
  • Structures the mess

deciding partner

  • Claims what users want
  • Never met them
  • A guess dressed as fact
AI is a superb partner for the thinking half of discovery and a dangerous one for the deciding half. The whole skill is keeping the line sharp.

Where AI earns its place

For the thinking half, a model is one of the best partners you can have, and it is worth being specific about why:

  • It widens the option space. Left alone, you explore the two or three solutions you already had in mind. Ask a model for ten more angles and most are useless, but two are things you would not have reached, and discovery lives on the options you did not already have.
  • It stress-tests an argument. Hand it your reasoning and ask where it breaks. A model is a tireless, ego-free devil’s advocate that will not get defensive when you push back. That is rare and useful.
  • It reframes. You are stuck seeing the problem one way. It offers three other framings, and one of them dissolves the thing you were stuck on.
  • It structures the mess. You have forty scattered observations. It helps you find the clusters and name the patterns. You still decide which patterns are real, but the sorting is real work it does well.

Every one of these keeps the human as the decider. The model expands and pressures and organizes your thinking. It does not conclude for you. Used this way it makes discovery faster and wider without touching the part that has to stay yours.

Where it quietly takes over

The danger is not that AI gives bad answers in discovery. It is that it gives confident, fluent, plausible answers to questions that can only be answered by contact with reality, and the confidence is indistinguishable from the confidence it shows when it is right.

Ask a model what your users struggle with and it will tell you, in clean prose, a list of pains. Some will be right. It has read enough that its priors are not worthless. But it has never met your users, watched them work, or seen where they actually get stuck. It is generating the most probable answer to “what do users of a thing like this struggle with,” which is a different question from “what do your users struggle with,” and the gap between those two questions is the entire job of discovery.

The failure mode is subtle because the output is useful-looking. A plausible list of user pains feels like progress. It reads like a research finding. And if you let it, it becomes one, sliding into your deck and your roadmap with the epistemic status of something you learned from users, when its actual status is something a model guessed. This is the AI slop failure mode wearing discovery clothes: more artifacts, more confident-sounding insights, no more actual contact with the truth.

Laundering: the specific sin to avoid

There is one move that does more damage than any other, and it deserves a name. Laundering is when a hypothesis a model generated gets cited later as if it were evidence a user gave you.

It happens gradually and without malice. On Monday you brainstorm with a model and it suggests users probably abandon the flow at the payment step. That is a hypothesis, a decent one. By Thursday it is in a doc as “users abandon at payment.” By the following week someone builds against it as a known fact. Nobody lied. At each step the claim just lost a little of its “we think” and gained a little “we know,” until a model’s guess is driving a build decision with the authority of research.

The defense is not to stop using AI in discovery. It is to keep the provenance of every claim ruthlessly visible. This hypothesis came from a model. This one came from a user session. This came from usage data. They are not the same kind of thing and they must never be allowed to look the same, because the whole value of discovery is knowing which of your beliefs are load-bearing evidence and which are plausible guesses waiting to be tested. The moment those two blur, you are making product like a feature factory, shipping against confident fiction and calling the motion progress.

Keeping the line sharp in practice

Concretely, a few habits keep thinking-partner from becoming deciding-partner:

  • Tag the source of every claim, always. Model-generated, user-sourced, data-sourced. If a belief in your discovery cannot say where it came from, it is not evidence yet, whatever its tone.
  • Never let a model summarize research into conclusions unsupervised. It will smooth away the uncertainty, drop the contradicting quote, and hand you a clean story. Clean is exactly what user research is not, and the mess it removes is where the real signal often hides.
  • Use it to generate hypotheses, then treat them as debts. Every model-suggested pain is a thing to validate, not a thing you know. It has earned a test, not a place in the roadmap.
  • Ask it to argue against your conclusion, not just for it. Its agreeableness is a trap. It will happily support whatever you are leaning toward. Aim its fluency at your own position and it becomes far more useful.

None of this is friction for its own sake. The point of the tags and the adversarial prompts is to preserve, through a fast and fluent process, the one distinction discovery exists to protect: the difference between what you have reason to believe and what you have merely been told plausibly. Lose that distinction and every decision downstream inherits a confidence it never earned, which is how a team ends up certain about a market it never actually looked at.

Where PaellaDoc fits

This is why in PaellaDoc provenance is not an afterthought, it is the structure. A hypothesis a model generated and a finding a user gave you are different kinds of nodes, and they stay different. You can use AI to widen and pressure your thinking as much as you want, because the system will not let a guess quietly acquire the authority of evidence. When a decision leans on a claim, you can see what the claim actually rests on: a user, a number, or a plausible sentence a model wrote.

Use AI for the half of discovery it is brilliant at. Let it widen your options, break your arguments, organize your mess. Just keep your hand on the decision, and keep the line between what you think and what you know bright enough that no fluent paragraph can smudge it. The model is a superb partner for thinking. It has never met your users. Do not let it pretend it has.

Frequently asked questions

Can AI do product discovery for you?

It can do the thinking half, not the deciding half. AI is superb at widening the option space, stress-testing an argument, reframing a problem, and structuring scattered observations. It is dangerous the moment it starts telling you what real users want, because it has never met them. Use it to expand and pressure your thinking; keep the decision, and the contact with reality, yours.

Can AI tell me what my users want?

No, though it will sound like it can. Ask a model what your users struggle with and it generates the most probable answer to “what do users of a thing like this struggle with,” which is a different question from what your users struggle with. That gap is the whole job of discovery. The answer is a plausible hypothesis, not a finding, until a real user confirms it.

What is laundering in AI-assisted discovery?

Laundering is when a hypothesis a model generated gets cited later as if it were evidence a user gave you. It happens gradually: a Monday brainstorm becomes a Thursday doc becomes a build decision the next week, losing a little “we think” and gaining a little “we know” at each step. Nobody lied, but a model’s guess ends up carrying the authority of research.

How do I keep AI hypotheses from turning into fake evidence?

Keep provenance ruthlessly visible. Tag every claim by source: model-generated, user-sourced, or data-sourced, and never let them look the same. Treat each model-suggested pain as a debt to validate, not a fact to build on. Never let a model summarize research into conclusions unsupervised, it smooths away the uncertainty and drops the contradicting quote. If a belief cannot say where it came from, it is not evidence yet.