Neither camp is right. Simulated respondents are excellent at one job and dangerous at another, and the line between the two is where most teams get lost.
Two positions dominate the conversation about synthetic users, and both are lazy. One says a language model prompted to act like your customer can replace user research. The other says the whole idea is a toy. The first sells software; the second sells reassurance. Neither survives contact with what these tools actually do.
Start with a distinction that gets blurred constantly. A synthetic user is a model pretending to be a respondent. An AI interviewer is a model talking to a real respondent. The second is a data-collection tool; we have run hundreds of such conversations alongside human interviews and found the findings converge, with the AI needing more conversations to get there because each one is thinner. The first is something else entirely, and it is the first we're talking about here.
Generating hypotheses. Ask a model to respond as a long-term subscriber facing a price increase and it will produce a plausible set of reactions: the ones a competent researcher would list in a planning session, plus a few they might have missed. It is fast, cheap and tireless at this. Used before a study, it widens the space of things worth asking about. Used during analysis, it can propose alternative explanations for a pattern that the team is too close to see.
That is real value. It is also the whole of the value, and it stops precisely where evidence begins.
A synthetic respondent cannot tell you what your customers do. It can tell you what text about customers like yours tends to say. The difference has four sharp edges.
Circularity. The model learned from the same articles, reports and forum threads your team already read. When it confirms your expectation, you have not found a second source. You have found the first source in a different voice.
Population validity. It is easy to prompt "a 34-year-old product manager in Berlin". It is impossible to know whether the response reflects that population or the model's average of everything ever written about it — and the two diverge most on exactly the segments you know least about, which are the ones you were hoping it would cover.
Preference versus behaviour. Even with real people, what they say they'd do and what they do are different measurements. A synthetic user gives you simulated stated preference — one step further from behaviour than the least reliable real method.
Confident fabrication. The output is fluent whether it is grounded or not, and fluency is the signal people use to judge credibility. A synthetic finding arrives dressed as a real one. Nothing in the format warns you.
Put those together and you get the rule. A synthetic user can generate the question. It cannot answer it. Treat its output as a list of hypotheses to test — and never as a data point in the count.
Prompt for range, not for answers: ask for the ten ways a segment might react, not for how it does react. Keep synthetic output in a separate column from real data, labelled, so it never gets averaged in by accident. Use it to design instruments — interview guides, survey options, edge cases — and then go and get real people. When a synthetic finding and a real finding disagree, the real one wins without a meeting. And when they agree, ask whether the model could have known this from the internet; if it could, agreement is not confirmation.
The question underneath all of this is the one the Institute exists to ask. When a machine can produce an answer instantly, how does a person decide whether the answer deserves belief? Synthetic users are a clean case, because the temptation is so visible: the output looks like evidence, costs nothing, and confirms what you hoped. That is not a research problem. It is a judgment problem, and no tool is going to solve it for you.
Four minutes. One real decision. Take the diagnostic →
Institute for Human Reasoning is a research and education organisation strengthening human reasoning for consequential decisions in a technology-shaped world.
Seventy years of research on evidence and decisions, in six findings that still embarrass every company calling itself data-driven.
ReadEvery company will say yes. The only measurement that means anything is where AI enters a decision — and whether anyone checks it when it disagrees with the boss.
Read