


Synthetic respondents are people simulated by an AI language model who answer questionnaires on behalf of a target group, based on given attributes such as age, income or region and without collecting data from real people.
Synthetic respondents are the basic method of synthetic research. They deliver answers in hours for which a survey needs weeks, and are therefore suited above all to the early phase of a decision.
Synthetic respondents make exploration faster: hypotheses, segments, an initial pre-selection. Teams can run through a long list of ideas, claims or tariff components, collect objections and pre-test questionnaires before any field budget is spent.
For interpreting the results, the yardstick is decisive. A language model reproduces what people have said and written. Even studies that attest a high level of agreement for synthetic respondents measure them against surveys, not against behaviour. In one study of 57 product surveys, for example, synthetic purchase intent reaches 90% of the test-retest reliability of human respondents (Maier et al. 2025); the benchmark was stated purchase intent on a scale. Whether people choose a new offer is shown by their behaviour. That is why the final variants go into a behavioural test with real people before the investment.
A consumer goods manufacturer is developing a refill system for cleaning products and has twelve wordings for the main promise. Synthetic respondents in three profiles rate the wordings, and the team drops eight and sharpens four. The four variants go into a behavioural test as realistic offer pages at an identical price. The synthetic ranking had variant B in front; on measured purchase intent, variant D is in front and B is in third place.
Synthetic consumers is usually a synonym. AI personas condense a segment into a conversational figure and serve dialogue rather than counting. A digital twin is additionally enriched with real data from a target group or person. Silicon sampling is the scientific origin of the method. Synthetic respondents differ from surveys of real people in that nobody is asked, and from a Painted Door Test in that no choice is observed.
Alongside good averages, studies show systematic deviations: artificially low variance and relationships that flip (Bisbee et al. 2024), overestimated well-known brands and underestimated lesser-known ones (Kaiser & Manewitsch 2024), and declining agreement with increasing cultural distance from the US (Atari et al. 2023). Synthetic respondents experience neither price nor alternatives nor consequences. For offers that do not yet exist, the model lacks the basis.
A behavioural test also has limits: it tests a few final variants, not dozens of ideas, and only partly explains the why. That is why the two methods complement each other in the process.
Maier et al. 2025: Across 57 product surveys with 9,300 human responses, synthetic purchase intent reaches 90% of the test-retest reliability of humans. The benchmark is stated purchase intent on a Likert scale, not observed behaviour. LLMs Reproduce Human Purchase Intent via Semantic Similarity Elicitation of Likert Ratings, arXiv 2510.08338. Source
Bisbee et al. 2024: Means of synthetic responses are close to a US election survey, but the variance is artificially low. Around 48% of the estimated relationships deviate significantly, 32% of them with the opposite sign. The same prompts produce different results over three months. Synthetic Replacements for Human Survey Data? The Perils of Large Language Models, Political Analysis 32(4). Source
Kaiser & Manewitsch (NIM) 2024: Compared with human respondents, AI answers deviate on 75% (soft drinks) and 80% (sportswear) of the questions. Well-known brands are overestimated, lesser-known ones underestimated, and answers are more positive and more uniform. Synthetische Befragte (Synthetic Respondents), Nuremberg Institute for Market Decisions. Source
Atari et al. 2023: The similarity between model responses and human responses declines with a country's cultural distance from the US (r = -0.70). The benchmark is international values surveys. Which Humans?, PsyArXiv Preprint. Source
Often well on averages, considerably worse on variance, niches, lesser-known brands and target groups outside the US. The yardstick is almost always a survey.
For exploration, collecting objections, pre-testing questionnaires and pre-selecting from many ideas or wordings.
Stated purchase intent can be simulated. Whether people actually choose a new offer is shown by behaviour, for example in a Painted Door Test.
Bring your decision question, and we will outline a possible test design.
You will speak with Daniel Putsche
Founder & CEO, 30 minutes
Read more