


Within a short time, synthetic research has gone from a niche topic to one of the most discussed topics in the industry. This overview sets out what it is suited for, how its quality is measured and at which point in a decision the behaviour of real people is decisive.
Large language models are trained on enormous amounts of text: forums, reviews, articles, survey results. If you ask such a model how a particular person would answer a question, it generates an answer that matches what people with a similar profile typically say in such texts. Several methods have emerged from this capability, and they come together under the term synthetic research.
The simplest form is synthetic respondents: a model answers a questionnaire on behalf of many simulated people with given characteristics such as age, income or region. AI personas go one step further and condense a segment into a character you can talk to. Digital twins are additionally enriched with real data about a target group or individual people, for example with earlier surveys, interviews or CRM data. Synthetic data in the narrower sense supplements or replaces missing values in existing datasets, for instance to fill up small subgroups.
What all these methods have in common: a language model reproduces what people have said and written. It condenses existing knowledge about a target group and makes it queryable. That is valuable for many tasks. But it is something different from observing what people do when they face a concrete offer with a price and alternatives.
Its value lies mainly at the start of a decision process. Synthetic respondents make exploration faster: formulating hypotheses, describing segments, collecting objections, checking questionnaires and discussion guides in advance, reducing a long list of ideas or claims to a shorter one. What used to cost weeks and a fieldwork budget can be run through in hours.
This works particularly well where a lot is known about a category: for incremental developments of established products, for well-known brands, for questions that have often been asked and answered in a similar form. Research, too, increasingly describes language models as collaborators in the research process, for example in drafting discussion guides, moderating interviews or condensing open-ended answers.
Two roles can be found in the market: synthetic methods as a pre-selection before surveying real people, and synthetic methods as a replacement for surveys. To assess them, what matters is which question is to be answered and what the result was measured against.
The quality of synthetic research is checked in validation studies. A result is compared with a reference measure, the so-called ground truth. In almost all known studies, this reference measure is a survey of real people: answers in the General Social Survey, election surveys, values surveys, conjoint studies or purchase intent asked on a scale.
This is methodologically sound and answers an important question: how well does a model replace a survey? It does not, however, answer the question of how well a model reflects behaviour. Even studies that attest a high level of agreement for synthetic respondents measure them against surveys, not against behaviour. If the survey itself reflects later behaviour only to a limited extent, a model that replicates the survey well inherits that gap.
For decisions about new offers, this is the core point. Research on the Say-Do Gap shows that stated purchase intent reflects sales of new products less well than of existing ones (Morwitz, Steckel & Gupta 2007). A model that matches stated purchase intent perfectly does not yet match who actually chooses.
Results vary widely depending on the task, and both sides of the debate have robust studies. On one side there are impressive figures: agents based on in-depth interviews reproduce answers in the General Social Survey 85% as accurately as the respondents themselves repeat their answers two weeks later (Park et al. 2024). Across 57 product surveys, synthetic purchase intent reaches 90% of the test-retest reliability of human respondents (Maier et al. 2025).
On the other side, studies show systematic deviations. In a study by the Nuremberg Institute for Market Decisions, AI answers deviated from the answers of real respondents for 75 to 80% of the questions; well-known brands were overestimated, less well-known ones underestimated (Kaiser & Manewitsch 2024). Averages can fit well while the spread is artificially low and relationships between characteristics flip: in a replication of a US election survey, around 48% of the estimated relationships deviated significantly (Bisbee et al. 2024). And agreement decreases with a country's cultural distance from the USA (Atari et al. 2023), which is relevant for target groups in Germany, France or the Nordics.
Both findings fit together. Where the model has a lot of material and is asked about averages, it performs well. Where differences, niches, new features or other cultural contexts are involved, it becomes less reliable. A study on willingness to pay shows this directly: for existing product features the values match conjoint results, for new features only after fine-tuning with earlier human data (Brand, Israeli & Ngwe 2023). What remains important in all cases: the benchmark was surveys.
The question is therefore not whether synthetic research is good or bad, but which evidence it provides for which decision. A helpful distinction is the one between incremental decisions in a familiar setting and decisions about new offers: a new tariff, a new pricing model, a product in a category in which the brand is not yet present.
For incremental questions, a model's data basis is dense, and a synthetic pre-selection saves time and budget. For new offers, this basis is missing: there are no texts about how people reacted to an offer that does not yet exist. In addition, synthetic respondents experience no friction. They have no budget, no alternatives in the next tab, no risk of regretting a bad purchase. They reflect language patterns, not trade-offs.
This does not only apply to innovation. For prices, tariffs, claims and brand presentation of existing products, too, the question is whether a change increases or lowers purchase or sign-up intent. A synthetic pre-selection can reduce the number of options. Which price or which claim holds up at the moment of the offer is shown by the measured choice of real people.
The more expensive and the newer the decision, the more the behaviour of real people counts. Synthetic respondents make exploration faster: hypotheses, segments, a first pre-selection. Whether people choose a new offer is shown by their behaviour. That is why the final variants go into a behavioural test with real people before the investment.
First: define its role in the process. Synthetic methods are suited to exploration, pre-selection and preparing studies. A decision about a new offer needs data from real people, and the question of whether they choose needs behavioural data.
Second: ask about the reference measure. Anyone presented with a validation should clarify what the model was measured against, in which category, in which country and whether existing or new products were involved. A high level of agreement with a survey is a statement about the survey.
Third: watch the spread and the niches. If all simulated people answer similarly or the best-known option comes out on top, that is often a feature of the model and not a characteristic of the target group.
Fourth: close the funnel deliberately. AI widens the funnel. At the end of the funnel is the behavioural test for the few variants that will actually receive investment. At Horizon, that is up to six variants, shown as realistic offers via Google and Meta ads to real people in their usual online environment, without a panel and without an incentive. Nothing is sold; anyone who chooses an offer learns afterwards that it is a test. What is measured is purchase or sign-up intent per variant.
An insurer is planning a new supplementary dental tariff and has formulated 20 possible benefit promises. Using synthetic respondents, the team runs the promises through three target group profiles, collects objections and sharpens the wording. After two days, four promises are on the shortlist; the synthetic ranking puts promise A clearly in front.
The four promises go into a behavioural test with real people as realistic offer pages at an identical price. After about four weeks, promise C leads on measured sign-up intent, with A in second place. The accompanying analysis shows one reason: A sounds convincing in the description but loses at the moment of the offer to C, which states a concrete waiting period. The synthetic pre-selection has served its purpose; the decision rests on behaviour.
Synthetic respondents: simulated answers to surveys, the basic method. Synthetic consumers: a term often used synonymously. Silicon sampling: the scientific origin, simulated samples from demographic profiles. Algorithmic fidelity: the measure of how well a model reproduces the answer patterns of groups.
AI personas: segments as conversational characters for exploration and objections. Digital twin: a model of a target group or person enriched with real data. Synthetic data and synthetic boosting: supplemented or filled-in values in existing datasets.
Variance compression and WEIRD bias: two known distortions, answers drifting towards the middle and a closeness to US patterns. Ground truth: the reference measure of a validation. AI concept screening and AI-generated concepts: their use before and within the innovation funnel. AI in market research and hybrid research: the wider framework in which synthetic methods, surveys and behaviour each have a clear role.
Synthetic research is developing quickly, and this overview describes the state of published research. Models trained on observed behaviour rather than on text and surveys are a different class of methods and are not assessed here. The statements refer to language models that simulate survey answers, and to decisions about new offers.
The behavioural test has limitations too. It measures whether and which variant people choose, but does not on its own explain why. It needs sufficient reach within the target group, a realistic offer and a metric defined in advance. It tests a few variants, not a list of 50 ideas. That is why the methods complement each other: synthetic exploration for breadth, surveys for motives, the behavioural test for the choice between the final variants.
Park et al. 2024: Agents based on interviews with 1,052 US citizens reproduce answers in the General Social Survey 85% as accurately as the people themselves repeat their answers two weeks later. The benchmark is survey answers and lab games, not purchasing behaviour. Generative Agent Simulations of 1,000 People, arXiv 2411.10109. Source
Maier et al. 2025: Across 57 product surveys with 9,300 human responses, synthetic purchase intent reaches 90% of the humans' test-retest reliability. The benchmark is purchase intent asked on a Likert scale, not observed behaviour. LLMs Reproduce Human Purchase Intent via Semantic Similarity Elicitation of Likert Ratings, arXiv 2510.08338. Source
Kaiser & Manewitsch (NIM) 2024: Compared with human respondents, AI answers deviate for 75% (soft drinks) and 80% (sportswear) of the questions. Well-known brands are overestimated, less well-known ones underestimated, and answers are more positive and more uniform. Synthetic Respondents (Synthetische Befragte), Nuremberg Institute for Market Decisions. Source
Bisbee et al. 2024: Averages of synthetic answers are close to a US election survey, but the spread is artificially low. Around 48% of the estimated relationships deviate significantly, 32% of them with the opposite sign. Identical prompts deliver different results over three months. Synthetic Replacements for Human Survey Data? The Perils of Large Language Models, Political Analysis 32(4). Source
Atari et al. 2023: The similarity between model answers and human answers decreases with a country's cultural distance from the USA (r = -0.70). The benchmark is international values surveys. Which Humans?, PsyArXiv preprint. Source
Brand, Israeli & Ngwe 2023: Language models plausibly reproduce willingness to pay for existing product features, but for new features only after fine-tuning with earlier human results. The benchmark is conjoint studies, i.e. surveys. Using LLMs for Market Research, Harvard Business School Working Paper 23-062. Source
For exploration, pre-selection and checking questionnaires, they can partly anticipate surveys. Studies show good agreement on averages, but deviations in spread, niches, new features and target groups outside the USA.
In almost all published studies, against surveys of real people, for example the General Social Survey or stated purchase intent. A high level of agreement therefore means that the model replicates the survey well, not necessarily behaviour.
For the early phase: forming hypotheses, describing segments, collecting objections and reducing a long list to a few candidates, especially in familiar categories.
A language model reproduces what people have said and written about existing products. For offers that do not yet exist, this basis is missing, and simulated respondents experience neither price nor alternatives nor consequences.
In its tests, Horizon measures the purchase or sign-up intent of real people in response to realistic offers. Synthetic methods can be useful beforehand to select the variants that go into the test.
Synthetic respondents are usually generated from demographic profiles. A digital twin is additionally enriched with real data about a target group or person and therefore reflects more precisely what is known about them.
Bring your decision question and your pre-selection, and we will outline a possible test design.
You will speak with Daniel Putsche
Founder & CEO, 30 minutes
Read more