


A digital twin in market research is an AI model of a target group or an individual person that is enriched with real data such as surveys, interviews or customer data and generates answers on behalf of that target group.
The term comes from industry, where a digital twin represents a machine in the computer. In market research, it stands for simulated consumers built on a real data basis.
Digital twins make existing knowledge about a target group queryable. Instead of searching through last year's study, a team can ask the twin new questions: how does the segment react to this wording? Which objections are to be expected? For exploration and pre-selection, this is fast and inexpensive, especially for categories and brands for which a lot of data is available.
A digital twin reflects what is already known about a target group. For offers that do not yet exist, it lacks this basis. That is where Horizon measures the purchase intent of real people, and nothing is sold. For incremental developments in a familiar setting, the data basis is denser; for a new tariff, a new pricing model or a new category, it is thinner.
The data basis usually consists of surveys and interviews. A public research dataset for digital twins covers 2,058 US respondents who answered around 500 questions over four waves (Toubia et al. 2025). Agents based on in-depth interviews with 1,052 people reproduce answers in the General Social Survey 85% as accurately as the people themselves repeat their answers two weeks later (Park et al. 2024).
Both studies show how far the technology has come with survey answers. Both measure against surveys and lab tasks. They do not answer how well a twin reflects the choice between real offers with price and alternatives.
A home appliance manufacturer has a digital twin of its core target group built from customer surveys of the last three years. For a new cordless kitchen machine, the team has the twin rate three positionings; all three receive similarly positive scores. In the subsequent behavioural test with real people, at the same price, the measured purchase intent of the strongest positioning is clearly above that of the weakest. The twin had no data on a product type that did not yet exist in the range.
Synthetic respondents are usually created only from demographic profiles, a digital twin additionally from real data. AI personas are often conversational characters without their own data basis. A model trained on observed behaviour rather than on surveys is a different class of method. A Painted Door Test does not simulate a target group, but observes which variant real people choose.
A twin is only as good as its data basis and inherits its biases. If that basis comes from surveys, it also inherits the gap between statement and behaviour. It ages with the data, and it knows no reaction to something that did not exist before. The behavioural test, in turn, only checks the variants it is given and explains motives only in part; for the breadth beforehand, a twin is useful.
Park et al. 2024: Agents based on interviews with 1,052 US citizens reproduce answers in the General Social Survey 85% as accurately as the people themselves repeat their answers two weeks later. The benchmark is survey answers and lab games, not purchasing behaviour. Generative Agent Simulations of 1,000 People, arXiv 2411.10109. Source
Toubia et al. 2025: Open dataset for digital twins: 2,058 US respondents answer around 500 questions over four waves. Twins are checked against repeated survey and experimental tasks. Twin-2K-500: A dataset for building digital twins of over 2,000 people based on their answers to over 500 questions, arXiv 2505.17479. Source
Often, yes, because it builds on real data about the target group. But it remains tied to this data and thus to what is already known.
For products for which there is no data, it lacks the basis. There, a behavioural test with real people provides the evidence.
In the known studies, against repeated surveys and lab tasks of the same people, not against purchasing behaviour.
Bring your decision question, and we will outline a possible test design.
You will speak with Daniel Putsche
Founder & CEO, 30 minutes
Read more