


Synthetic data in market research are artificially generated data points that supplement or replace real data collection, for example by filling up small subgroups, estimating missing values or fully simulating respondents.
Synthetic data is the umbrella term covering very different methods. For evaluation, it helps to keep them apart, because they have different strengths and risks.
When a study refers to synthetic data, different things can be meant. Imputation estimates individual missing answers based on existing data, a statistical method established for decades. Synthetic boosting fills up small subgroups of a real sample with modelled cases. Fully simulated respondents replace data collection entirely. The larger the synthetic share, the more the result depends on the model and the less on the people surveyed.
For decisions, the key question is therefore how much of the result is observed and how much is generated. Synthetic data supplements existing evidence; it does not create new evidence. For offers that do not yet exist, the data basis from which to supplement is missing. Whether people choose a new offer is shown by their behaviour.
A survey on a new mobile tariff has 1,000 respondents, including only 60 self-employed people. To be able to analyse this group, it is synthetically boosted to 300 cases. The boosted group shows a similar preference to the 60 real cases, only with a seemingly narrower confidence interval. The uncertainty is in fact that of the 60 people. For the decision on whether to launch a tariff for the self-employed, more observations of this group are needed, for example in a behavioural test with targeted outreach.
What share of the cases is observed, and what share is generated? On what data basis was the model trained or calibrated, and from which country and category does it come? Against what was the method validated, surveys or behaviour? Are confidence intervals reported on the basis of the real or the boosted number of cases?
If these questions can be answered, synthetic data can be put into proper context. For the decision on a new offer, the yardstick remains how real people behave when faced with this offer.
Synthetic respondents are a form of synthetic data in which the entire sample is simulated. Synthetic boosting supplements only subgroups. Synthetic data for data protection, such as anonymised copies of customer data, pursues a different goal and is not meant here. Behavioural data from a Painted Door Test is observed data: every measured decision comes from a real person.
Synthetic answers often show artificially low variance. In a replication of a US election survey, means were close to the original, but around 48% of the estimated relationships deviated significantly, and identical prompts produced different results over three months (Bisbee et al. 2024). Anyone using synthetic data should disclose what share is generated and not artificially shrink uncertainty. A behavioural test, for its part, only provides data on the variants and target groups tested, not on all conceivable ones.
Bisbee et al. 2024: Means of synthetic answers are close to a US election survey, but variance is artificially low. Around 48% of the estimated relationships deviate significantly, 32% of them with the opposite sign. Identical prompts produce different results over three months. Synthetic Replacements for Human Survey Data? The Perils of Large Language Models, Political Analysis 32(4). Source
Argyle et al. 2023: Coins the terms silicon sample and algorithmic fidelity: a language model prompted with the demographic backgrounds of real respondents reproduces the answer distributions of many subgroups from US election surveys well. The yardstick is survey data. Out of One, Many: Using Language Models to Simulate Human Samples, Political Analysis 31(3). Source
No. Synthetic respondents are a form of synthetic data in which the whole sample is simulated. Other forms supplement only individual values or subgroups.
For estimating individual missing values, for exploration and for first analyses of small groups, as long as it remains transparent what share is generated.
It increases the number of cases, not the information. The uncertainty remains that of the real observations.
Bring your decision question, and we will outline a possible test design.
You will speak with Daniel Putsche
Founder & CEO, 30 minutes
Read more