


Variance compression is the tendency of synthetic respondents to answer less diversely than real people, so that answers drift towards the middle and fringe opinions, niches and unusual reactions are under-represented in the results.
An average can be right while the distribution behind it is wrong. That is exactly what variance compression describes.
Language models generate likely answers. Likely means close to what occurs frequently in the training data. When a model is asked to play a thousand different people, the answers are therefore often more similar to one another than those of a thousand real people. Extreme agreement, clear rejection and the typical don't-know answers become rarer.
Bisbee and colleagues showed in 2024 that synthetic answers were close on average to a large US election study, but varied considerably less. Wang and colleagues found in 2025, in studies with 3,200 participants, that models portray demographic groups too flatly: the diversity within a group is lost.
Many decisions depend not on the average but on the edges. A new offer often lives on a small group of early buyers. A price increase fails because of the share that cancels. A niche brand grows with people who deviate from the mainstream. If the method smooths out precisely these edges, a concept that in truth polarises looks unremarkable.
A study by the NIM observed more positive answers, less variance and a preference for well-known brands among synthetic respondents. Niches and the new thus disappear into the middle.
A mobile network operator tests a tariff with no minimum contract term at a premium. Synthetic respondents give it an average of 6.1 out of 10 points, with almost all values between 5 and 7. A real survey gives a similar average of 5.8, but two camps: a quarter rate it 9 or 10, a third 3 or less.
For the decision, the quarter with high approval is the real information. Whether it also acts is shown by a behavioural test in which people see the tariff with its price.
Variance compression is not the same as a wrong average. It concerns the shape of the distribution. Related is the WEIRD bias, which describes whose patterns a model reproduces. Algorithmic fidelity asks how well distributions are matched overall. The acquiescence bias of real respondents is a separate effect that occurs in surveys with people.
With synthetic results, always look at variance and distribution, not just at averages. Deliberately take unusual concepts into the next step, even if they perform only moderately in screening. And check decisions that depend on minorities or early buyers with real people.
Newer methods try to increase variance artificially, for example through different prompting techniques. This improves the shape of the distribution, but does not create a new observation. A behavioural test also has limitations: it shows how many people act, not automatically why. For the why, qualitative research adds to it.
Bisbee et al. 2024: Averages of synthetic answers are close to the US election study ANES, but the answers vary less than in the real survey, regression coefficients often deviate considerably, and the same prompt produces markedly different results over three months. Synthetic Replacements for Human Survey Data? The Perils of Large Language Models, Political Analysis 32(4). Source
Wang, Morgenstern & Dickerson 2025: In studies with 3,200 participants from 16 demographic groups and four language models, the models portray groups in a distorted way and represent their internal diversity too flatly. Large language models that replace human participants can harmfully misportray and flatten identity groups, Nature Machine Intelligence. Source
NIM 2024/2025: GPT-4-based synthetic respondents deviate considerably from 500 real respondents on 75% (soft drinks) and 80% (sportswear) of questions respectively, overestimate well-known brands, underestimate lesser-known ones and answer more positively and with less variance. Synthetische Befragte, Forschungsvorhaben des Nürnberg Institut für Marktentscheidungen. Source
Partly, for example through calibration against real survey data. But that requires exactly the human data that was meant to be replaced, and the calibration only applies to similar questions.
No, it also affects ratings of concepts, prices and brands. Wherever the edges count, the effect influences the decision.
By conspicuously narrow distributions, missing don't-know answers and small differences between segments that would be clearer in real data.
Bring your decision question, and we will outline a possible test design.
You will speak with Daniel Putsche
Founder & CEO, 30 minutes
Read more