


Originally a critique of psychological samples, the term is now central to the question of whose voice a language model reflects.
In 2010, Henrich, Heine and Norenzayan coined the acronym WEIRD: Western, Educated, Industrialized, Rich, Democratic. They showed that many findings in behavioural research are based on samples from such societies, which in important respects are particularly unrepresentative of humanity.
Language models inherit this imbalance from their training data. Atari and colleagues found that model responses most closely resemble people from WEIRD societies and that the similarity decreases markedly with growing cultural distance from the USA. The relationship was strong: r = -0.70 across 49 countries.
From a European perspective, WEIRD at first sounds like ourselves. But the reference point is usually the USA. German, French, Spanish or Scandinavian consumers differ in price consciousness, trust in brands, attitudes to contracts and data protection. A model that matches US patterns better than European ones introduces a bias into a market entry in Spain or a tariff in Austria that is not visible in the result.
This is particularly relevant for multi-country decisions: if differences between markets appear too small, the choice easily falls on a single solution that does not hold up in individual countries.
In addition, not all groups within a country are equally well represented in text. Older people, rural regions or people who write little online leave fewer traces from which a model can learn.
How to deal with it: with synthetic results, ask what data the model relies on for your market. Use synthetic respondents for hypotheses about country differences, and check decisions on market entry and pricing in each country with real people.
A consumer goods manufacturer is evaluating a refill pack in Germany, France and Italy. Synthetic respondents show similar agreement of 58 to 61% in all three countries. A behavioural test with real people in all three markets, with the same offer and the same target group definition, reveals clear differences in measured purchase intent between the countries.
For the question of which market to launch in first, precisely this difference is the information.
WEIRD bias describes whose patterns a model reflects. Variance compression describes how uniformly it responds. The two often occur together. WEIRD bias differs from the sampling error of traditional surveys in that it does not decrease with more simulated cases.
The key study on language models is available as a preprint, and newer models are being deliberately trained on more languages and cultures. The extent may therefore change. The question of how well a model knows your specific market nonetheless remains valid.
Behavioural tests have their own limitation: they measure real people in the market where the ads run and cannot simply be transferred to other countries. That is why Horizon measures each market separately for multi-country questions.
Henrich, Heine & Norenzayan 2010: Coins the acronym WEIRD (Western, Educated, Industrialized, Rich, Democratic) and shows that samples from these societies are particularly unrepresentative of humanity as a whole in many areas of behavioural research. The weirdest people in the world?, Behavioral and Brain Sciences 33(2-3). Source
Atari et al. 2023: Language model responses most closely resemble people from WEIRD societies; the similarity decreases markedly with cultural distance from the USA (r = -0.70, 49 countries, World Values Survey data). Which Humans?, PsyArXiv Preprint. Source
Yes, because the reference point of many models is US patterns. Europe is WEIRD, but not identical to the USA.
Language changes the answers, but does not automatically make up for missing data on local buying habits.
With real people in each market, the same offer and the same target group definition, so that differences are due to the market and not to the set-up.
Bring your decision question, and we will outline a possible test design.
You will speak with Daniel Putsche
Founder & CEO, 30 minutes
Read more