


When a method is described as "valid" or "highly consistent", it is worth asking: consistent with what? The reference measure determines what a validation actually proves.
New methods, such as synthetic respondents or digital twins, are often described with agreement scores. These scores are only as meaningful as the reference measure against which they were measured. If the reference measure is a survey, a high level of agreement shows that the method replicates survey answers well. Whether people choose an offer has not yet been checked.
For decisions about new offers, prices or tariffs, this means: the question about the reference measure belongs at the start of every method assessment, not at the end.
A language model reproduces what people have said and written. Even studies that attest a high level of agreement for synthetic respondents measure them against surveys, not against behaviour. Two widely noted papers show this: in one, AI agents reach 86% of the consistency with which people repeat their own answers in a large social survey. In the other, synthetic answers to purchase intent questions reach around 90% of the test-retest reliability of human respondents. Both results are methodologically remarkable, and both use survey answers as the reference measure.
Survey answers themselves, however, deviate from behaviour: intentions are acted upon in about half of cases, and hypothetically stated willingness to pay is on average higher than the real one. A method that replicates surveys well inherits this deviation.
A team is assessing a method that simulates willingness to buy for three variants of a supplementary dental tariff. The provider's documentation shows a high level of agreement with an earlier survey. The team asks two questions: what was the agreement measured against, and were the tested offers already on the market?
The answer: measured against scale answers, for existing tariffs. For the new tariff, there is therefore no evidence based on behaviour. The team uses the method for pre-selection and checks the two remaining variants in a behavioural test.
Validity describes whether a method measures what it is supposed to measure; the reference measure is the benchmark against which this is checked. Reliability describes whether repeated measurements agree, and says nothing about whether the right thing is being measured. A method can be highly reliable and still be calibrated against the wrong reference measure.
Observed behaviour in the Painted Door Test is not a complete reference measure for the later market either: it shows the decision when faced with a realistic offer, not repeat purchases, returns or competitor reactions. The most honest benchmark for a method is therefore always the one that comes closest to your own decision.
Synthetic respondents make exploration faster: hypotheses, segments, a first pre-selection. Whether people choose a new offer is shown by their behaviour. That is why Horizon measures the purchase intent of real people for the final variants before the investment.
Park et al. 2024: AI agents based on interviews with 1,052 people reach 86% of the consistency with which the people repeat their own answers in the General Social Survey two weeks later. The benchmark is survey answers. LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals, arXiv 2411.10109. Source
Maier et al. 2025: Across 57 product surveys with 9,300 responses, synthetic answers reach around 90% of the test-retest reliability of human respondents. The benchmark is purchase intent asked on a scale, not behaviour. LLMs Reproduce Human Purchase Intent via Semantic Similarity Elicitation of Likert Ratings, arXiv 2510.08338. Source
Sheeran & Webb 2016: Intentions are translated into action in about half of cases. The Intention-Behavior Gap, Social and Personality Psychology Compass 10(9). Source
Schmidt & Bijmolt 2020: Across 77 studies and 115 effect sizes, hypothetically stated willingness to pay is on average 21% above the real one. Accurately measuring willingness to pay for consumer goods: a meta-analysis of the hypothetical bias, Journal of the Academy of Marketing Science 48. Source
What the agreement was measured against, a survey or behaviour, and whether the tested offers were already on the market.
No. For attitudes and motives, it is the right measure. For the question of whether people choose a new offer, observed behaviour is closer to the decision.
No. It shows the decision when faced with a realistic offer, not how things develop later in the market.
Bring your decision question, and we will outline a possible test design.
You will speak with Daniel Putsche
Founder & CEO, 30 minutes
Read more