Horizon
Home
/
Glossary
/
Ground Truth

Ground Truth

Ground truth is the reference measure against which the quality of a measurement or a model is checked; in market research, it determines whether a result is measured against survey answers or against observed behaviour.

Updated
September 28, 2026
· Horizon

When a method is described as "valid" or "highly consistent", it is worth asking: consistent with what? The reference measure determines what a validation actually proves.

Why this matters for your decision

New methods, such as synthetic respondents or digital twins, are often described with agreement scores. These scores are only as meaningful as the reference measure against which they were measured. If the reference measure is a survey, a high level of agreement shows that the method replicates survey answers well. Whether people choose an offer has not yet been checked.

For decisions about new offers, prices or tariffs, this means: the question about the reference measure belongs at the start of every method assessment, not at the end.

Survey or behaviour as the benchmark

A language model reproduces what people have said and written. Even studies that attest a high level of agreement for synthetic respondents measure them against surveys, not against behaviour. Two widely noted papers show this: in one, AI agents reach 86% of the consistency with which people repeat their own answers in a large social survey. In the other, synthetic answers to purchase intent questions reach around 90% of the test-retest reliability of human respondents. Both results are methodologically remarkable, and both use survey answers as the reference measure.

Survey answers themselves, however, deviate from behaviour: intentions are acted upon in about half of cases, and hypothetically stated willingness to pay is on average higher than the real one. A method that replicates surveys well inherits this deviation.

Example

A team is assessing a method that simulates willingness to buy for three variants of a supplementary dental tariff. The provider's documentation shows a high level of agreement with an earlier survey. The team asks two questions: what was the agreement measured against, and were the tested offers already on the market?

The answer: measured against scale answers, for existing tariffs. For the new tariff, there is therefore no evidence based on behaviour. The team uses the method for pre-selection and checks the two remaining variants in a behavioural test.

How it differs

Validity describes whether a method measures what it is supposed to measure; the reference measure is the benchmark against which this is checked. Reliability describes whether repeated measurements agree, and says nothing about whether the right thing is being measured. A method can be highly reliable and still be calibrated against the wrong reference measure.

Limitations

Observed behaviour in the Painted Door Test is not a complete reference measure for the later market either: it shows the decision when faced with a realistic offer, not repeat purchases, returns or competitor reactions. The most honest benchmark for a method is therefore always the one that comes closest to your own decision.

Synthetic respondents make exploration faster: hypotheses, segments, a first pre-selection. Whether people choose a new offer is shown by their behaviour. That is why Horizon measures the purchase intent of real people for the final variants before the investment.

Evidence

Park et al. 2024: AI agents based on interviews with 1,052 people reach 86% of the consistency with which the people repeat their own answers in the General Social Survey two weeks later. The benchmark is survey answers. LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals, arXiv 2411.10109. Source

Maier et al. 2025: Across 57 product surveys with 9,300 responses, synthetic answers reach around 90% of the test-retest reliability of human respondents. The benchmark is purchase intent asked on a scale, not behaviour. LLMs Reproduce Human Purchase Intent via Semantic Similarity Elicitation of Likert Ratings, arXiv 2510.08338. Source

Sheeran & Webb 2016: Intentions are translated into action in about half of cases. The Intention-Behavior Gap, Social and Personality Psychology Compass 10(9). Source

Schmidt & Bijmolt 2020: Across 77 studies and 115 effect sizes, hypothetically stated willingness to pay is on average 21% above the real one. Accurately measuring willingness to pay for consumer goods: a meta-analysis of the hypothetical bias, Journal of the Academy of Marketing Science 48. Source

Frequently asked questions

What should you ask a provider of synthetic data?

What the agreement was measured against, a survey or behaviour, and whether the tested offers were already on the market.

Is a survey a poor reference measure?

No. For attitudes and motives, it is the right measure. For the question of whether people choose a new offer, observed behaviour is closer to the decision.

Is behaviour in a test the same as sales?

No. It shows the decision when faced with a realistic offer, not how things develop later in the market.

Which benchmark do you want to measure your decision against?

Bring your decision question, and we will outline a possible test design.

Daniel Putsche

You will speak with Daniel Putsche
Founder & CEO, 30 minutes

Read more

Thank you, we have your request.

We will get back to you within one working day with suggested times.

Close

Demo · Video

A complete test, from design to data analysis

Play the demo

Chapters: test design · offer pages · live data · result

Book a call

First call

Book a call

30 minutes, your decision question, a possible test design. No preparation needed on your side.

By submitting, you agree to the processing of your details to arrange a call. Details in our privacy policy.

Request a call
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.
SAMPLE REPORTExample

Sample report: insurance

Data analysis · Test design · Metrics · Methodology

Sample report

Sample report: insurance

A complete results report with example values: research question, test design, purchase intent per variant and the data analysis.

12 pages, as a PDF to share internally
Free, download right after a short form
Note