


The term suggests reliability. What matters is what the fidelity is measured against.
Argyle and colleagues introduced the term in 2023, together with silicon sampling. A model is considered algorithmically faithful if its simulated answers reflect the distributions of a real survey for many subgroups, not just the average. Since then, algorithmic fidelity has been the common yardstick against which synthetic respondents are evaluated.
In 2024, Park and colleagues reported agents based on in-depth interviews that achieve 83 to 86% of the consistency with which people repeat their own answers after two weeks. In 2025, Maier and colleagues found 90% of the test-retest reliability of human respondents for purchase intent questions about personal care products. These are remarkable values.
All of these values compare models with surveys. Agreement with surveys is not agreement with behaviour. A model that perfectly matches what people state in a survey also inherits the well-known gap between statement and action. Meta-analyses show that a clear change in intention produces only a small-to-medium change in behaviour.
For you, this means: high algorithmic fidelity is a good argument for using synthetic respondents where you would otherwise have run a survey. It is not evidence that people will choose a new offer.
So for every reported agreement, ask about three things: what data it was compared against, in which category and in which country. Validation against a US opinion poll says little about a tariff in Germany.
A provider reports that its synthetic respondents agreed 88% with a survey on a new subscription-based travel cancellation insurance. In the survey, 42% said they would probably take out the insurance; the model arrives at 44%.
Both values describe stated intention. In the behavioural test with a realistic offer page, measured sign-up intent is clearly lower. Fidelity to the survey was high; the gap to behaviour remains.
Algorithmic fidelity is a form of validation against survey data. Ground truth in the sense of observed behaviour is a different yardstick. Variance compression and WEIRD bias describe typical deviations that reduce fidelity. Test-retest reliability measures how consistently people themselves answer, and often serves as an upper limit.
Fidelity depends on the question, model, prompt and target group. It holds better for topics with a large body of text than for new ones, and is better documented for US samples than for European ones. A high value in one category cannot be transferred to another.
The behavioural test also has limitations: it measures purchase intent in a realistic situation, not an actual purchase, because nothing is sold. Horizon measures where fidelity to surveys is not enough: on the question of whether real people choose a new offer.
Argyle et al. 2023: Introduces the terms silicon sample and algorithmic fidelity: a language model conditioned with socio-demographic profiles of real respondents from US surveys reproduces the answer distributions of many subgroups of these surveys well. The yardstick is survey data. Out of One, Many: Using Language Models to Simulate Human Samples, Political Analysis 31(3). Source
Park et al. 2024: Agents based on interviews with 1,052 Americans achieve 83 to 86% of the consistency with which the individuals repeat their own survey answers after two weeks; with demographics only, it is 74%. The yardstick is survey answers. LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals, arXiv. Source
Maier et al. 2025: Across 57 surveys on personal care products with 9,300 responses, simulated answers achieve 90% of the test-retest reliability of human respondents. The yardstick is purchase intent stated in the survey, not behaviour. LLMs Reproduce Human Purchase Intent via Semantic Similarity Elicitation of Likert Ratings, arXiv Preprint. Source
Yes, for replacing or supplementing surveys. For statements about behaviour, it is the wrong yardstick.
Because behavioural data on new offers is rarely available. That is exactly the gap a behavioural test closes.
The model is almost as consistent with the respondents as they are with themselves across two surveys. It remains a comparison of statements with statements.
Bring your decision question, and we will outline a possible test design.
You will speak with Daniel Putsche
Founder & CEO, 30 minutes
Read more