Horizon
Home
/
Glossary
/
Algorithmic Fidelity

Algorithmic Fidelity

Algorithmic fidelity is the degree to which a language model reproduces the answer distributions of specific population groups from existing human surveys when it is conditioned with matching personal profiles.

Updated
September 28, 2026
· Horizon

The term suggests reliability. What matters is what the fidelity is measured against.

What lies behind it

Argyle and colleagues introduced the term in 2023, together with silicon sampling. A model is considered algorithmically faithful if its simulated answers reflect the distributions of a real survey for many subgroups, not just the average. Since then, algorithmic fidelity has been the common yardstick against which synthetic respondents are evaluated.

In 2024, Park and colleagues reported agents based on in-depth interviews that achieve 83 to 86% of the consistency with which people repeat their own answers after two weeks. In 2025, Maier and colleagues found 90% of the test-retest reliability of human respondents for purchase intent questions about personal care products. These are remarkable values.

Why this matters for your decision

All of these values compare models with surveys. Agreement with surveys is not agreement with behaviour. A model that perfectly matches what people state in a survey also inherits the well-known gap between statement and action. Meta-analyses show that a clear change in intention produces only a small-to-medium change in behaviour.

For you, this means: high algorithmic fidelity is a good argument for using synthetic respondents where you would otherwise have run a survey. It is not evidence that people will choose a new offer.

So for every reported agreement, ask about three things: what data it was compared against, in which category and in which country. Validation against a US opinion poll says little about a tariff in Germany.

Example

A provider reports that its synthetic respondents agreed 88% with a survey on a new subscription-based travel cancellation insurance. In the survey, 42% said they would probably take out the insurance; the model arrives at 44%.

Both values describe stated intention. In the behavioural test with a realistic offer page, measured sign-up intent is clearly lower. Fidelity to the survey was high; the gap to behaviour remains.

How it differs

Algorithmic fidelity is a form of validation against survey data. Ground truth in the sense of observed behaviour is a different yardstick. Variance compression and WEIRD bias describe typical deviations that reduce fidelity. Test-retest reliability measures how consistently people themselves answer, and often serves as an upper limit.

Limitations

Fidelity depends on the question, model, prompt and target group. It holds better for topics with a large body of text than for new ones, and is better documented for US samples than for European ones. A high value in one category cannot be transferred to another.

The behavioural test also has limitations: it measures purchase intent in a realistic situation, not an actual purchase, because nothing is sold. Horizon measures where fidelity to surveys is not enough: on the question of whether real people choose a new offer.

Evidence

Argyle et al. 2023: Introduces the terms silicon sample and algorithmic fidelity: a language model conditioned with socio-demographic profiles of real respondents from US surveys reproduces the answer distributions of many subgroups of these surveys well. The yardstick is survey data. Out of One, Many: Using Language Models to Simulate Human Samples, Political Analysis 31(3). Source

Park et al. 2024: Agents based on interviews with 1,052 Americans achieve 83 to 86% of the consistency with which the individuals repeat their own survey answers after two weeks; with demographics only, it is 74%. The yardstick is survey answers. LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals, arXiv. Source

Maier et al. 2025: Across 57 surveys on personal care products with 9,300 responses, simulated answers achieve 90% of the test-retest reliability of human respondents. The yardstick is purchase intent stated in the survey, not behaviour. LLMs Reproduce Human Purchase Intent via Semantic Similarity Elicitation of Likert Ratings, arXiv Preprint. Source

Frequently asked questions

Is high algorithmic fidelity a mark of quality?

Yes, for replacing or supplementing surveys. For statements about behaviour, it is the wrong yardstick.

Why is validation not done against behaviour?

Because behavioural data on new offers is rarely available. That is exactly the gap a behavioural test closes.

What does 90% of test-retest reliability mean?

The model is almost as consistent with the respondents as they are with themselves across two surveys. It remains a comparison of statements with statements.

What do you measure the fidelity of your synthetic data against?

Bring your decision question, and we will outline a possible test design.

Daniel Putsche

You will speak with Daniel Putsche
Founder & CEO, 30 minutes

Read more

Thank you, we have your request.

We will get back to you within one working day with suggested times.

Close

Demo · Video

A complete test, from design to data analysis

Play the demo

Chapters: test design · offer pages · live data · result

Book a call

First call

Book a call

30 minutes, your decision question, a possible test design. No preparation needed on your side.

By submitting, you agree to the processing of your details to arrange a call. Details in our privacy policy.

Request a call
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.
SAMPLE REPORTExample

Sample report: insurance

Data analysis · Test design · Metrics · Methodology

Sample report

Sample report: insurance

A complete results report with example values: research question, test design, purchase intent per variant and the data analysis.

12 pages, as a PDF to share internally
Free, download right after a short form
Note