


AI concept screening is the pre-selection of product, tariff or communication concepts with language models that simulate respondents and rank concepts by the agreement they would be expected to receive in a survey, before real people are surveyed or tested.
If you have to pick three out of twenty concept ideas, a language model can produce a first ranking within hours. The question is what this ranking says and what it does not.
Concept screening with AI mainly changes the width of the funnel. Where five concepts used to go into a survey, dozens of variants of value propositions, tariff names or claims can now be pre-sorted. This saves time and budget in a phase in which mistakes are still cheap.
At the same time, a spurious result is easily created: a clean ranking with percentages looks like a finding, but it rests on a model that reproduces what people have said and written. For the decision on which concept gets budget, what counts is therefore which question the screening answers: which concepts are plausible and understandable? It can do that well. Which concept do people choose when it is a real offer with a price? Their behaviour shows that.
This does not only concern innovations. AI screening is now also used for claims, brand presentation, tariff names and price anchors for existing products, to reduce the number of variants before a survey or a test.
A language model receives personas, for example age, region and household situation, and rates each concept the way such a person would in a survey. Newer methods let the model answer freely and then translate the text into a scale. The result is a distribution of simulated agreement scores per concept.
A study across 57 surveys on personal care products shows how far this now goes: the simulated answers achieve a large part of the test-retest reliability of human respondents. The benchmark here is purchase intent stated in the survey, not measured behaviour.
An insurer has twelve concepts for a supplementary dental tariff: different benefit focuses, names and price anchors. An AI screening sorts out four concepts that seem incomprehensible to target groups over 50 and flags two concepts with the highest simulated agreement, 71 and 68 out of 100 points.
These two and a third, deliberately different concept then go into a behavioural test with real people. There it becomes clear which variant achieves the highest measured sign-up intent when price and alternatives are visible.
Classic concept screening surveys real people about concepts. AI concept screening simulates this survey. Both measure stated agreement, that is stated preference. A behavioural test such as the Painted Door Test, by contrast, measures whether people act in a realistic offer situation.
Screening differs from AI-generated concepts in that it does not create ideas but ranks existing ones. In practice, the two often run one after the other.
Language models tend towards the middle: they favour familiar patterns and tend to rate unusual concepts cautiously. In a study by the NIM, synthetic respondents overestimated well-known brands and underestimated less well-known ones. The very concept that breaks with habits can therefore drop out too early in screening. It is worth deliberately passing on a divergent concept.
The behavioural test has limits too: it tests a few mature variants, not dozens of raw ideas, and it needs a stimulus that looks like a real offer. That is why the two steps complement rather than replace each other.
AI widens the funnel. At the end of the funnel is the behavioural test for the few variants that are actually invested in. There, Horizon measures the purchase intent of real people with up to six variants, and nothing is sold.
Maier et al. 2025: Across 57 surveys on personal care products with 9,300 responses, simulated answers achieve 90 % of the test-retest reliability of human respondents. The benchmark is purchase intent stated in the survey, not behaviour. LLMs Reproduce Human Purchase Intent via Semantic Similarity Elicitation of Likert Ratings, arXiv Preprint. Source
Brand, Israeli & Ngwe 2025: A language model reproduces willingness to pay from conjoint studies for existing attributes; for new attributes, it needs fine-tuning with earlier human data from the same category, while data from other categories hardly helps. Using LLMs for Market Research, MSI Working Paper. Source
NIM 2024/2025: Synthetic respondents based on GPT-4 deviate clearly from 500 real respondents on 75 % (soft drinks) and 80 % (sportswear) of questions, overestimate well-known brands, underestimate less well-known ones and answer more positively and with less variance. Synthetic respondents, research project of the Nuremberg Institute for Market Decisions. Source
For an initial pre-selection, it can often come before a survey and shorten it. For the question of which concept people choose, it provides no evidence of its own, because it is calibrated on surveys.
Usually two to six. It makes sense to include a deliberately divergent concept alongside the front-runners, because models tend to underestimate the unfamiliar.
For a rough assessment of price anchors, yes. Willingness to pay for new offers remains a question of behaviour.
Bring your decision question, and we will outline a possible test design.
You will speak with Daniel Putsche
Founder & CEO, 30 minutes
Read more