


Testing AI-generated concepts means first pre-sorting product, tariff or claim ideas generated by language models, and then checking the few variants that are to be invested in against the behaviour of real people.
Generative AI has accelerated idea generation. Where a workshop used to deliver ten concepts, a model delivers a hundred in minutes. The bottleneck has thus shifted from the number of ideas to their evaluation.
AI widens the funnel. At the end of the funnel comes the behavioural test for the few variants that are actually invested in. In between lies the real work: selecting, from many plausible concepts, those that deserve to be tested.
Research shows that AI ideas can perform well. In a Wharton School study, 35 of the 40 ideas in the top ten percent of 400 product ideas came from GPT-4 (Girotra et al. 2023). The yardstick is worth noting: evaluation was based on stated purchase intent on a scale. The study therefore shows that AI ideas are convincing in surveys. Whether people choose them when price and alternatives are visible is a different question.
A second finding concerns diversity. Idea pools from a language model turn out more similar than those from groups of human participants; targeted prompting narrows the gap considerably (Meincke, Mollick & Terwiesch 2024). For pre-selection, this means: many concepts may be similar at their core. All the more important is that the final variants differ clearly in one dimension, so that a test can give a clear answer.
First: generate broadly, with AI and with people. Second: pre-sort, for example by feasibility, strategic fit, a survey-based concept screening or an AI concept screening. Third: design the two to six variants between which the investment decision will be made as realistic offers and check them in a behavioural test. Fourth: explore the why behind the result with a survey or interviews.
A consumer goods manufacturer has a language model generate 60 claim ideas for a new line of dishwasher tablets. The team pre-sorts them down to twelve, and a survey narrows them to four. In the Painted Door Test, each person sees an offer with exactly one of the four claims; price and pack are identical. Around 2,000 visitors per variant reach the offer page. The claim that ranked third in the survey achieves the highest measured purchase intent at 3.4 percent, the survey winner 2.5 percent.
AI concept screening evaluates concepts with a model, for example via synthetic respondents. That is a pre-selection step. Testing AI-generated concepts describes the entire process from generation to behavioural test. Synthetic respondents make exploration faster: hypotheses, segments, a first pre-selection. Whether people choose a new offer is shown by their behaviour.
The studies cited come from specific categories and target groups and cannot simply be transferred. The models are developing quickly, and statements about their strengths may look different in a year. A behavioural test also has limitations: it only checks the variants that are put into it. If a strong concept is discarded during pre-selection, the test cannot correct that.
Horizon checks the final variants in a Painted Door Test with real people in their familiar online environment, without a panel and without incentives. Nothing is sold.
Girotra, Meincke, Terwiesch & Ulrich 2023: Among the top 10% of 400 product ideas, 35 of 40 came from GPT-4. Quality was measured with stated purchase intent on a five-point scale. Ideas are Dimes a Dozen: Large Language Models for Idea Generation in Innovation, Working Paper (Wharton, SSRN 4526071). Source
Meincke, Mollick & Terwiesch 2024: Idea pools from GPT-4 are less diverse than those from groups of human participants; targeted prompting narrows the gap considerably. Prompting Diverse Ideas: Increasing AI Idea Variance, Working Paper (arXiv 2402.01727). Source
Morwitz, Steckel & Gupta 2007: Meta-analysis: stated purchase intent reflects later sales less well for new products than for existing ones. International Journal of Forecasting 23(3). Source
For pre-selection, that makes sense. Studies on agreement measure synthetic respondents against surveys, not against behaviour, which is why the final variants go into a behavioural test before the investment.
Two to six that differ clearly in one dimension, for example in the promise, the price or the claim.
That is for your legal department to clarify. Nothing is sold in the test, and anyone who chooses an offer learns afterwards that it is a test.
Bring your decision question, and we will outline a possible test design.
You will speak with Daniel Putsche
Founder & CEO, 30 minutes
Read more