


Evidence in the innovation process is the body of proof on which decisions about new offers, prices, tariffs and concepts rest, ideally measured before budget is committed rather than after the launch.
Most companies have an orderly process with stages and approvals. What often remains open is which kind of evidence is sufficient at which point, and whether it includes observed behaviour of real people.
Every innovation or pricing process has moments when money, time and attention are committed: approving development, choosing the tariff model, deciding on a price or a brand promise. The evidence on the table at these points determines the quality of the decision more than the number of workshops that came before.
Evidence can be roughly divided into three kinds. First, judgements: experience, internal alignment, expert opinion, analogies to earlier launches. Second, statements: surveys, concept tests, conjoint, interviews and focus groups, in other words what people say about their future behaviour. Third, behaviour: what people do when they face a concrete offer with a price and alternatives.
All three have their place. Judgements are fast and pool knowledge. Surveys explain motives, drivers and objections and help to rank many options. Behaviour answers the question of whether an offer is chosen. Evidence in the innovation process therefore does not mean replacing one method with another, but matching each question with the right kind of evidence and planning a behavioural proof point before the expensive decisions.
Why statements alone are rarely enough for the expensive decisions is well researched. A meta-analysis of 47 experiments shows that when intention is changed with medium to large strength (d = 0.66), behaviour changes only by a small to medium amount (d = 0.36) (Webb & Sheeran 2006). According to a review by the same authors, intentions are translated into action in about half of all cases (Sheeran & Webb 2016).
The same applies to prices. Across 77 studies, hypothetically stated willingness to pay is on average 21 per cent higher than real willingness to pay (Schmidt & Bijmolt 2020). And for new products in particular, stated purchase intent reflects later sales less well than for existing ones (Morwitz, Steckel & Gupta 2007). There is also a measurement effect: among customers who were surveyed, the link between intention and purchase is 58 per cent stronger than among those who were not, so the act of surveying itself changes later behaviour (Chandon, Morwitz & Reinartz 2005).
On the other side is the flop rate. Figures of 80 or 90 per cent failed products are often quoted. Castellion & Markham (2013) show that this figure is a myth: empirical studies since 1977 find a flop rate of 40 per cent or less. That is less dramatic, but still a lot when every launch involves development, tooling, packaging, distribution and marketing.
The gap between statement and behaviour has no fixed size and no fixed direction. It depends on category, price level, competition and situation. A blanket correction factor applied to survey results is therefore of little help. If you want to know how large the gap is for your own question, you have to observe behaviour.
The more expensive and the harder to reverse a decision is, the more weight the behavioural evidence deserves. This does not only concern new products: a price increase for existing customers, a new tariff, a new brand promise or the question of which of three equipment variants goes into series production are also decisions that can be checked against behaviour before implementation.
Timing determines the value. Measuring behaviour after the launch is valuable for optimisation, but it comes once the budget is committed. Behavioural evidence before the budget moves the moment of truth forward, into a phase in which variants are still cheap to change.
Finally, evidence changes the discussion in the committee. Where otherwise the most convincing presentation or the most senior opinion tips the balance, there is now a shared data point on the table. This makes decisions faster and traceable in hindsight, even if a launch later falls short of expectations.
The most effective lever is organisational, not methodological: a behavioural test is anchored as a gate before investment approval, a Behavioural Gate. This ensures that an expensive decision does not rest on judgements and statements alone.
For such a gate to hold, three things are defined before the test. The primary metric, meaning which step in the offer flow counts as measured purchase intent or sign-up intent. The sample size, so that a difference between variants can be detected with statistical reliability. And a stopping rule, so that nobody ends the test as soon as they like the result. Recording these points in writing beforehand makes the result traceable even for sceptics in the committee.
In practice, the role of behavioural evidence usually grows step by step: a first test on an upcoming decision, then a repeat on a second one, then an internal threshold above which a gate counts as passed. Many organisations already experiment without behavioural evidence being a fixed part of approval. The step from experiment to rule is the real shift.
Horizon runs the behavioural test as a Painted Door Test: realistic offer pages, served via Google and Meta ads to real people in their usual online environment, without a panel and without incentives, with up to six variants. Nothing is sold. Anyone who chooses an offer is then told transparently that it is a test. From the question to the data analysis takes around four weeks. The decision is made by those responsible, Horizon delivers the data analysis.
Ideation: breadth is what counts here. Workshops, trend analyses, customer interviews and AI-assisted idea generation provide raw material. Evidence in the narrower sense is not yet needed, but a clear question about which problem is to be solved is.
Concept and pre-selection: many ideas become few. Concept screening via survey, MaxDiff or expert judgement ranks quickly and cheaply. The result is a shortlist, not a basis for investment.
Business case and approval: this is where the largest budget is committed, and this is where the behavioural evidence belongs. The final variants, for example two promises, three prices or four concepts, are designed as realistic offers and checked against the behaviour of real people. The result feeds into the gate, together with cost, technology and strategy.
Development and launch: now it is about the how. MVP, pilot series, field test and A/B tests in the live system optimise implementation, usage and messaging. Sales and usage data show whether the assumptions hold and provide the basis for the next cycle.
Judgement and desk research: fast and cheap, suitable for a first assessment of a market. They do not replace measurement among customers.
Survey, concept test, conjoint and MaxDiff: strong for ranking many ideas or features, understanding motives and describing target groups. Choice-based methods reduce hypothetical bias, but remain statements without consequence.
AI-assisted methods: generative AI has greatly accelerated the early phases. Concepts, claims and variants are created in minutes, and synthetic respondents provide first assessments of segments and objections. AI widens the funnel. At the end of the funnel is the behavioural test for the few variants that actually receive investment.
A language model reproduces what people have said and written. Even studies that credit synthetic respondents with high agreement measure them against surveys, not against behaviour. For offers that do not yet exist, the data basis is also often missing. That is why observed behaviour remains the benchmark for the expensive decisions, while AI makes the work before them faster and broader.
Painted Door Test and Pretotyping: measure purchase or sign-up intent on a realistic offer before the product exists or the price is in the market. Suitable for the few final variants before investment.
A/B test in the live system: optimises what already exists when all variants can be delivered. It answers the question of how, the Painted Door Test the question of whether.
Pre-order and MVP: bring real payments or a working minimal product into play. They provide strong evidence, but require a delivery commitment or development effort.
A home appliance manufacturer has four concepts for a cordless kitchen machine on its shortlist. A concept test with a survey has cut the list from twelve to four, and all four achieve high approval scores. Before approving the tooling costs, management asks for behavioural evidence.
In the Painted Door Test, each person sees exactly one concept as a realistic offer with a price. Around 2,500 visitors per variant reach the offer page. The primary metric is a click on “Pre-order now” followed by entering an email address. Concept B reaches 3.1 per cent, concept A 2.0 per cent, C and D are lower. The difference between B and A is statistically reliable. In the concept test, A and B were almost level, with A even slightly ahead.
The committee now has a data point that shows which concept is chosen when price and alternatives are visible. Whether B goes into series production is still its own decision, together with cost, technology and strategy.
Starting point: the flop rate of new products shows how much is at stake, and why the often quoted 80 per cent is not supported by evidence. Evidence-based innovation describes the principle of basing decisions on the best available evidence.
Process and governance: the Stage-Gate process divides development into stages and approvals. The Behavioural Gate anchors a behavioural test before them. The go/no-go decision is made at the gate. Decision Intelligence puts the decision before the data set.
Pre-selection and content: concept screening pre-sorts many ideas. Value proposition testing checks which promise is chosen. Test and Learn describes the practice of preparing big bets through small tests.
Methods: Pretotyping, demand validation, pre-launch test, Smoke Test, concept test, pre-order test and Minimum Viable Product differ in whether a product already exists, whether payment is made and what is measured.
Measurement and statistics: primary metric, statistical significance, sample size, probability to be best, stopping rule and benchmark ensure that a test result is reliable and correctly interpreted.
Behavioural evidence before the launch is no substitute for the market. A Painted Door Test measures the response to an offer in a specific channel at a specific point in time. It does not reflect repeat purchase, satisfaction after use, competitor reactions or effects in physical retail.
The quality of the result depends on the test design. If variants differ in more than the variable being tested, it is unclear what explains the difference. Samples that are too small or a stop that comes too early produce false winners. That is why the primary metric, sample and stopping rule are set before the start.
Not every decision needs a behavioural test. For cheap, easily reversible decisions or products with a very low price, the effort is often out of proportion to the risk. And the why behind a result is usually better explained by a survey or an interview than by the test itself. Together, both give the fuller picture.
Webb & Sheeran 2006: 47 experiments: a medium to large change in intention (d = 0.66) produces only a small to medium change in behaviour (d = 0.36). Does changing behavioral intentions engender behavior change? A meta-analysis of the experimental evidence, Psychological Bulletin. Source
Schmidt & Bijmolt 2020: 77 studies, 115 effect sizes: hypothetically stated willingness to pay is on average 21% higher than real willingness to pay. Accurately measuring willingness to pay for consumer goods: a meta-analysis of the hypothetical bias, Journal of the Academy of Marketing Science. Source
Morwitz, Steckel & Gupta 2007: Meta-analysis: stated purchase intent reflects later sales less well for new products than for existing ones. International Journal of Forecasting 23(3). Source
Chandon, Morwitz & Reinartz 2005: Among surveyed customers, the link between intention and purchase is 58% stronger than among those not surveyed: the act of surveying itself inflates validity. Journal of Marketing 69(2). Source
Castellion & Markham 2013: The widespread assumption that 80% or more of new products fail is a myth; empirical studies since 1977 find a flop rate of 40% or less. Perspective: New Product Failure Rates: Influence of Argumentum ad Populum and Self-Interest, Journal of Product Innovation Management. Source
Judgements, statements from surveys and observed behaviour. All three have their place; for expensive, hard-to-reverse decisions, at least one piece of evidence should come from observed behaviour.
No. Surveys explain motives and rank many options, the behavioural test checks the few final variants. The two complement each other.
Before the approval that commits the largest budget: usually before development, tooling costs, a tariff change or a price change. At that point variants are still cheap to change.
No. Prices, tariffs, claims, brand promises and equipment variants of existing products can equally be checked against behaviour in advance.
Around four weeks from the question to the data analysis, with up to six variants. Nothing is sold, and anyone who chooses an offer is then informed transparently.
Bring your decision question, and we will outline a possible test design.
You will speak with Daniel Putsche
Founder & CEO, 30 minutes
Read more