


Go/no-go sounds binary, but in practice it is often a weighing of cost, technology, strategy and demand. The weakest piece of evidence in it is frequently demand.
For cost and technology, there are usually robust figures. For demand, surveys and analogies are often all that is available, and that is exactly where the greatest risk lies: surveyed purchase intent reflects sales less well for new products than for existing ones (Morwitz, Steckel & Gupta 2007), and hypothetically stated willingness to pay is on average above real willingness to pay (Schmidt & Bijmolt 2020). A go based on such statements can look well-founded and still stand on shaky ground.
A go holds when at least one piece of demand evidence comes from observed behaviour and it is clear in advance what a result means. This includes a primary metric, a baseline for comparison, such as an existing variant or an alternative, and a stopping rule. Without these definitions, every result gets interpreted to fit after the fact.
A no-go is also a result. Stopping a variant early saves budget for the better one. And "no clear winner" is information too: then other criteria decide, or the variants are sharpened and tested again.
It has proven useful to agree the criteria in writing before the test starts: which variants are in contention, which step in the offer flow counts as measured purchase or sign-up intent, which comparison variant is measured against, and how large a difference must be for it to change the decision. Also recording what happens with a close result avoids long discussions in the committee. This applies to new products as much as to price changes, tariff changes or a new brand promise.
A bank is planning a current account with a travel insurance package. The committee has agreed: go if the package variant achieves higher measured sign-up intent in the Painted Door Test than the variant without the package at a lower account fee, and the gap is statistically robust. Around 3,500 visitors per variant reach the offer page. The package variant achieves 1.9 percent, the variant with the lower account fee 2.0 percent. No robust difference: the committee decides against the additional effort for the bonus system.
The stage gate is where the go/no-go is made. The behavioural gate is a criterion before it that provides the behavioural evidence. Statistical significance and the probability of being the best variant are tools for putting a test result into context for the decision. They do not replace the decision.
A behavioural test provides a data point on demand, not on feasibility or profitability. A go from the test is no guarantee of success in the market: repeat purchase, competition and sales come later. For inexpensive, easily reversible decisions, a dedicated test is often not worthwhile.
For go/no-go, Horizon provides the data analysis from a Painted Door Test with real people. The decision is made by those responsible.
Morwitz, Steckel & Gupta 2007: Meta-analysis: surveyed purchase intent reflects later sales less well for new products than for existing ones. International Journal of Forecasting 23(3). Source
Schmidt & Bijmolt 2020: 77 studies, 115 effect sizes: hypothetically stated willingness to pay is on average 21 % above real willingness to pay. Accurately measuring willingness to pay for consumer goods: a meta-analysis of the hypothetical bias, Journal of the Academy of Marketing Science. Source
At least one piece of demand evidence from observed behaviour, assessed against a metric and baseline defined in advance, plus cost, technology and strategy.
That is a valid result. Then other criteria decide, or the variants are sharpened and tested again.
The responsible committee. The test provides a data point, not a decision.
Bring your decision question, and we will outline a possible test design.
You will speak with Daniel Putsche
Founder & CEO, 30 minutes
Read more