


Few decisions affect revenue and margin as directly as price, and few are as hard to reverse. This page explains how pricing decisions can be prepared, which methods exist and where behavioural data makes the difference.
Every pricing decision involves the same trade-off: a higher price earns more per unit sold but costs demand. A lower price wins demand but gives away margin. Where the best point lies depends on how strongly the target group reacts to price, and that reaction is exactly what is unknown before market entry.
On top of that, a price never stands alone. It works together with the scope of the offer, with how it is presented on the offer page, with other tiers in the range and with the reference prices people have in mind. That is why the question is rarely just which number goes on the price tag, but which offer is chosen at which price.
Pricing decisions are not limited to innovations. They range from pricing a new product to raising an existing tariff to the question of whether a third tier, a subscription or a different payment interval improves demand. For insurers, banks, energy suppliers, telecommunications providers, home appliance manufacturers, consumer goods and retail, these questions are part of everyday business.
New price: at what price does a new product, a new tariff or a new service go to market? There is no sales history of your own, and reference prices from competitors often fit only partly because the offer differs.
Price increase: how much demand does a higher price for an existing offer cost? The current price is the reference point here, and the increase is hard to reverse without losing credibility.
Price structure: how many tiers does an offer need, and how do they influence each other? Whether a premium tier strengthens the middle tier or a basic tier pulls customers down only becomes clear in the set.
Pricing model and payment interval: one-off price or subscription, monthly or annual, base price plus usage or flat fee. Such questions change how a price is read, even when the total stays the same.
Price presentation: does a price sit above a threshold, does an ending in 9 have an effect, how is a surcharge for an add-on presented? These questions also have measurable consequences for demand.
Each of these questions requires its own test design. If you want to answer several of them at once, for example which product and at which price, it is better to split them across several consecutive tests so that each result can be attributed to one cause.
Most pricing methods ask people what they would pay. Research has shown for decades that these answers are on average higher than what people actually pay. A meta-analysis of 77 studies with 115 effect sizes finds an average overestimation of willingness to pay of 21 percent for consumer goods (Schmidt & Bijmolt 2020). The bias is larger for higher-value products, larger for indirect methods such as conjoint than for direct questions, and larger in designs where the same person sees several conditions than in designs with only one condition per person.
Older meta-analyses from environmental and resource economics point in similar directions, with partly larger values: across 28 studies, the median ratio of hypothetical to real willingness to pay is 1.35, with a strongly right-skewed distribution (Murphy et al. 2005). Across 29 studies, hypothetical statements overstate real values on average by a factor of about 3 (List & Gallet 2001).
The studies also show that commitment helps. At the point of sale, the incentive-compatible BDM procedure produced lower willingness to pay than non-binding procedures (Wertenbroch & Skiera 2002). A comparison of several methods with real purchases found that binding procedures pass the tests, and that hypothetically biased methods can also lead to the right pricing decision in individual cases (Miller et al. 2011). The honest reading is therefore not that surveys are wrong, but that their deviation in any individual case is unknown.
How strongly demand reacts to price is itself substantial. A meta-analysis of 1,851 price elasticities from 81 studies finds an average of minus 2.62 (Bijmolt, van Heerde & Pieters 2005). On average, sales volume falls by a good two and a half percent when the price rises by one percent. Even small pricing errors therefore have noticeable consequences.
If you set a price on the basis of stated willingness to pay, you may be planning with demand that does not exist at that price. The result is business cases that are too optimistic, launch prices that have to be cut afterwards, or price increases that cost more customers than planned.
Conversely, cautious planning can leave money on the table. If the target group is less price-sensitive than assumed, a price that is too low often stays in place for years, because a later increase is harder to push through than a higher starting price.
Both errors are far cheaper to avoid before market entry than afterwards. The more expensive and the harder to reverse a pricing decision is, the more it pays to base it on observed behaviour rather than on assessments alone.
Van Westendorp uses four price questions to ask which prices are considered too cheap, cheap, expensive or too expensive, and derives an acceptable price range from them. Quick and inexpensive, but without demand per price.
Gabor-Granger asks, one price point after another, whether people would buy, and delivers a stated demand and revenue curve. Because each person sees several prices, the first price acts as an anchor and strategic answering comes into play.
Conjoint analysis has respondents choose between product profiles that differ in several attributes, including price. It is strong at weighting many attributes against each other. MaxDiff shows which attributes matter most, but not what they may cost.
The BDM auction and incentive-compatible conjoint make the answer binding and thereby reduce hypothetical bias. They require a product that can actually be handed out.
Price A/B tests in your own shop measure real purchases, but require that the product is already being sold and that different prices can be served live.
The behavioural test before market entry, for example as a Painted Door Test, shows real people a realistic offer page with exactly one price in their familiar online environment and measures purchase intent per price. It does not need a finished product, and nothing is sold.
The methods do not exclude each other, they answer different questions. A sensible sequence uses surveys where they are strong and behaviour where the decision is made: surveys and conjoint narrow down early which attributes count and within which range a price should be discussed. The two to six prices that are seriously in contention at the end are compared on behaviour.
A robust price comparison based on behaviour follows a few rules. Each person sees exactly one price, otherwise the first price seen becomes the anchor for all others. Everything except the price stays identical: product, service, page, images, copy, target group and channel. The price is tied to a clearly named service, and the price information is easy to read, so that the reaction is to the price and not to a lack of clarity.
At Horizon, the price is already shown in the ad copy, so that people see the same price from first contact through to the offer page. Brand search terms do not belong in a price test, because people who search specifically for the brand bring a higher willingness to pay. It is clarified in advance below which price an offer is no longer viable for the company, so that no price is tested that could not be implemented anyway. Pricing questions also use a stricter decision threshold than other test types, because pricing decisions are harder to reverse.
From the question to the data analysis, such a test takes around four weeks. The result is a data analysis: measured purchase intent per price, the differences between the variants and how certain these differences are. Which price is chosen is decided by the company, often together with its own margin and sales data.
A home appliance manufacturer wants to launch a cordless kitchen machine. A Van Westendorp survey yields an acceptable range of €300 to €480. Product management favours €449, sales €349.
The behavioural test runs four variants: €349, €399, €449 and €479, and each person sees only one price. Measured purchase intent is almost level at €349 and €399, falls by around a third at €449 and by slightly more than half at €479.
Calculated with the company's own unit costs, €399 delivers the highest contribution margin per visitor. In the survey, €449 was still in the middle of the acceptable range, while behaviour already shows a clear decline there. The survey narrowed down the range correctly, and only behaviour made the final choice between the prices possible.
A behavioural test also has limits, and they are part of an honest assessment. It measures purchase intent, not a purchase with payment. It compares only the price points tested, and values between them are interpolated. It shows the reaction of people who see the offer for the first time, not cancellations among existing customers or the reaction of competitors.
Not every pricing question fits an offer page. If a product is mainly chosen on the shelf next to competitors, if the decision is made on a comparison portal, or if the item is so inexpensive that nobody buys it individually online, an offer page does not reflect the decision. There, retail data, shelf tests or surveys are often the better choice. And if a product is already on the market with robust sales data on exactly this question, that data is usually the first source.
Surveys and conjoint analyses remain valuable. They explain why a price is perceived as too high, which attributes justify it and how different segments react. Behavioural data does not replace these insights, it complements them with the question of what people actually do at a specific price.
Willingness to pay: the highest price a person is willing to pay, and why stated values are higher than real ones.
Price testing and testing a price increase: how prices are compared on behaviour and how much demand an increase costs.
Price elasticity and price threshold: how strongly demand reacts to price and where it reacts abruptly.
Price tiers: how good-better-best models work and why tiers are tested as a set.
Van Westendorp, Gabor-Granger, conjoint analysis, MaxDiff and BDM auction: the most important survey and experimental methods with their strengths and limits.
Anchoring effect: why a price seen first influences all subsequent assessments (Ariely, Loewenstein & Prelec 2003).
Schmidt & Bijmolt 2020: 77 studies, 115 effect sizes: hypothetical willingness to pay is on average 21% above real willingness to pay. Indirect methods overestimate more than direct ones, within-subject designs more than between-subject designs, higher-value products more than inexpensive ones. Accurately measuring willingness to pay for consumer goods: a meta-analysis of the hypothetical bias, Journal of the Academy of Marketing Science 48(3). Source
Murphy et al. 2005: 28 studies: the median ratio of hypothetical to real willingness to pay is 1.35, and the distribution is strongly right-skewed. A Meta-analysis of Hypothetical Bias in Stated Preference Valuation, Environmental and Resource Economics 30(3). Source
Wertenbroch & Skiera 2002: In three studies, the incentive-compatible BDM procedure yields lower willingness to pay than non-incentive-compatible procedures (open question, double-bounded contingent valuation); the difference is due to the binding nature, not to cognitive effort. Measuring Consumers' Willingness to Pay at the Point of Purchase, Journal of Marketing Research 39(2). Source
Miller et al. 2011: Comparison of open question, choice-based conjoint, BDM and incentive-compatible conjoint with real purchases: BDM and incentive-compatible conjoint pass the tests; hypothetically biased methods can also lead to the right pricing decision in individual cases. How Should Consumers' Willingness to Pay Be Measured? An Empirical Comparison of State-of-the-Art Approaches, Journal of Marketing Research 48(1). Source
Bijmolt, van Heerde & Pieters 2005: Meta-analysis of 1,851 price elasticities from 81 studies: the average price elasticity is minus 2.62. New Empirical Generalizations on the Determinants of Price Elasticity, Journal of Marketing Research 42(2). Source
They provide good indications of the range and the reasons, but on average they overestimate willingness to pay. A meta-analysis finds a mean overestimation of 21 percent for consumer goods, with wide variation between studies.
No. Surveys and conjoint narrow down and explain. The behavioural test compares the few prices that are in contention at the end on behaviour.
At Horizon, up to six variants. Each person sees exactly one price, everything else stays identical.
No. There is no contract and no payment. Anyone who chooses an offer is then told transparently that it is a test.
For products chosen on the shelf next to competitors, for products sold through comparison portals and for very inexpensive items that nobody buys individually online. Other methods are more informative there.
From the question to the data analysis, around four weeks.
Bring your pricing question, and we will outline a possible test design.
You will speak with Daniel Putsche
Founder & CEO, 30 minutes
Read more