


A benchmark is a reference value from comparable earlier measurements against which a new result is placed in context; in a behavioural test, it shows whether a variant is not only better than the others but also strong in absolute terms.
A variant test tells you which variant is ahead. On its own, it does not answer whether the best variant is also good. That requires an external yardstick.
Suppose one of three concepts wins clearly. That can mean it is a strong concept, or that it is the least weak one. For an investment decision, this difference is central. Especially for questions about market demand at a given price, an absolute assessment therefore belongs alongside the variant comparison.
A benchmark, however, is only as good as its comparison group. A value from a different category, a different country or a different channel can make a variant look stronger or weaker than it is.
The comparison is made per variant, not per test, against a comparison group that matches on fixed attributes: test type, product category, price positioning, country or region, brand and channel. Numerator and denominator must be the same as for your own result, that is, the same primary metric relative to the visitors of the offer page.
If the exact comparison group is not sufficient, softer attributes are relaxed in a fixed order, and each relaxation is stated next to the value. The basis is always given, as the number of comparable variants. If the basis is thin, the comparison is presented as a tendency. Horizon shares benchmark values exclusively with clients as part of a test, not publicly.
An energy supplier tests three variants of a new heat pump rental model. Variant C is ahead by a clear margin. Compared with variants of other rental models from the same industry, the same country and the same channel, C is above the median but below the top quarter.
The assessment: C is the right choice among the variants tested and solid in the industry comparison, but not an outlier at the top. The team decides to test the pricing model before launch as well.
A benchmark is not a target and not an industry average from publications, but a comparison with measurements using the same method. Values from A/B tests in a live shop, from ad statistics or from surveys are not comparable, because they have different numerators and denominators. The variant comparison within a test remains the core result; the benchmark places it in context.
For new markets, new categories or rare price points, there is often no sufficient comparison group. In that case, no market-specific value is promised; instead, the result is read on its own, supplemented by the client's own data. Nor does a benchmark translate into sales: it shows how strong a variant is compared with similar tests, not how much it will sell in the market.
No. Benchmarks are built per variant and shared only with clients as part of a test, because public averages without a suitable comparison group are misleading.
Because a test can contain up to six variants with different prices or promises. The comparison has to fit the individual variant.
Then the result is read on its own, and the assessment is given honestly as a tendency or not at all.
Bring your decision question, and we will outline a possible test design.
You will speak with Daniel Putsche
Founder & CEO, 30 minutes
Read more