


A test produces many numbers: ad clicks, time on page, scroll depth, sign-ups. For a result to be robust, it must be clear in advance which of them counts.
If you only decide after the test which number counts, you will almost always find one by which a variant wins. With every additional measure considered, the probability rises that a difference occurs purely by chance. A primary metric set in advance protects against this and makes results comparable across tests.
For the decision, this means: management, product and insights talk about the same number, and the discussion is about what the result means, not about which metric to choose.
The rule is: the primary metric is the last binding click before the reveal. The closer a click is to a purchase decision, the fewer curiosity clicks it contains and the stronger the signal.
Single-stage setup: the offer page contains the buy button directly, and its click is the primary metric. Two-stage setup: the offer page is followed by a product detail page; the primary metric there is the click on "Add to cart". The first click ("Learn more") is then only an intermediate signal. Whether a two-stage test is run depends on whether enough visitors can be reached within the test period.
The denominator is always the number of unique visitors to the first offer page. Ad impressions and reach are planning figures, because they depend on the delivery of the ad platform, not on the variant. The email sign-up after the reveal is a secondary signal: it measures interest in news, not purchase intent.
A test compares three value propositions for a cordless kitchen machine, in two stages. Variant A achieves the most clicks on "Learn more" (11% of visitors), variant B the most clicks on "Add to cart" (2.1% compared with 1.6% for A).
Since the primary metric had been set in advance as the add-to-cart click, B is ahead. A sparks more curiosity, B triggers more purchase intent. Both findings are valuable, but only one decides.
Secondary metrics, such as ad clicks, the intermediate signal or the email sign-up, help with diagnosis but do not decide. The click-through rate of an ad measures advertising impact, not the offer. The primary metric is also not the same as a business goal: revenue or contribution margin only arise in the market and are calculated on the basis of the test result.
A single metric condenses behaviour heavily. It says which variant triggers more purchase intent, but not why. For that, it is worth looking at the intermediate signals and, where necessary, a supplementary qualitative survey.
In addition: the further down the journey the primary metric sits, the fewer signals occur and the more visitors the test needs. The choice is therefore a trade-off between signal strength and reach, and it is documented in the test design.
Simmons, Nelson & Simonsohn 2011: Simulation: anyone who checks from 10 observations per group onwards after every further 10 and stops at p < 0.05 obtains 14.3% instead of 5% false positive results; when checking after every observation, 22.1%. False-Positive Psychology, Psychological Science 22(11). Source
All of them are analysed, but the decision is based on one. Anyone who treats several metrics as equal increases the risk of mistaking a chance difference for a real one.
No. It happens after the person knows that it is a test, and therefore measures interest in news.
To the unique visitors of the first offer page, not to ad impressions.
Bring your decision question, and we will outline a possible test design.
You will speak with Daniel Putsche
Founder & CEO, 30 minutes
Read more