


Statistical significance means that an observed difference between variants is large enough that, assuming there were no real difference, it would only rarely arise by chance; a threshold of p below 0.05 is commonly used for this.
In variant tests, significance determines whether a difference is considered reliable. It is an important tool, but it is often overrated.
Two variants almost never show exactly the same values. The question is whether the gap points to a real difference in behaviour or lies within the range of chance. A significance test answers exactly this question and protects you from investing on the strength of a random fluctuation.
What it does not answer: how large and how important the difference is. The American Statistical Association states explicitly that a p-value measures neither the size of an effect nor the importance of a result, and that business decisions should not depend on a threshold alone.
Two offer pages for a cordless kitchen machine, 2,000 visitors each. Variant A: 80 clicks on "Add to basket" (4.0%). Variant B: 100 clicks (5.0%). The gap looks clear, a quarter more purchase intent.
The two-sample test for proportions gives a z-value of about 1.53 and a p-value of about 0.13. At the usual threshold of 0.05, the difference is not significant: with this sample, a gap of this size also arises without a real difference in about 13 out of 100 cases.
With 5,000 visitors each and the same proportions (200 versus 250 clicks), the p-value is about 0.016. The same gap is now significant, because the larger sample narrows the margin for chance.
Significance on the primary metric is one of several conditions for a winner. In addition, the probability to be best must reach a threshold set in advance, and the lead must be consistent with the upstream steps, meaning it must not be an effect of the ads. All values are recalculated from the raw counts. Results are classified as confirmed, directional or inconclusive.
The confidence interval shows the range in which the true value plausibly lies, and thus also the size of the difference. The probability to be best is a Bayesian quantity and answers a different question: how certain is it that a variant is ahead? Practical relevance is a question of substance: is a difference of this size worthwhile for the business?
Significance depends heavily on the sample. With very large samples almost any difference becomes significant, with small ones hardly any. Checking repeatedly during the test and stopping at the first significant value considerably increases the rate of false positives; that is why a stopping rule is part of it.
Even a cleanly significant result in the behavioural test says nothing about effects the test does not reflect, such as repeat purchase or competitor reactions. And not significant does not mean "no difference", but rather: the data is not sufficient to establish a difference with confidence.
Wasserstein & Lazar 2016: The American Statistical Association states that a p-value measures neither the size of an effect nor its importance, and that business decisions should not depend solely on whether a p-value falls below a threshold. The ASA Statement on p-Values: Context, Process, and Purpose, The American Statistician 70(2). Source
Simmons, Nelson & Simonsohn 2011: Simulation: checking from 10 observations per group onwards after every further 10 and stopping at p < 0.05 yields 14.3% instead of 5% false positives; checking after every observation yields 22.1%. False-Positive Psychology, Psychological Science 22(11). Source
No. Significance says that a difference is probably not due to chance. How large it is is shown by the confidence interval.
The data is not sufficient to establish a difference with confidence. A small difference or a tie can still be a valid insight.
Because it answers only one question. Horizon also checks the probability to be best and whether the lead is consistent across the flow.
Bring your decision question, and we will outline a possible test design.
You will speak with Daniel Putsche
Founder & CEO, 30 minutes
Read more