Horizon
Home
/
Glossary
/
Statistical Significance

Statistical Significance

Statistical significance means that an observed difference between variants is large enough that, assuming there were no real difference, it would only rarely arise by chance; a threshold of p below 0.05 is commonly used for this.

Updated
September 28, 2026
· Horizon

In variant tests, significance determines whether a difference is considered reliable. It is an important tool, but it is often overrated.

Why this matters for your decision

Two variants almost never show exactly the same values. The question is whether the gap points to a real difference in behaviour or lies within the range of chance. A significance test answers exactly this question and protects you from investing on the strength of a random fluctuation.

What it does not answer: how large and how important the difference is. The American Statistical Association states explicitly that a p-value measures neither the size of an effect nor the importance of a result, and that business decisions should not depend on a threshold alone.

Example: calculation with two variants

Two offer pages for a cordless kitchen machine, 2,000 visitors each. Variant A: 80 clicks on "Add to basket" (4.0%). Variant B: 100 clicks (5.0%). The gap looks clear, a quarter more purchase intent.

The two-sample test for proportions gives a z-value of about 1.53 and a p-value of about 0.13. At the usual threshold of 0.05, the difference is not significant: with this sample, a gap of this size also arises without a real difference in about 13 out of 100 cases.

With 5,000 visitors each and the same proportions (200 versus 250 clicks), the p-value is about 0.016. The same gap is now significant, because the larger sample narrows the margin for chance.

How Horizon uses significance

Significance on the primary metric is one of several conditions for a winner. In addition, the probability to be best must reach a threshold set in advance, and the lead must be consistent with the upstream steps, meaning it must not be an effect of the ads. All values are recalculated from the raw counts. Results are classified as confirmed, directional or inconclusive.

How it differs

The confidence interval shows the range in which the true value plausibly lies, and thus also the size of the difference. The probability to be best is a Bayesian quantity and answers a different question: how certain is it that a variant is ahead? Practical relevance is a question of substance: is a difference of this size worthwhile for the business?

Limitations

Significance depends heavily on the sample. With very large samples almost any difference becomes significant, with small ones hardly any. Checking repeatedly during the test and stopping at the first significant value considerably increases the rate of false positives; that is why a stopping rule is part of it.

Even a cleanly significant result in the behavioural test says nothing about effects the test does not reflect, such as repeat purchase or competitor reactions. And not significant does not mean "no difference", but rather: the data is not sufficient to establish a difference with confidence.

Evidence

Wasserstein & Lazar 2016: The American Statistical Association states that a p-value measures neither the size of an effect nor its importance, and that business decisions should not depend solely on whether a p-value falls below a threshold. The ASA Statement on p-Values: Context, Process, and Purpose, The American Statistician 70(2). Source

Simmons, Nelson & Simonsohn 2011: Simulation: checking from 10 observations per group onwards after every further 10 and stopping at p < 0.05 yields 14.3% instead of 5% false positives; checking after every observation yields 22.1%. False-Positive Psychology, Psychological Science 22(11). Source

Frequently asked questions

Does significant mean that the difference is large?

No. Significance says that a difference is probably not due to chance. How large it is is shown by the confidence interval.

What does a non-significant result mean?

The data is not sufficient to establish a difference with confidence. A small difference or a tie can still be a valid insight.

Why is significance alone not enough for a winner?

Because it answers only one question. Horizon also checks the probability to be best and whether the lead is consistent across the flow.

When is the difference reliable for your question?

Bring your decision question, and we will outline a possible test design.

Daniel Putsche

You will speak with Daniel Putsche
Founder & CEO, 30 minutes

Read more

Thank you, we have your request.

We will get back to you within one working day with suggested times.

Close

Demo · Video

A complete test, from design to data analysis

Play the demo

Chapters: test design · offer pages · live data · result

Book a call

First call

Book a call

30 minutes, your decision question, a possible test design. No preparation needed on your side.

By submitting, you agree to the processing of your details to arrange a call. Details in our privacy policy.

Request a call
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.
SAMPLE REPORTExample

Sample report: insurance

Data analysis · Test design · Metrics · Methodology

Sample report

Sample report: insurance

A complete results report with example values: research question, test design, purchase intent per variant and the data analysis.

12 pages, as a PDF to share internally
Free, download right after a short form
Note