


The significance test asks whether a difference could be due to chance. P(best) asks directly how certain it is that a variant is ahead. For decisions between several variants, this is often the more intuitive question.
With three to six variants, a pairwise significance test is not enough to answer the question decision makers actually ask: which variant is best, and how certain is that? P(best) answers exactly this question in a number that can be read without a degree in statistics: 92% means that, based on the data so far, the variant is ahead with a probability of 92%.
Because the number is so intuitive, it needs a threshold set in advance. Otherwise it becomes an argument for whichever variant you preferred anyway.
For each variant, a probability distribution of the true share is formed from visitors and clicks; a beta distribution is common. Random values are drawn from these distributions a great many times, and it is counted how often each variant has the highest value. The share of these draws is its P(best). The values of all variants add up to 100%.
Two offer pages with 2,000 visitors each: variant A 80 clicks (4.0%), variant B 100 clicks (5.0%). P(best) for B is about 94%. The classical test, by contrast, gives a p-value of about 0.13, so not significant.
The two numbers do not contradict each other; they answer different questions. B is probably ahead, but the difference is not yet established firmly enough to call it confirmed. With 5,000 visitors each and the same shares, P(best) rises to over 99%, and the difference becomes significant.
The threshold is set before the start for each test type: 90% for demand, product, feature, value proposition, brand and target group tests with the same price, 95% for price tests and all tests in which prices differ between variants. A winner only counts as confirmed if P(best) reaches the threshold and the difference on the primary metric is significant. If only one of the two conditions is met, the result is directional.
The p-value describes how surprising the data would be if there were no difference; P(best) describes how likely a variant is to be the best. The confidence interval shows the size of a difference, P(best) only the ranking. A variant can be the best with high probability and still be only narrowly ahead.
P(best) says nothing about whether the lead is large enough to pay off commercially, and nothing about whether the best variant is good in absolute terms. That question needs a comparison against comparable tests. With very few signals, P(best) reacts sensitively to individual clicks and should not decide on its own. And as with the significance test: anyone who keeps checking and stops the first time the threshold is crossed overestimates the certainty.
No. P(best) is the probability that a variant is ahead. Significance describes how unlikely the data would be without a real difference.
Pricing decisions are hard to reverse and more costly if they are wrong. That is why Horizon requires 95% instead of 90% there.
Then the result is inconclusive or directional. That too is a valid finding, for example that two variants are equally strong.
Bring your decision question, and we will outline a possible test design.
You will speak with Daniel Putsche
Founder & CEO, 30 minutes
Read more