


Running tests invite you to check. If you stop at the first favourable interim result, you overestimate the certainty of your result. A stopping rule separates planning from analysis.
Interim results fluctuate, especially in the first few days. Checking daily and stopping as soon as a difference looks significant considerably increases the rate of false positives. A widely cited simulation shows: if you check after every ten additional observations per group and stop at the first p-value below 0.05, the share of false positive results rises from 5 % to 14.3 %.
For the decision, this means: a winner that only emerged through clever stopping will not hold up in the market. The stopping rule makes the result defensible, including towards colleagues who expected a different outcome.
A good stopping rule answers four questions before the start: which metric decides? From what level of certainty does a variant count as ahead? What minimum data base is needed before an early end can be considered? And when does the test end at the latest, even without a clear result?
In addition, there are abort conditions in case a test does not work technically, for example if ads are not delivered or the cost per visitor is well above plan.
Before the start, two rules are written into the test design. The stopping rule: an early end is possible if the probability of being the best variant reaches the threshold for the test type (90 %, or 95 % for price differences) and the data base is sufficient; the planned maximum number of visitors ends the test in any case. The winner rule: a significant difference on the primary metric, P(best) above the threshold, and a lead that is consistent with the upstream steps.
The measurement period also counts: a test should cover a full week including a weekend wherever possible, because behaviour changes at weekends. Technical aborts are not decided before the third day.
Three tariff pages, each with 2,200 visitors after four days: 4.0 %, 5.0 % and 4.3 %. The middle variant is ahead with a P(best) of around 83 %. The 90 % threshold has not been reached, and the measurement period does not yet include a weekend.
According to the rule, the test keeps running, even though the interim result looks tempting. Without a rule, the temptation to declare the leading variant the winner at this point would be great.
The stopping rule is not the same as sample size planning: planning defines how many visitors are targeted, the stopping rule defines when the result is read and the test ends. Sequential testing methods are a statistical way of building repeated checking into the analysis from the outset.
No rule replaces judgement about the quality of the measurement period. More visitors do not compensate for a skewed period. A stopping rule can also only prevent deciding too early. It does not guarantee a clear-cut result: if a test ends without a clear winner, that is a valid result, for example that two variants are equally strong.
Simmons, Nelson & Simonsohn 2011: Simulation: checking from 10 observations per group after every 10 more and stopping at p < 0.05 yields 14.3 % instead of 5 % false positive results; checking after every observation yields 22.1 %. False-Positive Psychology, Psychological Science 22(11). Source
Wasserstein & Lazar 2016: The American Statistical Association states: a p-value measures neither the size of an effect nor its importance, and business decisions should not depend solely on whether a p-value falls below a threshold. The ASA Statement on p-Values: Context, Process, and Purpose, The American Statistician 70(2). Source
Yes, to check the technology and delivery. But decisions are made only according to the rule defined in advance.
Because repeated checking considerably increases the rate of false positives. An early interim result is often a random spike.
It is recorded in the test design before the start and agreed with the client. Deviations are documented in advance.
Bring your decision question, and we will outline a possible test design.
You will speak with Daniel Putsche
Founder & CEO, 30 minutes
Read more