Horizon
Home
/
Glossary
/
Stopping Rule

Stopping Rule

A stopping rule defines, before a test starts, under which conditions it will end and when a result counts as decided; it prevents a test from being stopped exactly when the numbers happen to look right.

Updated
September 28, 2026
· Horizon

Running tests invite you to check. If you stop at the first favourable interim result, you overestimate the certainty of your result. A stopping rule separates planning from analysis.

Why this matters for your decision

Interim results fluctuate, especially in the first few days. Checking daily and stopping as soon as a difference looks significant considerably increases the rate of false positives. A widely cited simulation shows: if you check after every ten additional observations per group and stop at the first p-value below 0.05, the share of false positive results rises from 5 % to 14.3 %.

For the decision, this means: a winner that only emerged through clever stopping will not hold up in the market. The stopping rule makes the result defensible, including towards colleagues who expected a different outcome.

What a stopping rule defines

A good stopping rule answers four questions before the start: which metric decides? From what level of certainty does a variant count as ahead? What minimum data base is needed before an early end can be considered? And when does the test end at the latest, even without a clear result?

In addition, there are abort conditions in case a test does not work technically, for example if ads are not delivered or the cost per visitor is well above plan.

How Horizon implements it

Before the start, two rules are written into the test design. The stopping rule: an early end is possible if the probability of being the best variant reaches the threshold for the test type (90 %, or 95 % for price differences) and the data base is sufficient; the planned maximum number of visitors ends the test in any case. The winner rule: a significant difference on the primary metric, P(best) above the threshold, and a lead that is consistent with the upstream steps.

The measurement period also counts: a test should cover a full week including a weekend wherever possible, because behaviour changes at weekends. Technical aborts are not decided before the third day.

Example

Three tariff pages, each with 2,200 visitors after four days: 4.0 %, 5.0 % and 4.3 %. The middle variant is ahead with a P(best) of around 83 %. The 90 % threshold has not been reached, and the measurement period does not yet include a weekend.

According to the rule, the test keeps running, even though the interim result looks tempting. Without a rule, the temptation to declare the leading variant the winner at this point would be great.

How it differs

The stopping rule is not the same as sample size planning: planning defines how many visitors are targeted, the stopping rule defines when the result is read and the test ends. Sequential testing methods are a statistical way of building repeated checking into the analysis from the outset.

Limitations

No rule replaces judgement about the quality of the measurement period. More visitors do not compensate for a skewed period. A stopping rule can also only prevent deciding too early. It does not guarantee a clear-cut result: if a test ends without a clear winner, that is a valid result, for example that two variants are equally strong.

Evidence

Simmons, Nelson & Simonsohn 2011: Simulation: checking from 10 observations per group after every 10 more and stopping at p < 0.05 yields 14.3 % instead of 5 % false positive results; checking after every observation yields 22.1 %. False-Positive Psychology, Psychological Science 22(11). Source

Wasserstein & Lazar 2016: The American Statistical Association states: a p-value measures neither the size of an effect nor its importance, and business decisions should not depend solely on whether a p-value falls below a threshold. The ASA Statement on p-Values: Context, Process, and Purpose, The American Statistician 70(2). Source

Frequently asked questions

Is it allowed to look at a running test?

Yes, to check the technology and delivery. But decisions are made only according to the rule defined in advance.

Why not simply stop as soon as the result is significant?

Because repeated checking considerably increases the rate of false positives. An early interim result is often a random spike.

Who defines the rule?

It is recorded in the test design before the start and agreed with the client. Deviations are documented in advance.

When is your test decided?

Bring your decision question, and we will outline a possible test design.

Daniel Putsche

You will speak with Daniel Putsche
Founder & CEO, 30 minutes

Read more

Thank you, we have your request.

We will get back to you within one working day with suggested times.

Close

Demo · Video

A complete test, from design to data analysis

Play the demo

Chapters: test design · offer pages · live data · result

Book a call

First call

Book a call

30 minutes, your decision question, a possible test design. No preparation needed on your side.

By submitting, you agree to the processing of your details to arrange a call. Details in our privacy policy.

Request a call
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.
SAMPLE REPORTExample

Sample report: insurance

Data analysis · Test design · Metrics · Methodology

Sample report

Sample report: insurance

A complete results report with example values: research question, test design, purchase intent per variant and the data analysis.

12 pages, as a PDF to share internally
Free, download right after a short form
Note