Skip to content

Experimentation & CRO

Did the Test Really Win? Significance and Peeking

Alessandro ScuottoPublished on 4 min read

The question that closes every A/B test seems simple: did the variant win? The honest answer is harder than the platform makes it look, because the way you look at a result can manufacture winners that do not exist. In this article I explain what statistical significance really means, why peeking is the error that kills more tests than it seems, and how I decide whether a test won without fooling myself.

What significance says, and what it does not

Statistical significance answers a precise question: how likely is it that the observed difference between variant and original is down to chance, rather than a real effect? It does not say how large the effect is, nor how important it is for the business. A test can be statistically significant and produce a gain so small it is not worth implementing, and it can be promising but not yet significant because the sample is too small.

The number the platform shows is reliable only under one condition: that the sample has reached the size decided before starting. On a still-growing sample, significance computed moment by moment does not mean what it appears to.

Peeking: the error that manufactures false winners

Peeking is checking results every day and stopping the test the moment the variant seems to win. It looks harmless, and it is instead the most costly error in experimentation. Every time you look at a partial result and decide whether to stop, you increase the chance of declaring a winner that, run to the end, would have returned to parity. It is not bad luck: it is mathematics. The more chances you give yourself to stop on a peak, the more often you will stop on a random one.

How you actually avoid it, without stopping looking at the data

The fix is not to ignore the test until it ends, which is unrealistic. It is to decide the sample size or minimum duration in advance, and commit to not making decisions before then. When the platform allows it, you use statistical methods designed for sequential reading, which account for the fact that you are looking multiple times and correct the threshold accordingly. As an Optimizely Experimentation Certified Core Strategist this is one of the points where I most often correct the instinct of someone who wants to stop a test because "it is already clear how it ends."

Another temptation to govern is post-hoc segmentation: slicing results by device, channel or country until you find a winning segment. The more cuts you make, the higher the chance of finding a significant one by pure chance. Segments should be decided before the test, in limited number and tied to a hypothesis.

What counts as a win

A test won when the primary metric decided in advance exceeded the agreed threshold, on the full sample, without the guardrail metrics getting worse. Any other definition, built after seeing the data, is a way of telling yourself the story you prefer. This discipline starts upstream, in the way you write the hypothesis: if you want to see how I tie metric and threshold from the formulation itself, I cover it in the article on how to write an A/B test hypothesis that gets results. The full picture of the method is in the guide to CRO and experimentation.

Get a CRO consultation

Share

Let's talk about your project

Found this useful? Let's see how it applies to your context.

All fields are required.

Experimentation & CRO

A/B Test Hypotheses That Get Results

How to write an A/B test hypothesis that gets results: from behavioural basis to expected metric, the format I use so I don't waste traffic.

4 min read
Experimentation & CRO

+57% Conversion: Anatomy of a CRO Program

The anatomy of the experimentation program behind the +57% conversion lift on Whirlpool EMEA: how it started, what we tested and what worked.

3 min read