Skip to content

Experimentation & CRO

CRO & Experimentation: The Complete A/B Test Guide

Alessandro ScuottoPublished on 8 min read

In most companies I have worked with, "doing CRO" means running an A/B test now and then on a button or a headline and hoping something happens. That is not experimentation; it is an isolated attempt dressed up as method. When I built the program that took the Whirlpool EMEA e-commerce to a +57% lift in conversion, the difference was not one lucky test: it was a process that generated good hypotheses one after another, with a statistical discipline that stopped us from fooling ourselves about the results. This guide collects that method in full, from the point where a hypothesis is born to the moment you decide, honestly, whether a test actually won.

What an experimentation program is, and why it is not a list of tests

CRO, conversion rate optimization, is often reduced to a tactic: change a colour, move a call to action, measure the click-through. An experimentation program is something else. It is a cycle that repeats: observe behaviour, form a hypothesis, test, read the result, and let the learning feed the next hypothesis. The single test matters less than the cycle that produces it.

In my work on Whirlpool EMEA that cycle had a fixed cadence: weekly backlog review, new tests launched when a traffic slot opened up, results read against thresholds set in advance, not once the numbers were in. Without this structure, a company ends up testing what is convenient to test, usually small visual details, instead of what real user behaviour suggests testing. The difference between one isolated winning test and a cumulative +57% almost always lives here: in the cadence, not in the brilliance of a single test.

Where a hypothesis worth testing comes from

A test is worth only as much as the hypothesis that generates it. Too many CRO backlogs are lists of "nice ideas" with no behavioural basis: variants get tested because a competitor uses them, not because a real friction in the user's journey was understood. My approach starts from a different place, and it predates my work in marketing: an MSc in Cyberpsychology, where I began asking why people behave the way they do online long before I tried to solve it with a test. That is the base on which I build every test hypothesis I launch, and I go deeper into it in the guide on online consumer psychology.

A solid hypothesis follows a simple structure: what I observe in real behaviour (session data, heatmaps, drop-off in a funnel), what change I believe addresses it, and which metric I expect to move, by how much. Not "let's try a green button," but "users abandon checkout at the shipping-cost step because the cost appears late: showing it earlier reduces abandonment at that step." That is the difference between a test that informs the next one and a test that ends up in a forgotten report. I have collected the full format in a guide on how to write an A/B test hypothesis that gets results.

Prioritising the backlog: not every test is worth the same time

With more test ideas than resources to run them (the norm, not the exception), prioritisation decides what actually happens. There are established frameworks: some weigh expected impact, confidence in the hypothesis and ease of implementation; others weigh the page's intrinsic potential, the importance of the traffic passing through it and how easy it is to test. Neither is the "right" one in the abstract: it depends on how much traffic you have, how costly it is to implement a variant, and how much your organisation tolerates tests that fail.

What matters, regardless of the framework chosen, is having an explicit, shared criterion, so the backlog does not fill up with the tests that are easiest to explain in a call instead of the ones most likely to move the numbers. On Whirlpool EMEA many of the biggest wins came not from the most-visited pages, but from the funnel steps where traffic was qualified and the observed friction was strongest.

Designing the experiment: metrics, randomisation, segments

A well-designed test decides before launch, not after, what counts as a win. You need one primary metric, the one that decides whether the test won or lost, and a small set of guardrail metrics that check you are not optimising one thing at the expense of another (more button clicks but lower margin per order, say). If you watch ten metrics hoping one turns out significant, sooner or later one will by pure chance: that is not a result, it is statistical noise dressed up as insight.

Segmentation deserves the same caution. It is tempting to slice results by device, traffic channel or country until you find a "winning" segment: but the more cuts you make, the higher the chance of finding a significant one by accident. Segments should be decided before the test, in limited number, and tied to a hypothesis (for example, if you believe friction hits mobile and desktop differently), not generated by mining the data after the test ends.

Statistical significance and the error that kills more tests than it seems

Peeking, checking results every day and stopping the test the moment the variant seems to win, is probably the most common and most costly error in experimentation. Every time you look at a partial result and decide whether to stop, you increase the chance of declaring a winner that, run to the end, would have returned to parity. Statistical significance computed on a still-growing sample does not mean what it appears to mean. I go deeper into this in the article on statistical significance and peeking.

The practical fix is not to stop looking at the data (that is normal) but to decide the sample size or minimum duration in advance, and to use, when the platform allows it, statistical methods designed for sequential reading rather than a single end-of-test check. In my work as an Optimizely Experimentation Certified Core Strategist this is one of the points where I most often correct the instinct of someone who wants to stop a test "because it is already clear how it ends."

Optimizely in practice: what the right platform changes

I use Optimizely every day as an experimentation platform, and it is the one on which I built the Whirlpool EMEA program. What a mature platform changes is not the ability to launch a variant (free tools do that too) but the QA before launch (checking that the variant behaves well across browsers and devices before exposing it to real traffic), the statistical engine that correctly handles the peeking problem above, and the ability to connect feature flags and A/B tests in the same tool, so a gradual product rollout and a conversion experiment speak the same language instead of living in separate systems.

This matters most in an enterprise context, where a badly implemented test costs more than lost time: it costs stakeholder trust in the experimentation program as a whole, and it is that trust that lets you keep testing even when a test loses.

Experimentation culture matters more than the single winning test

A CRO program that survives a year, and not just a quarter, needs something beyond the platform: a cadence shared with stakeholders, a common language about what "winning" a test means, and the acceptance that most tests, even well-designed ones, do not produce a clear winner. If every lost test is read as a failure of the program instead of as a learning, the organisation stops testing exactly when it would start learning the most.

Get a CRO consultation

Where to start if your first program begins today

If you are building an experimentation program from scratch, order matters more than tools. First you define how hypotheses are written and where they come from: real behaviour data, not isolated intuitions. Then you set an explicit prioritisation criterion, even a simple one, as long as it is shared and applied consistently. Only then do you choose the platform, because a powerful tool used without statistical discipline just produces false winners faster.

The +57% lift in conversion on Whirlpool EMEA did not come from an isolated stroke of genius, but from months of this cycle repeated without shortcuts: grounded hypotheses, a prioritised backlog, tests read to the planned end rather than to the point where the number looked best. It is a repeatable method, not a lucky break, and it is what I bring to every experimentation program I run for the enterprise clients I work with. If you want to see how it applies to your e-commerce or digital product, you will find an overview of the CRO and experimentation services I offer, or you can write to me directly.

Share

Let's talk about your project

Found this useful? Let's see how it applies to your context.

All fields are required.

Experimentation & CRO

A/B Test Hypotheses That Get Results

How to write an A/B test hypothesis that gets results: from behavioural basis to expected metric, the format I use so I don't waste traffic.

4 min read
Experimentation & CRO

+57% Conversion: Anatomy of a CRO Program

The anatomy of the experimentation program behind the +57% conversion lift on Whirlpool EMEA: how it started, what we tested and what worked.

3 min read