Incrementality testing vs A/B testing: the difference, and when to use each
An A/B test compares two versions of something, such as two ads or two landing pages, to see which performs better, and both groups see advertising. An incrementality test, or lift test, compares a group that sees advertising with a group that sees none, to measure whether the advertising produced sales that would not have happened otherwise.
The two are often confused because both split an audience or a market into groups. The difference is what the control group receives. In an A/B test it receives another version. In an incrementality test it receives nothing. This page defines each test, compares them on six dimensions, works through a lift calculation step by step, shows how lift is calculated, and sets out when each is the better choice and how the two fit into one testing programme.
A/B testing, defined
An A/B test splits an audience into two groups and shows each a different version of one thing: an ad creative, a headline, a bid strategy, a landing page, an email subject line. The version with the better result on the chosen metric wins.
Most ad platforms and site-testing tools run A/B tests natively. They are quick to set up and can return results in days when traffic is high.
A multivariate test extends the same idea to several elements at once, for example three headlines crossed with two images. It needs more traffic, because each combination needs enough exposure to compare.
Both groups are exposed to marketing, so an A/B test tells you which version works better. It does not tell you whether running the activity at all produced extra sales, because neither group represents the case with no activity.
Incrementality testing, defined
An incrementality test measures the effect of the activity itself. One group is exposed to a channel or campaign, and a comparable group is not. The difference in outcomes between the two is the incremental effect, often called lift.
The control group can be defined by person, as in a platform conversion lift study, or by region, as in a geo test. A geo test changes spend in one group of regions, adding budget or holding some back, while a matched set of comparable regions carries on untouched.
Because one group sees nothing, the test answers whether the channel earns its budget: would these sales have happened without the spend.
A third design switches spend on and off over time in the same market and compares the periods. It is simple to run, and it is more exposed to seasonality and other changes that happen at the same time, so it suits channels where regional or audience splits are not possible.
The differences, dimension by dimension
The question. An A/B test asks which version performs better. An incrementality test asks whether the activity produces sales that would not otherwise happen.
The control group. In an A/B test the control sees another version. In an incrementality test the control sees no advertising from the channel tested.
What it can return. An A/B test always produces a winner between two versions, even if both add nothing. An incrementality test can return zero, showing that the channel added nothing.
Scope. An A/B test works inside a channel. An incrementality test measures a channel or campaign as a whole.
Duration. An A/B test can finish in days with enough traffic. An incrementality test usually needs several weeks.
Cost. An A/B test costs little beyond setup. An incrementality test withholds or adds spend for its duration.
How to calculate incrementality
The basic calculation compares the two groups over the test period, after adjusting for how they compared before it started.
Lift is the difference in outcome between the exposed and control groups, divided by the control group's outcome. If the exposed group produced 1,100 orders and the adjusted control 1,000, lift is 10%.
Incremental conversions are the difference itself, here 100 orders.
Incremental return on ad spend divides the incremental revenue by the spend that produced it. It is usually lower than the return an ad platform reports, because platform figures also count sales that would have happened anyway.
These figures are illustrative. A real analysis also reports how certain the estimate is, and the smallest effect the test was designed to detect.
Why a winning A/B test can still lose money
An A/B test compares two versions of the same activity, so it finds the better of the two. If neither version produces extra sales, the winner is still declared, and the channel keeps its budget.
This is the most common gap in a testing programme built only on A/B tests: a steady stream of improvements inside channels whose overall contribution has never been measured. An incrementality test fills that gap by comparing the channel against no channel.
A common example is retargeting. An A/B test can show that one retargeting creative beats another. It cannot show whether the people being retargeted would have returned and bought anyway, which for many businesses is the larger question. A holdout on the retargeting audience answers it.
How long an incrementality test needs
Across 123 of our own geo experiments, 27% of tests under two weeks reached significance, 11 tests in that band, against 71% of tests running four to six weeks, 34 tests. Past six weeks the rate stops improving, and spend level separated nothing.
An A/B test can be read much sooner, because it compares two versions within the same traffic and the difference between versions is often larger than the effect of a whole channel.
What each test needs to be valid
An A/B test needs groups that are split evenly and see the versions over the same period, enough traffic to detect a difference, and one change at a time.
An incrementality test needs a control group that is genuinely comparable, enough weeks for the effect to show, and a clean window without overlapping launches or promotions. For platform lift studies, it needs the platform to link exposure to purchase for each person, which is harder when purchases happen offline or across devices.
The most frequent mistakes are the same for both: stopping a test as soon as it looks significant, changing budgets or creatives in the middle of the test, and choosing the success metric after seeing the results. Each one makes a result look stronger than the data supports.
When an A/B test is the better choice
An A/B test is the better choice for decisions inside a channel whose value is already established: which creative to run, which audience to target, which landing page converts better. It is fast and cheap, and it is the right tool for continuous optimisation.
It is also the better choice when traffic is high and the difference between versions is likely to be large, because it returns an answer in days rather than weeks.
When an incrementality test is the better choice
An incrementality test is the better choice when the question is whether a channel or campaign deserves its budget at all: before scaling a new channel, when reported returns look too good, when a large line item has never been measured, and when finance asks what a channel really adds.
It is also the way to check a platform's reported return. If a channel reports a strong return and an incrementality test finds a much smaller one, the difference is the share of sales the platform was crediting that would have happened anyway.
What neither test does
Neither test covers the whole marketing mix. Each measures one thing in one period, and results date as spend levels and seasons change.
An incrementality test measures a channel at one spend level. It does not show how the effect would change with more or less budget. That needs repeated tests or a marketing mix model calibrated with the test result.
An A/B test ranks versions of an activity. It does not show whether the activity itself should run.
And both need enough volume. A small business with few weekly conversions may not have the traffic for a meaningful A/B test or enough regions for a clean geo test, and results from underpowered tests are often unreliable.
How to choose
Choose by the question.
Which version works better: run an A/B test.
Whether the channel produces extra sales at all: run an incrementality test.
For a testing programme: use incrementality tests to decide which channels deserve budget, A/B tests to improve performance inside those channels, and a marketing mix model to set the split across them.
Cassandra is a marketing mix modeling platform that runs geo experiments in the same system as the model, using the same method family as GeoLift. A test result enters the model as a constraint, so it improves the estimates for other channels too.
Reads sit at campaign level, in the planning layer. A/B testing of creatives and landing pages stays with the ad platforms and site tools built for it.
A price is published.
Adding incrementality tests to an A/B programme
Most teams already run A/B tests and add incrementality tests later, starting with the largest channel or the one whose reported returns are least trusted.
The main change is calendar planning. An incrementality test needs a protected window of several weeks, so it has to be agreed with the teams that run promotions and launches.
Record each test in the same place: the question, the design, the window and the result. Over time that record shows which channels have been measured directly and which still rely on reported figures.
A useful order for the first year is one incrementality test per quarter on the largest channels, with A/B tests running continuously inside every channel. By the end of the year the biggest budget lines have a direct measurement behind them.
Questions
What does incrementality testing mean?
Measuring the extra sales an activity produces by comparing a group that saw it with a comparable group that did not. The difference is the incremental effect.
How do I calculate incrementality?
Subtract the control group's outcome from the exposed group's, after adjusting for how they compared before the test. Lift is that difference divided by the control outcome, and incremental return divides incremental revenue by spend.
Is a lift test the same as an A/B test?
No. An A/B test compares two versions and both groups see advertising. A lift test compares advertising with none, so it can show whether a channel adds anything at all.
How long does an incrementality test need to run?
Across 123 of our own geo experiments, 27% of tests under two weeks reached significance against 71% of tests running four to six weeks, and spend level did not separate the readable tests from the unreadable ones.
Can an A/B test measure incrementality?
Only if one of the two arms receives no advertising, at which point it is an incrementality test. Two versions of an ad compared against each other cannot show whether either one adds sales.