Geo experiments vs holdout tests: how each design works, and which one fits your business

A geo experiment measures incremental effect by changing spend in some regions and comparing them with similar regions left unchanged. A holdout test withholds advertising from a group and compares it with a group that saw it. A geo holdout is both at once: a geo experiment that pauses spend in the test regions.

The terms overlap, which is why comparisons of the two are often confusing. The useful distinction is the unit being split. A geo design splits by region, so it works wherever sales can be counted by region. An audience holdout splits by person, so it needs exposure and purchase to be linked for each individual. This page defines each design, compares them on six dimensions, covers how long a test needs and what it costs, and sets out when each is the better choice.

Geo experiments, defined

A geo experiment, or geo test, splits a market into regions such as cities, designated market areas or postcodes. Spend in one channel is changed in a group of test regions, while a matched group of comparable regions carries on untouched as the comparison. The difference in outcomes between the two groups, measured against how they tracked each other before the test, is the channel's incremental effect.

The change can go either way. A scale-up test adds budget in the test regions to see what extra spend produces. A geo holdout, sometimes called a go-dark test, pauses or reduces spend there to see what is lost without it.

The comparison regions are chosen because their historical outcomes moved in step with the test regions, so they stand in for what the test regions would have done without the change.

Two analysis approaches are common. Matched-market testing pairs each test region with one or a few similar regions. Synthetic control builds the comparison from a weighted blend of many untouched regions, chosen so that the blend follows the test regions closely before the test starts. Open-source tools such as Meta's GeoLift use the second approach.

Holdout tests, defined

A holdout test withholds advertising from a group and compares its outcome with a group that could see the advertising. The group can be defined by region, which makes it a geo holdout, or by person, which makes it an audience holdout.

An audience holdout excludes a share of an addressable audience from a channel's ads. Ad platforms offer this as a built-in study, often called a conversion lift or brand lift study, where the platform sets up the exclusion and reports the result.

Because the unexposed people are drawn from the same audience as the exposed ones, the comparison is direct. It relies on the platform being able to link each person's exposure to their later purchase.

The differences, dimension by dimension

The unit split. A geo experiment splits regions. An audience holdout splits people.

What it needs. A geo experiment needs sales that can be counted by region and enough regions to build a matched comparison. An audience holdout needs exposure and purchase linked at the level of the individual.

Channels covered. A geo experiment covers offline sales, television, connected TV, audio, out-of-home and any channel bought by region. An audience holdout covers addressable digital channels where the platform can control who sees ads.

Who runs it. A geo experiment is usually designed and read by the advertiser or an independent party. A platform audience holdout is designed, executed and reported by the platform being evaluated.

Cost to the business. A geo experiment affects the test regions only. An audience holdout affects the withheld share of the audience.

Contamination. A geo experiment can be affected by people who cross region boundaries. An audience holdout can be affected by people reached through other channels or other devices.

The question that decides which design you can use

Can each person's exposure be linked to their purchase, reliably, across the whole path?

Where it can, an audience holdout is available and often cheaper, because the platform handles the setup. Where it cannot, an audience holdout cannot measure the effect, however carefully it is run.

The link fails more often than teams expect. In-store and offline purchases break it, as do connected TV and audio. Shared devices, consent limits and long consideration cycles weaken it. In each of those cases a geo design is the one that works, because it only needs sales by region.

An example: a retailer that sells mostly in stores cannot use a platform audience holdout to measure a video campaign, because the platform cannot see who later bought in a shop. Pausing the campaign in a set of cities and comparing store sales with similar cities answers the question directly.

How long a test needs to run

Both designs need a window long enough for the effect to show. Across 123 of our own geo experiments, 27% of tests under two weeks reached significance against 71% of tests running four to six weeks, and spend level separated nothing.

In practice, plan for four to six weeks, and keep the window clean: no new launches, promotions or budget changes in the test regions while it runs. A shorter test saves little money and usually returns no readable result.

What a geo test costs in revenue

The usual objection to a holdout of any kind is the revenue at risk while spend is paused. Across 15 of our own geo experiments that returned a clear read, the median share of revenue exposed was 0.7%, and nine in ten exposed under 2.7%.

The share is small because a test does not need the whole market. It needs enough regions to measure a difference, and the comparison regions carry on as normal. A scale-up test puts no existing revenue at risk at all, since it adds spend.

Who reads the result

A platform-run holdout is designed, executed and reported by the party whose budget is being evaluated. That is a feature of the arrangement, and it is the reason many teams also run independent tests for their largest channels.

A geo experiment designed and read in-house, or by an independent party, keeps the design choices visible: which regions were chosen and why, how long it ran, and what effect it was able to detect. Writing those down before the test starts is what makes the result credible to people who were not involved.

The detectable effect deserves particular care. A test sized to detect only a large effect will report no effect for a channel with a moderate one, and that result is easy to misread as proof the channel does nothing. Stating the smallest effect the design could detect lets a reader interpret either outcome.

When a holdout test is the better choice

An audience holdout is the better choice for an addressable digital channel where the platform can link exposure to purchase, where the purchase happens online, and where a fast answer matters more than an independent one. It is quick to set up and needs no regional structure.

When a geo experiment is the better choice

A geo experiment is the better choice when sales happen offline or across several channels, when the channel is television, audio, out-of-home or connected TV, when identity cannot be joined reliably, and when the result has to stand up to scrutiny from outside the channel team.

It needs a business with regional structure. A company selling in a single small market may not have enough comparable regions to build a matched group.

What neither design does

Each test measures one channel, at one spend level, in one period. It does not show how the effect changes at a different spend level, and it needs repeating when conditions change.

Neither design covers the whole mix at once. Testing every channel in turn takes longer than most planning cycles, which is why teams use test results to calibrate a marketing mix model that covers the rest.

And neither creates a comparison group where the business has none. A single-market company with no regional structure and no addressable audience cannot run either design cleanly.

How to choose

Choose by what can be separated.

If exposure and purchase can be linked person by person, and the result does not need to be independent of the platform, an audience holdout is the quicker option.

If the purchase happens offline, the channel is bought by region, or identity cannot be joined, use a geo experiment.

If the result will decide a large budget, run the design that someone other than the channel's seller can check, and write down the regions, the length and the detectable effect before it starts.

How Cassandra runs geo experiments

Cassandra is a marketing mix modeling platform that designs and runs geo experiments in the same system as the model, using the same method family as GeoLift. Comparison regions are matched on how closely their history tracks the test regions before the test begins.

Each result enters the model as a constraint, so a test on one channel also improves the estimates for the rest of the mix. Reads sit at campaign level, in the planning layer.

A price is published.

Moving from platform tests to independent tests

Most teams start with the lift studies their ad platforms offer, because they are free and quick. Moving to independent geo tests usually begins with the largest channel, where the stakes justify a result the platform does not grade.

The groundwork is a clean history of sales by region, which is also what a marketing mix model needs. With that in place, the first geo test mostly involves choosing regions, agreeing the window with the teams that run promotions, and recording the design before launch.

Platform studies can continue alongside. Comparing the two on the same channel is itself useful evidence about how far the platform's own figures can be relied on.

Questions

What is a geo holdout test?

A geo experiment that pauses or reduces spend in a group of test regions and compares their sales with similar regions where spend continued. The difference shows what the paused spend was producing.

What is a holdout test?

A test that withholds advertising from a group, defined by region or by person, and compares that group's outcome with a group that could see the advertising. The difference is the incremental effect.

What are geo experiments?

Tests that change spend in one group of regions and compare the outcome with a matched group of regions left unchanged. They measure incremental effect without user-level tracking, so they work for offline and broadcast channels.

How much revenue does a geo test put at risk?

Across 15 of our own geo experiments that returned a clear read, the median share of revenue exposed was 0.7% and nine in ten exposed under 2.7%. A scale-up test adds spend, so it puts no existing revenue at risk.

Are platform-run lift studies reliable?

They are useful and quick. They are designed, run and reported by the platform whose budget is being evaluated, which is why many teams confirm the results for their largest channels with an independent test.