Testing incrementality without hurting live performance

The short answer

Capping what gets put at risk replaces splitting the account in half. A right-sized geo test, designed and sized in Cassandra, changes spend in one group of regions, adding budget or holding some back, while a matched set of comparable regions carries on untouched. What keeps the cost small is how little of the business sits in the treated group.

Applies to

EcommerceCharity/Non-profitAgencyBrand
Testing incrementality without hurting live performanceregionscomparisontreatedONE GROUP MOVES, THE REST DO NOT
Spend changes in one group of regions while a matched set carries on untouched, and the gap between them is the read.

Running an incrementality test without turning off a channel

The standard proposal is a ten percent account split, and it arrives in the same quarter the number that will drop gets judged. The method is invisible from where leadership sits, so a statistically valid test and a bad month look identical in the deck presented. Suspicion that platform numbers overstate the channel is exactly why the test gets requested, and that is exactly what makes the downside impossible to pre-argue.

The causal answer arrives with the exposure bounded and written down before anyone approves it, so the dip becomes a decision the business signed off on rather than a result someone has to defend after the fact. A design where only the treated regions change and the comparison regions carry on untouched puts the difference between two groups at stake, rather than the absence of a channel. The finding lands in weeks, on a schedule set against the account's own reporting calendar.

Geo testing when campaigns are deliberately unrestricted

Campaigns run broad on purpose, with no geographic or demographic limits, because the algorithm performs better left alone. Every test design on offer starts by forcing regions onto them, which changes the thing being measured and does it during the strongest quarter on record. The cost of that interference lands immediately on the account, while the learning is speculative and arrives later.

The experiments the account's own spend history already ran become readable, because regional variation nobody intended is still variation a model can use, and reading it costs nothing and disturbs nothing. A shortlist of the questions that history can already answer separates from the ones that genuinely need a live test. When a live test is the only route, it gets sized against the current run rate rather than a textbook percentage.

Testing when media commitments are contractual rather than tactical

The position is fixed by an annual agreement, so creating test conditions does not mean reallocating budget, it means opening a second supplier and paying more for the same reach. The increase is immediate, attributable and visible to the people who approved the original deal, while the answer bought with it is a maybe. So the question reopens every planning cycle, the current split gets noted as probably not optimal for reach, and nothing moves, because the cost of finding out is certain and the benefit is not.

The read from variation that already exists inside the account's own history settles the question without renegotiating anything or notifying anyone. The size of the gap replaces a directional opinion, which is what makes the conversation with the supplier possible at all. It arrives before the next commitment window, not after it closes.

Testing a dominant channel that cannot afford to be paused

One channel funds the year, and every test design on offer begins by asking the account to hold back a slice of the thing it cannot afford to lose. A single point of that channel is a material sum in the plan, the economic climate has made the board less tolerant of experiments, and the people who would approve it have watched income targets slip for reasons that had nothing to do with measurement. So the question stays open indefinitely, because the safe version of the answer has never been put in front of them.

A sequence sized to what the budget can genuinely absorb, in steps small enough that the first one is approved on its own merits without a debate about the whole programme, becomes possible. Each step gets designed to answer one question rather than several. And accumulated certainty builds up without ever having put the base at risk.

Running incrementality tests across a client book without breaking BAU

Holding a client's trading steady while still producing evidence their team will accept means the test has to survive a promo calendar nobody at the agency controls and a configuration nobody double-checked. A misconfigured test on client money is not a bad data point, it is the relationship, and once one goes wrong the doubt spreads across the rest of the book. Clients also get cold feet about holdouts precisely when the test would be most informative, which is peak.

A design coordinated against the client's own calendar lets the test and the promotions coexist instead of contaminating each other. Pre-flight verification before anything goes live is where labelling errors and stray geographies get caught rather than in the readout. The conversation in the readout stays about the finding rather than about competence, backed by a method their analysts can inspect.

What changes

Choosing between knowing and performing stops being necessary, because the size of the bet is set before the test starts and the worst case is a number already agreed to.

What this does not do

What decides whether a geo test reads is how long it runs, not how much it spends: across 123 of our experiments, tests shorter than two weeks returned a clear answer 27% of the time and tests run for four to six weeks did so 71% of the time. A test nobody can leave running for four weeks is a test worth delaying. Reads are at campaign level. There is no ad-set or audience-group depth, and this is a strategic layer rather than a day-to-day optimiser. Tests get designed with the account rather than proposed automatically: nothing here decides on the account's behalf which experiment to run next.

Who this is for

The teams this is written for are mid-market direct-to-consumer brands, retailers and publishers whose campaigns run broad and unrestricted or are locked into annual media agreements, and nonprofits whose largest channel is offline and funds the annual income target. It also applies to agencies that have to design and run an incrementality test on a client's own budget across a mixed-vertical book.

Questions

What is a geo incrementality test?

A geo incrementality test splits comparable regions into a treated group and a control group, changes spend in one and not the other, and reads the difference in outcomes. Because the two groups share the same seasonality and market conditions, the gap between them is the causal effect of the spend rather than a correlation with it.

How much revenue is exposed during a geo incrementality test?

Exposure is a design parameter rather than a fixed cost, because only the treated regions change while the matched comparison set carries on untouched. Across 15 of our own geo experiments that returned a clear read, the revenue exposed had a median of 0.7%, and nine in ten exposed under 2.7%. The number to agree internally is the maximum the business is willing to put at risk, not the expected loss.

Can incrementality be measured without running a test at all?

Sometimes, yes. Most accounts have unintended regional variation already sitting in their spend history, and where that variation is large enough and clean enough it works as an experiment that already happened. It will not answer every question, but it settles some of them at zero disruption and narrows what actually needs a live test.

How does an incrementality test run on a client account without disrupting trading?

By coordinating the design with the promotional calendar rather than around it, holding channel changes steady inside the test window, and verifying the configuration before launch rather than discovering errors in the readout. The risk to the account is the first thing engineered, not the last thing checked.

When does this not apply?

When the account does not have the spend or the history to separate signal from noise at regional level, when the decision needed is intraday or at ad-set level, or when what is actually needed is day-to-day campaign management rather than a strategic read. In those cases a geo test will return a wide interval that answers nothing.

The product behind it