Running an incrementality test without turning off a channel
The standard proposal is a ten percent account split, and it arrives in the same quarter the number that will drop gets judged. The method is invisible from where leadership sits, so a statistically valid test and a bad month look identical in the deck presented. Suspicion that platform numbers overstate the channel is exactly why the test gets requested, and that is exactly what makes the downside impossible to pre-argue.
The causal answer arrives with the exposure bounded and written down before anyone approves it, so the dip becomes a decision the business signed off on rather than a result someone has to defend after the fact. A design where only the treated regions change and the comparison regions carry on untouched puts the difference between two groups at stake, rather than the absence of a channel. The finding lands in weeks, on a schedule set against the account's own reporting calendar.
Geo testing when campaigns are deliberately unrestricted
Campaigns run broad on purpose, with no geographic or demographic limits, because the algorithm performs better left alone. Every test design on offer starts by forcing regions onto them, which changes the thing being measured and does it during the strongest quarter on record. The cost of that interference lands immediately on the account, while the learning is speculative and arrives later.
The experiments the account's own spend history already ran become readable, because regional variation nobody intended is still variation a model can use, and reading it costs nothing and disturbs nothing. A shortlist of the questions that history can already answer separates from the ones that genuinely need a live test. When a live test is the only route, it gets sized against the current run rate rather than a textbook percentage.
Testing when media commitments are contractual rather than tactical
The position is fixed by an annual agreement, so creating test conditions does not mean reallocating budget, it means opening a second supplier and paying more for the same reach. The increase is immediate, attributable and visible to the people who approved the original deal, while the answer bought with it is a maybe. So the question reopens every planning cycle, the current split gets noted as probably not optimal for reach, and nothing moves, because the cost of finding out is certain and the benefit is not.
The read from variation that already exists inside the account's own history settles the question without renegotiating anything or notifying anyone. The size of the gap replaces a directional opinion, which is what makes the conversation with the supplier possible at all. It arrives before the next commitment window, not after it closes.
Testing a dominant channel that cannot afford to be paused
One channel funds the year, and every test design on offer begins by asking the account to hold back a slice of the thing it cannot afford to lose. A single point of that channel is a material sum in the plan, the economic climate has made the board less tolerant of experiments, and the people who would approve it have watched income targets slip for reasons that had nothing to do with measurement. So the question stays open indefinitely, because the safe version of the answer has never been put in front of them.
A sequence sized to what the budget can genuinely absorb, in steps small enough that the first one is approved on its own merits without a debate about the whole programme, becomes possible. Each step gets designed to answer one question rather than several. And accumulated certainty builds up without ever having put the base at risk.
Running incrementality tests across a client book without breaking BAU
Holding a client's trading steady while still producing evidence their team will accept means the test has to survive a promo calendar nobody at the agency controls and a configuration nobody double-checked. A misconfigured test on client money is not a bad data point, it is the relationship, and once one goes wrong the doubt spreads across the rest of the book. Clients also get cold feet about holdouts precisely when the test would be most informative, which is peak.
A design coordinated against the client's own calendar lets the test and the promotions coexist instead of contaminating each other. Pre-flight verification before anything goes live is where labelling errors and stray geographies get caught rather than in the readout. The conversation in the readout stays about the finding rather than about competence, backed by a method their analysts can inspect.