Running incrementality tests around the academic calendar

The short answer

An incrementality test gets designed in Cassandra to close the exact gaps that broke a do-it-yourself attempt, with comparable regions selected and the test sized before launch, and each result carries into the next one instead of starting over. Timed against enrollment peaks and exam windows rather than a generic reporting calendar, one season's read becomes the next season's starting point.

Applies to

Universities/EducationBrand
Running incrementality tests around the academic calendarregionscomparisontreatedONE GROUP MOVES, THE REST DO NOT
Spend changes in one group of regions while a matched set carries on untouched, and the gap between them is the read.

Where this comes up

Running a second holdout test after the first one failed

The direct approach got tried once already: pause a channel in one region for a few weeks and read the difference, no outside design, no second look at the setup. The platform quietly redistributed the paused budget to other regions before the window closed, and the region chosen was small enough that even a clean result would barely move the needle, so weeks of discipline produced nothing usable. That failure is attached to whoever ran it, and the instinct going into a second attempt is to over-check every input, because a second inconclusive result would cost more than the first.

A design built to close exactly the gaps that broke the first attempt, with comparable regions selected from the account's own history and the test sized for what it can actually detect, gets verified before anything goes live. A result sized to actually answer the question replaces one the first attempt was too small to reach. A third run is no longer needed to trust a self-inflicted null result enough to walk away from it.

Re-opening the same budget question every planning cycle

One test answers one question, and by the time the next planning cycle starts, the answer has aged out and the same doubt about the same channel resurfaces, because nothing carried forward from the last round into this one. Each season a fresh budget effectively gets requested to re-litigate a question already answered once, and the pattern looks like a standing weakness rather than a one-time gap. Meanwhile the model that should hold that history still runs on last quarter's assumptions, because no one closed the loop between the result and the model that should have absorbed it.

A testing calendar that compounds instead of resetting means each season's read becomes a prior for the next one. One model keeps absorbing what the tests find, instead of a result and a model quietly drifting apart. And next year's budget conversation starts from a track record instead of a fresh request.

What changes

Starting over each season stops, because the read from the last test becomes the starting point for the next one instead of a question asked again from zero.

What this does not do

A designed test needs four uninterrupted weeks in a region more than it needs a large seasonal budget: across 123 of our experiments, four-week tests read 71% of the time against 27% under two weeks, and spend level did not separate them. A season shorter than the test window is the real constraint. Reads are at campaign level, not ad-set or audience-group depth, and this is a strategic layer rather than a day-to-day optimizer. The standing calendar is built together each cycle, not proposed automatically; nothing here decides unilaterally which test comes next.

Who this is for

Most relevant to education businesses that have already run a do-it-yourself pause test and gotten an inconclusive result, and to organic-heavy education brands running a mixed lead-generation and direct-enrollment funnel. It applies most where the same incrementality question keeps reopening each planning cycle instead of compounding into a track record, particularly around enrollment peaks and exam windows.

Questions

What is a designed geo incrementality experiment?

A designed geo experiment selects comparable regions from the business's own history and sizes the test before launch, how long it has to run and how small a change it can detect, rather than being assembled ad hoc. The comparison between the two groups is what produces a causal read, rather than a before-and-after look at one region alone.

Why does a self-run pause test often produce an inconclusive result?

Because small setup errors, such as not duplicating a paused campaign or a platform quietly redistributing the paused budget elsewhere, contaminate the comparison without being visible until the result comes back unreadable. Choosing a region too small to move the numbers compounds the problem, so weeks of discipline can still return nothing usable.

How do education businesses time incrementality tests around the academic calendar?

By working backward from enrollment peaks and exam-driven demand windows rather than forward from a generic reporting calendar, so the test result lands with enough runway to inform the decision it was meant to inform. A read that arrives after the peak answers a question that already closed.

When does this not apply?

When a channel or region does not have the spend or the history to separate a real effect from noise, when the decision needed is day-to-day campaign management, or when there is no runway left before the peak the result was meant to inform. In those cases a rushed test returns a wide interval that settles nothing.

What changes once incrementality testing becomes a standing program instead of a one-off?

Each season's result feeds directly into the model used for the next season's plan, instead of sitting alongside it as a separate finding that ages out. Budget conversations start from an accumulating track record instead of reopening the same unanswered question every cycle.

The product behind it