Running a second holdout test after the first one failed
The direct approach got tried once already: pause a channel in one region for a few weeks and read the difference, no outside design, no second look at the setup. The platform quietly redistributed the paused budget to other regions before the window closed, and the region chosen was small enough that even a clean result would barely move the needle, so weeks of discipline produced nothing usable. That failure is attached to whoever ran it, and the instinct going into a second attempt is to over-check every input, because a second inconclusive result would cost more than the first.
A design built to close exactly the gaps that broke the first attempt, with comparable regions selected from the account's own history and the test sized for what it can actually detect, gets verified before anything goes live. A result sized to actually answer the question replaces one the first attempt was too small to reach. A third run is no longer needed to trust a self-inflicted null result enough to walk away from it.
Re-opening the same budget question every planning cycle
One test answers one question, and by the time the next planning cycle starts, the answer has aged out and the same doubt about the same channel resurfaces, because nothing carried forward from the last round into this one. Each season a fresh budget effectively gets requested to re-litigate a question already answered once, and the pattern looks like a standing weakness rather than a one-time gap. Meanwhile the model that should hold that history still runs on last quarter's assumptions, because no one closed the loop between the result and the model that should have absorbed it.
A testing calendar that compounds instead of resetting means each season's read becomes a prior for the next one. One model keeps absorbing what the tests find, instead of a result and a model quietly drifting apart. And next year's budget conversation starts from a track record instead of a fresh request.