Marketing mix modeling vs incrementality testing: the differences, and how to use them together

Marketing mix modeling uses historical spend and sales data to estimate the contribution of every channel at once, and it suits budget allocation. Incrementality testing runs a controlled experiment, such as a geo test, to measure the causal effect of one channel over a set period. Most teams use both: tests calibrate the model.

The two methods answer different questions. A model covers the whole mix continuously, and its estimates come with uncertainty because they are fitted to observed data. A test covers one channel inside a window, and it measures cause and effect directly because it changes spend on purpose. This page explains how each works, compares them on six dimensions, sets out when each is the better choice, shows how long a test needs to run, and describes how the two work together.

Marketing mix modeling, defined

Marketing mix modeling, or MMM, fits a statistical model to weekly spend by channel, the outcome you care about, and the other factors that move it: seasonality, pricing, promotions and the economy. The model estimates how much each channel contributed, including delayed effects and the point at which extra spend starts to return less.

Because it works from aggregate data, it covers online and offline channels together and needs no user-level tracking. It answers portfolio questions: how to split a budget, where the next unit of spend returns most, and what a budget change is likely to do.

Its limit is that it learns from history. If a channel's spend rose in the same weeks demand rose, the model sees the pattern and cannot tell from the data alone which one drove the other.

Incrementality testing, defined

An incrementality test changes spend on purpose and measures what follows. In a geo test, spend in one channel is increased or held back in one group of regions while a matched set of comparable regions carries on untouched. The difference in outcomes between the two groups is the channel's incremental effect.

Other designs split audiences instead of regions, or switch spend on and off over time. All of them rely on a comparison group that did not receive the change.

Because the change is deliberate, the result measures cause and effect for that channel, at that spend level, in that period. It answers a validation question: did this spend produce sales that would not have happened otherwise.

Designing a test means three decisions made before it starts: which regions or audiences receive the change, how long it runs, and the smallest effect the design is able to detect. Writing the third down in advance is what makes a result that shows no effect interpretable afterwards.

The differences, dimension by dimension

Scope. A model covers every channel at once. A test covers one channel, or one campaign, at a time.

Method. A model infers effects from historical variation in spend. A test creates variation and observes the result.

What the answer means. A model returns an estimate with a range around it. A test returns a measured effect for the conditions it ran in.

Timing. A model can run on existing history today and is refreshed on a planning cycle. A test needs a window of several weeks that has not started yet.

Coverage of external factors. A model accounts for seasonality, price and the economy explicitly. A test controls for them through its comparison group.

Scaling to the whole budget. A model reports diminishing returns and supports reallocation across channels. A test measures one point and does not show how the effect changes at a different spend level.

Why a model benefits from a test

Every model is fitted to observed history, so it can mistake timing for effect. If spend and demand rise together, a more flexible model fits that pattern more closely without resolving it.

A test resolves it for one channel, because the spend change is independent of demand. The measured effect then becomes a constraint the model has to satisfy: whatever it concludes about the rest of the mix, it must reproduce the tested channel's result. Because channels are estimated together, that one anchored point also improves the estimates for the others.

The practical effect is that a small number of well-chosen tests improves the whole model, so a team does not need to test every channel.

Why a test benefits from a model

A test answers one question. With eight or ten channels, testing each in turn takes longer than the planning cycle the results are meant to inform, and running several at once risks overlapping comparison groups.

A result also dates. It describes one channel at one spend level in one season. A model carries the test result forward, combines it with the rest of the history, and shows how the channel is likely to respond at other spend levels.

A model also helps choose what to test next: the channels where its estimates are least certain and where the most budget is at stake.

How long an incrementality test needs to run

The most common reason a test fails to produce a readable result is a window that is too short.

Across 123 of our own geo experiments, 27% of tests under two weeks reached significance, 11 tests in that band. Tests running four to six weeks reached it 71% of the time, 34 tests. Past six weeks the rate stops improving, and spend level separated nothing.

In practice this means planning a testing programme around protected four-to-six-week windows. A team that can keep a window clean, with no overlapping launches or budget changes in the test regions, gets a readable result at a modest spend level.

Region selection matters as much as length. The comparison regions need a history of moving in step with the test regions before the test begins, and the choice should be made from that history before anyone sees results.

When an incrementality test on its own is the better choice

A test alone is the better choice when one channel dominates the budget and one decision is pending. If the question is whether the largest line item is worth its spend, a test answers it directly and faster than building a model.

It is also the better choice when a specific figure is disputed, for example when a platform's reported return is questioned inside the business. An experiment gives a direct answer to that one number.

And it is the only option when a channel has too little history for a model to read. A test creates its own variation, so it needs no history.

When a marketing mix model on its own is the better choice

A model alone is the better choice when no valid comparison group can be built. A business operating in one market, or selling through channels with no regional structure, may be unable to run a clean test, and a model with its uncertainty stated is the available instrument.

It is also the better choice when the question is allocation across the whole mix, which depends on how channels interact and on diminishing returns, and when the decision is due before a test window could finish.

What neither method does

Neither reads below campaign level for planning decisions, and neither replaces day-to-day optimisation inside a channel.

Neither makes a thin history thick. A model with too little data returns a wide range, which is the accurate result for that data. A test with too short a window returns no readable result, as the figures above show.

And a test result describes the conditions it ran in. It needs re-running when the spend level, the season or the channel's role changes substantially.

How to choose

Start with the method that answers the decision already due, then add the other.

If you need a budget split across the mix for a planning cycle, start with a marketing mix model. A test cannot answer a portfolio question.

If you need to confirm or challenge the value of one channel, start with an incrementality test. It gives a direct measurement of that channel.

For a lasting measurement programme, run both: the model to allocate and to choose what to test, and tests to calibrate the model.

How Cassandra runs both

Cassandra is a marketing mix modeling platform that runs the model and the experiments in one system. Its production engine is Bayesian.

Geo experiments are designed in the same place the model lives, using the same method family as GeoLift, and a measured result enters the model as a constraint. Reads sit at campaign level, in the planning layer.

A price is published.

Starting a combined programme

The first cycle is mostly data work. A clean, reconciled history of spend and outcomes is what the model needs, and it is also what makes it possible to pick comparable regions for a test. Teams that already have it move quickly.

The usual sequence is to fit a first model, use its least certain estimates to choose the first test, run the test over a protected window, and feed the result back into the model. Each later cycle is cheaper, because each test improves the model and each model refresh points to the next test worth running.

Keep a written record of every test: the channel, the regions, the window, the spend change and the result. That record is what lets a later team trust the calibration.

Questions

What are incrementality tests in marketing?

Experiments that measure the extra sales a channel produces by changing its spend for one group, such as a set of regions, and comparing the outcome with a similar group that did not receive the change.

Is incrementality testing better than marketing mix modeling?

They answer different questions. A test measures one channel directly inside a window. A model estimates every channel continuously. Most measurement programmes use tests to calibrate the model.

How long does an incrementality test need to run?

Across 123 of our own geo experiments, 27% of tests under two weeks reached significance, 11 tests in that band, against 71% of tests running four to six weeks, 34 tests. Past six weeks the rate stops improving, and spend level did not separate the readable tests from the unreadable ones.

Do you need to test every channel?

No. A measured effect on one channel becomes a constraint the model has to satisfy, which improves the estimates for the rest of the mix. A few well-chosen tests calibrate the whole model.

Can a test replace a marketing mix model?

For one channel and one decision, often yes. For an allocation across the whole mix, no, because sequential tests take longer than a planning cycle and do not show how channels interact.