Hierarchical marketing mix modeling: what it is, how partial pooling works, and when you need it
A hierarchical marketing mix model estimates effects for several markets, regions or product lines in one model and links them statistically, so that units with little data borrow strength from the group. The technique is called partial pooling. The alternatives are one combined model for everything, or a separate model for each unit.
The choice matters most for businesses that sell in many markets or across many product lines, where some units have thin data. A combined model hides the differences between markets. Separate models can struggle where a market has little history. A hierarchical model sits between the two. This page explains the three approaches, works through an example of partial pooling, compares the approaches on five dimensions, and sets out when it helps, where it can go wrong, and when separate models are the better choice.
Three ways to model many markets
Complete pooling. One model for all markets combined. It uses all the data at once, so it is stable, and it assumes every market responds the same way. Real differences between markets disappear into one national figure.
No pooling. A separate model for each market. Each market keeps its own estimates, and each model has only that market's data. Small markets, or channels with little spend in a market, can return unstable or implausible results.
Partial pooling. One hierarchical model in which each market has its own estimates, and those estimates are linked through a shared group-level distribution. Markets with plenty of data are driven mostly by their own data. Markets with little data are pulled toward the group.
The same logic applies to product lines, stores or customer segments, and a hierarchy can have several levels, for example regions within countries.
Modeling at a finer level also adds data. A national model with three years of weekly data has about 150 observations. The same period across twenty regions gives far more variation to learn from, provided the regional data is clean and spend actually varied between regions.
Partial pooling, explained
Partial pooling shrinks each unit's estimate toward the group in proportion to how little its own data supports it. The model learns two things at once: what each market's effect looks like, and how much markets tend to differ from each other.
An illustrative example: a retailer runs television in twelve regions. Ten have three years of steady data, and two opened last year. In separate models, the two new regions return wide or odd estimates. In a hierarchical model, their estimates start near what the other ten suggest, and move away from that as their own data accumulates.
Technically, the group-level distribution has its own parameters, often called hyperpriors, which describe the average effect across markets and how widely markets spread around it. Both are estimated along with each market's effect.
If markets turn out to be similar, the model pools strongly. If they differ a lot, it pools weakly and behaves more like separate models. The amount of pooling is estimated from the data.
The differences, dimension by dimension
Handling thin data. Separate models handle it poorly. A combined model hides it. A hierarchical model borrows from the group.
Local differences. Separate models keep them fully. A combined model loses them. A hierarchical model keeps them where the data supports them.
Complexity. Separate models are simple to build and explain. A hierarchical model is harder to specify, slower to fit and harder to diagnose.
Assumptions. A hierarchical model assumes the markets belong to a common group. If one market is genuinely different, pooling can pull it toward a group it does not resemble.
Framework. Partial pooling is native to Bayesian inference, where the group-level distribution acts as a prior for each unit. Frequentist equivalents exist as mixed-effects models.
How much history a channel needs
A common belief is that small channels come back empty because their budget share is too small. Our own fleet points somewhere else.
Across 138 advertisers, 881 models, a channel with two years of history returned a readable result about 97% of the time whatever its share of the budget, and 98.7% of the time for channels taking under 2%. These models fit each market separately. With enough history behind a channel, its share of budget did not decide whether it could be read.
In practice, that makes history the first thing to check. Where a market or channel has a long, varied history, a separate model can read it. Where it has little, pooling or an experiment is what adds the missing information.
Where partial pooling can go wrong
Pooling helps when the units really are alike. It can hurt when they are not.
If one market has different competitors, a different price position or a different channel mix, pooling pulls its estimates toward markets it does not resemble. The result looks stable and is wrong for that market.
The grouping itself is a modeling choice. Grouping by country, by region or by product type gives different answers, and the choice should follow how the business actually varies. Agreeing it with the people who know the markets is a sensible safeguard.
Two checks catch most problems. Compare each unit's pooled estimate with what a separate model says, and look closely at any unit the hierarchy has moved a long way. And check how much the model is pooling overall: very strong pooling across markets that the business knows to be different is a warning sign.
How to tell whether you need one
Three questions settle most cases.
How many units, and how much data each? A few large markets with long histories rarely need pooling. Dozens of regions or product lines, several with short or sparse histories, often do.
Do the units behave alike? If markets share channels, pricing and customer behaviour, pooling is reasonable. If they differ substantially, separate models or a looser hierarchy are safer.
What decision is the model for? A national budget split may need only one combined model. Allocating spend between regions needs estimates for each region, which is where hierarchy earns its complexity.
When a hierarchical model is the better choice
A hierarchical model is the better choice when a business has many markets, regions or product lines, when several of them have thin data, when the units behave broadly alike, and when decisions need estimates for each unit. It is also useful for geo-level models, where regional data adds variation that a national model cannot use.
When separate models are the better choice
Separate models are the better choice when each market has a long and varied history, when markets differ in ways pooling would blur, and when simplicity and speed matter. They are easier to explain, faster to fit and easier to check, and one market's data problems stay in that market.
They also suit businesses whose markets are managed by separate teams with separate budgets, where each team needs a model it can own and explain.
Where a single market or channel is short of history, an experiment is often a more direct fix than pooling: it measures that channel's effect instead of borrowing an estimate from other markets.
What neither approach fixes
Neither approach makes a model causal. Both learn from historical variation, so both can mistake a channel that moved with demand for one that drove it. Experiments supply that separation.
Neither creates data. Pooling borrows information from other units, and that information is only as good as the similarity between them.
And a hierarchical model is harder to maintain. It needs someone who can specify, fit and diagnose it, and a model nobody can maintain loses its value quickly. Fitting times also grow with the number of units, which matters for teams that refit often.
How to choose
A few markets, long histories, different behaviour: separate models.
Many markets or product lines, thin data in several, similar behaviour: a hierarchical model.
One channel or market short of history: run an experiment on it, whichever structure you use.
Either way: check how much history each channel has before blaming the budget, and agree the grouping with the people who know the markets.
Cassandra, the marketing mix modeling platform, runs a Bayesian production engine, and Bayesian inference is the framework partial pooling is native to. The engine does not pool, though. It fits one model per market rather than pooling markets into a single hierarchy, so a multi-market business runs one project per market.
What it does for a thin channel is bring in evidence from outside the model. Geo experiments run in the same system, using the same method family as GeoLift, and a measured result constrains the model. Reads sit at campaign level, in the planning layer.
A price is published.
Moving to a hierarchical structure
The main work is data preparation. A hierarchical model needs every unit's data in one consistent structure: the same weeks, the same channel definitions and the same outcome measure in every market. Aligning those is usually most of the effort.
Start with the grouping. Decide which units belong together, and why, before fitting anything. Then compare the hierarchical results with the separate models already in use, market by market, and investigate any market where the two differ sharply. That is where pooling is either helping most or blurring a real difference.
Keep the separate models running until the comparison is understood, and document the grouping and the reasons for it so a later team can revisit them.
Questions
What is a hierarchical mixed effect model?
A model in which each group, such as a market or product line, has its own parameters, and those parameters are linked through a shared distribution. In marketing mix modeling this is how partial pooling is implemented.
Can you provide an example of a hierarchical model?
A retailer modeling television across twelve regions in one model, where each region has its own television effect and new regions with little data borrow from the others until their own history builds up.
Why do small channels come back empty in a marketing mix model?
Usually because the channel has too little history or too little variation in its spend. In our fleet, channels with two years of history read about 97% of the time whatever their share of budget.
Is a hierarchical model always better?
No. It helps when units are similar and some have thin data. When markets differ substantially or each has a long history, separate models are simpler and can be more accurate.
Does a hierarchical model need Bayesian methods?
Partial pooling is most natural in Bayesian inference, where the group-level distribution acts as a prior. Frequentist mixed-effects models offer a similar structure.