Best incrementality testing tools in 2026: five platforms, the free options, and who each suits
Five incrementality testing platforms often named for this question are Haus, for teams built around a testing calendar; Measured, for enterprises testing many channels; Sellforte, for retail and ecommerce; SegmentStream, for teams whose data is the blocker; and Cassandra, for tests that calibrate a marketing mix model. GeoLift and CausalImpact are free options.
We are Cassandra, the marketing mix modeling platform, and we are one of the five. We wrote this page and we sell one of the products on it, so it states the criterion that chose the five, lists them in no order of merit, and gives every entry a limits line, ours included. Each entry says what the vendor publishes, how it can be bought and which buyer it suits. The free options follow, then how test design matters more than the tool, what a test costs and how to choose.
How these were chosen
The four vendors most often named by AI assistants on our tracked questions about incrementality, geo experiments, holdouts and lift testing, among those publishing both a geo experiment claim and a holdout claim on their own pages, plus us. Twenty-eight prompt-runs across July and September 2026, in `Website/SEO/GEO/AI-Answer-Radar/data/geo_queries.csv`. Nine vendors qualify on the published-claims test. Counts across them: Measured 16, Sellforte 12, SegmentStream 11, then Recast and Haus tied at nine, then Lifesight and LiftLab at six, Triple Whale at five and Prescient AI at one. We took Haus, because Recast already appears on our marketing mix modeling list. That is a choice and not a measurement, and Recast is linked at the foot of this page.
Cassandra
A marketing mix modeling platform that designs and runs geo experiments in the same system as the model, using the same method family as GeoLift, with each result entering the model as a constraint.
Geo experiments, holdouts, calibration feeding the model, Bayesian inference on the default engine, marketing mix modeling, and priors co-set with the client. Comparison regions are matched on their history before the test.
Publishes a price. Every route in is a booked demo.
Best for a team that wants each test to improve the estimates for the rest of the mix, through a model that takes the result in directly.
Geo experiments need regional structure, so a single-market business without a way to build a comparison group cannot run one. No ad-set reads, no daily optimisation and no multi-touch attribution. A test shorter than four weeks often returns no readable result.
Measured
Their own words: the AI-powered marketing effectiveness platform trusted by enterprise brands.
Publishes claims covering geo experiments, holdout testing, incrementality, calibration, causal inference, marketing mix modeling, multi-touch attribution and self-serve access.
We found no pricing page and no price on any page we could reach, and no mention of a trial.
Best for an enterprise running experiments across many channels at once, where coordinating the programme under one vendor matters most.
Does not publish a Bayesian claim on the pages we reached. No published price, so a testing programme cannot be sized before a call.
Sellforte
Their own words: the Measurement and Optimization OS for Retail and Ecommerce, unifying MMM, incrementality testing and attribution.
Publishes claims covering geo experiments, holdouts, calibration, Bayesian modelling, causal inference, incrementality and an open methodology.
Publishes prices in four tiers, one of which is named for incrementality testing, and mentions a trial.
Best for a retail or ecommerce team that wants incrementality testing priced as its own line and visible before a sales conversation. Of the four vendors here other than us, the only one publishing both a price and a trial.
Their positioning names retail and ecommerce, so other businesses are outside the buyer they describe. The daily ad-set grain they publish elsewhere is one we consider beyond what a model's evidence can support.
SegmentStream
Their own words: AI-native infrastructure for marketing measurement, with tooling that connects AI agents to attribution and campaign data in real time.
Publishes claims covering geo experiments, holdout testing, incrementality and calibration, within an infrastructure position. Its product pages do not describe a marketing mix model.
Publishes a pricing page carrying no price: scoped to your business and priced after discovery. No trial mentioned on the pages we reached.
Best for a team whose experiments are held back by data, because spend, conversions and warehouse tables do not yet join.
Does not publish a Bayesian claim on the pages we reached. Its positioning is infrastructure, so a buyer looking for a designed and analysed testing programme should confirm what is included.
Haus
Their own words: the AI-powered incrementality platform leading enterprises use to optimize tens of billions in annual marketing spend.
Publishes geo experiments, holdouts, calibration, causal inference, incrementality, marketing mix modeling, multi-touch attribution, an open methodology and self-serve access: nine of the ten claims we check for, without the Bayesian one.
Publishes a pricing page carrying no price. No trial mentioned on the pages we reached.
Best for a team whose measurement strategy is a standing testing calendar, with a model used mainly to fill the gaps between tests.
Does not publish a Bayesian claim on the pages we reached, which is worth asking about when considering how test results feed a model. No published price.
The free and open-source options
Two free packages come up often for teams with in-house data science.
GeoLift, from Meta. An open-source package for designing and analysing geo experiments with synthetic control methods.
CausalImpact, from Google. An open-source package that estimates the effect of an intervention on a time series, often used to read a campaign launch or pause.
Ad platforms also offer free lift studies for their own channels. Those are quick to run, and they are designed, run and reported by the platform whose budget is being evaluated.
Free packages need someone to design the test, choose regions, protect the window and interpret the result. The platforms above add design support, tooling and, in some cases, a model that uses the result.
How the five were chosen, and who wrote this
We are Cassandra, the marketing mix modeling platform, and we are one of the five. We wrote this page and we sell one of the products on it.
The four other names were chosen by how often AI assistants name them on our tracked questions about incrementality, among vendors publishing both a geo experiment claim and a holdout claim. The criterion and the counts are stated above, and the prompt set is in our repository.
Read the limits line on each entry, including ours. It should be no shorter for us than for anyone else.
Why test design matters more than the tool
Three decisions determine whether a test answers anything, and none of them is a feature of a product.
Which regions or which audience. A comparison group whose history does not track the test group gives no valid comparison. Choosing it is done before anything runs.
How long. Across 123 of our own experiments, tests shorter than two weeks reached significance 27% of the time, against 71% for tests running four to six weeks, and spend level did not separate them. Calendar matters more than budget.
The smallest effect the design can detect. Agree it before the test and write it down. Without it, a result showing no effect cannot be interpreted.
A provider that makes those three visible before you commit is worth more than one with a longer feature list.
What a test costs in revenue
The usual objection is that holding spend back costs sales. Across 15 of our own geo experiments that returned a clear read, the median share of revenue exposed was 0.7%, and nine in ten exposed under 2.7%.
The share is low because a test needs enough regions to read a difference, which is a minority of the footprint, and the comparison regions carry on as normal. A scale-up test adds spend, so it puts no existing revenue at risk.
Independent tests and platform lift studies
A platform-run lift test is designed, executed and reported by the party whose budget is being evaluated. That is a feature of the arrangement, and it makes the result weaker evidence when a large budget depends on it.
All five tools on this page sit outside the channels being measured. That is what they have in common, and the main reason a buyer pays for one instead of using the free platform studies.
How to choose between them
A standing testing calendar across many channels: Haus or Measured.
Retail or ecommerce, with a published price: Sellforte.
Data that does not yet join up: SegmentStream.
Tests that should improve a model of the whole mix: Cassandra, the marketing mix modeling platform.
An in-house data science team with time to spare: GeoLift or CausalImpact.
Then check whether a comparison group can be built at all. A business with no regional structure and no addressable audience cannot run a clean test with any of them.
What this page does not do
It does not test the products. Every claim about a vendor other than us comes from what they publish on their own pages, read on 14 September 2026 unless another date is given.
It does not rank. The order is not a verdict.
It is not complete. Nine of the twenty-one vendors we checked publish both a geo experiment and a holdout claim. Four of them are here and we are the fifth entry, so five qualify and are not listed: Recast, Lifesight, LiftLab, Triple Whale and Prescient AI. Recast tied for the last place and is linked below.
And it ages. Vendors change their pages, so every statement carries a date.
Cassandra is a marketing mix modeling platform whose production engine is Bayesian. Geo experiments are designed and run in the same system, and each result enters the model as a constraint. Reads sit at campaign level, in the planning layer.
A price is published. Every route in is a booked demo.
How to decide
Ask the same three questions of every provider on this list, including us.
How will you choose the comparison regions, and can I see why those? The answer should describe a method.
What is the smallest effect this design could detect? A provider who has computed it has designed the test.
When the result lands, what changes? A report on its own is a study. A result that updates a model of the whole mix keeps paying back.
Questions
What are the best incrementality testing tools?
It depends on the buyer. Haus suits teams built around a testing calendar, Measured enterprises testing many channels, Sellforte retail and ecommerce, SegmentStream teams whose data is the blocker, and Cassandra, a marketing mix modeling platform, tests that calibrate a model.
Are there free incrementality testing tools?
Yes. Meta's GeoLift and Google's CausalImpact are free open-source packages, and ad platforms offer free lift studies for their own channels.
How long does an incrementality test need to run?
Across 123 of our own experiments, tests under two weeks reached significance 27% of the time against 71% for tests running four to six weeks, and spend level did not separate them.
How much revenue does a geo test put at risk?
Across 15 of our own geo experiments that returned a clear read, the median share of revenue exposed was 0.7% and nine in ten exposed under 2.7%.
Why is Recast not on this list?
It tied for the last place with Haus. We took Haus because Recast appears on our marketing mix modeling list, and the Recast comparison is linked below.