Analytics
Incrementality testing with the tools you already have
Geo holdouts, audience holdouts and pause tests you can run inside Google Ads, Meta and your email platform, with the sizing maths before you spend.
By CartKernel · Published
An incrementality test answers one question: how much of this revenue would have arrived anyway. No attribution model can answer it, because attribution divides credit among the touchpoints that happened, and incrementality needs the world where the touchpoint did not happen. You do not need a vendor lift study or a beta programme to get that world. You need a group of customers or a set of regions that you deliberately stop advertising to, a period long enough to see the effect, and the discipline not to peek and stop early.
This is the practical version. Three designs, the sizing arithmetic, and how to read a result without fooling yourself.
What a holdout actually buys you
Every reported conversion belongs to one of two populations. Some buyers were going to purchase regardless, and the ad simply appeared on the path. Others bought because the ad did something: reminded them, introduced the brand, closed a comparison. Reported return on ad spend counts both. Incremental return counts only the second group.
The gap is largest where targeting is most precise, because precise targeting finds people who were already close to buying. Branded search, remarketing lists and abandoned cart audiences all sit in that zone. That is not a fault in the channel, it is a measurement problem, and incrementality exists as a concept because averages across a whole account hide it.
Design one: a geographic holdout
Best for anything you cannot switch off per person: Performance Max, broad prospecting, television, out of home, influencer pushes.
- Split your market into regions the platform can target and exclude. In Canada that is usually provinces or metropolitan areas; in the United States, designated market areas or states.
- Rank regions by revenue over the last twelve months and pair them so each pair has similar revenue, seasonality and product mix. Assign one of each pair to test and one to control, alternating.
- Turn the campaign off completely in the control regions. Partial reductions are hard to read.
- Hold the change for at least four weeks, longer if your purchase consideration cycle is long.
- Measure total store revenue per region from platform order data, not from ad platform reporting. The whole point is to see what happened outside the ad account.
The read is the difference in revenue between test and control regions across the period, indexed to the same regions during a matched pre-period. Where you want a confidence interval rather than a difference of two numbers, Google’s open-source CausalImpact package builds a counterfactual for the control series and reports the effect with credible intervals.
Two cautions. Holding out a region means giving up its revenue for the duration, so start with a control group that is around ten to twenty percent of the market. And regions leak: shoppers travel, and shipping addresses do not always match where the ad was seen. Leakage biases the result toward finding no effect, which makes a positive result more trustworthy and a null result less conclusive.
Design two: an audience holdout
Best for remarketing, customer match lists, email and SMS, where the platform can exclude a specific set of people.
- Email and SMS. Most platforms let you build a randomly sampled segment. Create a permanent holdout of five to ten percent of the list, excluded from every campaign and flow for a quarter, then compare revenue per subscriber against the sent group. This is the cleanest test any store can run, and it costs almost nothing.
- Paid remarketing. Split your remarketing audience by user ID hash or by a random custom audience, exclude one half in the campaign, and compare purchase rates. Platform-native experiment tools do this for you where available.
- Abandoned cart flows. Hold out a share of abandoners from the flow entirely. Recovered revenue reported by the flow always includes people who were returning to the cart anyway. The holdout tells you how many.
Keep the holdout stable for the whole test. Rotating members in and out contaminates both groups.
Design three: the structured pause
Best when a full experiment is not practical and you still need a signal. Pause the campaign, keep everything else constant, and read the change in total orders against a forecast built from the previous eight to twelve weeks.
This is the weakest design and it is still better than reading the platform’s own conversion column. It is only credible when nothing else changes: no promotion, no product launch, no seasonal turn, no competing campaign scaling in the same week. Write down the confounders you know about before you start, and abandon the read honestly if one of them fires.
A structured pause is the usual first test for branded search, which is where the argument between marketing and finance tends to start. Whether an ecommerce brand should bid on its own name sets out the factors that decide it, and the pause test is how you settle it for your store rather than in general.
Size the test before you run it
Most incrementality tests fail because they were never capable of detecting the effect they were looking for. Do this arithmetic first.
You need a baseline conversion count, an effect size you care about, and a duration. Say the control group generates 400 orders in a typical four-week window. Detecting a five percent difference in a count that size is beyond reach; the natural week-to-week variation is larger than the effect. Detecting a twenty percent difference is plausible. If the smallest effect worth acting on is smaller than the smallest effect the test can see, either extend the duration, enlarge the groups, or accept that this channel cannot be measured this way at your current volume.
A useful rule of thumb: the noise in an order count scales with the square root of the count, so quadrupling the test period halves the relative noise. A store doing a few hundred orders a month should plan for six to eight week tests and should test only the channels large enough to matter.
Reading the result without deceiving yourself
- Pre-register the decision. Write down, before the test starts, what result leads to what action. “If incremental return is below break-even we cut the budget by half” is a decision. “We will see what the data says” is not.
- Do not stop early. Checking daily and stopping when the numbers look good is the fastest way to manufacture a false positive.
- Use one revenue source. Store order data for both groups. Mixing platform-reported conversions for the test group with store revenue for the control group produces nonsense.
- Convert the lift into an incremental cost per acquisition. Spend divided by incremental orders, not reported orders. Compare that against your customer acquisition cost target, and against the first-order contribution margin.
- Expect a range, not a point. The honest output is “somewhere between a modest and a substantial effect”, and that is enough to act on when the whole range sits above or below break-even.
What to test, in order
- Branded search. The largest measured gap in most accounts, and the easiest pause test.
- Remarketing. High reported return, audiences that were already shopping. Ecommerce remarketing is worth running; the question is how much of it and at what bid.
- Email and SMS flows. The permanent holdout should be standing policy, not a one-off project.
- Prospecting and Performance Max. Geo holdout, longest duration, largest budget at stake. Read it alongside performance max spending on brand, because a campaign eating branded queries reports the same inflation a branded campaign does.
Where incrementality sits in the reporting stack
It does not replace daily reporting. It calibrates it. Run a test, learn that a channel’s true contribution is some fraction of its reported contribution, and apply that fraction as a standing adjustment until the next test. Meanwhile, watch total marketing efficiency at the account level, because the sum of channel-level reported returns will always exceed the store’s actual performance. The marketing efficiency ratio is the check that catches drift between tests, and the MER calculator will give you the current figure from total revenue and total spend.
Re-test each major channel once or twice a year, and after any structural change: a new campaign type, a large budget shift, a repositioning. Incrementality is a property of the current mix, not a constant. When a test result contradicts the dashboard, the test wins, and attribution reporting should be adjusted to reflect it rather than the other way round.