A platform dashboard can tell you how many conversions it attributed to a campaign. It cannot, on its own, tell you how many of those conversions would have happened anyway, from people who were already going to buy, search for your brand, or sign up regardless of whether they saw the ad. Incrementality testing is the discipline of estimating that difference directly, by comparing an exposed group against a comparable unexposed group, rather than relying on attribution rules alone.
Attribution versus causality
Attribution is a bookkeeping exercise: a platform or analytics tool applies a rule (last click, first click, data-driven, or another model) to decide which touchpoint gets credit for a conversion. Those rules are useful for operational reporting, but they do not involve a controlled comparison, so they cannot prove that the marketing activity caused the conversion rather than simply being present near it. Incrementality testing, by contrast, asks a causal question directly: if we had not run this activity, would this outcome still have happened? Answering that requires a comparison group that did not receive the activity, as similar as possible to the group that did.
Two feasible approaches: randomized and geo holdouts
Randomized holdouts
A randomized holdout (sometimes called a ghost ad or PSA-style holdout depending on the platform) withholds a specific ad or campaign from a randomly selected portion of your eligible audience, while everyone else is eligible to see it. Because the holdout group is chosen randomly from the same pool, it should be similar to the exposed group on average, which makes the comparison more trustworthy. Some ad platforms offer built-in conversion lift or holdout tools; check current official documentation for the specific platform you are using, since availability, minimum spend, and sample size requirements vary and change.
Geo holdouts
A geo holdout pauses or reduces marketing activity in a defined set of geographic regions while running it normally elsewhere, then compares the change in business outcomes between the two sets of regions over the test period. This approach is often more practical for smaller businesses or channels without a built-in randomization tool, but it depends on finding regions that are genuinely comparable in size, seasonality, and existing demand, which takes care to set up well.
| Factor | Randomized holdout | Geo holdout |
|---|---|---|
| Comparability of groups | Usually strong, since assignment is random | Depends on choosing genuinely similar regions |
| Feasibility for small businesses | Depends on platform tool availability | Often more accessible, no special platform tool required |
| Main risk | Contamination if holdout users are reached another way | Confounding regional events (weather, local promotions, competitor activity) |
| Typical output | Lift percentage with a confidence interval | Difference-in-differences estimate with a stated margin of uncertainty |
Feasibility and contamination
Before committing to a test, assess whether you can realistically maintain a clean holdout. Contamination happens when people assigned to the 'unexposed' group are reached anyway, for example through another channel targeting the same audience, through organic search for your brand, or through a geo holdout region that overlaps with a national campaign you forgot was still running elsewhere. Contamination tends to understate true lift, because it makes the two groups look more similar than they actually are. Also assess statistical feasibility: a very small audience or a very low baseline conversion rate may not generate enough data in a reasonable test window to produce a trustworthy estimate, which is a real limitation worth stating rather than hiding.
Defining outcomes and reporting uncertainty
Decide in advance what outcome you are measuring, such as purchases, qualified leads, or signups, and over what window, since measuring too short a window can miss delayed conversions and measuring too long a window invites other factors to contaminate the comparison. Report your result as a range reflecting the uncertainty in the test, for example 'this campaign likely drove between 8% and 18% incremental lift at a 90% confidence level,' rather than a single precise number presented as settled fact. Smaller sample sizes produce wider ranges; that is a feature of honest reporting, not a flaw in the test.
A test-design worksheet
The deliverable for this lesson is a worksheet to complete before launching any incrementality test.
- Define the specific business outcome being tested (purchases, qualified leads, signups) and its measurement window
- Choose randomized holdout or geo holdout based on available tools, audience size, and feasibility
- Identify and document contamination risks specific to this test (overlapping channels, overlapping regions)
- Estimate whether your audience or region size is large enough to produce a meaningful result in the test window
- Decide the confidence level and acceptable range width before seeing any results, to avoid moving the goalposts
- Plan to interpret lift alongside cost and contribution margin, not platform-reported ROAS alone
- Document assumptions and limitations in the final report, including anything that could have confounded the result
Common mistakes
- Treating platform attribution numbers as proof of causal impact without ever running a holdout.
- Choosing a holdout region that differs meaningfully in size, seasonality, or existing demand from the active regions.
- Reporting a single precise lift number without disclosing the uncertainty range behind it.
- Ending a test early because early results look favorable, before the planned sample size is reached.
- Comparing lift directly to platform ROAS instead of to the actual cost and contribution margin of the campaign.
When this is not the right tactic
Incrementality testing is not worth the setup cost for very small campaigns or very low-value test questions, where the business impact of being wrong is minor. It is also impractical when you cannot realistically isolate a comparison group, for example a single-location local business with no meaningful way to split audiences or geographies. In those cases, simpler before-and-after comparisons, read with appropriate caution about their limitations, may be the more proportionate approach, with incrementality testing reserved for higher-stakes budget decisions.
Where to go next
Read the marketing mix modeling lesson to understand a complementary, time-series approach to estimating channel contribution when you have enough historical data, and the marginal economics lesson to see how lift estimates should feed into scaling decisions rather than average ROAS alone.



