Imagine pausing your biggest ad campaign for a month and finding out revenue barely moved. Now imagine the opposite: pausing a campaign you thought was mediocre and watching sales fall off a cliff.
Both happen more often than most marketers would like to admit, and they reveal the central problem with modern ad measurement: platform-reported ROAS (return on ad spend) tells you what your ads touched, not what they caused. In an era where AI-driven campaigns like Performance Max and Advantage+ decide their own targeting and grade their own homework, that distinction has never mattered more.
Incrementality testing is how you close the gap. It is the practice of measuring what your advertising actually adds, the revenue that would not have happened without it. Here is how we think about it, and how to run your first test without a data science team.
Why Platform ROAS Overstates Reality
Ad platforms attribute a conversion to themselves whenever they can claim a touchpoint. Someone sees your retargeting ad after already deciding to buy? Attributed. Someone searches your brand name and clicks the ad above the organic listing they would have clicked anyway? Attributed.
None of this is fraud. It is just attribution doing what attribution does: assigning credit for observed conversions, with no ability to know what would have happened in the ad’s absence.
The problem compounds with automation. AI campaign types are explicitly designed to find the cheapest conversions available, and the cheapest conversions are usually people who were close to buying anyway. Your existing customers, your branded searchers, your abandoned-cart warm audiences. The result is reported ROAS that can be several multiples higher than true incremental return.
Incrementality, Defined Simply
Incrementality asks one question: of the conversions my ads get credit for, how many would have happened anyway?
The answer comes from a controlled comparison. You expose one group to advertising and withhold it from a comparable group, then measure the difference in outcomes. The difference is your incremental lift, the only number that tells you whether an ad dollar created revenue or just took credit for it.
Three Ways to Test, From Simplest to Most Rigorous
1. The Pause Test
The bluntest instrument: turn a campaign or channel off for two to four weeks and watch total revenue, not platform-attributed revenue. If you pause a campaign reporting 20% of your revenue and total sales dip 3%, you have learned something important.
Pause tests are free and simple, but noisy. Seasonality, promotions, and general variance can muddy the read. They work best for big, binary questions: does this channel matter at all?
2. Geo Holdout Testing
The workhorse for ecommerce brands. Split your market into matched groups of regions, keep ads running in one group, reduce or pause them in the other, and compare revenue between the groups over the test window.
Because both groups experience the same seasonality and promotions, geo tests isolate the advertising effect far better than a simple pause. They require enough geographic sales volume to read a signal, and a test window of typically four to six weeks, but no user-level data at all, which also makes them privacy-proof.
3. Platform Lift Studies and Always-On Holdouts
Meta and Google both offer conversion lift studies that randomize at the user level, exposing some users and holding out others. These are statistically clean and worth running when your spend qualifies. The tradeoff: the platform still administers the test, so we treat these as one input rather than the final word, and pair them with our own geo tests.
How Often Should You Test?
Our recommendation matches what leading measurement practitioners advise: run full-channel holdouts on your highest-spend platforms at least twice a year, ideally quarterly. Signals drift. Creative changes, competition shifts, and an incrementality number from eighteen months ago describes a campaign that no longer exists.
The cadence matters double if you feed the results into media mix modeling. MMM is only as strong as the real-world lift signals used to calibrate it.
What to Do With the Results
Incrementality results should change budgets. That is the entire point. A few patterns we see repeatedly:
- Branded search often shows low incrementality. Test it, and if lift is minimal, that budget usually works harder in prospecting.
- Retargeting is usually less incremental than reported. Warm audiences convert with or without a fifth reminder ad. Cap frequency and reallocate.
- Upper-funnel spend is usually more incremental than reported. Awareness campaigns look weak in click-based attribution but frequently show strong lift in holdout tests. This is where attribution systematically underinvests.
- Recalibrate your targets. If a channel’s true incremental ROAS is half its reported ROAS, adjust the target you manage it to, not just your expectations.
Common Mistakes That Ruin Tests
Keep an eye out for these: running tests during promotions or peak seasonality unless the test is specifically about that period, calling results early because the first week looked dramatic, testing regions that are too small to produce signal, and changing creative or offers mid-test. A clean test changes one variable: ad exposure.
Measurement Is a Competitive Advantage
Most of your competitors are still making budget decisions on platform-reported ROAS. Every distortion in those numbers, and there are many, leads them to overspend on low-lift tactics and underspend on what actually grows the business. A brand that knows its true incremental returns is playing a different game.
Frequently Asked Questions
What is the difference between ROAS and incrementality?
ROAS measures revenue attributed to ads based on touchpoints. Incrementality measures revenue that would not have happened without the ads. A campaign can report high ROAS while adding little incremental revenue if it mostly reaches people who would have bought anyway.
How long should an incrementality test run?
Geo holdout tests typically need four to six weeks: enough time to accumulate statistical signal and to let delayed conversions play out. Shorter windows produce noisy reads, especially for products with longer consideration cycles.
How much budget do I need to run incrementality tests?
Less than most brands assume. A pause test costs nothing. A geo holdout test costs only the measurement effort plus the revenue risk in held-out regions, which is typically a small fraction of monthly spend. The bigger requirement is sales volume: you need enough geographic transaction density to read a difference between test and control groups.
Want to know what your ad spend is actually driving? Book a strategy call with the Strat88 team and we will design an incrementality testing plan for your biggest channels.