While building a marketing mix model for a subscription app client, we were asking: how much revenue does Google Ads actually bring in? Google’s own reports looked good, but we know not to blindly trust them. So we ran a geo holdout test.
We paused Google App Campaigns (UAC) in 28 randomly chosen US states for six weeks and kept them running in the other 23. We couldn’t detect a meaningful drop in first-purchase revenue. The test was sensitive enough to catch any lift of 7% or more, so at this spend level, Google wasn’t driving a big share of revenue.
Here is how we ran the test, what it showed, and what we decided.
Why can’t platform ROAS tell you if Google Ads works?
Platform ROAS counts every purchase the ads touched. App Campaigns are built to find the people most likely to buy, and many of them would have bought anyway. So the better the targeting, the better the reported ROAS looks, whether or not the ads changed anything.
In other words, platform ROAS answers “which conversions did Google touch?”, not “which conversions would have disappeared without Google?” Only an incrementality test answers the second question, and that is the one budget decisions should depend on.
How we designed the geo holdout test
Google was paused in 28 states and kept running in 23. We randomized by US state, and included all 51 (the 50 states plus DC). Revenue varies a lot between states, so we used stratified random assignment: states were grouped by revenue level, then assigned at random within each group, so both sides got a similar mix of big and small markets.
Before starting, we checked that both groups had been following similar revenue trends over the five months before the test. They had, so any gap that opened up during the test could be read as the effect of the pause.
| Unit | US state, stratified random assignment |
|---|---|
| Groups | 28 states paused, 23 states running |
| Baseline | About five months before the test |
| Test | Six weeks with Google App Campaigns fully paused in the paused states |
| After | Ads back on everywhere, used as a second check |
| Outcome | Daily first-purchase revenue per state (lifetime and subscription) |
What did the holdout test show?
Both groups declined by a similar amount.
During the six weeks, revenue fell in both groups, by around a fifth. Something market-wide or seasonal pulled everything down at once, and when the test ended, both groups recovered together.
The paused states fell slightly further. That small extra drop is the estimated effect of pausing Google: roughly 3–4% of first-purchase revenue, depending on how you measure it. It was not statistically significant (p = 0.50). A gap this size would show up about half the time when you split the states at random with no ad pause at all.
What does a null result actually prove?
A null result is only worth something if the test could have caught a real effect, and ours was designed to do so. The smallest effect it could reliably detect was about 7% of baseline revenue (at 80% power and the usual 5% significance level). If Google had been adding 7% or more to first-purchase revenue, we would very likely have seen it. We didn’t.
So this doesn’t prove Google’s impact is zero, but it is good evidence that Google wasn’t driving a big revenue lift at this spend level.
It is also one experiment. Our power analysis says that if Google had been adding at least 7%, a test like this would catch it about 80% of the time, which still leaves a chance that we missed a real effect. Repeating the test would give stronger evidence, though some chance is always involved.
Did Google reach break-even ROAS?
What makes the result usable is the break-even math. We took the most generous reading the test allows, a true lift of exactly 7%, and set the revenue it represents against the ad spend saved in the paused states. Even at that upper bound, Google’s incremental ROAS stayed below break-even on first purchases. For the ads to break even, the true lift would have to be well above anything the test left room for.
That is the part worth copying: work out your break-even lift before the test, and make sure the test is big enough to detect it. Then “not significant” stops being a shrug and becomes an answer.
How do you avoid overstating a geo test result?
A geo test can be set up well and still give a misleading answer if it’s analyzed the wrong way. Here is the trap. The data had thousands of rows: 51 states, each with a revenue number for every day. Run a standard regression on all of them and you treat every state-day as an independent piece of evidence. It isn’t. A state’s revenue on Tuesday is mostly the same state behaving the same way as on Monday. Counting both as separate evidence overstates how much you know, and makes the result look far more certain than it is.
We randomized at state level, so we analyzed at state level, with a reshuffling test (randomization inference):
- Calculate the real difference-in-differences: the change in average daily revenue from before to during the test, paused states minus running states.
- Shuffle which 28 states are labeled “paused”, keeping every revenue number where it is, and recalculate. Repeat 2,000 times.
- See where the real result falls among the 2,000 made-up ones.
If the real result sits comfortably inside the range randomness produces on its own, you can’t tell it apart from noise. Ours sat right in the middle. The method makes no assumptions about how revenue is distributed.
Small methodological detail. Big impact on the conclusion.
A second check: did the paused states catch up when ads came back?
If pausing Google really caused revenue to drop, the paused states should fall behind the others while ads are off, then catch up once ads are back on. A seasonal coincidence has no reason to follow that schedule. So we ran the same reshuffling test, comparing the test period with the weeks after ads resumed. Again, the change was well within what chance alone produces.
| Test | Compares | p-value |
|---|---|---|
| Reshuffling test | Test period vs before | 0.50 |
| Reshuffling test | Test period vs after ads came back | 0.56 |
| State-level paired t-test | Test period vs before | 0.68 |
All three agree: the difference is well within what random assignment alone produces. Significant would mean below 0.05.
What did we decide?
We reduced Google spend. We did not switch it off. “No detectable lift” isn’t the same as “no lift”, and the test has limits:
- First purchases only. Renewals and longer-term value weren’t measured.
- One window. It ran once, in spring. Advertising can work differently in other seasons.
- 51 units. That is a modest sample for this kind of test. The reshuffling method is the most defensible option, but it doesn’t make the sample bigger.
So we cut Google spend rather than turning it off.
One more detail: the account had one of its best reported Android ROAS periods while this test ran. Platform ROAS and incrementality tell different stories, and it’s worth looking at both.
Lessons for your own incrementality test
- Platform ROAS shows what the platform touched, not what it caused. We saw the same thing when we tested brand traffic in App Campaigns.
- Size the test against your break-even. Know how big a lift the channel needs to pay off, and design a test that can detect it.
- Analyze at the level you randomized. Randomized by state means analyzed by state.
- Use geo tests to calibrate everything else. They don’t depend on user-level tracking, which is why we recommend them in our 2026 app tracking guide.
- Check that the difference disappears when the ads come back on. Real effects follow the schedule. Coincidences don’t.
FAQ
What is an incrementality test?
An incrementality test measures how much of a result an ad channel actually caused, by comparing a group exposed to the ads with a group that isn’t. Platform reports count every conversion the ads touched. An incrementality test counts only the ones that would have disappeared without them. Where you can’t pause a channel, marketing mix modeling estimates the same thing from historical data.
What is a geo holdout test?
A geo holdout pauses ads in a random set of regions while they keep running elsewhere, then compares revenue between the two. Here that meant 28 of 51 US states for six weeks, against five months of baseline. It needs no user-level tracking, so ATT and cookie loss don’t affect it.
Why is platform ROAS different from incremental ROAS?
Platform ROAS credits the ads with every conversion they touched, including people who would have bought anyway. Incremental ROAS counts only the revenue the ads added. In this test, the account had one of its best reported Android ROAS periods while the holdout found no detectable lift.
What does a minimum detectable effect tell you?
It is the smallest true effect a test can reliably catch. Ours was about 7% of baseline revenue. If that is below the lift your channel needs to break even, a null result means the channel’s incremental ROAS is probably below break-even. If it is above, the test can’t answer the question, and you need a longer or larger one.
Next up
Same app. Same channel. The UK instead of the US. Very different result. We’ll write that one up next.
The team behind the test
Chinmay KulkarniData Scientist · designed the test and ran the analysis
Markus SeppamApp & Business Pro · ran the Google campaigns and the holdout