Demand & Paid Media · established evidence
Geo-Experiments: The Randomized Controlled Trial for Local Advertising
A geo-experiment is a randomized controlled trial for advertising. Instead of tracking individual people, it divides a market into non-overlapping geographic regions, randomly assigns some regions to keep running ads and others to pause them, and reads the difference in outcomes between the two groups. Because the regions were assigned at random, that difference is a causal estimate of what the advertising caused, not a correlation the platform happened to report. The method was formalized by two Google statisticians, Jon Vaver and Jim Koehler, in a 2011 paper that framed geo experiments as a systematic way to measure true ad effectiveness and to inform bidding, budgeting, and campaign-design decisions. For a local or multi-location business the appeal is specific: it needs no individual-level tracking, no cookies, and no in-house data-science team, only enough separable regions and the discipline to hold one group out. This article explains why the design works and how to read its results.
The number your ad platform reports is not the number you want
An advertising dashboard reports a correlation. It records that a set of people saw an ad and that some of them later converted, and it credits the ad with those conversions. What it cannot see is the counterfactual: how many of those same people would have converted anyway, with no ad at all. The reported figure and the causal figure are different quantities, and for most campaigns the reported one is larger.
The size of that gap is not a matter of opinion. In a set of three controlled experiments, Randall Lewis, Justin Rao, and David Reiley documented what they called activity bias: the surprising degree to which a person's online behaviors are correlated in time. People who happen to be browsing are also more likely to be searching, clicking, and buying in the same window, advertisement or not. Any measurement that compares exposed users to a naively chosen unexposed group inherits that correlation and reads it as ad effect. The result is a systematic overestimate of what advertising did.
This is the problem a geo-experiment is built to solve. Rather than trusting the platform's attribution, it manufactures a genuine control group and reads the difference directly.
What a geo-experiment actually is
The design borrows its logic directly from the clinical trial. In medicine, you cannot know what a drug did to a single patient, because you never observe that patient both treated and untreated. You solve the problem at the level of groups: randomly assign patients to a treatment arm and a control arm, and the random assignment makes the two groups statistically comparable in everything except the treatment. The difference in outcomes is then attributable to the treatment.
Geography as the unit of randomization
A geo-experiment applies that same logic, but the unit that gets randomized is a place, not a person. Vaver and Koehler's method partitions a market into non-overlapping geographic areas, cities, metros, or designated market regions, and randomly assigns them to conditions. One set of regions keeps the advertising running. The other set has it turned off or held down for the duration of the test.
Because the assignment is random and the regions do not overlap, the control regions become a credible estimate of what would have happened in the treated regions absent the advertising. Subtract the control outcome from the treated outcome and you have the incremental lift: the sales, calls, or bookings the advertising actually caused, above the baseline the business would have earned anyway.
A pre-period to calibrate, a test period to measure
In practice the design uses a pre-experiment period to learn the normal relationship between the regions, then compares the test period against the level that relationship predicts. Vaver and Koehler framed the output not as a vanity number but as an input to decisions: how to set bids, how to size budgets, and how to design the next campaign. The experiment is a measurement instrument, not a report card.
Why geography measures lift without tracking a single person
The quiet advantage of the geo design is that it never needs to identify anyone. It reads outcomes at the level of the region, total conversions in region A versus total conversions in region B, so it does not depend on cookies, device graphs, or individual-level identity resolution at all. The randomization does the work that individual tracking is usually asked, and often fails, to do.
That property has become more valuable, not less, as the tracking ground has shifted underneath advertisers. Google spent six years building its Privacy Sandbox as a replacement for the third-party cookie, then terminated the initiative in 2025, retiring the remaining APIs and leaving third-party cookies in Chrome with no removal timeline. The lesson for a measurement strategy is that platform-level tracking policy is not a stable foundation to build on. A method whose validity comes from randomizing regions rather than from following users is insulated from that volatility by construction.
The evidence that reported lift and real lift diverge
The reason to go to this trouble is that holdout experiments repeatedly find less lift than dashboards claim, and sometimes find none. The clearest demonstration in the paid-search literature is a large-scale randomized field experiment run at eBay by Thomas Blake, Chris Nosko, and Steven Tadelis. Advertising on the company's own branded keywords produced no measurable short-term incremental benefit: the users clicking those ads were people already headed to eBay, who would have arrived through organic results or direct navigation regardless.
For non-brand keywords the picture was more mixed and more instructive. New and infrequent users were positively influenced by the ads, but frequent users, whose purchases the ads did not change, absorbed most of the spend, dragging the average return negative. A single blended dashboard number would have hidden that entirely. Only a design that separated the causal effect from the correlational one could surface it, which is exactly what a geo-holdout does at the regional level.
This is a single-firm study, and its specific magnitudes should not be transplanted onto a med-spa or a plumbing company as if they were constants. What generalizes is the mechanism: the more an ad is shown to people who already intended to buy, the more attribution overstates it, and branded and retargeting inventory is where that overlap is highest.
Cheaper cousins: ghost ads and the cost of causal measurement
Geo-experiments are not the only causal instrument, and they are not always the cheapest. The older gold standard, the public-service-announcement holdout, showed the control group an unrelated charity ad in place of the real one so the two groups were treated identically except for the message. It works, but it is expensive, because the advertiser pays to serve placebo impressions.
Garrett Johnson, Randall Lewis, and Elmar Nubbemeyer introduced a lower-cost alternative they called ghost ads: rather than buying placebo placements, the system records the counterfactual impressions a control user would have been served, and compares against those. On a retargeting campaign, their method measured a lift of 17.2 percent in site visits and 10.5 percent in purchases while working natively with real-time ad-delivery optimization and at a fraction of the holdout cost.
The point for a local advertiser is not that one method wins. It is that a whole toolkit exists for answering the causal question, and that geo-experiments occupy the position in that toolkit where the business runs across separable regions and wants an answer that owes nothing to individual tracking.
What a geo-experiment can and cannot tell you
The design has real preconditions, and acknowledging them is part of using it well. It needs enough non-overlapping regions to randomize, which means a genuinely single-location business cannot run a clean version of it. It needs the regions to be reasonably separable, so that turning ads off in one place does not simply push demand into another. It needs a test window long enough to accumulate signal, because the incremental effect is usually a modest fraction of total sales and modest effects require patience and volume to detect.
It is also worth being precise about what the wider industry claims here versus what is established. Practitioner analyses frequently report that measured incremental ROAS runs 30 to 70 percent below platform-reported ROAS, with branded search the worst offender. That direction is consistent with the peer-reviewed findings above, but the specific percentage ranges come from unaudited vendor case studies, not academic replication, and should be treated as directional rather than as a number to quote as fact. The correct posture is to measure your own gap, not to import someone else's.
Reading the result
A geo-experiment returns an incremental figure that will almost always be lower than the platform's attributed figure. That is the finding, not a failure of the test. The disciplined response is to treat the incremental number as the real one and to re-plan bids and budgets against it, which is precisely the decision use Vaver and Koehler had in mind.
It also explains the growing practitioner preference for a blunt anchor metric. Marketing Efficiency Ratio, total revenue divided by total marketing spend, cannot be gamed by shifting attribution credit between channels, because it never assigns channel-level credit in the first place. It is an industry construct rather than an academic one, but its rationale is the same activity-bias and attribution-gaming logic the field-experiment literature documents. A geo-experiment tells you the causal contribution of a specific lever; a blended ratio keeps you honest about the whole. Neither is a naked platform-reported ROAS, and that is the point.
Who should run one
The design fits a particular shape of business. If you operate across several markets, a multi-location med-spa group, a home-services company covering many metros, a franchise system, you already have the raw material a geo-experiment needs, and you are also the kind of advertiser for whom a persistent attribution gap compounds into real money over a year of spend.
You do not need a statistics department to benefit from the method. You need someone to design the test so the regions are comparable, hold the discipline of the holdout, and read the result without flattering it. That measurement practice is what decides what to keep spending on and what to cut.
The evidence
Key findings, with their sources
-
Geo experiments randomize non-overlapping geographic regions into treatment and control ad conditions to measure true ad effectiveness without individual-level tracking, explicitly to inform bidding, budgeting, and campaign design.
established Vaver, J. & Koehler, J., "Measuring Ad Effectiveness Using Geo Experiments", Google Inc., 2011.
-
Across three controlled experiments, correlation among a user's online behaviors (activity bias) led observational methods to systematically overestimate the effects of advertising.
established Lewis, R. A., Rao, J. M. & Reiley, D. H., "Here, There, and Everywhere: Correlated Online Behaviors Can Lead to Overestimates of the Effects of Advertising", WWW '11, 2011.
-
Paid search ads on branded keywords produced no measurable short-term incremental benefit in a large-scale randomized field experiment at eBay; for non-brand keywords, frequent users absorbed most spend and drove average returns negative.
established Blake, T., Nosko, C. & Tadelis, S., "Consumer Heterogeneity and Paid Search Effectiveness: A Large-Scale Field Experiment", Econometrica 83(1), 2015.
-
A ghost-ad measurement of a retargeting campaign found a lift of 17.2% in site visits and 10.5% in purchases, at a fraction of the cost of a traditional PSA holdout.
established Johnson, G. A., Lewis, R. A. & Nubbemeyer, E. I., "Ghost Ads: Improving the Economics of Measuring Online Ad Effectiveness", Journal of Marketing Research 54(6), 2017.
-
Practitioner analyses report measured incremental ROAS running roughly 30 to 70 percent below platform-reported ROAS, with branded search the widest divergence.
contested Industry practitioner synthesis (Prescient AI, Eightx, MHI, layerfive), 2025-2026; directionally consistent with the peer-reviewed field-experiment literature but not independently audited.
Calibration
What is proven, what is promising, what is unproven
| Evidence tier | Tactics | What the evidence says |
|---|---|---|
| established | Geo-holdout tests, activity-bias correction, field-experiment incrementality, ghost-ad counterfactuals | Vaver-Koehler 2011; Lewis-Rao-Reiley 2011; Blake-Nosko-Tadelis 2015; Johnson-Lewis-Nubbemeyer 2017 (peer-reviewed and primary-source) |
| emerging | Marketing Efficiency Ratio (blended ROAS) as the anti-gaming anchor metric | Industry practitioner consensus, grounded in the attribution-gaming and activity-bias literature but not itself an academic construct |
| contested | Quoting a specific iROAS gap (for example, 30 to 70 percent below platform ROAS) as a fixed figure | Unaudited vendor case studies; direction supported by peer-reviewed work, magnitude not independently replicated |
Reference
Glossary
- Geo-experiment
- A randomized controlled trial for advertising in which non-overlapping geographic regions, not individuals, are randomly assigned to run ads or hold them out, so the difference in outcomes is a causal estimate of ad lift.
- Incrementality
- The additional outcomes advertising actually caused, above the baseline a business would have earned with no ad. It is the difference between the treated and control groups, not the total the platform attributes.
- Incremental ROAS (iROAS)
- Return on ad spend calculated only on the incremental, causally measured revenue, rather than on all revenue an attribution model credits to the channel.
- Activity bias
- The tendency for a user's online behaviors to be correlated in time, which makes exposed users look more active than a naive control group and inflates observational estimates of ad effect.
- Counterfactual (holdout)
- The outcome that would have occurred without the advertising. A control group, whether a region, a PSA arm, or a ghost-ad record, is a way to observe it.
- Ghost ads
- A low-cost incrementality method that records the ad impressions a control user would have been served, rather than paying to serve placebo ads, then compares treated outcomes against that recorded counterfactual.
- Marketing Efficiency Ratio (MER)
- Total revenue divided by total marketing spend, a blended figure that cannot be gamed by shifting attribution credit between channels because it never assigns channel-level credit.
Straight answers
Frequently asked questions
What is a geo-experiment in advertising?
It is a randomized controlled trial run on places instead of people. A market is split into non-overlapping regions that are randomly assigned to keep running ads or to pause them, and the difference in outcomes between the two groups estimates the lift the advertising actually caused. The method was formalized by Vaver and Koehler at Google in 2011.
Do I need cookies or individual tracking to run a geo holdout test?
No. The test reads outcomes at the level of the region, total conversions in the ad regions versus the holdout regions, so it never needs to identify individuals. That is a large part of its appeal now that third-party tracking has become unstable, with Google terminating its Privacy Sandbox replacement plan in 2025.
Do I need a data-science team to do this?
You need statistical care, but not an in-house department. The hard parts are choosing comparable regions, holding the discipline of the holdout for the full test window, and reading the result without inflating it. That is a measurement practice a specialist can run for you; the design itself is decades old and well documented.
How is a geo-experiment different from the ROAS my ad platform reports?
The platform reports a correlation: people who saw an ad and later converted. It cannot see how many would have converted anyway. A geo-experiment builds a genuine control group so it can. The two numbers are different quantities, and field experiments repeatedly find the causal one is lower, sometimes near zero for branded search, as the eBay study by Blake, Nosko, and Tadelis showed.
How many locations do I need to run one?
Enough non-overlapping, reasonably separable regions to randomize into two comparable groups, which means a true single-location business cannot run a clean geo-experiment. Multi-location businesses, franchises, and companies covering several metros are the natural fit, and they are also where an unmeasured attribution gap compounds most over a year of spend.
Provenance
Sources
- Vaver, J. & Koehler, J., "Measuring Ad Effectiveness Using Geo Experiments", Google Inc., 2011 (established)research.google
- Lewis, R. A., Rao, J. M. & Reiley, D. H., "Here, There, and Everywhere: Correlated Online Behaviors Can Lead to Overestimates of the Effects of Advertising", Proceedings of WWW '11, 2011 (established)doi.org
- Blake, T., Nosko, C. & Tadelis, S., "Consumer Heterogeneity and Paid Search Effectiveness: A Large-Scale Field Experiment", Econometrica, 83(1), 155-174, 2015 (established)doi.org
- Johnson, G. A., Lewis, R. A. & Nubbemeyer, E. I., "Ghost Ads: Improving the Economics of Measuring Online Ad Effectiveness", Journal of Marketing Research, 54(6), 867-884, 2017 (established)doi.org
- Google, "Next steps for Privacy Sandbox and tracking protections in Chrome", privacysandbox.google.com/blog, 2025 (established, recent industry-documented event)
- Industry practitioner synthesis on the iROAS gap (Prescient AI, Eightx, MHI, layerfive), 2025-2026 (contested, unaudited vendor case studies)
Every figure above is attributed to a real, dated source and tagged with its evidence tier. Where a claim could not be verified to a primary source, it is not stated as fact.