Measurement & Honesty · established evidence

Incrementality Testing With Geo Experiments: Proving a Campaign Worked Without a User-Level Pixel

Last reviewed 2026-07-20. Written by Chandranshu Kumar, Founder, Raveneye Global. · 11 min read

Incrementality testing answers the only question a paid campaign is really being asked: did the spend cause sales that would not have happened anyway, or did the platform simply take credit for customers who were already coming? For years, that answer depended on a user-level pixel that followed individuals across the web. That pixel is degrading, both because privacy changes have removed much of the signal and because the observational methods built on it systematically disagree with what controlled experiments show. Geo experiments offer a way through. Instead of tracking people, they hold out matched geographic markets and use a synthetic control, a weighted blend of untreated regions, to model what would have happened without the campaign. The gap between the treated markets and that modeled baseline is the incremental lift. The method is established and used across the industry; the vendor-reported accuracy figures are honest but early. This piece explains how it works, what the evidence supports, and where it stops.

The measurement you actually want is a counterfactual

Almost every number a business sees about its advertising describes a correlation: this many people who saw the ad went on to buy. The number a business should care about is a counterfactual: how many of those buyers would not have purchased if the ad had never run. The first is what happened alongside the campaign. The second is what the campaign caused. They are not the same quantity, and the difference is not academic.

The reason the two diverge is selection. Ad systems are built to show impressions to the people most likely to convert, and retargeting and branded search in particular reach audiences who were already on their way to buying. An attribution model that credits those conversions to the ad is measuring the platform's targeting skill, not its persuasive effect. Isolating the causal share requires a comparison group that did not receive the campaign but is otherwise like the group that did.

This is the same logic that underpins a randomized controlled trial in medicine. The treatment effect is defined as the difference between a treated population and an equivalent untreated one. The whole discipline of incrementality testing is the effort to construct that untreated comparison honestly when the unit being treated is a market rather than a patient.

Why observational attribution keeps getting it wrong

The strongest evidence that correlation-based measurement misleads comes from inside one of the platforms that sells it. Gordon, Zettelmeyer, Bhargava and Chapsky ran fifteen large-scale randomized field experiments at Facebook and compared the true causal lift each experiment produced against what standard observational methods, the same family that underpins most multi-touch attribution and lookalike modeling, would have estimated from the identical data.

The observational estimates frequently landed in the wrong direction or the wrong magnitude, and conditioning on rich demographic and behavioral covariates did not reliably fix the gap. This matters because it means the failure is not a data-quality problem that better tracking would solve. It is structural. The methods answer the wrong question, so feeding them more individual-level signal does not make the answer right, it makes a confident wrong answer.

That finding reframes the privacy conversation entirely. If the pixel-based estimate was already systematically biased, then losing the pixel is less a catastrophe than an occasion to adopt the method that was more trustworthy all along.

The pixel was fraying before privacy finished it

Two forces are pulling user-level tracking apart at once. The first is regulatory and platform-driven. Apple's App Tracking Transparency, introduced with iOS 14.5 in April 2021, required apps to ask permission before tracking users across other companies' apps and sites, and industry measurement reports put the opt-out rate high enough that a large share of the mobile signal simply stopped flowing. Pixel-based revenue attribution for e-commerce, by the same vendor reporting, fell from capturing most conversions to capturing a clear minority.

The second force is the deprecation and partial replacement of third-party cookies, which removed much of the cross-site scaffolding that desktop attribution relied on. The industry response has been to move toward server-side and consented signals and, more importantly for this discussion, toward aggregate methods that never depended on identifying individuals in the first place: marketing mix modeling and geo experiments.

The opt-out and attribution-loss percentages that circulate come from ad-tech vendors with a commercial interest in the narrative, not from independent peer review, so they are best read as directional rather than precise. What is not in dispute is the direction and the dates: individual-level measurement is materially weaker than it was, and it is not coming back.

How a geo experiment works

A geo experiment abandons the individual as the unit of measurement and adopts the geographic market instead. The advertiser divides its footprint into regions, runs the campaign in some of them, deliberately withholds it in others, and measures the difference in outcomes between the two sets. Because the holdout regions never saw the campaign, they stand in for the counterfactual: what the treated regions would have done anyway.

Matched markets and the holdout

The credibility of the design rests on how comparable the treated and control regions are before the campaign starts. If the two sets moved together historically, the control set is a reasonable forecast of the treated set in the absence of any intervention. This is why geo experiments are typically planned around a pre-period of shared history rather than assembled at random on the day a campaign launches.

Holding out a market has a real cost: the advertiser accepts that some regions will receive no campaign for the duration of the test. That cost is the price of a trustworthy answer, and it is why geo tests are scoped to a defined window and a lift worth detecting rather than run indefinitely.

The synthetic control

Rather than pairing each treated market with a single similar one, the synthetic control method builds the comparison as a weighted combination of untreated markets, chosen so that the blend tracks the treated market's pre-campaign history as closely as possible. The synthetic control is then projected forward through the campaign window to estimate what the treated market would have done without the ads. The difference between the observed outcome and that projection is the estimated incremental lift.

Meta's open-source GeoLift library is the most widely used implementation of this method for advertising, and its public methodology documents the synthetic-control construction, the power analysis for choosing test markets, and the inference procedure for judging whether an observed lift is distinguishable from noise.

Reading the lift

The output of a well-run geo test is an estimate of incremental conversions or revenue, with an interval around it and a stated confidence that the effect is real rather than chance. Divided by the spend, it yields an incremental return on ad spend, a figure that cannot be inflated by a platform crediting itself for pre-existing demand, because the holdout markets already subtract that demand out.

What the evidence says about geo methods

The method itself is established: synthetic-control geo experimentation is used across the advertising industry and is the design that both platform measurement teams and independent practitioners reach for when a pixel is unavailable or untrusted. The comparative performance numbers, however, deserve a more careful label.

An independent head-to-head simulation study by Recast compared several open-source geo-testing tools and reported that GeoLift's coverage, the share of simulations in which its interval contained the true effect, came closest to the 95 percent target at roughly 92 to 95 percent, and that its false-positive rate, the share of simulations in which it detected an effect that was not there, was the lowest of the tools compared at roughly 3 to 5 percent. Those are encouraging results, and they are the kind of validation the field needs.

They are also a single vendor-run simulation rather than a peer-reviewed benchmark. The right reading is that the numbers support the method's reliability in the tested conditions without settling it as a law. A geo test is a strong instrument; it is not an oracle, and its accuracy on any given account depends on how well the pre-period matched, how large the true effect is relative to market noise, and whether the test had the statistical power to see it.

Geo experiments and marketing mix modeling are complements, not rivals

Geo experiments belong to a broader family of aggregate, privacy-durable measurement methods that survived the collapse of user-level tracking precisely because they never used it. Marketing mix modeling is the other member of that family. Rather than following individuals, it regresses an outcome time series such as sales on marketing time series, with adstock and saturation transforms to model carryover and diminishing returns, and recovers each channel's contribution statistically.

The clearest signal of how seriously the industry now takes these methods is that both of the largest ad platforms have open-sourced their internal versions: Google's Meridian, built on a geo-level Bayesian hierarchical media-mix model, and Meta's Robyn. That the companies whose businesses depend on advertising being seen to work have published the aggregate methods that grade advertising more skeptically than a pixel does is itself evidence of where credible measurement is heading.

The two methods answer different questions and are strongest together. A geo experiment gives a clean causal read on one campaign or channel over one window. A mix model gives a continuous, whole-account picture across channels and time, calibrated by the experiments. Neither replaces the other, and a serious measurement practice runs experiments to anchor the model rather than choosing between them.

The traps that make a test lie

A geo experiment is only as honest as its discipline. The catalog of ways controlled experiments mislead is well documented by Kohavi, Tang and Xu, drawing on the experience of running more than twenty thousand experiments a year across Google, LinkedIn and Microsoft. The most common is peeking: checking results before the test reaches its planned sample and stopping early the moment they look significant, which manufactures false positives because random noise will eventually cross a threshold if you keep looking. Sample-ratio mismatch, novelty and primacy effects, and misread interactions round out the list.

For a single local business there is a harder, quieter constraint underneath all of these: statistical power. Standard sample-size math at conventional significance and power levels, for a minimum detectable effect in the single-digit percentages, calls for tens to hundreds of thousands of observations. At the traffic a single-location business realistically sees, detecting even a sizeable relative lift can require many months, well past the point where the campaign or the business has already changed. This is a structural reason, not a moral one, that owner-operated firms have so often been sold vanity metrics instead of causal ones.

The resolution is not a shortcut. It is choosing tests scaled to the question, accepting that some effects are too small to prove at a single site in a useful window, and, where possible, pooling evidence across many similar businesses to recover the statistical power that no single one of them has. An experiment you were not powered to read is not evidence; it is a coin flip wearing a lab coat.

What a geo experiment cannot tell you

The method has real limits. A geo test measures the aggregate effect of what ran in the treated markets during the window; it does not attribute that effect to a specific creative, keyword or audience unless the design isolates one. It assumes the holdout markets were not contaminated by spillover, which cross-border media and word of mouth can violate. It measures the period it ran, not next quarter, and a lift that holds in spring may not hold in a promotion-heavy December.

It also sits beside a genuinely unsolved measurement frontier. Whether and why an AI answer engine cites a given business is a second attribution problem being invented from scratch, and there is no agreed sampling methodology for it yet, in part because generative systems are non-deterministic by construction. Geo experiments are the mature end of the measurement world; share-of-answer measurement is the immature end. A serious practice states which is which rather than lending the maturity of one to the other.

None of this is a reason to prefer the confident wrong number. It is the reason to prefer the bounded one, and to keep re-running it as accounts and platforms change.

The evidence

Key findings, with their sources

  • In an independent head-to-head simulation, GeoLift's interval coverage came closest to the 95% target at roughly 92 to 95 percent, and its false-positive rate was the lowest among the open-source geo-testing tools compared at roughly 3 to 5 percent.

    emerging Recast Research, "Open-Source Geo-Experiment Tools: A Head-to-Head Simulation Study" (research.getrecast.com); facebookincubator/GeoLift (GitHub).

  • Across fifteen large-scale randomized field experiments at Facebook, standard observational attribution methods frequently produced lift estimates in the wrong direction or of the wrong magnitude versus the randomized ground truth, even after conditioning on rich demographic and behavioral covariates.

    established Gordon, B.R., Zettelmeyer, F., Bhargava, N., Chapsky, D. (2019), "A Comparison of Approaches to Advertising Measurement: Evidence from Big Field Experiments at Facebook," Marketing Science 38(2):193-225.

  • After Apple's App Tracking Transparency (iOS 14.5, April 2021), industry reporting put e-commerce pixel-based revenue attribution as falling from capturing most conversions (roughly 80 to 95 percent) to a clear minority (roughly 60 to 70 percent), pushing the industry toward server-side and aggregate methods.

    contested Industry measurement reports (AppsFlyer opt-in-rate study; PubMatic ad-spend shift data), summarized in marketing-industry press. Point estimates are vendor-reported and should be read as directional.

  • Both of the largest ad platforms have open-sourced their internal marketing mix modeling methodology: Google's Meridian (geo-level Bayesian hierarchical media-mix model) and Meta's Robyn.

    established Google, "Meridian" (business.google.com); Tueller, N. et al. (2024), "Packaging Up Media Mix Modeling: An Introduction to Robyn's Open-Source Approach," arXiv:2403.14674; facebookexperimental/Robyn (GitHub).

  • Peeking, checking a test before it reaches its planned sample and stopping early on an apparent significant result, is a cataloged cause of false positives, documented from experience running more than 20,000 controlled experiments a year at Google, LinkedIn and Microsoft.

    established Kohavi, R., Tang, D., Xu, Y. (2020), Trustworthy Online Controlled Experiments: A Practical Guide to A/B Testing, Cambridge University Press.

  • Standard power analysis (95% significance, 80% power, a single-digit-percent minimum detectable effect) typically requires tens to hundreds of thousands of observations, so at realistic single-location traffic, detecting even a sizeable relative lift can take many months.

    established Industry synthesis of standard power-analysis methodology (Analytics-Toolkit.com; Statsig, "Power Analysis for A/B Testing"), consistent with the general statistics of hypothesis testing.

Calibration

What is proven, what is promising, what is unproven

Evidence tierTacticsWhat the evidence says
EstablishedHold out matched markets and read the difference; synthetic-control geo experiments; anchor a marketing mix model with experiments; treat observational last-touch attribution as biased.Gordon et al. (2019) on observational vs randomized divergence; Kohavi et al. (2020) on experiment traps; GeoLift, Meridian and Robyn as widely used, open-sourced aggregate methods.
EmergingRely on a specific geo tool's reported coverage and false-positive rate; extend mix modeling down to a single-location business.Recast's GeoLift coverage (92 to 95 percent) and false-positive (3 to 5 percent) figures are a single vendor-run simulation, not peer-reviewed; MMM fit at MSME scale is not yet well evidenced.
ContestedQuote precise signal-loss percentages from the cookie and ATT era; treat any AI-answer "visibility score" as a settled measurement.iOS ATT loss figures are vendor-reported and directional; there is no agreed sampling methodology for share-of-answer, and generative engines are non-deterministic by construction.

Reference

Glossary

Incrementality
The share of outcomes a campaign actually caused, measured against a counterfactual of what would have happened without it. Distinct from attributed conversions, which only correlate with the ad.
Geo experiment
A test that treats geographic markets, not individuals, as the unit of measurement: run a campaign in some regions, withhold it in matched others, and read the difference as lift.
Synthetic control
A weighted blend of untreated markets built to track a treated market's pre-campaign history, then projected forward to estimate what the treated market would have done without the campaign.
Holdout
A group, here a set of markets, deliberately not shown a campaign so it can serve as the untreated comparison against which lift is measured.
Marketing mix modeling
A statistical method that regresses an outcome time series on marketing time series, with carryover and saturation transforms, to estimate each channel's contribution without tracking individuals.
Incremental ROAS
Return on ad spend computed from the incremental sales a campaign caused, rather than from all sales a platform credits to itself, so pre-existing demand is subtracted out.

Straight answers

Frequently asked questions

What is incrementality testing?

It is measuring how much of a campaign's results the campaign actually caused, rather than how many conversions happened alongside it. The test constructs a comparison that did not receive the campaign, a holdout, and reads the difference as the incremental lift, which is the honest basis for an incremental return on ad spend.

How do geo experiments measure ad impact without a user-level pixel?

By changing the unit of measurement from the person to the market. The campaign runs in some regions and is withheld in matched others, and a synthetic control, a weighted blend of the untreated regions, models what the treated regions would have done anyway. The gap is the lift. No individual is tracked, so the method is durable to privacy changes.

Is a geo experiment better than the ROAS my ad platform reports?

It answers a different and harder question. A platform reports the conversions it can associate with its own ads, which includes buyers who were already coming. A geo test subtracts that pre-existing demand out by comparing against markets that saw no campaign. Platform ROAS is useful for operating a campaign day to day; a geo test is what tells you whether the spend caused anything.

Can a small single-location business run a valid geo test?

Sometimes, and honestly not always. Detecting a real effect requires statistical power, and at low traffic a small lift can be impossible to distinguish from noise within a useful window. The responsible answer is to scope the test to a lift large enough to see, accept that some questions cannot be proven at a single site, and, where possible, pool evidence across similar businesses.

Do geo experiments replace marketing mix modeling?

No. They are complements. A geo experiment gives a clean causal read on one campaign over one window; a mix model gives a continuous whole-account picture across channels and time. The strongest practice uses experiments to anchor and calibrate the model rather than choosing one over the other.

What can a geo experiment not tell me?

It measures the aggregate effect of everything that ran in the treated markets during the window, not the contribution of a single creative or keyword unless the design isolates one. It assumes no spillover between markets, and it describes the period it ran, not future seasons. It is a strong, bounded instrument, not a universal one.

Provenance

Sources

  1. Gordon, B.R., Zettelmeyer, F., Bhargava, N., Chapsky, D. (2019), "A Comparison of Approaches to Advertising Measurement: Evidence from Big Field Experiments at Facebook," Marketing Science 38(2):193-225 (established)
  2. facebookincubator/GeoLift (GitHub), open-source synthetic-control geo-experimentation library and methodology documentation (established: method)
  3. Recast Research, "Open-Source Geo-Experiment Tools: A Head-to-Head Simulation Study" (research.getrecast.com) (emerging: comparative figures are a single vendor-run simulation, not peer-reviewed)
  4. Kohavi, R., Tang, D., Xu, Y. (2020), Trustworthy Online Controlled Experiments: A Practical Guide to A/B Testing, Cambridge University Press (established)
  5. Google, "Meridian" open-source Bayesian marketing mix model (business.google.com); Sun, Wang, Jin et al. geo-level Bayesian Hierarchical Media Mix Modeling lineage (established)
  6. Tueller, N. et al. (2024), "Packaging Up Media Mix Modeling: An Introduction to Robyn's Open-Source Approach," arXiv:2403.14674; facebookexperimental/Robyn (GitHub) (established)arxiv.org
  7. AppsFlyer opt-in-rate study and PubMatic ad-spend shift data on iOS 14.5 App Tracking Transparency, summarized in marketing-industry press (contested: vendor-reported point estimates, treat as directional)
  8. Analytics-Toolkit.com and Statsig, "Power Analysis for A/B Testing," industry synthesis of standard power-analysis methodology (established as statistical math; illustrative figures are industry-sourced)

Every figure above is attributed to a real, dated source and tagged with its evidence tier. Where a claim could not be verified to a primary source, it is not stated as fact.

What this means for your ad budget

If the real measure of a campaign is what it caused and not what a platform claims, then the number on your ad dashboard is answering the wrong question, and no amount of extra tracking fixes that. What fixes it is a standing measurement layer that runs geo-lift and holdout tests on your spend, reports incremental return instead of self-graded ROAS, and re-runs as your accounts change. That is exactly what the Incrementality and Measurement Retainer does: a measurement seam that sits on top of your paid media and reports what the spend actually caused.

service Incrementality & Measurement Retainer A standing measurement layer over your paid media: experiment design, geo-lift and holdout tests, and a monthly incremental read reviewed by a specialist. Your ad spend stays yours, paid straight to the platforms, never marked up. See how it works

Start free with a Machine-Readiness Score, a specialist-reviewed read of where you stand across search and AI answers. No guaranteed number, and no obligation.