Demand & Paid Media · established evidence

The Incrementality Gap: What a Decade of Field Experiments Says About Paid Search ROI

Last reviewed 2026-07-20. Written by Chandranshu Kumar, Founder, Raveneye Global. · 11 min read

Paid search incrementality is the share of a campaign's measured return that the advertising actually caused, as opposed to sales the business would have won anyway. A decade of randomized field experiments points to an uncomfortable conclusion: much of the return platforms report on paid search is not incremental to the advertiser. In a large experiment at eBay, ads on the company's own brand keywords produced no measurable short-term lift, and across non-brand terms the spend was absorbed mostly by frequent buyers whose purchases were unaffected. Separate research explains why the dashboards read high in the first place, a statistical artifact called activity bias, and later methods made the clean counterfactual cheap enough to run at scale. The consequence is a persistent gap between platform-reported return and true incremental return. The evidence is strong, though much of it comes from single firms, so the correct response is to measure lift directly rather than to assume any universal number.

A decade of experiments, one uncomfortable finding

Nearly every account dashboard answers the same question the same way: it reports how much revenue occurred while the advertising ran and attributes that revenue to the advertising. This is a correlational measurement dressed as a causal one. The question an owner actually needs answered is different and harder: of the sales the platform is claiming, how many would not have happened without the ad?

That difference has a name in the economics literature. Incrementality is the portion of measured return that the advertising genuinely caused, the counterfactual difference between the world with the campaign and an identical world without it. Between 2011 and 2017, a small set of large-scale field experiments, run inside real advertisers and published in peer-reviewed venues, measured that quantity directly rather than inferring it from a dashboard. Their results do not say advertising never works. They say the number platforms report is systematically higher than the number the advertising earned, and that the size of the gap is predictable.

The eBay experiment: when branded search returned nothing

The most cited result comes from a randomized experiment eBay ran on its own paid search. When the researchers switched off ads on eBay's branded keywords, the terms containing the word "eBay" itself, they found no measurable short-term change in sales. The traffic simply arrived through organic and direct routes instead. On those terms, the paid clicks had been buying visits the company was going to receive anyway.

The non-brand story split two ways, and for a small advertiser, was more instructive. New and infrequent users were positively influenced by the ads. Frequent users, whose purchasing was largely unaffected by advertising, absorbed most of the spend, and because they were the bulk of exposures, the average return across the tested campaigns came out negative. The lesson is not that paid search fails. It is that who the ad reaches determines whether it is incremental, and reaching people already on their way to you is the most expensive way to buy customers you already had.

Why branded search is the sharpest case

Branded search is the clearest example because exposure and intent are almost perfectly correlated: a person typing your name has already chosen you. An ad served into that moment reads brilliantly in the dashboard and frequently causes very little. This is exactly the condition under which platform-reported return and true incremental return diverge most, which is why brand-defense bidding should be treated as a measurable hypothesis, not a default line item.

Activity bias: why observational ad metrics read high

If experiments keep finding smaller effects than dashboards report, the natural question is why the dashboards are wrong in a consistent direction. A 2011 paper answered it with a mechanism the authors called activity bias. Across three controlled experiments, they showed that a user's online behaviors are surprisingly correlated in time: a person who happens to be browsing at a given moment is also more likely to be searching, clicking, and buying at that same moment, advertisement or no advertisement.

That correlation quietly inflates any measurement that compares people who saw an ad against people who did not, because the exposed group was more active to begin with. The ad gets credited for behavior that the underlying burst of activity would have produced on its own. Activity bias is not a flaw in a particular platform's reporting; it is a property of observational data itself, and it is the statistical reason last-click and view-through attribution tend to overstate causal contribution.

Ghost ads: making the clean counterfactual affordable

The clean way to remove activity bias is a randomized holdout: withhold the ad from a comparable control group and compare. Historically that was expensive, because the classic method ran public-service-announcement placements in the control cells, spending real money to show unrelated ads simply to hold the slot. Small advertisers could not justify the cost, so almost nobody measured lift and almost everybody trusted the dashboard.

A 2017 method called ghost ads changed the economics. Instead of buying placebo placements, the system records the counterfactual: it logs which control-group users would have been shown the ad had they not been held out, then compares outcomes between those would-have-seen-it users and the ones who actually saw it. It works natively with real-time ad delivery and costs a fraction of the older approach. On the retargeting campaign the authors studied, the method measured a genuine causal lift of 17.2 percent in site visits and 10.5 percent in purchases, a real effect, cleanly isolated, at a cost an ordinary advertiser could bear. The significance for a local business is that measuring lift is no longer a luxury reserved for platforms and enterprises.

The incrementality gap in industry practice

The academic results describe a mechanism. Practitioners who run holdouts for a living report a magnitude, and here the evidence changes tier. Measurement vendors publishing in 2025 and 2026 commonly find that measured incremental return on ad spend, incremental ROAS, runs well below the ROAS a platform reports about its own work, with branded search and retargeting cited as the channels where the two diverge most.

These figures are directionally consistent with the peer-reviewed findings above, and they describe the same mechanism the eBay and activity-bias studies isolated. They should be read for what they are: unaudited vendor and practitioner analyses, not independent academic replication. We cite the specific percentage ranges as industry-reported rather than as established fact, and we do not attach a single universal number to any account we have not measured. The defensible statement is qualitative and strong: for most advertisers the gap is real and often large, and its size for a given business is an empirical question, not a constant.

Why the gap is a property of the measurement, not a scandal

None of this means the platforms are lying. A dashboard reporting return is answering the question it was built to answer, how much revenue occurred while the campaign ran, and it does that arithmetic correctly. The problem is that the arithmetic is not the causal question, and the two answers separate by an amount that depends on one variable: how correlated the ad exposure is with pre-existing intent.

That single variable organizes the whole literature. Where exposure and intent are tightly coupled, branded search, retargeting people who already visited, the reported number can be almost entirely non-incremental. Where the ad genuinely reaches new demand, prospecting to people who did not know you, more of the reported return is real. The last-click model, by handing full credit to the final touch before a sale, is structurally biased toward the tightly-coupled, least-incremental cases. Understanding the gap is therefore not a reason to stop advertising; it is a reason to stop treating platform-reported ROAS as if it were a measurement of cause.

What actually measures lift: holdouts and geo experiments

If dashboards cannot answer the causal question, something has to. The established answer is experimentation. Geo experiments, randomizing non-overlapping geographic regions into advertising and control conditions, provide a systematic causal method for measuring true ad effectiveness without any individual-level tracking, and were explicitly designed to inform bidding, budgeting, and campaign decisions. For a multi-location business, whole cities or regions can be held back as controls and the difference read as lift.

Where geographic holdouts do not fit, audience holdouts and ghost-ad style counterfactuals do the same job at the user level. The common thread across every credible method is a control group and a pre-committed test design: the window, the held-back cohort, and the success threshold agreed before the test runs, so the result is a genuine read rather than a story told after the fact. Paired with a blended efficiency anchor across the whole account, this is what separates a number you can act on from a number that merely looks healthy.

Reading the evidence

The strongest results here rest on single firms. The eBay experiment is one company's paid search; the ghost-ad magnitudes come from one retargeting campaign; the industry percentages come from vendor case studies. That is a real limit, and it is why the correct claim is a mechanism, not a universal constant. What generalizes is the causal logic, heterogeneous consumer response, activity bias, the exposure-intent correlation, not any particular percentage.

Read that way, the finding is durable and actionable at once. A decade of experiments establishes that platform-reported return overstates causal contribution in a predictable direction, and that the tools to measure the truth are now cheap enough for an ordinary advertiser to use. The dashboard is not the enemy and advertising is not a waste. The defensible position is simply to measure lift instead of assuming it, and to size the gap for your own account rather than borrowing someone else's number.

The evidence

Key findings, with their sources

  • In a large randomized field experiment, ads on eBay's own branded keywords produced no measurable short-term incremental benefit; across non-brand terms, frequent buyers whose purchases were unaffected absorbed most of the spend, yielding negative average returns.

    established Blake, Nosko & Tadelis, "Consumer Heterogeneity and Paid Search Effectiveness: A Large-Scale Field Experiment", Econometrica, 83(1), 2015.

  • Observational estimates of advertising effectiveness are systematically biased upward by "activity bias": a user's online behaviors are correlated in time, so the exposed group was already more active, demonstrated across three controlled experiments.

    established Lewis, Rao & Reiley, "Here, There, and Everywhere: Correlated Online Behaviors Can Lead to Overestimates of the Effects of Advertising", WWW '11, 2011.

  • The "ghost ads" method measures incrementality by recording the counterfactual impressions a control group would have seen, at a fraction of the cost of PSA holdouts; on the retargeting campaign studied it isolated a causal lift of 17.2% in site visits and 10.5% in purchases.

    established Johnson, Lewis & Nubbemeyer, "Ghost Ads: Improving the Economics of Measuring Online Ad Effectiveness", Journal of Marketing Research, 54(6), 2017.

  • Practitioner and vendor analyses report that measured incremental ROAS commonly runs well below platform-reported ROAS, with branded search and retargeting the channels where the two diverge most.

    contested Synthesized from measurement-vendor and practitioner analyses, 2025-2026 (unaudited; directionally consistent with the peer-reviewed field experiments, not an independent replication).

  • Geo experiments, randomizing non-overlapping regions into treatment and control, provide a systematic causal method for measuring true ad effectiveness without individual-level tracking, designed to inform bidding, budgeting, and campaign decisions.

    established Vaver & Koehler, "Measuring Ad Effectiveness Using Geo Experiments", Google Inc., 2011.

Calibration

What is proven, what is promising, what is unproven

Evidence tierTacticsWhat the evidence says
EstablishedThe brand-keyword null result, activity bias, the ghost-ad method, and geo experiments as a causal toolkit.Peer-reviewed field experiments and published methodology (Blake-Nosko-Tadelis 2015; Lewis-Rao-Reiley 2011; Johnson-Lewis-Nubbemeyer 2017; Vaver-Koehler 2011).
EmergingBlended Marketing Efficiency Ratio (MER) as an anti-attribution-gaming anchor across the whole account.Industry-originated metric whose rationale follows directly from the activity-bias and attribution literature, not yet an academic construct.
ContestedAny specific "iROAS runs 30 to 70 percent below platform ROAS" figure, or a fixed branded-search waste percentage.Unaudited vendor and practitioner case studies; directionally consistent with the established findings but not independently replicated. Cite as industry-reported, never as a constant.

Reference

Glossary

Incrementality
The portion of measured return an advertisement genuinely caused: the counterfactual difference between the world with the campaign and an identical world without it.
Incremental ROAS (iROAS)
Return on ad spend calculated only against the revenue a campaign actually caused, measured against a held-back control, rather than against every sale that occurred while the ad ran.
Activity bias
The upward bias in observational ad measurement caused by the time-correlation of a user's online behaviors: the exposed group was already more active, so the ad is credited for activity it did not cause.
Ghost ads
A measurement method that records which control-group users would have been shown an ad, enabling a clean incrementality read without buying placebo placements.
Geo experiment
A causal test that randomizes non-overlapping geographic regions into advertising and control conditions, reading the difference as true lift without individual-level tracking.

Straight answers

Frequently asked questions

What is paid search incrementality?

It is the share of a paid-search campaign's return that the advertising actually caused, as distinct from sales the business would have won through organic, direct, or repeat channels anyway. The field experiments in this article measure that quantity directly instead of inferring it from a platform dashboard.

Does this mean Google Ads does not work?

No. The evidence says platform-reported return overstates causal contribution in a predictable direction, worst where the ad reaches people who were already coming to you. Advertising that reaches genuinely new demand can be strongly incremental. The point is to measure which of your spend is which, not to stop advertising.

Why does branded search show up as the least incremental?

Because a person searching your name has already chosen you, so exposure and intent are almost perfectly correlated. An ad served into that moment reads well in the dashboard and often causes very little additional purchasing, which is exactly the eBay result. Brand-defense bidding should be treated as a testable hypothesis, not an automatic line item.

How would a small business actually measure incremental return?

With a control group and a pre-committed design: a geo holdout that withholds advertising from whole regions, an audience holdout at the user level, or a ghost-ad style counterfactual. The window, the held-back cohort, and the success threshold are agreed before the test runs, and the whole account is anchored to a blended efficiency figure no single channel can inflate.

Are the "iROAS runs 30 to 70 percent below platform ROAS" figures reliable?

Treat them as directional, not exact. They come from unaudited vendor and practitioner case studies and are consistent with the peer-reviewed mechanism, but they are not independent academic replication. The defensible claim is that the gap is real and often large; its precise size for any given account is an empirical question to be measured, not assumed.

Provenance

Sources

  1. Blake, T., Nosko, C. & Tadelis, S., "Consumer Heterogeneity and Paid Search Effectiveness: A Large-Scale Field Experiment", Econometrica, 83(1), 2015 (established)doi.org
  2. Lewis, R. A., Rao, J. M. & Reiley, D. H., "Here, There, and Everywhere: Correlated Online Behaviors Can Lead to Overestimates of the Effects of Advertising", WWW '11, 2011 (established)doi.org
  3. Johnson, G. A., Lewis, R. A. & Nubbemeyer, E. I., "Ghost Ads: Improving the Economics of Measuring Online Ad Effectiveness", Journal of Marketing Research, 54(6), 2017 (established)doi.org
  4. Vaver, J. & Koehler, J., "Measuring Ad Effectiveness Using Geo Experiments", Google Inc., 2011 (established)research.google
  5. Measurement-vendor and practitioner analyses of the iROAS-to-platform-ROAS gap, 2025-2026 (contested / industry-reported, pending independent replication)

Every figure above is attributed to a real, dated source and tagged with its evidence tier. Where a claim could not be verified to a primary source, it is not stated as fact.

What this means for your ad budget

The evidence points to one operational question most businesses cannot answer: of everything you spent last month, how much actually caused a sale that would not have happened anyway? Platform dashboards cannot answer that about their own work. A standing measurement layer, controlled holdouts and geo-lift tests read against a blended efficiency figure, can. That is what the Incrementality and Measurement Retainer runs, so your budget follows the spend that genuinely earns new customers instead of the channels best at claiming credit.

service Incrementality & Measurement Retainer An experiment-led measurement layer over your paid media: geo-lift and holdout tests that measure true incremental return, reported with confidence intervals against a blended MER, with your ad spend paid straight to the platforms and never marked up. If the leak is inside one Google Ads account, the Paid Search Conversion Audit finds where the click-to-customer path breaks first. See how it works

Start free with a Machine-Readiness Score, a specialist-reviewed read of where you stand across search and AI answers. No guaranteed number, and no obligation.