Conversion Science · established evidence

Why Your Google Ads Dashboard Overstates Its Own Impact: Lessons From a Randomized Field Experiment at eBay

Last reviewed 2026-07-20. Written by Chandranshu Kumar, Founder, Raveneye Global. · 11 min read

The Google Ads attribution report tells you how many conversions followed an ad click. It does not tell you how many of those conversions would have happened without the ad, and that gap is where the number inflates. In 2015 three economists ran large randomized experiments at eBay and compared the platform-style, observational estimate of paid-search return against the true causal lift measured by turning ads off for randomly chosen groups. The observational estimate was substantially larger than the real effect, because ad clicks concentrate among people who already intended to buy. Brand-keyword ads, the ones bidding on a company's own name, produced no measurable short-term incremental benefit at all once measured this way. The lesson is not that paid search never works. It is that a dashboard built on last-click attribution measures correlation and reports it as cause.

The dashboard answers the wrong question

A Google Ads report is a record of clicks that were followed by conversions. When a buyer clicks an ad and later purchases, the conversion is assigned to that click, and the account's return on ad spend is computed from the assigned conversions. The arithmetic is correct. The inference most owners draw from it is not.

The number every account is judged on answers the question, how many sales came after an ad click? The question a budget owner actually needs answered is different: how many sales happened because of the ad that would not have happened otherwise? The first is a matter of counting. The second is a matter of causation, and causation cannot be read off an observational report no matter how granular the report is.

This is the specific failure the eBay experiment was built to expose. It is also why measurement scientists insist that a controlled experiment, not a richer attribution model, is the only instrument that recovers the causal number.

What Blake, Nosko and Tadelis actually did at eBay

In "Consumer Heterogeneity and Paid Search Effectiveness: A Large-Scale Field Experiment," published in Econometrica in 2015, Thomas Blake, Chris Nosko and Steven Tadelis used eBay's own scale to run the test that observational data cannot substitute for. Rather than infer the effect of paid search from historical clicks and conversions, they randomized. For selected keywords and geographies, paid-search ads were switched off for randomly assigned groups of users while remaining on for others, creating a genuine treatment and control.

With that design in place, the counterfactual is observed rather than assumed. The purchases that still occurred in the groups with ads turned off are, by construction, the purchases that did not need the ad. The incremental effect of paid search is the difference between the groups, and nothing about a user's prior intent can contaminate it, because assignment to treatment or control was random.

The authors then compared this experimental estimate against the kind of non-experimental estimate a standard attribution report produces. The observational method credited paid search with far more value than the randomized experiment could find. The divergence was not noise. It was the structural bias that any last-click or click-attributed report carries.

Brand keyword bidding: paying to be found by people already looking for you

The sharpest result concerned brand keywords, the search terms that contain the advertiser's own name. Intuitively these look like cheap, high-converting inventory, and in a dashboard they are: someone searching for "eBay" who clicks an eBay ad and buys will be recorded as a conversion driven by the ad.

The experiment showed this to be an accounting illusion for a well-known brand. Brand-keyword ads produced no measurable incremental short-term benefit once tested against a randomized control. The people clicking those ads were, in the main, people already navigating to the site. The ad intercepted demand that the organic listing directly beneath it would have captured for free. The conversions were real; the incremental conversions caused by the ad were, within the study, not distinguishable from zero.

This finding needs a boundary drawn around it, addressed in full below. eBay is one of the most recognized commerce brands in the world, and the null result is strongest precisely where a brand is already the destination a searcher has in mind. It does not license a blanket claim that every business should abandon brand-keyword ads. It does mean that the value a brand-term campaign reports is the most likely of all paid tactics to be borrowed from demand you already owned.

Why last click attribution bias inflates the number

The mechanism is selection, not fraud. Advertising platforms are extremely good at showing ads to people who are likely to convert. That is the product working as designed. But it means the population that sees and clicks an ad is not a random sample of the market; it is enriched with people who were already close to buying.

When a report attributes their conversions to the click, it conflates two effects that only an experiment can separate. One is persuasion: the ad genuinely moved someone who would not otherwise have bought. The other is selection: the ad was simply present at the moment a committed buyer was going to act anyway. Observational attribution sums the two and prints the total as if it were all persuasion. The eBay experiment removed the selection component by randomization, and a large part of the reported return went with it.

This is why adding more touchpoints does not fix the problem. Multi touch attribution distributes credit across several clicks rather than only the last, which changes how the inflated total is divided but not the fact that the total itself is inflated by selection. A model that only ever sees users who were shown ads cannot, by construction, observe what those users would have done with no ads at all. Only a holdout can.

The gold standard is an experiment, not a better attribution model

The eBay study sits inside a broader consensus in the measurement-science literature: the credible way to know whether an intervention caused an outcome is to run a randomized controlled experiment with a holdout, then check the experiment itself for the errors that quietly invalidate results.

Ron Kohavi, Diane Tang and Ya Xu, drawing on more than twenty thousand controlled experiments run annually at Microsoft, codify this discipline in their 2020 practitioner text. Two of their warnings are directly relevant to anyone tempted to trust a flattering ad report.

A surprising win is usually an error

The authors promote "Twyman's law," the heuristic that any figure that looks surprising or too good is usually wrong, into a working rule. A paid-search report showing an implausibly high return is exactly the kind of surprising number the law targets, and the eBay experiment is what happens when someone finally checks. The professional reflex is not to celebrate the number but to ask what artifact produced it.

Even real experiments fail silently without checks

Trustworthy experimentation is not free just because you randomized. A Sample Ratio Mismatch, where the actual split between treatment and control deviates from the intended ratio, trips a diagnostic check in roughly six percent of experiments at Microsoft scale and invalidates the result regardless of the lift observed. Separately, "peeking" at a running test and stopping when it looks significant inflates the false-positive rate well above its nominal level, above forty percent under aggressive peeking in analyses of production systems. The point for a business owner is that incrementality has to be measured by people who apply these checks, not asserted by a dashboard that applies none.

Reading incremental ROAS and MER

If the platform number overstates impact, what should a budget owner look at instead? Two more reliable instruments, neither of which a single-campaign dashboard shows by default.

The first is incremental ROAS, sometimes written iROAS. Where ordinary return on ad spend divides all attributed revenue by cost, incremental ROAS divides only the revenue the ads actually caused, established by a holdout, by cost. It is almost always a lower number than the platform figure, and it is the more accurate one. The eBay result is, in effect, a demonstration that for some campaigns, brand terms in particular, incremental ROAS can be dramatically below the reported ROAS.

The second is the media efficiency ratio, or MER, which is total revenue over total advertising spend across the whole account rather than per campaign. MER is deliberately blunt. It cannot be gamed by one campaign claiming credit for another's demand, because it never tries to attribute at the campaign level. Watched over time against spend changes, it exposes the situation the eBay study warns about: spend rises, the per-campaign dashboards stay green, and blended revenue barely moves, which is the signature of paying for demand you already had.

Neither instrument is a silver bullet, and reliable measurement usually combines a blended top-line view with periodic holdout tests on the campaigns large enough to justify them. What both instruments share is that they refuse to treat a click that preceded a sale as proof that the click caused the sale.

What this study proves, and what it does not

The eBay experiment is among the strongest causal-inference results available in marketing measurement: a real randomized field experiment at large scale, published in a top-tier peer-reviewed economics journal. Its core finding, that observational attribution inflates paid-search ROI because clicks select for pre-existing intent, is established and has been influential precisely because the method leaves little room to argue with the direction of the bias.

Its limits deserve the same scrutiny. The experiment measured one advertiser, eBay, a brand with enormous unaided awareness, and the starkest null result is for the case where that awareness matters most, brand keywords. A smaller or less-known business bidding on its own name may be defending against competitors' ads or reaching people who genuinely would not have found it otherwise, and the incremental value there is an empirical question for that business, not a foregone conclusion. The study also measured short-term effects; long-run brand consequences of turning ads off are outside its window.

The correct reading is therefore narrow and durable at once. Narrow, because the exact magnitude does not transfer to your account. Durable, because the mechanism does: any report that counts conversions after clicks, without a holdout, will overstate the causal contribution of advertising to people who were already going to convert. The remedy is not to distrust all advertising. It is to measure it the way the evidence says it must be measured.

The evidence

Key findings, with their sources

  • A large-scale randomized field experiment at eBay found that non-experimental, attribution-style estimates of paid-search return are inflated relative to the true causal lift, because ad clicks correlate with pre-existing purchase intent.

    established Blake, T., Nosko, C. & Tadelis, S., "Consumer Heterogeneity and Paid Search Effectiveness: A Large-Scale Field Experiment", Econometrica 83(1), 2015, pp. 155-174 (NBER Working Paper No. 20171).

  • Brand-keyword ads, bidding on the advertiser's own name, showed no measurable incremental short-term benefit once effectiveness was measured experimentally rather than by attribution.

    established Blake, Nosko & Tadelis, Econometrica 83(1), 2015 (randomized experiments at eBay).

  • Trustworthy causal measurement requires a controlled experiment plus a defined checklist, not a p-value alone; the framework is built on more than 20,000 controlled experiments run annually at Microsoft, and treats a surprising result as most likely an error ("Twyman's law").

    established Kohavi, R., Tang, D. & Xu, Y., "Trustworthy Online Controlled Experiments: A Practical Guide to A/B Testing", Cambridge University Press, 2020.

  • A Sample Ratio Mismatch trips a diagnostic check in roughly 6% of controlled experiments at large scale and invalidates the result regardless of the observed lift, so even a randomized test can mislead without integrity checks.

    established Fabijan et al., "Diagnosing Sample Ratio Mismatch in Online Controlled Experiments", KDD '19 (ACM SIGKDD, 2019); Kohavi et al., 2020.

  • Continuously monitoring a fixed-horizon A/B test and stopping when it looks significant ("peeking") can push the realized false-positive rate well above the nominal 5%, above 40% under aggressive peeking in analyses of production systems.

    established Johari, R., Koomen, P., Pekelis, L. & Walsh, D., "Peeking at A/B Tests: Why It Matters, and What to Do About It", KDD '17 (ACM SIGKDD, 2017).

Calibration

What is proven, what is promising, what is unproven

Evidence tierTacticsWhat the evidence says
establishedObservational/last-click attribution inflates paid-search ROI; the causal number requires a randomized holdoutBlake, Nosko & Tadelis, Econometrica 2015 (eBay field experiment); Kohavi, Tang & Xu, 2020
establishedBrand-keyword ads for a highly recognized brand showed no measurable short-term incremental liftBlake, Nosko & Tadelis, Econometrica 2015
contestedWhether the brand-keyword null generalizes to small or low-awareness businesses (competitive defense, genuine reach)Bounded by the single-advertiser, short-term design of the eBay study; an empirical question per account, not settled by it

Reference

Glossary

Incrementality
The share of an outcome caused by an intervention that would not have happened without it. For advertising, the conversions the ads actually produced, over and above what would have occurred with no ads.
Counterfactual
What would have happened in the absence of the intervention. In a randomized experiment it is observed directly in the control group; in an attribution report it is only assumed.
Last-click attribution
A reporting rule that assigns a conversion to the final ad click preceding it. It counts clicks that came before sales and cannot distinguish persuasion from the ad merely being present for a committed buyer.
Brand keyword
A search term containing the advertiser's own name or brand. Ads on these terms often report high conversion rates because searchers already intend to reach that business.
Incremental ROAS (iROAS)
Return on ad spend computed only on the revenue the ads causally produced, established by a holdout, rather than all attributed revenue. Almost always lower than the platform-reported figure.
Media efficiency ratio (MER)
Total revenue divided by total advertising spend across the whole account, rather than per campaign. A blunt, hard-to-game top-line check on whether added spend is actually growing the business.
Randomized field experiment
A test run in a live market where units are randomly assigned to see or not see the intervention, so that differences in outcome can be attributed to it rather than to pre-existing differences between groups.

Straight answers

Frequently asked questions

Does this mean I should turn off my Google Ads?

No. The eBay study does not show that paid search is worthless; it shows that the platform report overstates its causal contribution, most severely for brand keywords at a well-known brand. The practical move is to measure incremental value, through a holdout or a blended view, rather than to trust the attributed number, and then keep the spend that earns real incremental customers and cut the spend that only re-books demand you already owned.

What is the difference between ROAS and incremental ROAS?

Standard return on ad spend divides all revenue attributed to the ads by their cost. Incremental ROAS divides only the revenue the ads actually caused, established by comparing a group shown the ads against a randomly held-out group that was not, by that same cost. Incremental ROAS is almost always the lower and the more accurate number, because the standard figure includes sales that would have happened anyway.

Does the eBay finding mean brand keyword bidding is always a waste?

Not universally. The null result is strongest for eBay, a brand searchers already intend to reach, where the paid ad largely intercepts a click the organic listing would have won for free. A smaller or less-known business may use brand terms to defend against competitors bidding on its name, or to reach people who would not otherwise have found it. Whether that spend is incremental is an empirical question for your account, best answered with a brand-term holdout test rather than assumed either way.

What is MER and why does it matter here?

MER, the media efficiency ratio, is total revenue divided by total advertising spend across the whole account. It matters because it cannot be inflated by one campaign claiming credit for another campaign's demand. Watched over time against changes in spend, it reveals the exact pattern the eBay study warns about: spend rising while blended revenue barely moves, which means the added budget is buying demand you already had.

How can a small business measure paid search incrementality without eBay's scale?

You do not need eBay's scale to run a holdout. Practical options include geographic holdouts, where ads run in some regions and pause in comparable ones, brand-term on/off tests over defined windows, and watching MER against spend changes. The discipline that matters is the same one the measurement literature insists on: define a control that never saw the ads, check the test for basic errors like an uneven split, and resist stopping the moment the result looks good.

Provenance

Sources

  1. Blake, T., Nosko, C. & Tadelis, S., "Consumer Heterogeneity and Paid Search Effectiveness: A Large-Scale Field Experiment", Econometrica, 83(1), 2015, pp. 155-174 (NBER Working Paper No. 20171) (established)
  2. Kohavi, R., Tang, D. & Xu, Y., "Trustworthy Online Controlled Experiments: A Practical Guide to A/B Testing", Cambridge University Press, 2020, ISBN 9781108724265 (established)cambridge.org
  3. Fabijan, A., Gupchup, J., Gupta, S., Omhover, J., Qin, W., Vermeer, L. & Dmitriev, P., "Diagnosing Sample Ratio Mismatch in Online Controlled Experiments", KDD '19, ACM SIGKDD, 2019 (established)
  4. Johari, R., Koomen, P., Pekelis, L. & Walsh, D., "Peeking at A/B Tests: Why It Matters, and What to Do About It", KDD '17, ACM SIGKDD, 2017 (established)

Every figure above is attributed to a real, dated source and tagged with its evidence tier. Where a claim could not be verified to a primary source, it is not stated as fact.

What this means for your ad budget

If your Google Ads report is the only number telling you the ads work, you are looking at the exact figure the eBay experiment showed is inflated. The practical question is not whether to advertise, it is how much of your reported return is real incremental business and how much is demand you already owned, especially on branded search. A Paid Media Diagnostic reads every dollar already spent, separates genuine incremental value from last-click credit, and hands you a ranked list of fixes with the money each one recovers, before you commit another dollar or any management retainer.

diagnostic Paid Media Diagnostic A fixed-scope, specialist read of your existing ad spend that recalculates real return past last click using blended measurement and incremental ROAS, then sizes and sequences the leaks. Diagnosis, not management: no bids are changed, and your ad spend is never marked up. See how it works

Start free with a Machine-Readiness Score, a specialist-reviewed read of where you stand across search, AI answers, the map pack, and reputation. No guaranteed number, and no obligation.