Demand & Paid Media · established evidence
Ghost Ads and the Hidden Cost of Knowing What Your Advertising Actually Does
Ghost ads are a measurement technique, not an ad format. They answer the question every advertiser should ask and few can: of everything an ad campaign appeared to achieve, how much did the ads actually cause. The old way of answering it was to run a controlled experiment in which a randomly chosen control group was shown a placebo, usually a public-service announcement, so its behavior could be compared with the treated group. That worked, but it forced the advertiser to buy worthless impressions at real cost. The ghost-ad method removed that cost by recording the ad each control user would have been shown, the counterfactual impression, without paying to display anything. The result was causal measurement at a fraction of the price, which is precisely why it matters for a business spending a few thousand dollars a month rather than a few million.
The question a dashboard cannot answer
An advertising platform reports what happened while its ads were running. It records impressions, clicks, and the sales that followed, and it attributes those sales to itself. What it does not report, because it structurally cannot, is what would have happened if the ads had never run. That difference, between the world with the ad and the world without it, is the only thing that tells you whether the money did any work. Researchers call it incremental effect, or incrementality.
The gap between the two is not academic hair-splitting. Retargeting and branded search look strong on a platform report for a specific reason: they reach people who were already heading toward a purchase. The ad appears, the sale follows, and the platform claims the credit, even though the customer had largely decided already. To know what the advertising truly moved, you have to compare buyers who saw the ad against otherwise identical buyers who did not. Constructing that comparison is the entire problem, and it is harder than it sounds.
Why you cannot simply compare the exposed and the unexposed
The intuitive shortcut is to compare people who saw an ad with people who did not, and treat the difference as the ad's effect. This is wrong, and it is wrong in a predictable direction. The people an ad system chooses to show ads to, and the people who happen to be online and active enough to be shown them, are systematically different from everyone else, before any ad is served.
Lewis, Rao, and Reiley demonstrated this across three controlled experiments in 2011, naming the phenomenon activity bias: a person browsing the web at a given moment is, at that same moment, more likely to search, click, and buy, ad or no ad, simply because they are active. Because ad exposure correlates with that underlying activity, observational estimates of ad effectiveness are biased upward, sometimes enough to manufacture a large apparent effect where the true effect is near zero. The comparison group has to be built by randomization, not by who happened to see the ad.
The old fix worked, and it was expensive
The clean solution borrowed from clinical trials: randomize users into a treatment group that receives the real ad and a control group that receives a placebo, then compare outcomes. Because assignment is random, the two groups are statistically identical apart from the ad, so any difference in purchases is caused by the ad. In online advertising the placebo was usually a public-service announcement, a charity or awareness message shown in the same slot the real ad would have occupied.
The method is sound but it carries a real bill. To run it, the advertiser buys and serves impressions to the control group that generate no commercial value, purely so the experiment has a comparison. On top of that direct waste, a public-service-announcement design fights the ad system's own optimization: modern delivery is a real-time auction that constantly re-targets and re-prices, and forcing a static placebo into that machinery is both costly and technically awkward. For a large advertiser, the price of certainty was tolerable. For a small one, it was often prohibitive, which meant the businesses least able to waste budget were the ones most locked out of knowing whether their budget worked.
The ghost-ad insight: record the impression you would have shown
In 2017, Johnson, Lewis, and Nubbemeyer proposed a change that sounds small and is not. Instead of buying placebo impressions for the control group, the ad system simply records, for each control user, the moment at which it would have shown the advertiser's ad, and logs that counterfactual impression while displaying whatever ad wins the auction for a different advertiser instead. No wasted spend, no placebo to purchase, and the comparison group is defined by exactly the same targeting and auction logic that selected the treated group.
This matters for two reasons the authors make explicit. First, cost: recording a would-be impression is close to free, so the dominant expense of the old placebo design disappears. Second, fidelity: because the ghost impression is generated by the live delivery system rather than bolted on beside it, the method works natively with real-time optimization instead of against it, and it identifies a cleaner treatment and control population than intent-to-treat holdout designs that hold out whole audiences in advance.
What "counterfactual" means here in plain terms
A counterfactual is simply the road not taken, recorded. The ad system knows, at auction time, that this particular control user is one it would have served the advertiser's ad to. It notes that fact, then serves something else. Later, the advertiser compares purchases among the users who actually saw the ad against purchases among the users who would have seen it but did not. The would-have-seen group is the control, and building it costs almost nothing.
The evidence that it works
The ghost-ad paper did not only propose the design, it demonstrated it. Applied to a retargeting campaign, the method measured a lift of 17.2 percent in site visits and 10.5 percent in purchases attributable to the ads, results obtained without the cost structure of a public-service-announcement holdout. The contribution is twofold: a credible causal estimate of what the campaign actually caused, and a demonstration that such an estimate can be produced cheaply enough to be routine rather than exceptional.
The significance is less any single percentage than the economics behind it. Once causal measurement stops requiring a bespoke, budget-burning experiment, it becomes something an ordinary advertiser can afford to run on cadence, which is the difference between measuring your advertising once as a curiosity and measuring it continuously as a discipline.
What the same literature says about paid search
Ghost ads sit inside a decade of field experiments that all point the same way: platform-reported returns tend to overstate causal contribution, and the size of the overstatement grows with how correlated the ad exposure is with intent the buyer already had. Blake, Nosko, and Tadelis found, in a large randomized experiment at eBay, that paid search ads on the company's own branded keywords produced no measurable short-term incremental benefit, and that even for non-brand terms the returns were dragged negative by frequent users who would have purchased regardless. The brand-defense line item that looks efficient on a dashboard may be paying for clicks that cost nothing to win organically.
For advertisers who cannot instrument user-level ghost ads, a complementary tool exists. Vaver and Koehler's geo-experiment methodology randomizes non-overlapping geographic regions into treatment and control ad conditions, producing a causal read of lift without any individual-level tracking. It is the accessible cousin of the ghost-ad idea: same logic of a randomized counterfactual, implemented at the level of markets rather than users, and explicitly designed to inform bidding and budgeting decisions.
What the evidence supports: a cheaper method, not a free lunch
Two caveats are worth stating plainly. First, the ghost-ad technique in its original form depends on the ad platform's cooperation, because only the delivery system can log the counterfactual impression at auction time. An independent advertiser cannot always run the pure method unaided, which is part of why geo experiments and holdout designs remain the practical instruments for many smaller campaigns.
Second, the widely repeated industry claim that measured incremental return runs some 30 to 70 percent below platform-reported return should be treated as directional, not settled. That range comes from vendor and practitioner analyses rather than peer-reviewed replication. What is established is the mechanism the range describes: activity bias and heterogeneous consumer response are real, documented, and point consistently toward overstatement. The specific number for any one business has to be measured, not assumed, which is the entire reason a measurement practice exists at all.
The through-line from all of this is not that advertising fails. It is that what an ad caused and what a platform reports are different quantities, that the difference is knowable, and that knowing it is now affordable in a way it was not a decade ago. For a business spending real money every month, that shift is the difference between steering by a number designed to flatter the channel and steering by a number designed to be true.
The evidence
Key findings, with their sources
-
Applied to a retargeting campaign, the ghost-ad method measured a 17.2% lift in site visits and a 10.5% lift in purchases attributable to the ads.
established Johnson, Lewis & Nubbemeyer, "Ghost Ads: Improving the Economics of Measuring Online Ad Effectiveness", Journal of Marketing Research, 54(6), 2017.
-
Recording counterfactual (would-be) impressions lets advertisers measure incrementality at a fraction of the cost of public-service-announcement holdout experiments, while working natively with real-time ad delivery.
established Johnson, Lewis & Nubbemeyer, "Ghost Ads", Journal of Marketing Research, 54(6), 2017.
-
Observational estimates of ad effectiveness are systematically biased upward by activity bias, because a user active enough to be shown an ad is, at that moment, more likely to search, click, and buy regardless of the ad.
established Lewis, Rao & Reiley, "Here, There, and Everywhere: Correlated Online Behaviors Can Lead to Overestimates of the Effects of Advertising", WWW '11, 2011.
-
In a large-scale randomized field experiment, paid search ads on branded keywords produced no measurable short-term incremental benefit, and non-brand returns were dragged negative by frequent users who would have purchased anyway.
established Blake, Nosko & Tadelis, "Consumer Heterogeneity and Paid Search Effectiveness: A Large-Scale Field Experiment", Econometrica, 83(1), 2015.
-
Geo experiments, which randomize non-overlapping regions into treatment and control ad conditions, provide a causal measure of ad lift without individual-level tracking.
established Vaver & Koehler, "Measuring Ad Effectiveness Using Geo Experiments", Google Inc., 2011.
-
Industry analyses report that measured incremental ROAS often runs roughly 30 to 70 percent below platform-reported ROAS, with branded search cited as the largest divergence.
contested Practitioner and vendor analyses (Prescient AI, Eightx, MHI Growth Engine, layerfive.com), 2025-2026, directionally consistent with the peer-reviewed field-experiment literature but not independently replicated.
Calibration
What is proven, what is promising, what is unproven
| Evidence tier | Tactics | What the evidence says |
|---|---|---|
| established | Ghost-ad counterfactual measurement; PSA/holdout randomized experiments; geo experiments; activity-bias correction; branded-search incrementality caution | Peer-reviewed field experiments and methodology papers (Johnson-Lewis-Nubbemeyer 2017; Lewis-Rao-Reiley 2011; Blake-Nosko-Tadelis 2015; Vaver-Koehler 2011). |
| contested | Citing a specific "iROAS runs 30-70% below reported ROAS" figure as fact | Vendor and practitioner blog synthesis (2025-2026); mechanism is well-evidenced, the specific range is not peer-reviewed and needs primary data. |
Reference
Glossary
- Incrementality
- The portion of an outcome (visits, purchases) that an ad actually caused, measured as the difference between the world with the ad and the world without it, rather than everything that happened while the ad ran.
- Ghost ad
- A measurement technique in which the ad system records the impression a control user would have been shown, the counterfactual, without paying to display it, so a clean comparison group can be built at almost no cost.
- Counterfactual
- The road not taken, recorded. In ad measurement, the ad a control user would have received had they been in the treated group, logged so their behavior can serve as an honest baseline.
- PSA holdout
- The older randomized design in which a control group is shown a placebo, typically a public-service announcement, in the ad slot, so its behavior can be compared with users who saw the real ad. Effective but costly, because the placebo impressions are bought and serve no commercial purpose.
- Activity bias
- The upward distortion in observational ad estimates caused by the fact that users active enough to be served ads are, at that moment, already more likely to search, click, and buy.
- Geo experiment
- A causal measurement design that randomizes non-overlapping geographic regions into ad and no-ad conditions, measuring lift at the level of markets rather than individuals and requiring no user-level tracking.
Straight answers
Frequently asked questions
What are ghost ads?
Ghost ads are a way to measure what an advertising campaign actually caused. Instead of buying placebo impressions for a control group, the ad system records the impression each control user would have been shown, the counterfactual, and compares the two groups. It produces a causal estimate of lift at a fraction of the cost of a traditional holdout experiment.
How is a ghost ad different from a normal A/B test or holdout?
A classic holdout shows a control group a placebo ad, which costs money to serve and fights the platform's real-time optimization. A ghost ad shows the control group nothing extra; it merely logs the impression they would have received. The comparison is defined by the same targeting and auction that selected the treated group, so it is both cheaper and cleaner.
Why does the platform-reported ROAS overstate what my ads did?
Because a platform records what happened while its ads ran and attributes nearby sales to itself, including sales the customer had already decided to make. Field experiments show this overstatement is real and largest for retargeting and branded search, where the ad reaches people who were already coming to you. The figure that matters is what the ad caused, not what happened alongside it.
Can a small business actually measure incrementality affordably?
Yes, in principle, which is the point of the ghost-ad economics. The pure ghost-ad method usually needs platform cooperation, but geo experiments apply the same randomized-counterfactual logic at the level of markets and need no individual-level tracking, making causal measurement feasible on a modest budget. The specific gap for any one account has to be measured, not assumed.
Is the "30 to 70 percent below reported ROAS" figure reliable?
Treat it as directional. The mechanism behind it, activity bias and heterogeneous consumer response, is well documented in peer-reviewed research. The specific percentage range comes from vendor and practitioner analyses rather than independent replication, so it is a reasonable prior to test, not a fact to quote as your own number.
Provenance
Sources
- Johnson, G. A., Lewis, R. A. & Nubbemeyer, E. I., "Ghost Ads: Improving the Economics of Measuring Online Ad Effectiveness", Journal of Marketing Research, 54(6), 867-884, 2017 (established)doi.org
- Lewis, R. A., Rao, J. M. & Reiley, D. H., "Here, There, and Everywhere: Correlated Online Behaviors Can Lead to Overestimates of the Effects of Advertising", Proceedings of WWW '11, 2011 (established)doi.org
- Blake, T., Nosko, C. & Tadelis, S., "Consumer Heterogeneity and Paid Search Effectiveness: A Large-Scale Field Experiment", Econometrica, 83(1), 155-174, 2015 (established)doi.org
- Vaver, J. & Koehler, J., "Measuring Ad Effectiveness Using Geo Experiments", Google Inc., 2011 (established)research.google
- Practitioner/vendor incrementality analyses (Prescient AI, Eightx, MHI Growth Engine, layerfive.com), 2025-2026 (contested, industry-reported, not independently replicated)
Every figure above is attributed to a real, dated source and tagged with its evidence tier. Where a claim could not be verified to a primary source, it is not stated as fact.