Measurement & Honesty · established evidence
Why "Half My Advertising Is Wasted" Is a Myth Marketers Still Believe
"Half my advertising is wasted; I just don't know which half." The line is almost always pinned on the department-store merchant John Wanamaker, and it has become the marketing industry's founding parable about measurement. There is a problem: no one can verify he ever said it. The earliest documented match is a secondhand 1919 account, and the same sentiment has separately been attributed to William Lever and William Wrigley. So the quote is very likely a myth. The frustration underneath it is not. For most of the last century the answer to which half is working was genuinely unknowable, because you can measure what you spent but not the customers you would have won anyway. That has changed. Randomized field experiments and modern causal methods can now estimate real advertising lift, and they routinely disagree with the observational reports most businesses still trust. This piece separates the apocryphal quote from the measurable problem.
The quote almost no one actually said
Start with the sentence itself, because its own history is the first lesson. When a researcher at Quote Investigator traced "one-half the money I spend for advertising is wasted, and I do not know which half," the trail did not lead cleanly back to Wanamaker at all. The earliest solid match is a secondhand 1919 account, published years after the words were supposedly spoken, and the identical sentiment has been credited elsewhere to the soap magnate William Lever and the chewing-gum magnate William Wrigley.
The attribution, in other words, is unverified. That is a fitting origin for the industry's favorite line about measurement: the parable everyone repeats to prove that advertising cannot be measured is itself a claim no one has been able to verify. It is repeated because it feels true, which is exactly the failure mode it warns against.
The real problem the parable named
Debunking the quote is not the same as dismissing it. Whoever first said it, the sentence endured for a century because it named a genuine and hard condition. An advertiser can always see what was spent. What no invoice reveals is the counterfactual: how many of the customers who arrived after a campaign would have arrived without it. Sales that follow advertising are not the same as sales caused by advertising, and for most of the last hundred years there was no rigorous way to separate the two.
This is why the reports businesses read can be so misleading. A dashboard that credits a channel with every conversion that touched it is measuring correlation, not cause. It answers "which sales came after this ad" when the owner is really asking "which sales would not have happened without it." Those are different questions, and the gap between them is precisely the half Wanamaker's ghost could not find.
What multi-touch attribution gets wrong when you test it against reality
The strongest evidence that this gap is real, and not a rhetorical flourish, comes from putting standard attribution to the test against a randomized experiment. In 2019 Brett Gordon, Florian Zettelmeyer, Neha Bhargava and Dan Chapsky published exactly that test in Marketing Science. Working with fifteen large-scale randomized field experiments run at Facebook, spanning more than 500 million user-experiment observations and over 1.6 billion ad impressions, they had a rare thing: the true causal lift of the advertising, established by random assignment.
They then asked what the usual observational methods, the family of techniques that underpins most multi-touch attribution and lookalike modeling, would have concluded from the same data. The answer was uncomfortable. Even after conditioning on rich demographic and behavioral covariates, the observational estimates frequently came out in the wrong direction or of the wrong magnitude compared with the experimental ground truth. The convenient number and the true number were often not close.
The implication for an owner reading a platform dashboard is direct. Multi-touch attribution is not a neutral window onto what worked; it is a model whose assumptions can fail quietly, and when a randomized experiment is available to check it, it often does. The number is not wasted because advertising does not work. It is unreliable because the method used to score it can be wrong.
Why the old tracking got shakier before it got better
The measurement problem did not stand still. As individual-level tracking eroded, the ground under conventional attribution shifted too. Apple's App Tracking Transparency, introduced with iOS 14.5 in April 2021, required apps to ask permission before tracking users across other apps and sites. Industry reports put the opt-out rate high, with figures around 75 percent of iOS users declining, and pixel-based revenue attribution reportedly falling from something like 80 to 95 percent capture down to 60 to 70 percent.
Those specific percentages come from ad-tech vendors with a commercial interest in the story, so they are best read as directional rather than exact. The direction, though, is not in dispute: individual-level attribution got less reliable at the same moment more businesses were leaning on it. That pressure is what pushed the measurement conversation back toward aggregate, experimental methods that never depended on following individuals in the first place.
Marketing mix modeling, geo-experiments, and what actually works
Here the story turns constructive. The tools that can actually answer Wanamaker's question exist today, and they do not require a surveillance pixel on every buyer.
The older of the two is marketing mix modeling. Rather than following individuals, it treats the problem as statistical causal inference: it regresses a sales time series on the marketing time series, with adjustments for the lagged carryover and diminishing returns of spend. Because it never touches user-level identifiers, it survived the collapse of cookie and device tracking largely intact. Its credibility is reinforced by an unusual fact: both Google and Meta have open-sourced their internal methodologies, Google's Meridian and Meta's Robyn, which is a strong signal that the two largest ad platforms now treat mix modeling as the serious fallback once individual tracking degrades.
The sharper instrument is the geo-experiment. Instead of holding out individual users, it holds out whole matched geographic markets and uses a synthetic-control method to estimate what the treated markets would have done untouched. Meta's open-source GeoLift library is a common implementation; in one independent simulation study its coverage landed at 92 to 95 percent, closest to the 95 percent target, with a false-positive rate of 3 to 5 percent, the lowest among the open-source geo-testing tools compared. Those comparative numbers come from a vendor-run simulation rather than a peer-reviewed paper, so treat them as industry-grade evidence, not settled science. The underlying method, measuring incremental lift without a user-level pixel, is well established and used across the industry.
The scale floor: why one local business cannot just run the experiment
There is a catch here, and skipping it would repeat the original sin. Rigorous measurement has a scale floor, and most owner-operated local businesses fall beneath it. This is a structural fact, not a failure of effort.
Standard power calculations, the arithmetic that decides how much data a valid test needs, typically demand tens to hundreds of thousands of observations to detect the modest effect sizes that matter, on the order of 50,000 to 500,000 users for a conventional test. At the traffic a single-location business realistically sees, roughly a thousand visitors a month, detecting even a 20 percent relative lift can take more than seven months, and a 10 percent lift more than two and a half years. By the time such a test could conclude, the campaign, the season, and often the business itself have already changed.
This is the structural, non-moral reason local owners have so often been sold vanity metrics: the clean experiment that works for a platform is genuinely out of reach for one small firm. The resolution is not a shortcut but aggregation, pooling data across many similar businesses to recover the statistical power no single one of them has, and treating any single-client claim skeptically until pooled evidence backs it. That is an empirical-generalization posture, borrowed from marketing science, applied to local visibility.
The industry has corrected itself before
The pattern in this article is older than any one metric. The industry repeatedly adopts a convenient, computable number, treats it as truth for years, and only corrects course once someone runs the harder experiment or the field disciplines itself.
The public-relations industry did exactly that when it renounced Advertising Value Equivalency. Its own voluntary standard, the Barcelona Principles, revised most recently in 2020, rejects AVE outright in favor of outcome-based measurement that is transparent, consistent, and valid. Regulators reinforce the same line: the US Federal Trade Commission's Endorsement Guides require that testimonials reflect honest experience and that any material connection be clearly disclosed, a live framework with an active enforcement record. And the academic literature is candid about its own limits, with a documented, long-standing shortage of replication studies in marketing that should make any confident "this always works" claim suspect.
The same discipline applies to any new metric, including a disciplined visibility score: state the method, disclose the limits, and treat each reading as another step in an ongoing correction, not a final answer that ends the argument.
The new Wanamaker problem: the AI answer
There is a fresh version of the old frustration, arriving just as the classic one becomes tractable. Buyers increasingly ask an answer engine, ChatGPT, Perplexity, Gemini, Google AI Overviews, and act on the synthesized reply. Whether and why one of those engines names a given business is a new measurement problem being invented almost from scratch.
It is genuinely unsettled. There is no standardized, agreed methodology for measuring "share of answer." Generative engines are non-deterministic, so the same question can return different answers across runs; they personalize; and they are not fully observable from outside. Any "AI visibility" number is therefore a sample-based estimate whose reliability depends entirely on the sampling method behind it. The right response is transparency about that immaturity, not a claim that it is already solved.
So the Wanamaker question returns in a new form. Buyers now ask whether you appear at all in the answer they act on, a more basic question than which half of the spend works. The first step is the same as it always was. Measure where you actually stand, disclose how you measured it, and skip dressing a convenient number up as certainty.
The evidence
Key findings, with their sources
-
Across 15 randomized field experiments (500M+ user-experiment observations, 1.6B ad impressions), standard observational attribution methods frequently produced lift estimates in the wrong direction or of the wrong magnitude versus the experimental ground truth, even after conditioning on rich covariates.
established Gordon, Zettelmeyer, Bhargava & Chapsky, "A Comparison of Approaches to Advertising Measurement: Evidence from Big Field Experiments at Facebook", Marketing Science 38(2):193-225, 2019.
-
The "half my advertising is wasted" line has no verified original source; the earliest documented match is a secondhand 1919 account, and the same sentiment is also attributed to William Lever and William Wrigley.
established Quote Investigator, "One-Half the Money I Spend for Advertising Is Wasted, But I Have Never Been Able To Decide Which Half", 2022.
-
At roughly 1,000 visitors per month, detecting a 20% relative lift in a conventional A/B test can take 7+ months and a 10% lift 31+ months; standard tests typically require 50,000 to 500,000 users.
established Industry synthesis of standard power-analysis methodology (Analytics-Toolkit.com; Statsig, "Power Analysis for A/B Testing"); the underlying power math is not in dispute, the traffic figures are industry illustrations.
-
After iOS 14.5 App Tracking Transparency (April 2021), industry reports put the opt-out rate near 75% of iOS users and pixel-based revenue attribution falling from 80-95% capture to 60-70%.
emerging Ad-tech vendor reports (AppsFlyer opt-in study; PubMatic ad-spend shift data), summarized in marketing-industry press; vendor-sourced, treat as directional.
-
In an independent simulation study, Meta's open-source GeoLift showed coverage of 92-95% (closest to the 95% target) and a false-positive rate of 3-5%, lowest among the open-source geo-testing tools compared.
emerging facebookincubator/GeoLift (GitHub); Recast Research, "Open-Source Geo-Experiment Tools: A Head-to-Head Simulation Study"; vendor-run simulation, not peer-reviewed.
-
The Barcelona Principles 3.0 (2020) are the PR/communications industry's own ratified standard rejecting Advertising Value Equivalency in favor of transparent, outcome-based measurement.
established AMEC, Barcelona Principles 3.0, 2020, amecorg.com.
-
The FTC Endorsement Guides (16 CFR Part 255) require honest testimonials and clear disclosure of any material connection; the FTC's 2025 enforcement report cites 150+ actions on deceptive endorsement practices.
established Federal Trade Commission, 16 CFR Part 255, "Guides Concerning the Use of Endorsements and Testimonials in Advertising".
Calibration
What is proven, what is promising, what is unproven
| Evidence tier | Tactics | What the evidence says |
|---|---|---|
| established | The Wanamaker "wasted half" quote is apocryphal | Quote Investigator (2022): earliest match is a secondhand 1919 account; also attributed to Lever and Wrigley. |
| established | Observational attribution diverges from randomized-experiment ground truth | Gordon, Zettelmeyer, Bhargava & Chapsky (2019), Marketing Science, 15 Facebook field experiments. |
| established | Marketing mix models and geo-experiments measure lift without user-level tracking | Standard MMM literature; Google Meridian and Meta Robyn (open-sourced); Meta GeoLift synthetic-control library. |
| emerging | How well those rigorous methods perform at single-location MSME scale | Power-analysis math shows a structural scale floor; GeoLift comparative numbers are from a vendor simulation, not a peer-reviewed study. |
| emerging | A standardized way to measure AI-answer visibility (share of answer) | GEO founding paper (Aggarwal et al., 2024) exists, but engines are non-deterministic and there is no agreed sampling standard yet. |
Reference
Glossary
- The Wanamaker line
- The saying "half my advertising is wasted; I just don't know which half," popularly credited to John Wanamaker but with no verified original source.
- Counterfactual
- What would have happened without the advertising. Measuring true effect means comparing outcomes against this unseen baseline, not just counting sales that followed a campaign.
- Observational attribution
- Estimating advertising effect from data on who saw ads and who converted, without random assignment. Multi-touch attribution belongs to this family, and it can misstate true lift.
- Marketing mix modeling
- A causal-inference method that regresses a sales time series on marketing time series (with carryover and saturation adjustments) instead of tracking individuals.
- Geo-experiment
- A test that holds out matched geographic markets rather than individual users and uses a synthetic-control estimate to measure incremental lift without a user-level pixel.
- Advertising Value Equivalency (AVE)
- A discredited PR metric that priced earned coverage as if it were paid advertising space, formally rejected by the Barcelona Principles.
- How often a business is named inside AI answers to buyer questions. There is no standardized methodology for measuring it yet, so any figure is a sample-based estimate.
Straight answers
Frequently asked questions
Did John Wanamaker actually say "half my advertising is wasted"?
There is no verified evidence that he did. Quote Investigator traced the line to a secondhand 1919 account and found the same sentiment attributed to William Lever and William Wrigley as well. The attribution is unverified, which makes it a myth repeated because it feels true, not because it is sourced.
If the quote is a myth, is the problem it describes real?
Yes. The underlying difficulty is genuine: you can see what you spent, but not the customers you would have won without spending. Separating sales that followed advertising from sales caused by advertising is a real, hard problem, and for most of the last century there was no rigorous way to do it.
Can you actually measure which advertising is working now?
Better than ever, using randomized field experiments and modern causal methods. Gordon and colleagues (2019) showed that experiments and standard observational attribution often disagree, which is why the reliable tools are marketing mix modeling and geo-experiments rather than a dashboard that credits every touch.
Can a single local business run these experiments?
Usually not on its own. Statistical power requirements put valid tests out of reach at typical local traffic; detecting a modest lift can take many months or years. The practical path is aggregation, pooling many similar businesses to recover the power no single one has, and being skeptical of any single-client claim until pooled evidence supports it.
Does this apply to AI answers too?
It is the newest version of the same problem. There is no agreed method for measuring "share of answer," and generative engines are non-deterministic, so any AI-visibility number is a sample-based estimate that is only as good as its disclosed sampling method. The correct posture is transparency about that immaturity, not false certainty.
Provenance
Sources
- Gordon, B.R., Zettelmeyer, F., Bhargava, N., Chapsky, D., "A Comparison of Approaches to Advertising Measurement: Evidence from Big Field Experiments at Facebook", Marketing Science 38(2):193-225, 2019 (established)
- Quote Investigator, "One-Half the Money I Spend for Advertising Is Wasted, But I Have Never Been Able To Decide Which Half", 2022 (established: attribution is unverified)quoteinvestigator.com
- Kohavi, R., Tang, D., Xu, Y., Trustworthy Online Controlled Experiments: A Practical Guide to A/B Testing, Cambridge University Press, 2020 (established)
- Industry synthesis of power-analysis methodology (Analytics-Toolkit.com; Statsig, "Power Analysis for A/B Testing") (established math, industry-illustrative figures)
- Wikipedia, "Marketing mix modeling", summarizing the standard MMM literature (established)en.wikipedia.org
- Google, Meridian open-source marketing mix model; Meta, facebookexperimental/Robyn; Tueller et al., "Packaging Up Media Mix Modeling", arXiv:2403.14674, 2024 (established that both exist and are open-sourced; emerging as to single-location fit)arxiv.org
- facebookincubator/GeoLift (GitHub); Recast Research, "Open-Source Geo-Experiment Tools: A Head-to-Head Simulation Study" (established method; industry-grade, not peer-reviewed, comparative numbers)github.com
- Ad-tech vendor reports on iOS 14.5 App Tracking Transparency (AppsFlyer opt-in study; PubMatic ad-spend shift data), 2021 onward (established policy event; directional, vendor-sourced magnitudes)
- AMEC, Barcelona Principles 3.0, 2020 (established)amecorg.com
- Federal Trade Commission, 16 CFR Part 255, "Guides Concerning the Use of Endorsements and Testimonials in Advertising" (established)ecfr.gov
- Evanschitzky, H., Baumgarth, C., Hubbard, R., Armstrong, J.S., "Replication Research in Marketing Revisited: A Note on a Disturbing Trend", Journal of Business Research, 2007 (established)
- Aggarwal, P. et al., "GEO: Generative Engine Optimization", arXiv:2311.09735, KDD 2024 (emerging)arxiv.org
Every figure above is attributed to a real, dated source and tagged with its evidence tier. Where a claim could not be verified to a primary source, it is not stated as fact.