Measurement & Honesty · established evidence
Does Marketing Science Replicate? A Close Look at the Evidence Behind "What Works"
Does marketing science replicate? Partly, and knowing which part is the whole discipline. Marketing has a documented, decades-old replication gap: it publishes far fewer replication studies than the natural sciences, so many of its published findings have simply never been retested. Worse, a large share of popular marketing advice borrows from social psychology, where several celebrated effects, priming among them, failed to reproduce when researchers ran them again at scale. Yet the picture is not uniform. A handful of marketing regularities, such as the double jeopardy law and the buyer-acquisition patterns behind "How Brands Grow," are among the most replicated findings in the field. The conclusion is not that marketing knows nothing. It is that "what works" is a mix of durable laws and untested folklore wearing the same confident voice, and buyers deserve to know which is which before they pay for it.
The replication gap marketing rarely mentions about itself
Replication, rerunning a study to see whether its result holds, is how a science tells durable knowledge from a lucky first result. It is unglamorous and it is the load-bearing wall of the whole enterprise. A finding that no one has ever reproduced is not yet knowledge; it is a hypothesis that got published.
By that standard marketing has a problem its own scholars named long ago. In a 2007 analysis in the Journal of Business Research, Evanschitzky, Baumgarth, Hubbard and Armstrong revisited the state of replication in the discipline and documented what they called a disturbing trend: marketing publishes markedly fewer replication studies than the natural sciences, and the share was not improving over time. The consequence is structural. If a field rewards novel findings and rarely funds or publishes the boring work of retesting them, most of its accepted results are, quite literally, untested past their first appearance.
This is not an accusation of fraud. It is a description of an incentive system. The point for a business owner is narrower and more useful: when a marketer cites a study to justify a tactic, the relevant question is "has anyone ever reproduced it," not just "is the study real," and for a large fraction of marketing claims the answer is no.
The replication crisis, and why priming is the cautionary tale
Marketing does not live alone. A great deal of its persuasion advice, especially the neuromarketing and behavioral-science genre, is imported from academic social psychology, and that field spent the last decade in a public reckoning now known as the replication crisis. Large coordinated efforts to rerun well-known psychology experiments found that a substantial portion did not reproduce, or reproduced with effects far smaller than the headline studies claimed.
The instructive case is priming, the idea that a subtle cue can meaningfully shift later behavior without a person noticing. Priming effects have been invoked constantly in marketing-psychology claims, and several prominent ones suffered high-profile replication failures, with real-world effect sizes proving far less reliable than the original lab findings suggested. That does not mean every priming claim is false. It means the category earned its skepticism, and any confident marketing tactic built on a single striking priming study is standing on ground the psychologists themselves stopped trusting.
The lesson transfers cleanly. When a pitch leans on "studies show that people subconsciously..." the correct reflex is to ask which studies, how large, and whether they held up when someone tried again. Often the trail ends at one memorable experiment that never survived a serious retest.
What does survive replication: the laws that keep holding
The record does not stop at the failures. Some marketing findings are among the most replicated regularities in all of social science, and they deserve to be trusted precisely because they were retested across categories and decades.
The double jeopardy law
Brands with smaller market share suffer twice over: they have fewer buyers, and those buyers are, on average, slightly less loyal. First observed by McPhee in 1963 and generalized to brand purchasing by Ehrenberg and colleagues, the double jeopardy pattern has been replicated across packaged goods, banking, insurance and newer categories. It is a genuine empirical law, not a slogan. Its practical meaning for a small local business is oddly comforting: when a larger competitor appears to be both more chosen and more loved, that gap is a predictable statistical consequence of size, not proof that you are doing something wrong.
Mental and physical availability
The empirical program behind Byron Sharp's How Brands Grow and the Ehrenberg-Bass Institute argues that brands grow mainly by being easy to bring to mind in a buying moment and easy to find and buy, and mainly by acquiring more buyers rather than deepening loyalty. These are framed by the Institute as patterns replicated across many categories over decades, and within the marketing-science literature they are treated as established. The caveat worth stating plainly: extending these laws to AI-mediated, answer-engine discovery specifically is new and untested, so the principle is durable while the application to being cited by ChatGPT or Google AI Overviews is an inference we should hold loosely.
What happens when a plausible metric meets a real experiment
The clearest way to see the difference between folklore and finding is to take a measurement everyone trusts and test it against an experiment. Gordon, Zettelmeyer, Bhargava and Chapsky did exactly that with advertising attribution. Using 15 large-scale randomized field experiments at Facebook, drawn from more than 500 million user-experiment observations and 1.6 billion ad impressions, they compared the true causal lift an ad produced against what standard observational attribution methods, the machinery behind most multi-touch attribution, would have estimated from the same data.
The observational methods frequently got it wrong, sometimes in the wrong direction or the wrong magnitude, even after controlling for rich demographic and behavioral data. In other words, a widely sold, computationally convenient way of measuring "what works" did not survive contact with a randomized experiment. This is the same failure mode as the replication gap, just at the level of a metric rather than a paper: a plausible number gets adopted at scale, gets treated as truth, and only gets corrected when someone runs the harder test.
The industry has a habit of retiring its own bad metrics
There is a hopeful side to this story, and it is the tradition of self-correction. Marketing and its adjacent industries do, periodically, admit a favorite metric was never valid and vote to retire it.
The public-relations industry did this openly. For years the standard way to value coverage was Advertising Value Equivalency, a made-up number that priced editorial coverage as though it were paid advertising. In 2010, and again in the revised Barcelona Principles of 2015 and 2020, the industry's own measurement body formally rejected AVE in favor of outcome-based measurement, calling instead for metrics that are transparent, consistent and valid. An entire field looked at its flagship number and called it a vanity metric to its face.
Even the industry's founding parable about measurement turns out to be unverified. The famous line, "half my advertising is wasted, I just don't know which half," is popularly pinned on the retailer John Wanamaker, yet the earliest documented match is a secondhand 1919 account and the same sentiment has been attributed to several other men. The one story marketing tells about the honesty of measurement cannot itself be sourced. That is not a reason for cynicism. It is a reason to treat confident attribution, of quotes and of results, as a claim to be checked rather than a fact to be repeated.
How to read a marketing claim like a scientist
None of this requires a statistics degree. It requires a short, repeatable set of questions that separate a durable finding from a confident guess.
- Ask whether it has ever been replicated. A single striking study is a starting point, not a conclusion. The findings you can lean on, double jeopardy, buyer-acquisition patterns, were retested across many categories and decades.
- Weigh the source data against the claim. A pattern derived from hundreds of large national brands may simply not transfer to a single-location business, and an honest presenter says so rather than smoothing it over.
- Distrust the round, memorable number. When a figure like "accounts for 83% of the variance" comes from one industry analysis of a few dozen cases, it is a single result, not a settled law, and should be labeled that way.
- Prefer experiments to observation. A randomized or geo-based test that withholds treatment from a control tells you what an intervention caused; a dashboard that correlates spend with outcomes usually cannot.
- Watch for the vanity metric. If a number is easy to compute and always goes up, be suspicious. The metrics worth acting on are the ones that can also deliver bad news.
What this means for buying marketing services
The takeaway is not nihilism about marketing. It is a posture. Some of what a good marketer knows is genuine, replicated science; some of it is untested convention repeated with a straight face; and the two are usually delivered in the same confident tone. A firm worth hiring is one that can tell you, without prompting, which of its claims rests on replicated evidence, which rests on a single study, and which rests on nothing firmer than industry habit.
The Machine-Readiness Score, a four-pillar read of where a business stands across classic search, the local map pack, AI answers and reputation, sits inside this corrective tradition rather than outside it: a published method, sourced figures, and clear labels on what is measured versus what is still an emerging estimate.
The evidence
Key findings, with their sources
-
Marketing has a documented, long-standing replication gap: the discipline publishes markedly fewer replication studies than the natural sciences, and the trend was not improving.
established Evanschitzky, H., Baumgarth, C., Hubbard, R. & Armstrong, J.S., "Replication Research in Marketing Revisited: A Note on a Disturbing Trend," Journal of Business Research, 2007.
-
Several prominent priming effects, a category often invoked in marketing-psychology claims, suffered high-profile replication failures, with real-world effect sizes proving far less reliable than early lab findings suggested.
established General psychology replication-crisis literature (Open Science Collaboration and follow-on priming-specific replication work), as summarized alongside Evanschitzky et al. (2007).
-
The double jeopardy law, smaller-share brands have both fewer buyers and slightly lower loyalty, has been replicated across categories including packaged goods, banking, insurance and newer categories.
established Ehrenberg, A.S.C., Goodhardt, G.J. & Barwise, T.P., "Double Jeopardy Revisited," Journal of Marketing, 54(3), 1990; McPhee, W. (1963), Formal Theories of Mass Behavior.
-
Standard observational attribution methods frequently produced advertising-effect estimates in the wrong direction or magnitude when checked against 15 randomized field experiments (500M+ user-experiment observations, 1.6B ad impressions).
established Gordon, B.R., Zettelmeyer, F., Bhargava, N. & Chapsky, D., "A Comparison of Approaches to Advertising Measurement: Evidence from Big Field Experiments at Facebook," Marketing Science, 38(2), 2019.
-
The PR industry formally rejected Advertising Value Equivalency as a vanity metric in favor of outcome-based measurement, in the Barcelona Principles (2010, revised 2015 and 2020).
established AMEC, Barcelona Principles 3.0 (2020), International Association for the Measurement and Evaluation of Communication.
-
The "half my advertising is wasted" line has no verified original source; the earliest documented match traces to a secondhand 1919 speech and the sentiment is attributed to several people, not confirmed to Wanamaker.
contested Quote Investigator (2022), "One-Half the Money I Spend for Advertising Is Wasted, But I Have Never Been Able To Decide Which Half."
-
A share-of-search analysis reporting it "accounted for around 83%" of market-share variance is a single industry study across roughly 30 cases, not a peer-reviewed law.
emerging Hankins, J., Share of Search Council research (myshareofsearch.com); Binet, L., summarized in Marketing Week.
Calibration
What is proven, what is promising, what is unproven
| Evidence tier | Tactics | What the evidence says |
|---|---|---|
| established | Double jeopardy; mental/physical availability and buyer acquisition (How Brands Grow); attribution needs experiments (Gordon et al.); AVE is a vanity metric (Barcelona Principles); marketing's replication gap itself. | Peer-reviewed or ratified-industry sources, retested across categories, decades, or large experiments. |
| emerging | Share of search as a leading indicator; applying brand-growth laws specifically to AI-answer / generative-engine discovery. | Genuine, named, published research or reasonable inference, but single studies or new applications not yet replicated at scale. |
| contested | The Wanamaker "half wasted" attribution; any tactic resting on a single unreplicated priming study; round, memorable statistics from one analysis. | Unverified attributions, findings that failed to reproduce, or single-analysis numbers repeated as settled law. |
Reference
Glossary
- Replication
- Rerunning a study to see whether its result holds. A finding that has never been reproduced is a hypothesis that got published, not settled knowledge.
- Replication crisis
- The finding, most publicized in psychology, that a substantial share of well-known experiments do not reproduce, or reproduce with much smaller effects than first claimed.
- Priming
- The idea that a subtle cue shifts later behavior without a person noticing. Widely used in marketing-psychology claims and the subject of prominent replication failures.
- Double jeopardy law
- A replicated pattern in which smaller-share brands have both fewer buyers and slightly lower average loyalty, purely as a consequence of size.
- Vanity metric
- A number that is easy to compute and tends only to rise, offering reassurance rather than a basis for a decision. AVE was a canonical example.
- Observational vs experimental
- Observational measurement correlates what happened; experimental measurement withholds treatment from a control to isolate what an intervention actually caused.
Straight answers
Frequently asked questions
Does marketing science replicate?
Partly. Marketing has a documented replication gap, it publishes far fewer replication studies than the natural sciences, so many findings have never been retested. But a few regularities, such as the double jeopardy law and the buyer-acquisition patterns behind How Brands Grow, are among the most replicated results in social science. The task is telling the durable laws apart from untested convention.
Is the replication crisis in psychology relevant to marketing?
Yes, because much marketing persuasion advice is borrowed from social psychology. Several celebrated effects, priming among them, failed to reproduce or shrank sharply when retested. Any marketing tactic built on a single striking psychology study inherits that fragility.
So is most marketing advice wrong?
No. "What works" is a mix of replicated laws and untested folklore delivered in the same confident tone. Some of it is genuine science; some has never been retested. The fair demand of a marketer is that they tell you which is which.
How can a business owner tell a durable finding from a guess?
Ask whether it has ever been replicated, weigh the source data against the claim, distrust round memorable numbers from a single analysis, prefer experiments to dashboards, and be suspicious of any metric that only ever goes up. These few questions separate most real findings from confident convention.
What does this mean for measuring my own marketing?
Insist on published method and clear labels. A trustworthy read tells you what is measured versus estimated, with sourced figures throughout. That is the standard behind the Machine-Readiness Score, which is built to sit inside marketing's tradition of correcting its own metrics rather than adding another unverifiable one.
Provenance
Sources
- Evanschitzky, H., Baumgarth, C., Hubbard, R. & Armstrong, J.S., "Replication Research in Marketing Revisited: A Note on a Disturbing Trend," Journal of Business Research, 60(4), 2007 (established)doi.org
- Open Science Collaboration and follow-on priming-specific replication work, general psychology replication-crisis literature (established that a replication crisis is documented)
- Ehrenberg, A.S.C., Goodhardt, G.J. & Barwise, T.P., "Double Jeopardy Revisited," Journal of Marketing, 54(3), 1990; McPhee, W., Formal Theories of Mass Behavior, 1963 (established)
- Sharp, B., How Brands Grow: What Marketers Don't Know, Oxford University Press, 2010; Ehrenberg-Bass Institute research program (established as a replicated pattern; AI-answer application new and untested)
- Gordon, B.R., Zettelmeyer, F., Bhargava, N. & Chapsky, D., "A Comparison of Approaches to Advertising Measurement: Evidence from Big Field Experiments at Facebook," Marketing Science, 38(2), 2019 (established)
- AMEC, Barcelona Principles 3.0, International Association for the Measurement and Evaluation of Communication, 2020 (established)
- Quote Investigator, "One-Half the Money I Spend for Advertising Is Wasted...", 2022 (contested / unverified attribution)
- Hankins, J., Share of Search Council (myshareofsearch.com); Binet, L., summarized in Marketing Week (emerging, single industry analysis)
Every figure above is attributed to a real, dated source and tagged with its evidence tier. Where a claim could not be verified to a primary source, it is not stated as fact.