Measurement & Honesty · established evidence
What Happens When You Test Multi-Touch Attribution Against Reality: The Facebook Field Experiments
Multi-touch attribution, and the last-click model that preceded it, is an observational method: it reads which ads a converting customer happened to touch and credits them, without ever asking what would have happened if the ad had never run. In 2019 a team led by Brett Gordon put that method on trial. Using fifteen large-scale randomized field experiments at Facebook, they compared the true causal lift an experiment measures against the lift the standard observational methods would have reported from the same data. The two frequently disagreed, sometimes in the wrong direction, sometimes by a wide margin, and the disagreement did not close even after conditioning on rich demographic and behavioral covariates. The lesson is not that attribution reports are useless, but that they answer a different question from the one owners think they are answering, and only a controlled experiment settles the difference.
The question attribution quietly avoids
Every attribution report answers a question of correlation: of the customers who converted, which marketing touches did they encounter along the way, and in what order. Last-click credits the final touch. Multi-touch attribution spreads the credit across several. Both are reading the path of people who already converted.
The question a business owner actually cares about is causal: of the money spent, how much bought a sale that would not have happened anyway. Those two questions have the same units and look like the same number on a dashboard, which is exactly why they are so easily confused. An ad shown to someone who was already going to buy gets full credit in an attribution model and zero credit in reality. The gap between the two is not a rounding error. It is the entire measurement problem.
For most of the industry's history that gap was assumed to be small, or was simply never tested, because testing it requires something most advertisers never had: a genuine control group of buyers who were deliberately not shown the ad. The Facebook field experiments are important because they built that control group at scale, fifteen times over, and measured the gap directly.
The experiment that put attribution on trial
Gordon, Zettelmeyer, Bhargava and Chapsky designed a study with a rare property: it could measure the same advertising effect two ways at once, on the same people, at the same time. Their 2019 paper in Marketing Science reports fifteen large-scale randomized field experiments run on Facebook's advertising platform, spanning more than 500 million user-experiment observations and 1.6 billion ad impressions.
In each experiment, users were randomly split into a test group that could see the advertiser's campaign and a control group that could not. The difference in outcomes between those two randomized groups is the ground truth, the real causal lift, because randomization is the one method that guarantees the two groups differ only in whether they saw the ad.
Then the researchers did something clever. From the very same population, they estimated the ad's effect a second way, using the standard observational methods that underpin most attribution and lookalike modeling. They fit those models with the rich individual-level data a platform like Facebook actually has, extensive demographic and behavioral covariates, and compared the observational answer to the experimental ground truth.
Why this design closes the usual excuse
The power of the design is that it removes the usual excuse. When an observational model disagrees with an experiment, defenders can always say the model simply lacked the right variables. Here the model had them: the data was Facebook's own, some of the richest individual-level data in commercial existence. If observational attribution were going to work anywhere, it would work here.
What the Facebook field experiments actually found
The observational methods frequently failed to recover the experimental ground truth. Across the fifteen studies, effect estimates from the standard observational approaches diverged from the randomized result, at times pointing in the wrong direction and at times off by a wide margin. Critically, the divergence persisted even after conditioning on the rich demographic and behavioral covariates available in the data.
That last point is the one to sit with. The failure was not a data problem that more variables would fix. It was structural. Observational methods have to reconstruct a counterfactual, what these buyers would have done unseen, from people who are systematically different: the people an ad system chose to show the ad to are, by design, the people most likely to convert. No amount of covariate adjustment fully undoes a selection that the ad platform itself engineered.
The authors' conclusion is measured, and worth quoting in spirit rather than overstating: observational methods can sometimes approximate experimental results, but they do so unreliably, and an advertiser cannot know in advance whether a given observational estimate is one of the good cases or one of the badly wrong ones. Unreliable in an unknown direction is, for decision-making, close to unusable.
Why last-click and multi-touch attribution inherit the same bias
Last-click attribution has an obvious flaw: it hands all the credit to the final touch and none to everything that built the demand. Multi-touch attribution was sold as the sophisticated fix, distributing credit across the whole path to purchase. But the Facebook result cuts underneath that debate, because both models share the same original sin. Both are observational. Both credit touches on the paths of people who converted, and neither constructs the control group that would reveal which of those conversions the ads actually caused.
This is why the two most over-credited tactics in almost every account are retargeting and branded search. Both deliberately reach people who are already close to buying. An observational model sees those people convert shortly after the touch and awards a glowing return. An experiment that holds a matched group back frequently finds much of that return was going to arrive regardless. The dashboard is not lying about the arithmetic it performed. It is answering the wrong question with impressive precision.
The practical implication is uncomfortable but freeing: the choice between last-click and a fancier multi-touch attribution model is a smaller decision than it appears. Refining the credit-splitting rules of an observational model does not make the model causal. Only an experiment does that.
What rigorous experiments demand, and where they break
If experiments are the standard, it is worth being equally clear about how hard they are to run well. The canonical practitioner text, Ron Kohavi, Diane Tang and Ya Xu's Trustworthy Online Controlled Experiments, distilled from experience running more than twenty thousand experiments a year at Google, LinkedIn and Microsoft, catalogues the ways controlled tests quietly produce false answers.
The most common is peeking: checking a test before it reaches its planned sample size and stopping the moment it looks significant, which inflates false positives. Others include sample ratio mismatch, where the split between groups is subtly broken, and novelty and primacy effects, where a change looks powerful only because it is new. A controlled experiment is the right tool, but a carelessly run experiment can mislead as confidently as the attribution model it was meant to check.
There is also a scale problem that matters enormously for smaller businesses. Standard statistical power requirements for a clean individual-level test, at conventional significance and power targets, typically call for tens of thousands to hundreds of thousands of users. At the traffic a single-location business realistically sees, an experiment sensitive enough to trust can take many months to conclude, by which point the campaign, the season or the business has already changed. This is not a moral failing on the owner's part. It is arithmetic, and it is why the answer for small firms is rarely "run your own fifteen experiments."
The methods that survived: mix modeling and geo experiments
The response to all this, across the serious end of the industry, has been a return to methods that measure causally without depending on tracking individual people. Two matter most.
Marketing mix modeling
Marketing mix modeling is a statistical method that regresses an outcome time series, such as sales, on marketing time series, with transforms for carryover and diminishing returns, rather than following individual users. That property is precisely why it outlived the collapse of cookie and device tracking. Tellingly, both of the largest ad platforms have open-sourced their internal mix-modeling methodology, Google as Meridian and Meta as Robyn, a strong signal that the industry now treats aggregate causal modeling as the credible fallback once individual attribution degrades. How well these models perform at the scale of a single small business, rather than a national brand, is still an open, emerging question.
Geo experiments and holdouts
The nearest small-scale cousin of the Facebook design is the geo experiment: hold back whole geographic markets as a control, keep the campaign running in matched markets, and measure the difference. Meta's open-source GeoLift library formalizes this with synthetic-control methods, and needs no user-level pixel at all. An independent simulation study reported GeoLift recovering the intended coverage most closely and with the lowest false-positive rate among the open-source geo tools it compared, though those specific comparative figures come from a vendor-run simulation rather than a peer-reviewed study, and should be read as industry-grade evidence rather than settled fact.
Why the signal got worse right when it mattered
The attribution problem did not stay still after 2019. Apple's App Tracking Transparency, introduced with iOS 14.5 in April 2021, sharply reduced the individual-level signal that observational attribution depends on. Industry and ad-tech vendors have reported figures such as a large majority of iOS users opting out of cross-app tracking and material drops in pixel-based attribution accuracy, though these point estimates come from parties with a commercial stake in the narrative and are best treated as directional rather than precise.
The direction, however, is not in dispute, and it sharpens the lesson from the Facebook experiments. Observational attribution was already unreliable when the tracking was rich. As the tracking thinned, the same models were asked to do the same reconstruction with less to work from. This is the structural reason the industry pivoted toward consented server-side signals and aggregate, experiment-led methods: not fashion, but the accumulation of evidence that the individual-tracking approach was fragile even at its peak.
A recurring pattern worth recognizing
The Facebook study is one episode in an older story: the marketing industry repeatedly adopts a convenient, computable metric, treats it as truth for years, and only corrects once someone runs the harder experiment. The industry's own founding parable about waste, the line that half of all advertising is wasted and no one knows which half, is popularly pinned on John Wanamaker but has no verified original source, a fitting emblem for a field whose measurement folklore is itself unverified.
The public relations industry offers the cleanest example of self-correction. Under the Barcelona Principles, an industry standard first ratified in 2010 and revised as Barcelona Principles 3.0 in 2020, communications professionals formally renounced Advertising Value Equivalency, a flattering vanity metric, in favor of outcome-based measurement that is transparent, consistent and valid. An entire industry voted to kill its own comfortable number.
This is the tradition a rigorous measurement practice belongs inside, not outside. The corrective posture is not to claim a metric that finally solves everything, but to disclose the method, separate the measurement from the money that could bias it, and never report a figure without the baseline behind it. In the United States that standard also has a legal edge: the Federal Trade Commission's Endorsement Guides in 16 CFR Part 255 require that testimonials be genuine and that material connections be disclosed, a reminder that inflated or unsubstantiated claims are not merely bad manners.
What this means for a business that is not Facebook
A single owner-operated business will never run fifteen randomized experiments across 1.6 billion impressions. That is the wrong takeaway. The right one is that the standard is set, and any measurement that ignores it should be read with suspicion. When a report claims a precise return from a channel, the question is simply: against what control? If the answer is "none, we credited the touches on the paths of people who converted," then the number is an attribution estimate, not a measured causal result, and the Facebook experiments tell you it can be wrong in a direction you cannot predict.
The workable path for a smaller business is not to abandon measurement but to change what counts as proof: prefer holdout and geo experiments where volume allows, hold the whole account to a blended efficiency figure that no single platform can inflate, report every result with its baseline and its confidence, and treat any dashboard ROAS as a claim to be checked rather than a fact to be banked. None of that guarantees a particular outcome. It guarantees that the number you steer by is one you can actually trust.
The evidence
Key findings, with their sources
-
Across 15 large-scale randomized field experiments at Facebook, standard observational attribution methods frequently diverged from the true experimental lift, sometimes in the wrong direction, even after conditioning on rich demographic and behavioral covariates.
established Gordon, Zettelmeyer, Bhargava & Chapsky, "A Comparison of Approaches to Advertising Measurement: Evidence from Big Field Experiments at Facebook", Marketing Science 38(2), 2019.
-
The study spanned more than 500 million user-experiment observations and 1.6 billion ad impressions, measuring causal lift and observational estimates on the same population.
established Gordon, Zettelmeyer, Bhargava & Chapsky, Marketing Science 38(2):193-225, 2019.
-
Peeking, sample ratio mismatch, and novelty and primacy effects are documented ways controlled experiments produce false answers, catalogued from running more than 20,000 experiments a year at Google, LinkedIn and Microsoft.
established Kohavi, Tang & Xu, "Trustworthy Online Controlled Experiments", Cambridge University Press, 2020.
-
Both of the largest ad platforms have open-sourced their internal marketing mix modeling methodology, Google as Meridian and Meta as Robyn, treating aggregate causal modeling as the credible fallback as individual attribution degrades.
established Google, Meridian (business.google.com); Tueller et al., "Packaging Up Media Mix Modeling", arXiv:2403.14674, 2024; facebookexperimental/Robyn.
-
In an independent simulation, Meta's open-source GeoLift recovered its intended coverage most closely (about 92 to 95 percent) and with the lowest false-positive rate (about 3 to 5 percent) among the open-source geo-experiment tools compared.
emerging facebookincubator/GeoLift (GitHub); Recast Research, "Open-Source Geo-Experiment Tools: A Head-to-Head Simulation Study".
-
The PR and communications industry formally rejected Advertising Value Equivalency as a vanity metric under the Barcelona Principles (2010, revised as 3.0 in 2020), in favor of transparent, valid, outcome-based measurement.
established AMEC, Barcelona Principles 3.0, 2020.
-
The line that "half my advertising is wasted, I just don't know which half", popularly attributed to John Wanamaker, has no verified original source.
contested Quote Investigator, 2022, "One-Half the Money I Spend for Advertising Is Wasted".
Calibration
What is proven, what is promising, what is unproven
| Evidence tier | Tactics | What the evidence says |
|---|---|---|
| established | Randomized field experiments (RCTs) as ground truth; treating observational attribution as an unreliable approximation; documented experiment failure modes (peeking, SRM). | Gordon et al., Marketing Science 2019; Kohavi, Tang & Xu, 2020. |
| established (method) / emerging (numbers) | Geo experiments and synthetic-control holdouts (geo-lift) to measure incremental lift without a user-level pixel. | facebookincubator/GeoLift; Recast Research vendor simulation (comparative figures not peer-reviewed). |
| established (exists) / emerging (MSME fit) | Marketing mix modeling on aggregate time series (Meridian, Robyn) as the tracking-independent fallback. | Google Meridian; Robyn, arXiv:2403.14674, 2024. Performance at single-location scale is unproven. |
| contested / directional | Citing specific post-iOS-ATT signal-loss magnitudes as if precise. | Ad-tech vendor reports; treat point estimates as directional, attribute to the named source. |
Reference
Glossary
- Multi-touch attribution (MTA)
- An observational method that distributes conversion credit across the several marketing touches a converting customer encountered. Because it reads the paths of people who already converted, it does not establish what the ads caused.
- Observational method
- Any estimate of an effect drawn from data that was not randomized, requiring the analyst to reconstruct a counterfactual from people who differ systematically from those exposed.
- Randomized field experiment (RCT)
- A test that randomly splits an audience into a group that can see an ad and a control group that cannot, so the two differ only in exposure. The difference in outcomes is the true causal lift.
- Ground truth
- The real causal effect, established here by randomization, against which an observational estimate can be judged right or wrong.
- Selection bias
- The distortion that arises because an ad system deliberately shows ads to the people most likely to convert, making exposed and unexposed groups non-comparable in ways covariates cannot fully fix.
- Incrementality (iROAS)
- The return calculated only on the revenue an ad genuinely caused, measured against a held-back control, as opposed to platform-reported ROAS that credits nearby conversions.
- Marketing mix modeling (MMM)
- A statistical method that regresses an outcome time series on marketing time series, with carryover and saturation transforms, to estimate effects without tracking individuals.
- Geo-lift
- A geographic holdout experiment: whole markets are held back as a control while matched markets keep running the campaign, using synthetic-control methods and no user-level pixel.
Straight answers
Frequently asked questions
What did the Facebook field experiments find about attribution?
That standard observational attribution methods frequently failed to recover the true causal lift measured by a randomized experiment on the same people. The estimates sometimes pointed in the wrong direction and sometimes missed by a wide margin, and the gap did not close even when the models were given rich demographic and behavioral data. The takeaway is that an attribution number can be wrong in a direction you cannot predict in advance.
Should I stop using multi-touch attribution?
Not necessarily, but you should stop treating it as a measured causal result. Multi-touch attribution is a useful descriptive picture of the paths customers take, and it can help with directional decisions. It is not proof that a channel caused sales, because it never constructs a control group. The reasonable use is to read it as a claim to be checked by an experiment, not a fact to bank. Refining the credit-splitting rules does not make an observational model causal.
What is the difference between attribution and incrementality?
Attribution asks which touches a converting customer encountered and assigns credit among them. Incrementality asks how much revenue an ad actually caused that would not have happened otherwise, measured against a control group that was deliberately not shown the ad. Attribution reads correlation on the paths of people who converted; incrementality measures causation. They can look like the same number on a dashboard and mean very different things.
If Facebook could not make observational attribution reliable with all its data, can better data fix it for me?
The evidence says no. The Facebook study used some of the richest individual-level data in commercial existence and the divergence from experimental truth persisted anyway. The failure is structural: an ad system chooses to show ads to the people most likely to convert, so the exposed and unexposed groups are not comparable, and no amount of extra covariates fully undoes a selection the platform engineered. More data does not repair a broken counterfactual; a control group does.
How can a small business measure honestly if it cannot run 15 experiments?
By changing what counts as proof rather than by imitating Facebook. Prefer geo or holdout experiments where volume allows, hold the whole account to a blended efficiency figure no single platform can inflate, report every result with its baseline and confidence, and treat any dashboard return as a claim to verify. Statistical power for a clean individual-level test is out of reach at small traffic, which is exactly why aggregate and experiment-led methods, not fancier attribution, are the reliable path.
Provenance
Sources
- Gordon, B.R., Zettelmeyer, F., Bhargava, N., Chapsky, D., "A Comparison of Approaches to Advertising Measurement: Evidence from Big Field Experiments at Facebook", Marketing Science, 38(2), 193-225, 2019 (established)
- Kohavi, R., Tang, D., Xu, Y., "Trustworthy Online Controlled Experiments: A Practical Guide to A/B Testing", Cambridge University Press, 2020 (established)doi.org
- Google, Meridian, open-source Bayesian marketing mix model (business.google.com); Google Research geo-level Bayesian Hierarchical Media Mix Modeling lineage (established that it exists and is open-sourced; emerging as to single-business fit)developers.google.com
- Tueller, N. et al., "Packaging Up Media Mix Modeling: An Introduction to Robyn's Open-Source Approach", arXiv:2403.14674, 2024; facebookexperimental/Robyn (established)arxiv.org
- facebookincubator/GeoLift (GitHub); Recast Research, "Open-Source Geo-Experiment Tools: A Head-to-Head Simulation Study" (established method; emerging/industry-grade comparative numbers)github.com
- AMEC, Barcelona Principles 3.0, 2020 (established industry standard)amecorg.com
- Federal Trade Commission, 16 CFR Part 255, "Guides Concerning the Use of Endorsements and Testimonials in Advertising" (established regulation)ecfr.gov
- Multiple ad-tech vendor reports on iOS 14.5 App Tracking Transparency signal loss, 2021 onward (established as a dated policy event; directional/contested point estimates)
- Quote Investigator, "One-Half the Money I Spend for Advertising Is Wasted, But I Have Never Been Able To Decide Which Half", 2022 (established that the Wanamaker attribution is unverified)quoteinvestigator.com
Every figure above is attributed to a real, dated source and tagged with its evidence tier. Where a claim could not be verified to a primary source, it is not stated as fact.