Demand & Paid Media · established evidence
Attribution Windows and the Illusion of the Last Click: A Synthesis
Last click attribution is a rule for splitting credit, not a measurement of cause. It assigns a conversion to whichever touchpoint the buyer clicked last, and every richer model that followed, first-click, linear, time-decay, position-based, data-driven, and full multi-touch attribution, is a different rule for splitting the same observed credit among the touchpoints a system happened to record. All of them answer the bookkeeping question, which recorded interactions preceded the sale, and none of them answer the causal question, would the sale have happened anyway. The distinction is not pedantic. Two large bodies of field-experiment evidence, the eBay branded-keyword null result and the activity-bias experiments, show that observational credit assignment systematically overstates what paid media actually caused, because the exposed users were already the ones most likely to convert. Attribution describes the path. Only a controlled counterfactual measures the lift.
What an attribution model actually is: a rule for splitting credit
An attribution model is a convention. It takes the set of interactions a measurement system recorded before a conversion, a paid click here, an organic visit there, an email open, a direct return, and applies a fixed rule for dividing the credit for that conversion among them. Last click hands all of it to the final recorded touch. First click hands all of it to the first. Linear splits it evenly. Time-decay weights recent touches more heavily. Position-based reserves large shares for the first and last touches and spreads the rest. Data-driven attribution and multi-touch attribution use the observed data to distribute fractional credit across the path.
What every one of these models has in common is more important than what separates them. They all operate on the same input, a log of interactions that were observed to co-occur with conversions, and they all produce the same kind of output, a division of credit that sums to the conversions that happened. Not one of them contains a step that estimates what would have happened in the absence of the advertising. That step, the counterfactual, is the entire content of the word "caused," and it is exactly the step that attribution models omit by construction.
This is why comparing attribution models against each other is a comparison of conventions, not a search for accuracy. A business that switches from last click to a data-driven model has not moved closer to the truth about causation. It has changed the accounting rule, which reshuffles credit between channels while leaving the underlying question, did the spend produce incremental customers, entirely untouched.
The attribution window: an arbitrary boundary that changes the story
Before any model can split credit, a system has to decide which interactions are even eligible to receive it. That decision is the attribution window, also called the lookback window, and it is set by a number: count clicks within one day, or seven, or thirty, or ninety, and count viewed-but-not-clicked impressions separately or not at all. The window is a policy choice, not a property of the buyer.
Because the window is a boundary drawn on a continuous stream of behavior, moving it moves the reported result without anything in the real world changing. Widen a click window from seven to thirty days and more prior touches become eligible, so a channel that tends to appear early in the path gains credit it did not have a moment before. Add a view-through window and display and video impressions start collecting conversions they were previously invisible to. The customers, the sales, and the advertising were identical in both readings. Only the accounting boundary moved.
The practical consequence is that two agencies auditing the same account with different window settings can produce materially different channel reports and both be internally correct. When a report leads with a channel ROAS, the window that produced it is doing quiet, load-bearing work, and a figure quoted without its window is not yet a fact.
Why the last click is an illusion
Last click is the default many platforms still report against, and its appeal is obvious: the final touch feels closest to the decision, so crediting it feels intuitive. The problem is that the final recorded touch is frequently the point where the buyer, having already decided, arrives to transact. A person who has resolved to book a med-spa treatment and searches the brand name to find the site is a conversion that was going to happen; the paid click that intercepts that search collects the credit for a sale it did not create.
This is not a modeling nuance that a better model corrects. It is a structural property of any credit rule applied to observational data, because the users who see and click ads are, on average, already the users most inclined to convert. The final section names the mechanism precisely, but the intuition is simple: the closer a touchpoint sits to an intent the buyer already had, the more credit it harvests and the less of that credit it earned.
The field experiment that broke the tie: eBay branded keywords
The cleanest evidence that attribution overstates paid-search impact does not come from a better model. It comes from turning the ads off at random and watching what changed. In a large-scale randomized field experiment, eBay suspended paid search across randomized regions and measured the effect on sales rather than inferring it from click logs.
On the company's own branded keywords the result was stark: the ads produced no measurable short-term incremental benefit. Buyers who searched for eBay and would have clicked the paid listing simply clicked the organic listing instead when the ad was gone, and arrived all the same. The paid click had been collecting credit, under every attribution model, for traffic that was already going to convert. For non-branded terms the picture was heterogeneous rather than uniformly null: new and infrequent users were positively influenced, but frequent users whose purchases were unaffected by ads absorbed most of the spend, dragging average returns negative.
The finding is a single-firm study and generalizes as a documented mechanism, not a universal constant; a med-spa or a home-services account is not eBay. What travels is the mechanism, that brand-term and high-intent paid clicks tend toward low incrementality precisely where last click credits them most, and that consumer response is heterogeneous enough that an account-level average hides who the spend actually moved.
Activity bias: why every observational model tilts the same way
If the eBay result showed that last click can be wrong, the activity-bias literature explains why the error runs in one predictable direction and why no purely observational model escapes it. Across three controlled experiments, researchers demonstrated that a user's online behaviors are correlated in time: someone who happens to be active online, browsing, searching, checking email, is simultaneously more likely to be exposed to an ad and more likely to take the measured action, whether or not the ad had any effect.
That correlation is fatal to observational comparison. When an analyst compares users who saw an ad against users who did not, the exposed group was already the more active, more purchase-inclined group before the ad entered the picture. The gap between the two groups is read as advertising effect, but much of it is pre-existing difference. The result is a systematic upward bias in observational estimates of advertising effectiveness.
Multi-touch attribution and data-driven attribution do not solve this
It is tempting to believe that a sophisticated multi-touch or data-driven model, by using more of the path, corrects the bias that last click introduces. It does not, because the input is still observational. Distributing fractional credit more cleverly across a set of touches that were themselves correlated with pre-existing intent produces a more elaborate description of the same biased data. The models differ in how they split credit; they are identical in that none of them observes the counterfactual world where the ad was absent. Comparing attribution models is therefore the wrong axis of debate. The real fault line is not last click versus multi-touch. It is observational credit assignment versus experimental measurement of lift.
Attribution answers a bookkeeping question. Incrementality answers a causal one
The synthesis is a category distinction. Attribution, in every form, is bookkeeping: given the conversions that occurred and the touches that preceded them, apportion the credit. It is a description of observed paths. Incrementality is a causal quantity: of the conversions attributed to a channel, how many would not have occurred without it. The first can be computed from a log. The second requires a comparison against a world in which the spend did not happen, which no log contains.
Treating an attributed number as a causal one is the error that quietly wastes budget. A channel can be credited with a large share of conversions under any model and cause almost none of them, which is the branded-search pattern the eBay experiment isolated. Reporting that reshuffles credit between channels, whether by changing the model or widening the window, feels like measurement but never crosses from description into causation. The gap between what a platform says it did and what it did is not a reporting bug to be tuned away; it is the difference between two questions.
The honest alternative: measure the counterfactual
If observational attribution cannot answer the causal question, the answer has to be produced by building a counterfactual on purpose, which is what the experimental literature does. Geo experiments randomize non-overlapping geographic regions into treatment and control ad conditions and read the difference in outcomes, giving a systematic causal method for measuring true ad effectiveness without individual-level tracking, explicitly designed to inform bidding, budgeting, and campaign decisions. Ghost-ad methods make the same counterfactual cheaper by recording the would-be impressions a control group should have seen, isolating incremental lift at a fraction of the cost of public-service-announcement holdouts; on the retargeting campaign originally studied, the method isolated a causal lift of 17.2 percent in site visits and 10.5 percent in purchases.
At the account level, a blunter instrument does useful defensive work. The Marketing Efficiency Ratio, total revenue divided by total marketing spend, is a blended figure that cannot be gamed by shifting attribution credit between channels, because it never assigns channel-level credit in the first place. It is an industry-originated anchor rather than an academic construct, but its rationale follows directly from the activity-bias and attribution literature: if channel-level numbers are the thing being manipulated by model and window choices, a whole-account figure sidesteps the manipulation. Used together, an experiment for causal reads and a blended anchor for account-level honesty, they answer the question attribution cannot.
A word of calibration on the numbers that circulate here. Practitioner and vendor analyses commonly report that measured incremental ROAS runs well below platform-reported ROAS, with branded search and retargeting the channels where the two diverge most. That direction is consistent with the peer-reviewed field experiments above, but the specific percentage ranges come from unaudited case studies, not independent replication, and should be cited as industry-reported, never as a constant.
What this means for reading an attribution report
None of this argues for abandoning attribution. A path log is genuinely useful for understanding how buyers move and for spotting broken tracking, and last click remains a defensible operational default for some decisions precisely because it is simple and stable. The argument is narrower and firmer: an attributed number is a description of credit, and it should be read as one, with its model and its window stated, and it should never be promoted to a claim about what the spend caused without a counterfactual behind it.
The operational discipline that follows is to hold two layers apart. Let attribution do the bookkeeping it is good at, and let a separate, experimental layer, geo tests, holdouts, and a blended efficiency anchor, carry every claim about incremental return. The moment a report blurs the two, presenting reshuffled credit as proof of causation, it has stopped measuring and started asserting.
The evidence
Key findings, with their sources
-
In a large-scale randomized field experiment, ads on eBay's own branded keywords produced no measurable short-term incremental benefit; on non-brand terms, frequent buyers whose purchases were unaffected absorbed most of the spend, yielding negative average returns.
established Blake, Nosko & Tadelis, "Consumer Heterogeneity and Paid Search Effectiveness: A Large-Scale Field Experiment", Econometrica, 83(1), 2015.
-
Observational estimates of advertising effectiveness are systematically biased upward by "activity bias": a user's online behaviors are correlated in time, so the ad-exposed group was already more active and more purchase-inclined, demonstrated across three controlled experiments.
established Lewis, Rao & Reiley, "Here, There, and Everywhere: Correlated Online Behaviors Can Lead to Overestimates of the Effects of Advertising", WWW '11, 2011.
-
The "ghost ads" method measures incrementality by recording the counterfactual impressions a control group would have seen, at a fraction of the cost of PSA holdouts; on the retargeting campaign studied it isolated a causal lift of 17.2% in site visits and 10.5% in purchases.
established Johnson, Lewis & Nubbemeyer, "Ghost Ads: Improving the Economics of Measuring Online Ad Effectiveness", Journal of Marketing Research, 54(6), 2017.
-
Geo experiments, randomizing non-overlapping regions into treatment and control, provide a systematic causal method for measuring true ad effectiveness without individual-level tracking, designed to inform bidding, budgeting, and campaign decisions.
established Vaver & Koehler, "Measuring Ad Effectiveness Using Geo Experiments", Google Inc., 2011.
-
Practitioner and vendor analyses report that measured incremental ROAS commonly runs well below platform-reported ROAS, with branded search and retargeting the channels where the two diverge most.
contested Synthesized from measurement-vendor and practitioner analyses, 2025-2026 (unaudited; directionally consistent with the peer-reviewed field experiments, not an independent replication).
Calibration
What is proven, what is promising, what is unproven
| Evidence tier | Tactics | What the evidence says |
|---|---|---|
| Established | Attribution as credit assignment rather than causal estimate; the eBay branded-keyword null result; activity bias as a systematic upward bias in observational estimates; geo experiments and ghost ads as counterfactual measurement. | Peer-reviewed field experiments and published methodology (Blake-Nosko-Tadelis 2015; Lewis-Rao-Reiley 2011; Johnson-Lewis-Nubbemeyer 2017; Vaver-Koehler 2011). |
| Emerging | Blended Marketing Efficiency Ratio (MER) as an account-level anchor that resists attribution gaming. | Industry-originated metric whose rationale follows directly from the activity-bias and attribution literature, not yet an academic construct. |
| Contested | Any specific "iROAS runs X percent below platform ROAS" figure, or a fixed branded-search waste percentage. | Unaudited vendor and practitioner case studies; directionally consistent with the established findings but not independently replicated. Cite as industry-reported, never as a constant. |
Reference
Glossary
- Last-click attribution
- A credit rule that assigns a conversion entirely to the final recorded interaction before it. Simple and stable, but it credits the touch nearest an intent the buyer may already have had.
- Attribution window (lookback window)
- The time boundary that decides which prior interactions are eligible for credit, for example a 7-day, 30-day, or 90-day click window. Moving the window changes the reported result without anything in the real world changing.
- Multi-touch attribution (MTA)
- A family of rules that distribute fractional credit across several touches in a path. More elaborate than last click, but still computed from observational data and still not a measure of causation.
- Data-driven attribution
- Attribution that uses the observed path data to assign fractional credit. It changes how credit is split, not whether the underlying number is causal.
- Incrementality
- The causal quantity attribution omits: of the conversions credited to a channel, how many would not have occurred without it. Requires a comparison against a world in which the spend did not happen.
- Activity bias
- The upward bias in observational ad-effect estimates that arises because a user's online behaviors are correlated in time, so the exposed group was already more active and more likely to convert before any ad.
- Marketing Efficiency Ratio (MER)
- Total revenue divided by total marketing spend, a blended account-level figure that cannot be gamed by shifting attribution credit between channels because it assigns no channel-level credit.
Straight answers
Frequently asked questions
Which attribution model is the most accurate?
The question assumes the models differ in accuracy, but they differ in convention. Last click, linear, time-decay, position-based, data-driven, and multi-touch attribution are all rules for splitting observed credit, and none of them estimates whether the advertising caused the sale. Choosing between them is choosing an accounting rule, not getting closer to the truth about causation. The accurate answer to "did this spend cause customers" comes from a controlled experiment, not a better attribution model.
Is last-click attribution wrong?
It is not wrong as bookkeeping; it correctly records which touch was last. It becomes misleading when the credit it assigns is read as impact. Because the last recorded touch is often where an already-decided buyer arrives to transact, last click tends to over-credit high-intent, brand, and retargeting clicks that would have converted anyway. That pattern is exactly what a randomized experiment at eBay isolated on branded keywords.
Does multi-touch attribution fix the problem?
No. Multi-touch and data-driven attribution use more of the path and split credit more finely, but the input is still observational, and the touches themselves are correlated with the buyer's pre-existing intent. A more detailed division of biased data is a more detailed description, not a causal measurement. The real dividing line is observational credit assignment versus experimental measurement of lift, not last click versus multi-touch.
What is an attribution window and does its length matter?
The attribution window, or lookback window, sets how far back before a conversion an interaction can still receive credit, and whether view-through impressions count. Its length matters a great deal: widening it lets earlier touches collect credit and can materially change a channel report while the underlying sales are identical. A channel figure quoted without stating its window is incomplete, because the window is doing quiet, load-bearing work.
How do you actually measure whether paid search works?
By building a counterfactual on purpose. Geo experiments randomize regions into ad and no-ad conditions and read the difference in outcomes; ghost-ad and holdout methods record the impressions a control group would have seen and measure the incremental lift. At the account level, a blended Marketing Efficiency Ratio resists the credit-shuffling that channel-level attribution invites. Together they answer the causal question attribution cannot.
Provenance
Sources
- Blake, T., Nosko, C. & Tadelis, S., "Consumer Heterogeneity and Paid Search Effectiveness: A Large-Scale Field Experiment", Econometrica, 83(1), 2015 (established)doi.org
- Lewis, R. A., Rao, J. M. & Reiley, D. H., "Here, There, and Everywhere: Correlated Online Behaviors Can Lead to Overestimates of the Effects of Advertising", WWW '11, 2011 (established)doi.org
- Johnson, G. A., Lewis, R. A. & Nubbemeyer, E. I., "Ghost Ads: Improving the Economics of Measuring Online Ad Effectiveness", Journal of Marketing Research, 54(6), 2017 (established)doi.org
- Vaver, J. & Koehler, J., "Measuring Ad Effectiveness Using Geo Experiments", Google Inc., 2011 (established)research.google
- Measurement-vendor and practitioner analyses of the iROAS-to-platform-ROAS gap, 2025-2026 (contested / industry-reported, pending independent replication)
Every figure above is attributed to a real, dated source and tagged with its evidence tier. Where a claim could not be verified to a primary source, it is not stated as fact.