Demand & Paid Media
The ROAS Illusion: Why Your Ad Platform's Return Number Isn't the Real One
Independent incrementality research shows platform-reported ROAS is structurally unreliable, not in the single direction most marketers assume.
Part of Demand & Paid Media in the Insights library.
Abstract
Every performance marketer has heard some version of the claim that platform-reported ROAS is inflated by a flat 20 to 60 percent. That number does not survive contact with the evidence. What the strongest research available actually shows, from a peer-reviewed 2019 Facebook field-experiment study through newer incrementality testing panels, is that platform attribution is unreliable in a way that depends on the channel and the funnel stage: branded search and retargeting are reliably and heavily overstated, while upper-funnel spend, and in at least one large rigorous panel Meta's own aggregate reporting, has been shown to run the opposite way. The finding here is not a percentage. It is that self-reported platform ROAS cannot be trusted at face value for budget decisions in either direction, and small businesses are structurally the ones least able to find out which way their own account is biased.
Procter & Gamble supplied the first widely cited real-world test of that trust back in 2017 and 2018, when Marc Pritchard cut roughly $200 million of digital ad spend, about 20 to 50 percent of spend on some digital channels, over nine months and found no discernible drop in business performance, with reach actually rising about 10 percent. Two years later, a peer-reviewed study in Marketing Science examined 15 large-scale Facebook field experiments covering more than 500 million user-experiment observations and found that the observational and attribution methods marketers normally rely on generally overestimated advertising effectiveness compared to true randomized results, though in some cases they significantly underestimated it. That finding, that the platform's own numbers can be wrong in either direction, has mostly been ignored in favor of a simpler, punchier story: platforms always inflate. What has actually changed since is not the underlying math but the emergence of a small, commercially motivated incrementality-testing industry, firms like Stella and Haus, that run geo-holdout and randomized tests for brands who can afford them and then publish benchmark numbers to sell more of that testing. Those numbers are useful directional evidence. They are not an independent census, and none of them agrees on a single figure.
What the evidence shows is not that marketers don't know the platform number is suspect. In a vendor-sponsored 2026 survey of 210 senior marketing leaders, 91 percent said platform-reported results are overstated to some degree, and two-thirds estimated at least 11 percent of budget is wasted to attribution and optimization lag. The gap is between that belief and what marketers actually do with it: more than 80 percent of the same group said they rely primarily on proxy signals rather than verified purchase data when optimizing campaigns day to day, and 35 percent admitted those proxy-based decisions don't hold up once reconciled against actual sales. A separate, independent 2020 survey found only 40 percent of US marketers and agencies required their media partners to use buyer-approved measurement tools at all, and that major walled-garden platforms typically did not comply even when asked. The pattern across both surveys is the same: distrust of the number, followed by continued use of the number, because the alternative, real incrementality testing, is expensive and slow, and the dashboard is right there.
If the number a platform reports for its own performance cannot be trusted at face value, then a business's whole picture of where its market's attention actually sits, and where a dollar spent there actually converts, is being drawn from a distorted map. A brand that trusts branded-search ROAS at face value is very likely spending against a channel that was already going to convert those customers anyway; the strongest available benchmark data puts that kind of bottom-funnel inflation at 5 to 10 times the true incremental return, though that specific multiplier comes from a single vendor and should be read as a directional warning rather than a fixed rate. A brand that cuts an upper-funnel or automated campaign because its reported ROAS looks weak may be cutting the one channel that was quietly building demand it never got credit for. The only way to correct that map is to look at where attention and intent actually move, independent of any one platform's self-interested accounting of its own performance, which is the entire reason a demand-side view of the attention terrain matters more than any single channel's dashboard.
The data, in one read
The Percentage That Doesn't Exist
Search for the gap between platform-reported ROAS and true incremental return and you will find a version of the same claim everywhere: platforms overstate results by 20 to 60 percent. It is a clean, quotable number, and it is not supported by any independent, cross-platform, peer-reviewed source. Every rigorous piece of evidence we could locate and verify shows something more complicated and more useful: the bias is channel- and funnel-stage-dependent, running heavily in one direction for some spend and, credibly, in the opposite direction for other spend. There is no single haircut you can apply to a reported ROAS number and trust the result.
That is not a comfortable finding for anyone selling a simple fix, and it is worth sitting with. Two of the sources that came closest to a flat 40 to 60 percent figure turned out, on inspection, to be a vendor's explicitly hypothetical sales example and an unattributable aggregation of ad-tech blog commentary with no named study behind it. Both were set aside. What survives is smaller, older in places, and more honest.
There is no single haircut you can apply to a reported ROAS number and trust the result.
What the One Peer-Reviewed Study Actually Found
The strongest evidence available is also the oldest and the least commercially motivated. In 2019, researchers working with Facebook published a study in Marketing Science, a peer-reviewed journal of the Institute for Operations Research and the Management Sciences, examining 15 large-scale field experiments covering more than 500 million user-experiment observations and 1.6 billion ad impressions. They compared the observational and attribution methods marketers normally use against results from true randomized controlled experiments, the gold standard for isolating what an ad actually caused. The finding was not a clean overstatement number. Commonly used attribution methods generally overestimated advertising effectiveness relative to the randomized-experiment truth, but in some cases they significantly underestimated it.
That single sentence is the most academically credible fact in this entire subject, and it already contradicts the flat-overstatement story before a single vendor benchmark enters the picture. The direction of the bias is not fixed. It depends on what is being measured and how.
The $200 Million Real-World Test
Before any of the modern incrementality-testing vendors existed, Procter & Gamble ran the largest natural experiment on record for this exact question. Over nine months in 2017 and 2018, Chief Brand Officer Marc Pritchard cut roughly $200 million of digital ad spend, about 20 to 50 percent of spend on some digital channels, largely in response to concerns about ad fraud, viewability, and walled-garden measurement. The result, corroborated across multiple trade publications at the time, was no discernible negative impact on P&G's business performance. Reach actually increased by about 10 percent, because the spend that was cut had been buying largely redundant or unmeasurable impressions rather than incremental audience.
This is not a controlled experiment with a clean causal design the way the 2019 Marketing Science study is. It is a real company making a real, large cut and living with the consequences in public. But it points the same direction as the peer-reviewed research: a meaningful share of reported digital ad performance was never doing the work its reported ROAS implied.
Where the Money Actually Goes Before It Reaches Anyone
Part of why platform-reported numbers are so unreliable is structural, not statistical. In December 2023, the Association of National Advertisers published a log-level analysis of 21 major advertisers, $123 million in ad spend and 35.5 billion impressions, tracing money through 12 participating supply-chain companies from the moment it entered a demand-side platform to the moment it reached a consumer. The finding: only 36 cents of every dollar effectively reached the consumer as working media. Roughly 29 percent of spend was consumed by transaction costs, the fees that sit between the advertiser and the audience, and about 35 percent went to inventory that was non-viewable, invalid, or otherwise unmeasurable.
That is not an incrementality problem, it is a supply-chain problem, but it compounds the measurement problem. A platform reporting ROAS on the dollar it received is reporting on a dollar that, on average across this study's sample, had already lost nearly two-thirds of its working power before it ever reached a person capable of converting.
Only 36 cents of every dollar effectively reached the consumer as working media.
The Reliable Distortion: Branded Search and Retargeting
Where the evidence does point toward a consistent, sizable overstatement is at the bottom of the funnel. Across 225 geo-based incrementality tests run between August 2024 and December 2025, the incrementality-testing platform Stella reported an overall median incremental ROAS of 2.31x, with branded search and retargeting showing inflation of 5 to 10 times against platform-reported numbers. In the channel-level breakdown from the same panel, Google Search branded spend came in at a median incremental ROAS of just 0.70x, meaning that on an incremental basis it was likely losing money even though the platform's own attributed ROAS almost certainly looked strong.
This is the one piece of quantified evidence in this entire subject that should be read carefully rather than dismissed: it comes from a single vendor that sells incrementality testing and has a direct commercial incentive to show a large, fixable gap, so it is labeled contested here rather than established. But the mechanism it describes, that branded search and retargeting largely capture customers who were already going to convert, is well understood and consistent with the peer-reviewed evidence above. Treat the specific multiplier as directional, not as gospel.
When the Platform Undercounts You
The more counterintuitive finding, and the one that most directly breaks the flat-overstatement narrative, comes from a different incrementality vendor looking at Meta specifically. Across 640 Meta geo-holdout experiments, Haus found that for every $100 in platform-attributed DTC revenue under a 7-day click attribution window, Meta actually generated $115 in incremental revenue in aggregate, a 15 percent platform understatement. In a separate panel of 46 Meta incrementality studies among ecommerce brands with $15 million to $100 million in annual revenue, Stella found Meta delivering 21 percent more incremental revenue than the platform itself reported.
Both findings come with important texture. Haus found that Advantage+ and other automated campaign types tended to over-report relative to manual campaigns, while mid- and upper-funnel campaigns tended to under-report, especially when measured on a DTC-only basis. Fully 32 percent of Meta's total measured incremental impact in that panel lifted retail and marketplace sales that a DTC-only ROAS calculation would miss entirely. Both of these findings are also vendor-sourced and commercially interested, so they carry the same contested label as the bottom-funnel numbers above. But they are directly relevant because they show the bias running in the opposite direction from the popular assumption, on the same platform, sometimes in the same campaign portfolio.
For every $100 in platform-attributed revenue, Meta actually generated $115 in incremental revenue, a 15 percent understatement.
Believing It's Broken, Still Running It the Same Way
Marketers are not naive about any of this in the abstract. A 2026 survey of 210 senior marketing leaders across brands and agencies found that 91 percent believe platform-reported results are overstated to some degree, and a third believe attribution and optimization lag wastes more than 26 percent of budget. This particular survey should be read with a caveat: live verification found it was sponsored content commissioned by Affinity Solutions through its own Outcomes Marketing Council, not independent research, and no sampling methodology or margin of error was disclosed. It is included here as a labeled, vendor-commissioned data point on marketer sentiment, not as validated measurement of the actual overstatement size.
What is more telling, and comes from an independent 2020 eMarketer analysis, is the gap between that stated distrust and actual behavior. Only 40 percent of US marketers and agencies required their media partners to use buyer-approved measurement tools, and major walled-garden platforms typically did not comply even when asked. Forty-four percent of marketing professionals surveyed separately rated social media among the hardest channels to attribute sales revenue to, and 63 percent said insufficient media-quality transparency would most likely affect their spend on Facebook specifically. The distrust is real and widespread. The day-to-day optimization decisions still mostly run on the number the platform hands you, because building or buying an independent measurement capability is expensive, and the dashboard is free.
Why This Hits Small and Midsize Businesses Hardest
Every rigorous source in this study, the 2019 Marketing Science RCT, the geo-holdout panels, the incrementality benchmarks, has one thing in common: it required spend and conversion volume far beyond what most small and midsize advertisers generate. Geo-holdout testing needs enough geographic markets and enough conversions per market to detect a statistically meaningful difference. Randomized controlled experiments at the scale of the Gordon et al. study require Facebook-level infrastructure. None of the sources reviewed for this study included an SMB-specific incrementality dataset for accounts spending in the sub-$2,500 to $10,000 monthly range that describes most small businesses.
That is a reasonable inference from what incrementality testing structurally requires, not a measured finding about SMB accounts specifically. A small business cannot run its own version of the P&G experiment or a Meta geo-holdout panel. It is left choosing between trusting a platform number that the best available evidence says is unreliable in an unpredictable direction, or spending money it does not have on testing infrastructure built for brands ten times its size. That asymmetry, not the size of the gap itself, is the real finding for a business this size.
What to Do With a Number You Can't Fully Trust
None of this means platform reporting is useless, and it does not mean the answer is to ignore ROAS entirely. It means treating a platform's self-reported number as one signal among several, weighted by what is now reasonably well established: be most skeptical of branded search and retargeting ROAS, since the mechanism for inflation there, capturing customers who were already converting, is well understood even where the exact multiplier is vendor-sourced. Be slower to cut upper-funnel, prospecting, or automated campaigns purely on a weak reported ROAS number, since at least one rigorous panel found that exact category understated on a large sample.
The deeper fix is not a better attribution model from inside any single platform, it is an independent view of where a market's attention actually sits and how it is moving, measured from outside the walled garden that has every incentive to report its own performance favorably. That is a different kind of map than a ROAS dashboard, and it is the one worth building next.
The evidence, in numbers
Key findings, dated and sourced
-
In 15 large-scale U.S. Facebook field experiments (500 million-plus user-experiment observations, 1.6 billion ad impressions), commonly used observational and attribution methods generally overestimated advertising effectiveness relative to true randomized-experiment results, though in some cases they significantly underestimated it.
established Marketing Science (INFORMS) / Northwestern & Facebook researchers, Gordon, Zettelmeyer, Bhargava, Chapsky, "A Comparison of Approaches to Advertising Measurement: Evidence from Big Field Experiments at Facebook" (2019)
-
Procter & Gamble cut roughly $200 million of digital ad spend (about 20 to 50 percent of spend on some digital channels) over nine months of 2017 and 2018, found no discernible negative impact on business performance, and saw reach increase about 10 percent.
established Procter & Gamble (Marc Pritchard, Chief Brand Officer), Corroborated across Adweek, The Drum, Marketing Week, and Marketing-Interactive reporting on P&G's digital spend cuts (2017-2018)
-
A log-level analysis of 21 major advertisers ($123 million ad spend, 35.5 billion impressions, 12 supply-chain companies participating) found only 36 cents of every dollar entering a demand-side platform effectively reaches the consumer as working media; transaction costs consumed about 29 percent of spend and non-viewable, invalid, or unmeasurable inventory consumed about 35 percent.
established Association of National Advertisers (ANA), "Programmatic Media Supply Chain Transparency Study" (2023-12)
-
In an October 2020 survey, 44 percent of marketing professionals worldwide rated social media among the hardest channels to attribute sales revenue to; a separate Integral Ad Science survey found insufficient media-quality transparency would most likely affect ad spend on Facebook for 63 percent of respondents (versus 31 percent for YouTube, 28 percent for Instagram); only 40 percent of US marketers and agencies required media partners to use buyer-approved measurement tools, and major walled gardens typically did not comply.
established eMarketer (citing DemandLab/Ascend2, Integral Ad Science, Advertiser Perceptions), "Inside walled gardens: The long-standing challenge of ad measurement and attribution" (2020-10)
-
Across 225 geo-based incrementality tests run August 2024 to December 2025, overall median incremental ROAS was 2.31x (interquartile range 1.36x to 3.24x), and the gap between platform-reported ROAS and true incremental ROAS often reaches 2 to 3x, with branded search and retargeting showing 5 to 10x inflation.
contested Stella (StellaHeyStella incrementality testing platform), "2025 DTC Digital Advertising Incrementality Benchmarks" (2026-05-26)
-
Median incremental ROAS by channel in the same 225-test panel: Tatari CTV 3.30x, Google Performance Max 2.98x, Meta 2.92x, Google YouTube 2.17x, Google Shopping 1.86x, Google Search non-branded 1.46x, TikTok 0.94x, Google Search branded 0.70x.
contested Stella (StellaHeyStella incrementality testing platform), "2025 DTC Digital Advertising Incrementality Benchmarks" (2026-05-26)
-
In a separate panel of 46 Meta incrementality studies (ecommerce brands, $15 million to $100 million annual revenue, January to April 2025), Meta delivered 21 percent more incremental revenue than the platform itself reported, the opposite direction from a flat 'platform overstates' assumption.
contested Stella (StellaHeyStella incrementality testing platform), "Incrementality Study: How Incremental is Meta Really?" (2025)
-
Across 640 Meta geo-holdout incrementality experiments, for every $100 in platform-attributed DTC revenue (7-day click window), Meta actually generated $115 in incremental revenue in aggregate, a 15 percent platform understatement; Advantage+ and automated campaigns tended to over-report versus manual campaigns, and mid- and upper-funnel campaigns tended to under-report, especially when measured DTC-only; 32 percent of Meta's total incremental impact lifted non-DTC sales a DTC-only ROAS calculation would miss entirely.
contested Haus, "Is Meta Incremental?" (2025-08-13)
-
In a survey of 210 senior marketing leaders across brands and agencies, 91 percent said platform-reported results are overstated to some degree; over 80 percent rely primarily on proxy signals rather than verified purchase data when optimizing; 35 percent say proxy-based optimization decisions don't hold up when reconciled against actual sales; two-thirds estimate at least 11 percent of budget is wasted to attribution and optimization lag, and a third believe the figure exceeds 26 percent.
contested Affinity Solutions / Outcomes Marketing Council, "Nearly 91% of marketers believe their platform results are overstated" (Marketing Dive sponsored content placement) (2026-05-18)
Learning outcomes
What this study teaches
- Stop treating platform ROAS as inflated by a fixed percentage. The best evidence shows the bias runs in different directions depending on the channel and the funnel stage, not as one uniform haircut.
- Be most skeptical of branded search and retargeting ROAS specifically. The mechanism for inflation there, crediting customers who were already going to convert, is well understood even though the sharpest multipliers come from a single interested vendor.
- Don't cut upper-funnel or automated campaigns on a weak reported ROAS number alone. At least one large, rigorous panel found that category understated its real incremental impact.
- Treat every incrementality benchmark number in circulation, including the ones in this study, as coming from vendors who sell incrementality testing. Read the direction as a useful signal and the exact multiplier with real caution.
- If you can't afford your own geo-holdout or randomized test, the better fallback is an independent view of where your market's attention actually sits, not a deeper trust in any single platform's own dashboard.
Honest limits
What this does not yet settle
- No independent, cross-platform, peer-reviewed census exists that computes a single reliable overstatement percentage applicable across all channels, verticals, and account sizes. Every rigorous source shows the bias is channel- and funnel-stage-dependent rather than a flat haircut.
- Almost all quantified incrementality-gap data available comes from ad-tech vendors that sell incrementality-testing products and therefore have a commercial incentive to demonstrate a large, fixable gap. Independent academic replication more recent than the 2019 Gordon et al. Facebook RCT study could not be located and verified in this research pass, and is flagged here as a gap rather than assumed.
- No SMB-specific (sub-$2,500 to $10,000 monthly spend) incrementality dataset was found. It is a reasonable inference from what geo-holdout and RCT testing structurally require, not a measured finding, that small businesses are the least able to detect which way the bias runs for their own account.
- No evidence connects AI-search or generative-answer adoption to the platform-ROAS-versus-incremental-ROAS gap. None should be assumed; this study does not draw that connection.
- The 91 percent marketer-belief figure measures perception via a vendor-commissioned survey, not a validated measurement of actual overstatement size, and its sampling methodology and margin of error were not disclosed.
This is a synthesis of dated, attributed evidence, not a census. The AI-answer layer in particular has no independent, Nielsen-grade measurement yet, so readings of it are directional and named as a frontier, never presented as settled.
Straight answers
Frequently asked questions
Is platform-reported ROAS always inflated?
No, and that is the central finding of this study. The one peer-reviewed source, a 2019 Marketing Science study of 15 Facebook field experiments covering more than 500 million user-experiment observations, found that commonly used attribution methods generally overestimated advertising effectiveness compared to true randomized results, but in some cases significantly underestimated it. The direction of the bias depends on the channel and the funnel stage, not a fixed rule.
How much does branded search and retargeting ROAS overstate the real return?
Stella's panel of 225 geo-based incrementality tests found branded search and retargeting inflated 5 to 10 times against platform-reported numbers, with Google Search branded spend coming in at a median incremental ROAS of just 0.70x, meaning it was likely losing money on an incremental basis. This figure is vendor-sourced from a company that sells incrementality testing, so the study labels it contested and treats the multiplier as directional rather than a fixed rate.
Can a platform ever understate its own ROAS?
Yes. Haus found that across 640 Meta geo-holdout experiments, for every 100 dollars in platform-attributed DTC revenue under a 7-day click window, Meta actually generated 115 dollars in incremental revenue, a 15 percent understatement. A separate Stella panel of 46 Meta studies found Meta delivering 21 percent more incremental revenue than the platform reported, with automated and upper-funnel campaigns most likely to be undercounted.
Why does this problem hit small and midsize businesses hardest?
Every rigorous source in this study, from the 2019 Facebook RCT to the geo-holdout panels, required spend and conversion volume far beyond what most small advertisers generate, and no SMB-specific incrementality dataset was found for accounts spending in the sub-2,500 to 10,000 dollar monthly range. That is stated as a reasonable inference from what incrementality testing structurally requires, not a measured finding about SMB accounts specifically, but it leaves smaller businesses choosing between an unreliable platform number and testing infrastructure built for brands ten times their size.
If platform ROAS cannot be trusted, what should a business do instead?
The study recommends treating platform-reported ROAS as one signal among several rather than ignoring it: be most skeptical of branded search and retargeting ROAS, and be slower to cut upper-funnel or automated campaigns on a weak reported number alone, since at least one rigorous panel found that category understated its real impact. Where a business cannot afford its own geo-holdout or randomized test, the study frames an independent view of where a market's attention actually sits as the better fallback than deeper trust in any single platform's dashboard.
Provenance
References
- Gordon, Zettelmeyer, Bhargava, Chapsky, "A Comparison of Approaches to Advertising Measurement: Evidence from Big Field Experiments at Facebook," Marketing Science (INFORMS), 2019. https://pubsonline.informs.org/doi/10.1287/mksc.2018.1135
- Adweek and corroborating trade press, reporting on Procter & Gamble's 2017-2018 digital ad spend cuts. https://www.adweek.com/brand-marketing/when-procter-gamble-cut-200-million-in-digital-ad-spend-its-marketing-became-10-more-effective/
- Association of National Advertisers, "Programmatic Media Supply Chain Transparency Study," December 2023. https://www.ana.net/miccontent/show/id/rr-2023-12-ana-programmatic-media-supply-chain-transparency-study
- eMarketer, "Inside walled gardens: The long-standing challenge of ad measurement and attribution," October 2020. https://www.emarketer.com/content/inside-walled-gardens-long-standing-challenge-of-ad-measurement-attribution
- Stella (StellaHeyStella), "2025 DTC Digital Advertising Incrementality Benchmarks," May 2026. https://www.stellaheystella.com/blog/2025-dtc-digital-advertising-incrementality-benchmarks
- Stella (StellaHeyStella), "Incrementality Study: How Incremental is Meta Really?", 2025. https://www.stellaheystella.com/blog/incrementality-study-how-incremental-is-meta-really
- Haus, "Is Meta Incremental?", August 2025. https://www.haus.io/blog/is-meta-incremental
- Affinity Solutions / Outcomes Marketing Council, sponsored content via Marketing Dive, May 2026. https://www.marketingdive.com/spons/nearly-91-of-marketers-believe-their-platform-results-are-overstated-here/816171/
Every measured figure is dated to its capture and tagged with an evidence tier. Every cited work is real and locatable. Where an engine could not be captured this round, it is named as uncaptured, not estimated. Small-sample readings are labelled as directional.