Conversion Science · emerging evidence
Three Companies, Three Core Web Vitals, Three Revenue Outcomes: What the LCP and INP Case Studies Actually Prove
Core Web Vitals are Google's field measurements of how a page loads, responds, and holds still for real users. Three widely cited case studies published on Google's web.dev show that improving them can move revenue: Rakuten 24 reported a 53.37 percent rise in revenue per visitor after cutting Largest Contentful Paint, Vodafone Italy reported an 8 percent sales increase after a 31 percent LCP improvement, and redBus reported a 7 percent sales increase after improving Interaction to Next Paint. Read carelessly, those numbers become a promise. Read closely, they are something narrower and more useful: single-company before-and-after records, each an existence proof that the link between Core Web Vitals and money is real, none of them a controlled experiment you can lift and apply to your own site as a fixed return. This article separates what the three case studies prove from what they cannot, and shows how to use them without overclaiming.
The three case studies, stated precisely
Google publishes a Core Web Vitals case-study series on web.dev in which named companies document a performance change and the business metric that moved alongside it. Three of these are cited far more than the rest, usually stripped of their context and reduced to a single impressive percentage. Stated in full, they are more instructive than the headline versions.
Rakuten 24, a Japanese online store, optimized Largest Contentful Paint and reported a 53.37 percent increase in revenue per visitor and a 33.13 percent increase in conversion rate over the measured period. Vodafone Italy improved its LCP by 31 percent and reported an 8 percent increase in sales. redBus, an Indian bus-ticketing platform, improved Interaction to Next Paint, the metric that captures how quickly a page answers a tap, and reported a 7 percent increase in sales.
Every number in that paragraph is real, publicly documented, and attributable. Each also comes from exactly one company, measured before and after its own work, on its own traffic. That single fact governs everything that follows.
What a Core Web Vitals case study can prove, and what it cannot
A before-and-after case study is a record of one thing that happened once. It can establish that an outcome is possible. It cannot establish how often that outcome occurs, how large it is on average, or whether the change being credited actually caused it. In the language of evidence, these are existence proofs, not effect sizes.
The distinction matters because the numbers invite a specific error. When someone reads that Rakuten 24 gained 53.37 percent revenue per visitor, the tempting inference is that faster LCP yields something in the neighborhood of a 50 percent lift. That inference treats a single observation as an elasticity, a stable ratio between an input and an output that holds across cases. Nothing in a single before-and-after record supports that leap.
No counterfactual
The core problem is the missing counterfactual. To know that a speed improvement caused the revenue change, you would need to know what Rakuten 24's revenue would have done over the same period without the change. A controlled experiment supplies that by holding back a randomized control group and comparing. A before-and-after case study has no control group. It compares a company to its own past, during which any number of other things also changed.
Confounds that travel with a redesign
Performance work rarely ships alone. Cutting LCP often means new image handling, cleaner markup, deferred scripts, and sometimes a redesigned template. Any of those can move conversion on its own. Seasonality, a marketing push, a pricing change, or a shift in traffic mix during the measurement window can move it too. The case study attributes the gain to speed because speed is what the team set out to fix, but the record cannot separate the speed effect from everything riding alongside it.
Why before-and-after inflates, in a direction we can name
The deeper point is not merely that a single case study is uncertain. It is that uncontrolled before-and-after measurement tends to overstate the effect of the thing being studied, and the mechanism is well documented in the measurement literature.
The strongest evidence comes from a large randomized field experiment at eBay. Blake, Nosko and Tadelis showed that observational estimates of paid-search return, the kind produced by ordinary attribution reporting, were inflated relative to the true causal effect measured experimentally, because ad clicks correlate with buyers who were already going to purchase. The general lesson generalizes past advertising: when the population you measure after an intervention is not a random slice, non-experimental estimates carry the pre-existing intent of that population, not the clean effect of the change. A company motivated enough to invest in a serious performance project, measured during the window it chose to report, is not a neutral sample.
Ron Kohavi, Diane Tang and Ya Xu, drawing on more than twenty thousand experiments a year at Microsoft, formalize the same caution as a working rule: a surprising result is usually an error until it survives a checklist of validity threats. Their canon exists precisely because uncontrolled or improperly controlled measurement produces confident numbers that do not replicate. A 53 percent figure from one company, with no control and no replication, is the archetype of a result their method would quarantine before trusting.
The large-N corroboration, and its own limits
If the three case studies were the only evidence, the fair verdict would be "interesting, unproven." They are not. A much larger study points the same way, which is why the direction of the effect is credible even though the magnitudes are not transferable.
"Milliseconds Make Millions," commissioned by Google and conducted by the agency 55 and Deloitte Digital, analyzed more than thirty million real user sessions across thirty-seven European and American brand sites. A 0.1 second improvement in mobile load speed was associated with an 8.4 percent increase in retail conversions and a 9.2 percent increase in average order value, with travel conversions rising 10.1 percent. That is a far stronger design than any single case study: many brands, many sessions, a consistent measured relationship.
It still is not a randomized controlled trial, and it was commissioned by a party with an interest in the finding, so it corroborates rather than proves. But it does the one thing the three case studies cannot do on their own: it establishes that the speed-to-revenue relationship holds across a population, not just in three flattering anecdotes. The case studies then become believable illustrations of a real pattern, which is exactly the weight they can bear and no more.
The tiers of evidence
Not all of the evidence in this article is equally strong, and treating it as if it were would repeat the error the article is about. It is worth sorting explicitly.
How to use these case studies without overclaiming
The case studies are useful the moment you stop asking them the wrong question. "How much will speed lift my revenue?" is unanswerable from this evidence. "Is it plausible that a slow page is costing me measurable revenue, and worth measuring on my own site?" is answered clearly: yes.
That reframes the work. The correct move is not to borrow Rakuten 24's 53 percent, or Vodafone's 8 percent, or redBus's 7 percent, and forecast it onto your business. It is to measure your own page against the thresholds Google actually reads, find where it fails, fix those specific faults, and observe your own before-and-after with clear eyes about its limits. Your own N of one is still an N of one, but it is measured on your traffic and your buyers, which makes it the only figure that describes you.
Google grades Core Web Vitals on field data at the seventy-fifth percentile of real visitors, against published thresholds: Largest Contentful Paint at or under 2.5 seconds, Interaction to Next Paint at or under 200 milliseconds, and Cumulative Layout Shift at or under 0.1. A page can look finished in the browser and still fail these checks for the slice of users on slower devices and networks. INP, the responsiveness metric redBus improved, is a common failure point because it captures the lag after a tap that a fast-looking page can still hide.
The rule this leaves you with
Three companies improved three Core Web Vitals and reported three revenue gains. The evidence is real, it is documented, and pointed in the same direction as a study of thirty million sessions. That is enough to justify taking page performance seriously as a revenue lever. It is not enough to promise a number, and any vendor who quotes you Rakuten 24's 53 percent as your expected return is misreading a case study, or hoping you will.
The position worth holding is a measured one: Core Web Vitals can move revenue, the size of the move is specific to your site and can only be found by measuring it, and the work that pays off is fixing the faults on your own page against the thresholds Google reads, not chasing someone else's percentage.
The evidence
Key findings, with their sources
-
Rakuten 24 reported a 53.37% increase in revenue per visitor and a 33.13% increase in conversion rate after optimizing Largest Contentful Paint.
emerging Google web.dev, Core Web Vitals case-study series (Rakuten 24), web.dev/case-studies. Single-company before/after case study, not a controlled experiment.
-
Vodafone Italy improved LCP by 31% and reported an 8% increase in sales.
emerging Google web.dev, Core Web Vitals case-study series (Vodafone Italy), web.dev/case-studies. Single-company before/after case study.
-
redBus improved Interaction to Next Paint (INP) and reported a 7% increase in sales.
emerging Google web.dev, Core Web Vitals case-study series (redBus), web.dev/case-studies. Single-company before/after case study.
-
A 0.1-second improvement in mobile load speed was associated with an 8.4% increase in retail conversions and a 9.2% increase in average order value, across more than 30 million sessions on 37 brand sites.
established Google, agency 55 & Deloitte Digital, "Milliseconds Make Millions", 2020. Large-N industry study, Google-commissioned; treat as corroborating, not fully independent.
-
Observational (attribution-style) estimates of return were inflated relative to the true causal effect measured experimentally, because clicks correlate with buyers already intending to purchase.
established Blake, Nosko & Tadelis, "Consumer Heterogeneity and Paid Search Effectiveness: A Large-Scale Field Experiment", Econometrica 83(1), 2015.
-
Uncontrolled measurement produces confident numbers that do not replicate; a surprising result is usually an error until it survives a validity checklist.
established Kohavi, Tang & Xu, "Trustworthy Online Controlled Experiments", Cambridge University Press, 2020 (built on 20,000+ experiments/year at Microsoft).
Calibration
What is proven, what is promising, what is unproven
| Evidence tier | Tactics | What the evidence says |
|---|---|---|
| Established | Page speed relates to revenue across a population of sites; uncontrolled before/after overstates effects; observational estimates carry pre-existing intent. | "Milliseconds Make Millions" (30M+ sessions); Blake, Nosko & Tadelis (Econometrica 2015); Kohavi, Tang & Xu (2020). |
| Emerging | A specific Core Web Vitals improvement produced a specific revenue gain at a named company (Rakuten 24, Vodafone Italy, redBus). | Google web.dev Core Web Vitals case-study series. Real and documented, but N=1 per case and not randomized experiments. |
| Contested / do not claim | That another company's reported percentage (for example 53.37%) is a reliable forecast of your own revenue lift. | No evidence supports transferring a single before/after figure as an elasticity; the measurement literature predicts it will not hold. |
Reference
Glossary
- Core Web Vitals
- Google's set of field measurements of real user experience: Largest Contentful Paint (loading), Interaction to Next Paint (responsiveness), and Cumulative Layout Shift (visual stability), graded at the 75th percentile of real visitors.
- Largest Contentful Paint (LCP)
- How long the largest visible element takes to render. Google's "good" threshold is 2.5 seconds or less at the 75th percentile.
- Interaction to Next Paint (INP)
- How quickly a page visibly responds after a user taps or clicks. Google's "good" threshold is 200 milliseconds or less. The metric redBus improved.
- Existence proof
- Evidence that an outcome is possible, shown by a single documented instance. It establishes that something can happen, not how often or how large the effect is on average.
- Elasticity
- A stable ratio between an input and an output that holds across cases. A single before-and-after result is not an elasticity and cannot be used as one.
- Counterfactual
- What would have happened without the change. A controlled experiment estimates it with a randomized control group; a before-and-after case study has none.
Straight answers
Frequently asked questions
What did the Rakuten 24, Vodafone Italy, and redBus case studies actually find?
Each documented a Core Web Vitals improvement alongside a business gain: Rakuten 24 reported a 53.37% rise in revenue per visitor after cutting LCP, Vodafone Italy an 8% sales increase after a 31% LCP improvement, and redBus a 7% sales increase after improving INP. All three are real and published by Google on web.dev, and all three come from a single company measuring itself before and after.
Do Core Web Vitals affect revenue?
The evidence supports that page performance and revenue are related. A study of more than 30 million sessions across 37 brands ("Milliseconds Make Millions") found a 0.1-second mobile speed improvement associated with an 8.4% retail conversion increase. The three case studies illustrate the same pattern at named companies. The direction is credible; the exact size depends on your site.
Why can't I use the 53% figure to forecast my own revenue lift?
Because it is one company's before-and-after result with no control group, measured during a window that also contained other changes and a self-selected audience. Treating a single observation as a fixed return ignores the missing counterfactual and the confounds that travel with a redesign. The measurement literature predicts that such transfers overstate the real effect.
What is a fair way to use these case studies?
As existence proofs. They justify measuring your own page against the thresholds Google reads and fixing what fails, then observing your own before-and-after with clear eyes about its limits. They do not justify borrowing another company's percentage as your expected outcome.
What are the Core Web Vitals thresholds Google grades against?
Google grades field data at the 75th percentile of real visitors: Largest Contentful Paint at or under 2.5 seconds, Interaction to Next Paint at or under 200 milliseconds, and Cumulative Layout Shift at or under 0.1. A page can look finished and still fail these for users on slower devices and connections.
Provenance
Sources
- Google, web.dev Core Web Vitals case-study series (Rakuten 24, Vodafone Italy, redBus), web.dev/case-studies (emerging, N=1 per case, not randomized experiments)
- Google, agency 55 & Deloitte Digital, "Milliseconds Make Millions", 2020 (established; large-N, Google-commissioned, corroborating rather than fully independent)web.dev
- Blake, T., Nosko, C. & Tadelis, S., "Consumer Heterogeneity and Paid Search Effectiveness: A Large-Scale Field Experiment", Econometrica 83(1), 2015, 155-174 (established, top-tier peer-reviewed field experiment)
- Kohavi, R., Tang, D. & Xu, Y., "Trustworthy Online Controlled Experiments: A Practical Guide to A/B Testing", Cambridge University Press, 2020 (established, practitioner-academic canon)cambridge.org
- Fabijan, A. et al., "Diagnosing Sample Ratio Mismatch in Online Controlled Experiments", KDD '19, doi:10.1145/3292500.3330722 (established, on why uncontrolled splits invalidate results)doi.org
Every figure above is attributed to a real, dated source and tagged with its evidence tier. Where a claim could not be verified to a primary source, it is not stated as fact.