Choice Science · established evidence
The Yelp Effect, Ten Years Later: What a Regression Discontinuity Actually Proved
The Yelp effect is the finding that a business rating on a review platform can move revenue by a measurable amount. The single most cited estimate comes from Michael Luca, who used a regression discontinuity design around Yelp's rating-rounding thresholds, matched to Washington State tax records, and found that a one-star increase in a restaurant's Yelp rating produces a 5 to 9 percent revenue increase. The number is real and causally identified, which is rare in marketing research. What is widely misread is its scope: the effect was found for independent restaurants only, chains showed no relationship, and the study measured restaurants in one state at one time. Treating "5 to 9 percent per star" as a universal law for every business is the mistake this article exists to correct.
What the Yelp effect actually is
In a 2011 Harvard Business School working paper, "Reviews, Reputation, and Revenue: The Case of Yelp.com," Michael Luca set out to answer a question that sounds simple and is not: does a better Yelp rating cause a business to earn more, or do good businesses simply earn both better ratings and more revenue for the same underlying reason?
The headline result is precise. A one-star increase in a restaurant's Yelp rating is associated with a 5 to 9 percent increase in revenue. That is the figure that entered the industry's vocabulary as "the Yelp effect." But the figure matters far less than the design behind it, because the design is what lets Luca say cause rather than correlation, and it is the design that also fences off exactly where the claim stops being true.
Why a regression discontinuity can claim cause, not just correlation
The obvious way to study ratings and revenue is to line up businesses by their star rating and see whether higher-rated ones earn more. They do, but that comparison proves nothing about cause. A restaurant with a genuinely better kitchen will tend to earn both a higher rating and higher revenue, so the two move together without one driving the other. This is the endogeneity problem, and it defeats almost every casual claim about reviews.
Luca's move was to exploit a quirk in how Yelp displays ratings. The site shows a rounded star rating in half-star increments, even though it computes the underlying average to far more decimal places. A restaurant with a true average of 3.24 and one with a true average of 3.26 are, on quality, essentially identical, yet the site rounds one down to 3.0 stars and the other up to 3.5. The displayed rating jumps at the rounding threshold while the actual quality does not.
That discontinuity is the natural experiment. Businesses sitting just on either side of a rounding cutoff are, on average, interchangeable in every respect except the star rating a buyer sees. Comparing revenue across that narrow boundary isolates the effect of the displayed rating itself, cleanly separated from underlying quality. To measure revenue, Luca matched the ratings data to Washington State Department of Revenue tax records rather than relying on self-reported or estimated figures. That is what earns the study its authority: a credible identification strategy joined to hard administrative revenue data.
Does a Yelp rating affect sales, and only where the study looked
The most important sentence in the paper is not the 5 to 9 percent. It is the finding that the effect is driven entirely by independent restaurants. Chain restaurants show no meaningful relationship between their Yelp rating and their revenue at all.
The reason is intuitive once stated. Buyers already hold strong prior beliefs about a chain. A person choosing a national franchise knows roughly what they will get before they read a single review, so the marginal star moves nothing. An independent restaurant has no such prior working for it. For the independent, the Yelp rating is often the only credible third-party signal a buyer has, so it carries the full weight of the decision. The revenue effect lives exactly where the information gap is widest.
Luca also documented that this dynamic is associated with a measurable decline in chain market share over the period studied, consistent with review platforms shifting attention toward independents that were previously hard to evaluate. That is a genuine and underappreciated part of the finding: review sites did not just reward good ratings, they partially re-leveled the field between independents and chains.
What the study does not prove
Because the number is quotable, it gets stretched far past what the evidence supports. Three limits deserve to be stated plainly.
It does not prove that more reviews, by themselves, equal more money
The study identifies the effect of a change in the displayed star rating at a rounding threshold. It is not a study of review volume, review velocity, or review recency, and it does not license the claim that adding reviews mechanically adds revenue. Those may matter, but they are separate questions the regression discontinuity was not built to answer, and importing the 5 to 9 percent figure to justify a volume campaign is a misuse of it.
It does not prove the star average is a reliable measure of quality
A higher rating causing more revenue is not the same as a higher rating meaning better quality. In a separate and sobering body of work, de Langhe, Fernbach and Lichtenstein examined 1,272 products across 120 categories and found that average user ratings did not converge with independent Consumer Reports quality scores, were frequently built on too few ratings to be statistically informative, and ran higher for pricier and premium-brand items independent of actual quality, even as buyers lean on the star average more heavily than on better cues. The rating is a powerful persuasion signal, as Luca shows. Whether it is an accurate quality signal is a different matter, and the evidence says it is weaker than buyers assume.
It does not prove the rating is a clean input you can simply raise
The rating a business carries is itself the product of strategy, not a neutral thermometer. Luca and Zervas, using Yelp's own filtered-review flags as a fraud proxy, found that fake reviews are more common for businesses with weak existing reputations and rise when a business faces more direct competition. Reputation manipulation is a rational, predictable response to competitive pressure. That does not weaken Luca's causal claim about the displayed rating, but it does mean the input side of the story is adversarial, which is precisely why raising a rating honestly is harder, and more valuable, than the headline number suggests.
The older correlational evidence agrees on direction, and on asymmetry
Luca's work did not appear in a vacuum. Five years earlier, Chevalier and Mayzlin compared relative book sales across Amazon.com and BarnesAndNoble.com and found that a one-star improvement in a book's average rating correlated with up to a 9.9 percent increase in relative sales. That study is correlational rather than causally identified, so it carries less weight on the question of cause, but its direction and rough magnitude are consistent with Luca's later, cleaner estimate.
Chevalier and Mayzlin added a finding that reframes how a business should think about its reputation: the impact of one-star reviews was larger in magnitude than the impact of five-star reviews. Negative information moved buyers more than an equivalent amount of positive information, an asymmetry consistent with loss aversion. Put together with Luca, the practical reading is that a bad rating is not the mirror image of a good one. The downside of a poor reputation is heavier than the upside of a strong one, which is an argument for defense, not just accumulation.
How the effect gets misread ten years on
The gap between what the study proved and how it is cited has a predictable shape. The common misreadings are worth naming directly.
- Treating "5 to 9 percent per star" as a universal constant. It was estimated for independent restaurants in Washington State. Applying the exact number to a med spa, a law firm, or a home-services contractor is an extrapolation, not a finding.
- Assuming the effect exists for every business type. It did not appear for chains at all, because buyers already held strong priors about them. Any business with a strong pre-existing brand should expect a smaller marginal effect than an unknown independent.
- Confusing this causal study with the many correlational ones. Most "reviews drive revenue" claims online rest on studies that cannot separate cause from quality. Luca's can. That distinction is the whole reason the paper is cited, and collapsing it flatters weaker evidence.
- Reading a persuasion effect as a quality guarantee. The rating moves buyers whether or not it accurately reflects quality, so a rise in stars is not proof the underlying service improved.
- Assuming the 2011 restaurant result transfers unchanged to today's AI-answer surfaces. That is a live extrapolation, examined below, not something the study established.
Does the effect transfer to the AI-answer era?
It is tempting to assume that if a star rating moved restaurant revenue on a list of Yelp results, it must move revenue even more now that an AI engine reads reviews and names a single business in its answer. The intuition is reasonable, and the underlying mechanics, a trusted third-party signal filling an information gap, plausibly still apply. But it should be labeled for what it is: an extrapolation, not a measured result.
The causal literature, Luca included, studied ranked lists and human browsing. It did not study single-answer generative surfaces, which did not exist when the data was collected. When an engine synthesizes one recommendation instead of returning ten, the same trust and social-proof mechanics may concentrate onto whichever business the engine names, raising the stakes of the rating rather than lowering them. That is a defensible hypothesis and a reason to take reputation seriously in the AI-answer era. It is not something the regression discontinuity proved, and anyone who claims it did is overreaching.
Reputation is now a regulated, adversarial system
One thing has changed unambiguously since 2011. Reputation is no longer an honor system. In October 2024 the Federal Trade Commission's Trade Regulation Rule on the Use of Consumer Reviews and Testimonials, 16 CFR Part 465, made fake, incentivized, insider, and suppressed reviews federal violations carrying civil penalties of up to 51,744 dollars each. At the same time, Google's own Search Quality Rater Guidelines name trust as the load-bearing member of its Experience, Expertise, Authoritativeness and Trust framework, the criterion the other three feed into.
Read alongside the academic record, the message is consistent. A rating genuinely moves revenue, especially for an independent with no brand to fall back on. That rating is a persuasion signal more than a quality one, it responds asymmetrically to bad news, and it now sits inside a compliance regime where the shortcuts are illegal. Building a reputation that earns the Yelp effect honestly, and can survive scrutiny, is a discipline, not a growth hack.
The evidence
Key findings, with their sources
-
A one-star increase in a restaurant's Yelp rating produces a 5 to 9 percent revenue increase, identified with a regression discontinuity design around Yelp's rounding thresholds matched to Washington State tax records.
established Luca, M., "Reviews, Reputation, and Revenue: The Case of Yelp.com," Harvard Business School Working Paper 12-016, 2011/2016.
-
The revenue effect is driven entirely by independent restaurants; chain restaurants show no meaningful rating-to-revenue relationship, and the dynamic is associated with a documented decline in chain market share.
established Luca, M., "Reviews, Reputation, and Revenue: The Case of Yelp.com," Harvard Business School Working Paper 12-016, 2011/2016.
-
A one-star improvement in a book's average rating correlated with up to a 9.9 percent increase in relative sales, and one-star reviews moved sales more in magnitude than five-star reviews (a loss-aversion-consistent asymmetry).
established Chevalier, J.A. & Mayzlin, D., "The Effect of Word of Mouth on Sales: Online Book Reviews," Journal of Marketing Research, 43(3), 2006.
-
Across 1,272 products in 120 categories, average user ratings did not converge with independent Consumer Reports quality scores and ran higher for pricier items independent of actual quality, yet buyers weight the star average heavily.
contested de Langhe, B., Fernbach, P.M. & Lichtenstein, D.R., "Navigating by the Stars," Journal of Consumer Research, 42(6), 2016.
-
Fake reviews are more common for businesses with weak existing reputations and rise when a business faces more direct competition, showing manipulation is a strategic response to competitive pressure.
established Luca, M. & Zervas, G., "Fake It Till You Make It: Reputation, Competition, and Yelp Review Fraud," Management Science, 62(12), 2016.
-
Since October 2024, fake, incentivized, insider, and suppressed reviews are federal violations carrying civil penalties of up to 51,744 dollars each.
established Federal Trade Commission, "Trade Regulation Rule on the Use of Consumer Reviews and Testimonials," 16 CFR Part 465, effective Oct 21, 2024.
Calibration
What is proven, what is promising, what is unproven
| Evidence tier | Tactics | What the evidence says |
|---|---|---|
| established | The 5 to 9 percent per-star revenue effect for independent restaurants; the asymmetry of negative over positive ratings | Causally identified in Luca 2011 via regression discontinuity on tax records; direction corroborated correlationally by Chevalier & Mayzlin 2006. |
| contested | Reading the star average as a reliable proxy for underlying quality | de Langhe et al. 2016 show the average rating is a weaker quality signal than buyers assume; it persuades more reliably than it measures. |
| emerging | Transferring the 2011 restaurant effect to other verticals and to single-answer AI surfaces | A reasonable extrapolation from the same trust mechanics, but not directly measured; the causal studies covered ranked lists and human browsing, not generative answers. |
Reference
Glossary
- Regression discontinuity design
- A quasi-experimental method that estimates a causal effect by comparing units just above and just below an arbitrary threshold, where they are near-identical except for the treatment the threshold assigns. Luca used Yelp's star-rounding cutoffs as the threshold.
- Yelp rounding threshold
- The point at which Yelp rounds a business's true underlying average rating up or down to the nearest half-star for display. Two businesses on either side are essentially equal in quality but show different star ratings.
- Endogeneity
- When a supposed cause and its outcome are both driven by a third unmeasured factor, so their correlation cannot be read as causation. Here, genuine quality raises both rating and revenue, which is what the discontinuity design was built to get around.
- Loss aversion
- The tendency for losses to loom larger than equivalent gains. In reviews, it appears as negative ratings moving buyers more than an equal amount of positive rating, as found by Chevalier & Mayzlin.
- E-E-A-T
- Experience, Expertise, Authoritativeness and Trust: the criteria in Google's Search Quality Rater Guidelines, with trust named as the load-bearing member the others feed into.
Straight answers
Frequently asked questions
Does a Yelp rating actually affect sales?
For independent restaurants, yes, and causally. Michael Luca's 2011 study used a regression discontinuity design around Yelp's rounding thresholds, matched to Washington State tax records, and found a one-star increase produced a 5 to 9 percent revenue increase. The design separates the effect of the displayed rating from underlying quality, which is why it can claim cause rather than correlation.
How much is one Yelp star worth?
The best causal estimate is 5 to 9 percent of revenue per star, but that figure was measured for independent restaurants in one US state. It is a strong reference point, not a universal constant, and applying the exact number to a different business type is an extrapolation the study did not make.
Does the 5 to 9 percent figure apply to my med spa or home-services business?
Not directly. The study measured restaurants. The underlying logic, that a third-party rating carries more weight when a buyer lacks a strong prior, plausibly extends to any independent local business with an information gap, which describes most med spas and home-services firms. But the direction likely transfers while the exact magnitude is unproven for those verticals.
Why did the effect not show up for chain restaurants?
Because buyers already hold strong prior beliefs about a chain and know roughly what they will get before reading any review, so the marginal star changes little. An independent has no such prior, so its rating carries the full weight of the decision. The revenue effect lives where the information gap is widest.
Is a higher star average proof of better quality?
No. Luca shows the rating moves buyers, but separate work by de Langhe, Fernbach and Lichtenstein found average ratings did not track independent quality scores across 1,272 products and ran higher for pricier items regardless of quality. The star average is a strong persuasion signal and a weaker quality signal than buyers assume.
Provenance
Sources
- Luca, M., "Reviews, Reputation, and Revenue: The Case of Yelp.com," Harvard Business School Working Paper 12-016, 2011/2016 (established)
- Chevalier, J.A. & Mayzlin, D., "The Effect of Word of Mouth on Sales: Online Book Reviews," Journal of Marketing Research, 43(3), 2006 (established)
- de Langhe, B., Fernbach, P.M. & Lichtenstein, D.R., "Navigating by the Stars: Investigating the Actual and Perceived Validity of Online User Ratings," Journal of Consumer Research, 42(6), 2016 (established; qualifies the star average as a quality signal)
- Luca, M. & Zervas, G., "Fake It Till You Make It: Reputation, Competition, and Yelp Review Fraud," Management Science, 62(12), 2016 (established)
- Federal Trade Commission, "Trade Regulation Rule on the Use of Consumer Reviews and Testimonials," 16 CFR Part 465, effective Oct 21, 2024 (established, binding US regulation)ecfr.gov
- Google, Search Quality Rater Guidelines, and "E-A-T gets an extra E for Experience," Google Search Central Blog, Dec 2022 (established, primary-source policy document)
Every figure above is attributed to a real, dated source and tagged with its evidence tier. Where a claim could not be verified to a primary source, it is not stated as fact.