Vertical Playbooks · established evidence

One Star, Five to Nine Percent: What the Ratings-Revenue Literature Says About High-Consideration Local Services

Last reviewed 2026-07-20. Written by Chandranshu Kumar, Founder, Raveneye Global. · 10 min read

How much do reviews affect revenue? The most credible answer comes from a single natural experiment. Michael Luca found that a one-star increase in a restaurant's Yelp rating produced a 5 to 9 percent increase in revenue, an effect concentrated entirely in independent restaurants, with no measurable effect for chains. That is the strongest causal estimate the literature offers, and it is a restaurant estimate. No equivalent large-sample study was located for home services, dental, or fitness. The position is therefore split in two: the mechanism, a synthesized reputation signal steering which business a buyer chooses, plausibly generalizes to any high-consideration local service; the specific 5 to 9 percent magnitude does not transfer to those verticals as a number. This piece reads the evidence for both claims, marks exactly where established findings stop and extrapolation begins, and treats the gap as a call for primary data rather than a figure to borrow.

How much do reviews affect revenue? Start with the one study that measured it

Most published claims about reviews and revenue are correlations dressed as causes. A business with more five-star reviews earns more, but higher-quality businesses also earn more and attract better reviews, so the raw association cannot separate the review's effect from the underlying quality it reflects. To measure the causal value of a rating point, you need a source of rating variation that is unrelated to true quality.

Michael Luca's study of Yelp supplied exactly that. Yelp displays a business's rating rounded to the nearest half-star, so two restaurants of nearly identical underlying quality can be pushed onto different sides of a rounding threshold and shown a visibly different rating for reasons that have nothing to do with the food. That rounding is a natural experiment. Comparing restaurants just above and just below each threshold isolates the effect of the displayed rating itself.

The result is the number in this article's title. A one-star increase in the displayed Yelp rating produced a 5 to 9 percent increase in revenue. It is, to date, the cleanest causal estimate of what a rating point is worth to an independent local business, and it is why the finding has become a reference point across local-commerce and platform-economics research.

Why the effect lands on independents, not chains

The most instructive part of Luca's finding is not the headline range but its boundary. The revenue effect was concentrated in independent restaurants. Chain-affiliated restaurants showed no measurable rating effect at all.

The proposed explanation is that consumers already hold a quality prior for a chain. A national brand has taught the buyer, through advertising and repeated experience, roughly what to expect before a single review is read, so the marginal review changes little. An independent has no such buffer. For an unfamiliar local business, the synthesized rating is close to the only quality signal the buyer has, so it carries the full weight of the decision. The study also found consumers respond more strongly to ratings backed by more visible and more numerous reviews, consistent with the rating mattering most precisely where it is the buyer's primary evidence.

For the businesses this publication serves, owner-operated and small-team local services rather than franchises of national brands, this boundary is the relevant half of the finding. It says the rating point is worth most to exactly the kind of firm that has no brand prior to fall back on. It does not say the 5 to 9 percent figure is theirs to claim.

The mechanism generalizes; the magnitude may not

There is a defensible reason to expect the underlying mechanism to extend beyond restaurants. The mechanism is not about food. It is about a buyer choosing an unfamiliar provider under uncertainty, using a reputation signal that a platform has already aggregated on their behalf. That structure describes how people now select a plumber, a dentist, or a gym as much as it describes how they pick a restaurant.

Adjacent research on provider selection points the same way. A large topic-modeling study of online health communities, covering 747 doctors and 105,032 reviews, found that narrative reviews measurably shifted which provider a patient chose, and that the content of those reviews predicted choice differentially. The synthesized reputation signal moved the decision, which is the same causal role the star rating plays in the Yelp result.

Two limits keep this from being a transfer of the number. First, that provider-choice study ran on a non-US platform, so its mechanism generalizes more safely than its magnitude. Second, and more fundamentally, a mechanism that operates in the same direction can operate at a very different strength. Booking a dental implant, hiring a contractor for a roof, and choosing a restaurant for Friday differ in price, risk, switching cost, and how many other signals the buyer consults. Any of those could make the revenue response to a rating point larger or smaller than the restaurant estimate. Direction is a reasonable inference. Size is not.

What we do not know: home services, dental, and fitness

This piece has a real gap at its core. In the research assembled for it, no equivalent large-sample causal study of ratings and revenue was located for home services, dental practices, or fitness. Extending the 5 to 9 percent figure to those verticals is extrapolation, and it should be labeled as such wherever it appears.

It matters to be specific about why the missing evidence cannot simply be replaced with the numbers already circulating in vertical marketing content.

Dental: widely repeated statistics that could not be traced

Figures such as "70 percent of patients use the internet to find a new dental practice" or "77 percent cite reviews as their first step" appear across many dental-marketing sites, usually attributed to a professional association. In this research pass those specific percentages could not be traced to a disclosed primary study, including a direct check of the association's own materials, which describe how to survey patients rather than publishing those national findings. They are treated here as unverified and are not cited as fact. The established provider-choice literature is used in their place.

Fitness: a genuine white space

The strongest fitness data describes membership and engagement, not discovery. A record 81 million Americans belonged to a gym or studio in 2025, roughly 26 percent of the population aged 6 and over, but membership figures say nothing about how a consumer searches for and selects a studio, or what a rating point is worth once they do. No primary source describing fitness discovery behavior specifically was found. This is an open question, not a settled one.

Home services: economics reported by the interested party

For home services the available acquisition-cost numbers, including the frequently cited premium of lead marketplaces over owned search, come from contractor anecdote and marketing-vendor content rather than an independent audited dataset. They are directionally suggestive and evidentially thin. Using them requires the caveat that they are industry-reported estimates, not measured figures.

Why the rating point may be worth more now than when it was measured

Luca's natural experiment predates the current discovery environment, and the intervening change runs in one direction: the reputation signal now feeds more than one gate at once. The same aggregate that a human buyer reads is also read by the systems that decide who the buyer sees at all.

The local map pack is the first of these. Local searchers click the local three-pack results far more often than the organic links beneath it, and businesses that appear in the pack draw materially more calls, direction requests, and clicks than comparable businesses that do not. Review count, rating, and recency are among the inputs that influence who occupies those positions, so a rating point does not only persuade the buyer who reaches the profile; it helps decide whether the buyer reaches it.

The generative answer layer is the second gate, and an emerging one. Consumer use of AI tools to find a local business has risen sharply in reported surveys, and AI local recommendations draw heavily on business-profile and review data as their source. A rating point therefore now compounds across surfaces that did not exist when it was first measured. This is a reason to expect the present-day value of reputation to sit at or above the original estimate, and simultaneously a reason not to attach a firmer number to that expectation than the evidence supports.

The signal moves buyers, but trust in the signal is itself moving

A finding that a rating point is valuable is not a license to treat reviews as a lever that pays without limit. The reliability buyers grant the signal is itself shifting, and not only upward.

General trust in online reviews has declined over several years, from a peak near 84 percent to roughly half of consumers in recent survey data. Trust in AI-chatbot output specifically sits lower still: a 2026 survey found only about 29 percent of US chatbot users trust the information a lot or some, even as chatbot use for information search has grown. The practical reading is that reputation work compounds where the signal is credible and consistent, and decays where it looks manufactured. It also underlines why the manufactured shortcut is off the table on evidence grounds as much as legal ones: a reputation the buyer no longer believes is not an asset, whatever its star average.

How to read a "reviews lift revenue by X percent" claim

Because the demand for a clean number is strong and the supply of rigorous ones is thin, the market is full of confident percentages. The useful skill for an owner is not memorizing a figure but grading the claim behind it. Three questions separate an established finding from a borrowed one.

  • Is it causal or correlational? A number from a natural experiment or randomized design (Luca's rounding thresholds) measures the effect of the rating. A number from "businesses with more reviews earn more" only shows that quality and reviews travel together.
  • Is it the right vertical? The 5 to 9 percent estimate is a restaurant estimate. A source that reports it as the value of a rating point for a dental practice or a gym has transplanted a number across a boundary the original study never tested.
  • Is the primary source locatable? An established statistic names an author, a venue, and a year you can check. A figure attributed only to an unnamed "study" or a trade association with no locatable document is a fabrication risk, and several of the most repeated vertical statistics fail this test.

What this means for a high-consideration local business

The evidence supports a precise and limited conclusion. For an independent, owner-operated local business, the kind with no national brand prior doing the persuading, a better reputation signal plausibly carries real revenue weight, and that weight may be larger today than when it was measured because the same signal now governs visibility across the map pack and AI answers as well as the buyer's own reading. The direction is well founded. The exact size, in any vertical other than restaurants, is not yet known and should not be quoted as though it were.

What follows from that is not a promise of a percentage. It is a case for building the reputation signal on real evidence and measuring the movement against a real baseline, so that the value in your own vertical becomes something you observe rather than something you borrow from a restaurant study. That is the difference between a marketing claim and a measured outcome.

The evidence

Key findings, with their sources

  • A one-star increase in a restaurant's Yelp rating produced a 5 to 9 percent increase in revenue.

    established Luca, M., "Reviews, Reputation, and Revenue: The Case of Yelp.com", Harvard Business School Working Paper 12-016, 2011 (rev. 2016).

  • The revenue effect was concentrated in independent restaurants; chain-affiliated restaurants showed no measurable rating effect, plausibly because consumers already hold a brand-based quality prior.

    established Luca, M., HBS Working Paper 12-016, 2011 (rev. 2016).

  • Narrative reviews measurably shifted which provider a patient chose, with review content differentially predicting choice, across 747 doctors and 105,032 reviews.

    established Zhang M, Sun Y, Zhao X, Wang L, Xiong J, "The Impact of Narrative Reviews on Patient E-doctor Choice in Online Health Communities", INQUIRY, 2023, PMID 37357728 (non-US platform; mechanism generalizes, magnitude may not).

  • Local searchers click the local three-pack results about 44 percent of the time versus 29 percent for organic, and pack businesses draw materially more calls, directions, and clicks than non-pack businesses.

    emerging Aggregated Google local-search behavior studies as reported via industry local-SEO research, 2025.

  • Reported consumer use of an AI tool to find a local-business recommendation rose to about 45 percent in the trailing year as of the 2026 survey, from 6 percent a year earlier.

    emerging BrightLocal, Local Consumer Review Survey 2026 (single-source year-over-year swing, worth independent verification).

  • About 29 percent of US chatbot users trust the information "a lot" or "some", and general trust in online reviews has declined from a peak near 84 percent to roughly half of consumers.

    established Pew Research Center, "Americans and AI 2026", June 17, 2026; BrightLocal Local Consumer Review Survey (trend).

  • A record 81 million Americans belonged to a gym or studio in 2025 (about 26 percent of the population aged 6 and over), a membership figure that does not describe discovery or selection behavior.

    established Health & Fitness Association (formerly IHRSA), 2025 US Health & Fitness Consumer Report.

Calibration

What is proven, what is promising, what is unproven

Evidence tierTacticsWhat the evidence says
EstablishedThe causal value of a rating point for independent restaurants (a one-star increase worth 5 to 9 percent of revenue), and the no-effect boundary for chains.Luca, HBS Working Paper 12-016, 2011 (rev. 2016), a quasi-experiment using Yelp half-star rounding.
Established (mechanism)That a synthesized reputation signal shifts which provider a buyer chooses in high-consideration, credence-like services.Zhang et al. 2023 (provider choice, non-US); local map-pack click and action behavior, 2025.
EmergingThat AI answer engines now use reputation and profile data as a local source, and that consumer AI discovery is rising, raising the present-day stakes of the signal.BrightLocal 2025 to 2026; Pew Research Center 2026 (single-vendor or fast-moving survey data).
Extrapolation, needs primary dataApplying the specific 5 to 9 percent magnitude to home services, dental, or fitness, where no equivalent causal study was located.No source. Flagged for RavenEye Visibility Corpus measurement against real client baselines.

Reference

Glossary

Natural experiment
A situation where something outside the researcher's control varies a factor almost at random, letting a cause be isolated without a lab trial. Yelp's half-star rounding is the natural experiment behind the ratings-revenue estimate.
Reputation buffer
The stock of prior belief a known brand carries, which makes any single review matter less. Chains have one; an unfamiliar independent business does not, which is why the rating point is worth more to the independent.
Credence good
A product or service whose quality the buyer cannot fully verify even after purchase (medical, legal, many home repairs). Buyers of credence goods lean harder on reputation signals, because direct verification is not available to them.
Extrapolation
Extending a finding beyond the setting it was measured in. Applying a restaurant estimate to a dental practice is extrapolation: it may be right in direction but is unproven in size.
Share of answer
How often a business is named across the surfaces that now synthesize local recommendations, including the map pack and AI answers, as distinct from where it ranks in a list of links.

Straight answers

Frequently asked questions

How much do reviews affect revenue?

The best causal estimate is Michael Luca's Yelp study, which found a one-star increase in rating raised an independent restaurant's revenue by 5 to 9 percent. That figure is specific to restaurants. For other local-service verticals the mechanism plausibly holds, but the exact percentage has not been measured, so it should not be quoted as though it applies everywhere.

Does a Yelp rating actually affect sales?

For independent restaurants, yes, and it was measured cleanly: the study used Yelp's half-star rounding as a natural experiment to isolate the effect of the displayed rating from the underlying quality it reflects. The same design has not been repeated for most other verticals, so the direction is well established while the size outside restaurants is not.

Do reviews matter as much for chains as for independents?

The evidence says no. In Luca's study the revenue effect was concentrated in independent restaurants, and chains showed no measurable rating effect. The likely reason is that consumers already hold a quality expectation for a national brand, so a single review changes little. An unfamiliar independent has no such buffer, which is why the rating point carries more weight there.

Does the 5 to 9 percent figure apply to my dental practice, home services business, or gym?

It should not be assumed to. That number was measured for restaurants only, and no equivalent large-sample study was found for dental, home services, or fitness. It is reasonable to expect reputation to matter in those verticals; attaching the restaurant percentage to them is not supported. The right move is to measure the effect against your own baseline rather than borrow a figure.

Are the "70 percent of dental patients" style statistics reliable?

Treat them with caution. Several widely repeated vertical statistics, including common dental figures, could not be traced to a disclosed primary study in this research pass. A trustworthy statistic names an author, a venue, and a year you can locate. When a figure is attributed only to an unnamed study or an association with no findable document, it is a fabrication risk and should not be repeated as fact.

Provenance

Sources

  1. Luca, M., "Reviews, Reputation, and Revenue: The Case of Yelp.com", Harvard Business School Working Paper No. 12-016, 2011 (rev. 2016) (established)hbs.edu
  2. Zhang M, Sun Y, Zhao X, Wang L, Xiong J, "The Impact of Narrative Reviews on Patient E-doctor Choice in Online Health Communities", INQUIRY, 2023, PMID 37357728 (established; non-US platform)pubmed.ncbi.nlm.nih.gov
  3. Federal Trade Commission, Trade Regulation Rule on the Use of Consumer Reviews and Testimonials, 16 CFR Part 465, effective Oct. 21, 2024 (established)ecfr.gov
  4. BrightLocal, Local Consumer Review Survey, 2024 and 2026 editions (emerging; industry-primary survey, flagged per figure)brightlocal.com
  5. Pew Research Center, "Americans and AI 2026: Chatbots, Smart Devices and Views on Impact", June 17, 2026 (established)pewresearch.org
  6. Health & Fitness Association (formerly IHRSA), 2025 US Health & Fitness Consumer Report (established; membership behavior, not discovery)healthandfitness.org
  7. Aggregated Google local-search behavior studies via industry local-SEO research, 2025 (established direction, emerging on precise magnitude)

Every figure above is attributed to a real, dated source and tagged with its evidence tier. Where a claim could not be verified to a primary source, it is not stated as fact.

What a rating point is worth in your vertical

The literature has a clear limit: the 5 to 9 percent figure was measured for restaurants, and the value in home services, dental, or fitness has to be observed, not borrowed. That starts with a reputation built the way the evidence rewards, owned and consistent profiles, a steady flow of real reviews, and a baseline you can measure against. A Reputation Foundation Sprint puts that footing in place as one coordinated build, and a Review Acquisition System Setup keeps genuine, compliant reviews flowing so the movement is real and yours to measure.

service Reputation Foundation Sprint A one-time, coordinated build that claims and cleans up every profile a buyer might find, stands up a compliant review-acquisition system, and installs a crisis plan, all measured against your Reputation and Sentiment baseline on day one. See how it works

Start free with a Machine-Readiness Score, a specialist-reviewed read of where your reputation stands across search and AI answers. No guaranteed number, and no obligation.