Measurement & Honesty · established evidence
Vanity Metrics, Decision Metrics: A Field Guide for Owners Who Do Not Have a Data Team
Vanity metrics are numbers that go up and tell you nothing you can act on: impressions, raw traffic, follower counts, an advertising-value-equivalent for your press coverage. Decision metrics are the few that would actually change what you do next week if they moved, and you can name the move. An owner-operator without a data team does not need a dashboard of forty figures. You need one decision metric per surface where buyers now find and choose you, classic search, the local map pack, AI answers, and reputation. This guide names those four, explains why the scale of a single-location business makes the usual tests unworkable, ties each metric to published measurement evidence, and separates which pillars rest on settled science from which, like AI answers, are still emerging. The goal is a short list you can read and act on, not a longer one that only looks like progress.
The metric you were sold is rarely the metric you can act on
The oldest complaint in marketing is that you cannot tell which half of your spending works. The line is usually put in John Wanamaker's mouth, though it has never been traced to him: the earliest documented match is a secondhand 1919 speech, and the same sentiment has been credited to William Lever and William Wrigley too. The industry's founding parable about measurement is itself unverified, which is a fitting place to start a discussion about the numbers you are handed.
The pattern the parable names has repeated for a century. A convenient, computable number gets adopted at scale, gets treated as truth for years, and only gets corrected once someone runs the harder experiment or an industry votes to discipline itself. The public-relations trade did exactly that when it renounced Advertising Value Equivalency, a metric that priced coverage as if it were bought advertising. A vanity metric, in short, is any number that is easy to produce, pleasant to report, and disconnected from a decision.
A decision metric passes a harder test. If it moves, you would do something different, and you can name what. It resists gaming, it points at an outcome rather than an output, and it survives being read by someone who did not want the answer. Everything below is built to that standard.
Why owner-operators get sold vanity metrics: a scale problem, not a moral one
There is a structural reason a single-location business gets handed impressions instead of proof, and it is not that agencies are lazy. Rigorous measurement has a scale floor most local businesses fall beneath.
Consider the tool everyone reaches for first, the A/B test. Standard sample-size math at 95 percent significance and 80 percent power, chasing a realistic 5 to 10 percent minimum detectable effect, typically calls for tens to hundreds of thousands of visitors. At the traffic a real med-spa or home-services site sees, roughly a thousand visitors a month, detecting even a 20 percent relative lift can take more than seven months, and a 10 percent lift more than two and a half years, well past the point where the season, the offer, or the business itself has already changed.
The same floor sits under the famous brand-building findings. The 60:40 split and the double-jeopardy law were derived from large, multi-brand, multi-decade datasets, not from single small firms. This is why owner-operators end up with vanity metrics: the firm-level experiment they would need is structurally out of reach, so they are sold the number that is easy to compute instead. The real resolution is not a shortcut. It is aggregation, pooling many similar businesses to recover statistical power no single one of them has, and choosing metrics that do not require a randomized trial to read.
The test of a decision metric: could it change what you do on Monday?
Before naming the four, it helps to state the discipline, because the same discipline is what a serious measurement standard already demands. When the communications industry rewrote its rules as the Barcelona Principles, the core instruction was to measure outcomes, not outputs, and to treat measurement as something that must be transparent, consistent, and valid rather than a flattering number.
Translated for an owner, a decision metric has four properties. It is actionable: a move in it has an obvious next step. It is honest: it is never reported naked, always with its method and its comparison. It is hard to inflate: you cannot juice it without doing the underlying work. And it is tied to a buyer, not to your own activity. Impressions fail every test. The four metrics below pass them.
One decision metric per Machine-Readiness Score pillar
The Machine-Readiness Score reads your visibility across four pillars, because those are the four places a buyer now decides: classic search, the local map pack, AI answers, and reputation. Each pillar has exactly one number worth acting on. Track these four and you have a truer picture than a forty-line dashboard gives.
Classic search: presence for the queries a ready buyer types
The vanity version is total sessions or total impressions, a number that rises with brand searches, bots, and irrelevant traffic. The decision metric is narrower and harder to fake: do you appear on the first page for the short list of queries a buyer types when they are ready to book, the service plus the city, the "near me", the specific procedure? That is a presence question, and presence is something you can read and fix directly.
Presence is also the metric to trust because the attributed-click numbers beneath it are shakier than they look. When researchers compared standard observational attribution against fifteen large randomized experiments, the observational estimates frequently pointed in the wrong direction or the wrong magnitude even with rich data. Counting whether you show up for the right questions is more reliable than trusting a platform's claim about which click caused a sale.
Local map pack: appearance for money queries, plus review velocity
For a local service business the three-pack often decides the trade. The vanity metric is total profile views. The decision metric is whether you appear in the pack for your money queries at all, paired with whether your review count is actually moving month over month. Both are legible without a data team, and both change what you do: if you are absent from the pack, the fix is a different project than if you are present but static.
AI answers: share of answer, as a disclosed sample
When someone asks an engine "who is the best contractor near me", it names a few businesses. The decision metric is your share of answer: across a fixed panel of the questions your buyers actually ask, how often are you named? This is the newest pillar and it must be stated carefully. There is no standardized, agreed method for measuring AI-search visibility yet. Generative engines are non-deterministic, they personalize, and they are not fully observable from outside, so any share-of-answer figure is a sample-based estimate whose worth depends entirely on a disclosed sampling method. Read as a disclosed sample, it is the most forward-looking metric you have. Read as a precise score, it is a vanity metric wearing a lab coat.
Reputation: recent, real review signal against the named competitor
The vanity metric is a lifetime star average that has not moved in two years. The decision metric is recent, genuine review signal, rating and velocity, read against the specific competitor an engine keeps naming above you. Two caveats are non-negotiable here. First, if you are smaller you will predictably look a little less loved, and that is a law, not a failing, which the next section explains. Second, the reviews must be real: US endorsement rules require testimonials to reflect honest experience and any material connection to be disclosed, and enforcement is active. A bought or faked review is not a shortcut, it is a liability.
The double-jeopardy trap in every reputation number
One statistical law deserves its own section because it quietly distorts how owners read their reputation pillar. Double jeopardy is one of the most replicated regularities in marketing science: brands with smaller market share have both fewer buyers and slightly lower loyalty among the buyers they do have. It has held across packaged goods, banking, insurance, and newer categories like streaming and ride-share.
Applied to your review counts, the lesson is not to panic when the larger competitor has more reviews and a marginally higher average. That is the expected pattern for a smaller player, not evidence you are worse. The decision this metric should drive is a steady program to raise real review velocity, not a despairing comparison of lifetime totals. Reading double jeopardy correctly is the difference between acting and flinching.
The traps even good metrics fall into
A decision metric can still be misread. The most common failure is peeking: checking a test or a trend the moment it looks like it is working and acting on it early. In large-scale industrial experimentation this is a catalogued way to convince yourself of an effect that is not there, alongside sample-ratio mismatch and novelty effects. For an owner the practical rule is to let a change run a fair, pre-decided window before you judge it, and to compare against a baseline rather than against your hopes.
The second trap is the naked number. A metric reported without its method and its comparison is not a decision metric, it is decoration. "Views up 30 percent" means nothing without the window, the source, and what it is up against. Every number in this guide is meant to be carried with its method attached, which is also the rule the Machine-Readiness Score itself is built on.
What to stop tracking
The corollary of four decision metrics is a longer list of numbers you can safely ignore. None of these should drive a decision on their own.
- Total impressions or total ad impressions, which rise with irrelevant reach.
- Raw website sessions, which mix bots, brand searches, and idle curiosity with buyers.
- Follower and like counts, which almost never move booked business.
- An advertising-value-equivalent for press coverage, a metric its own industry formally abandoned.
- Lifetime star averages that have not moved in a year, which hide recent trajectory.
- Any single-session A/B test result read early, which the statistics say you cannot yet trust.
The limit of this guide
Three of the four decision metrics rest on settled ground: the scale floor under small-business testing, the double-jeopardy law, the abandonment of vanity metrics like the advertising-value-equivalent, and the endorsement rules are all established. The fourth, share of answer, is genuinely emerging, and the correct posture is transparency about that immaturity rather than a claim to have solved it. The Machine-Readiness Score follows that same corrective tradition: a small set of measured numbers, each reported with its method, which is the most an owner without a data team should ever act on.
The evidence
Key findings, with their sources
-
Detecting even a 20 percent relative lift at roughly 1,000 visitors a month can take more than seven months, and a 10 percent lift more than two and a half years, because standard A/B tests need tens to hundreds of thousands of users for adequate power.
established Industry synthesis of standard power-analysis methodology (Analytics-Toolkit.com; Statsig, "Power Analysis for A/B Testing"), consistent with the statistics of hypothesis testing.
-
The public-relations industry formally rejected Advertising Value Equivalency in favor of outcome-based measurement, calling for measurement to be transparent, consistent, and valid rather than reliant on vanity metrics.
established AMEC, Barcelona Principles 3.0 (2020), amecorg.com.
-
Brands with lower market share have both fewer buyers and slightly lower average loyalty, a "double jeopardy" pattern replicated across packaged goods, banking, insurance, streaming, and ride-share.
established Ehrenberg, Goodhardt & Barwise, "Double Jeopardy Revisited," Journal of Marketing 54(3), 1990; McPhee (1963).
-
US endorsement rules require testimonials to reflect honest experience and any material connection between advertiser and endorser to be disclosed, with active 2025-2026 enforcement.
established Federal Trade Commission, 16 CFR Part 255, "Guides Concerning the Use of Endorsements and Testimonials in Advertising."
-
Checking test results before the planned sample size is reached ("peeking") is a documented way to stop early on a false positive, catalogued from experience running 20,000+ experiments a year at Google, LinkedIn, and Microsoft.
established Kohavi, Tang & Xu, Trustworthy Online Controlled Experiments, Cambridge University Press, 2020.
-
Standard observational attribution methods frequently produced effect estimates in the wrong direction or magnitude when checked against fifteen large randomized field experiments, even after conditioning on rich data.
established Gordon, Zettelmeyer, Bhargava & Chapsky, "A Comparison of Approaches to Advertising Measurement," Marketing Science 38(2), 2019.
-
There is no standardized, agreed method for measuring share of answer; generative engines are non-deterministic and not fully observable, so any AI-visibility figure is a sample-based estimate whose reliability depends on disclosed sampling.
emerging Inference from Aggarwal et al., "GEO: Generative Engine Optimization," arXiv:2311.09735, KDD 2024, plus documented LLM non-determinism.
-
The "half my advertising is wasted" line popularly attributed to John Wanamaker has no verified original source; the earliest documented match is a secondhand 1919 speech.
established Quote Investigator (2022), "One-Half the Money I Spend for Advertising Is Wasted."
Calibration
What is proven, what is promising, what is unproven
| Evidence tier | Tactics | What the evidence says |
|---|---|---|
| established | The scale-floor case for choosing read-and-fix decision metrics over firm-level A/B tests; the classic-search, map-pack, and reputation decision metrics; the double-jeopardy read of review counts; the no-naked-metric and honest-endorsement discipline. | Power-analysis math; Kohavi et al. (2020); Gordon et al. (2019); Ehrenberg et al. (1990); AMEC Barcelona Principles 3.0; FTC 16 CFR 255. |
| emerging | The AI-answers decision metric (share of answer), read strictly as a disclosed sample, not a precise score. | Aggarwal et al. (2024) as the founding academic framework; the field has no agreed sampling standard and the engines are non-deterministic. |
Reference
Glossary
- Vanity metric
- A number that is easy to produce and pleasant to report but disconnected from any decision, such as impressions, raw traffic, or follower counts.
- Decision metric
- A number that would change what you do next if it moved, is reported with its method, resists gaming, and points at a buyer outcome rather than your own activity.
- Statistical power
- The probability a test detects a real effect. Small sites rarely have the traffic to reach adequate power, which is why firm-level A/B testing is usually infeasible for them.
- Advertising Value Equivalency (AVE)
- A discredited PR metric that priced earned coverage as if it were paid advertising. Its own industry abandoned it in the Barcelona Principles.
- Double jeopardy
- The replicated law that smaller-share brands have both fewer buyers and slightly lower loyalty, so a smaller business will predictably show fewer reviews and a marginally lower average.
- Across a fixed panel of buyer questions, how often an engine names your business inside its synthesized answer. A disclosed sample, not a precise score, given today's non-deterministic engines.
- Peeking
- Checking a test or trend before its planned window ends and acting on a result that looks significant but is not, a documented cause of false positives.
Straight answers
Frequently asked questions
What is a vanity metric?
A number that is easy to produce and flattering to report but disconnected from any decision. Impressions, raw sessions, follower counts, and an advertising-value-equivalent for press are the classic examples: they can rise without any change in booked business.
What is a decision metric?
A number that would change what you do next if it moved, and you can name the move. It is reported with its method and comparison, it is hard to inflate without doing the underlying work, and it is tied to a buyer rather than to your own activity.
Which marketing metrics should a small business without a data team actually track?
Four, one per surface where buyers decide: presence on page one for the queries a ready buyer types (classic search); appearance in the map pack for your money queries plus review velocity (local); share of answer read as a disclosed sample (AI answers); and recent, real review signal against the competitor engines keep naming (reputation).
Why can I not just run an A/B test like the big platforms do?
Statistical power. Standard tests need tens to hundreds of thousands of visitors to detect a realistic effect. At about a thousand visitors a month, detecting a 20 percent lift can take over seven months, by which point the season or offer has already changed. The alternative is choosing metrics you can read and fix directly, and pooling data across many similar businesses.
My bigger competitor has far more reviews. Am I doing something wrong?
Probably not. Double jeopardy is a well-replicated law: smaller-share businesses predictably have fewer reviews and a slightly lower average. The right response is a steady program to raise real review velocity, not a despairing comparison of lifetime totals, and never bought or faked reviews, which US endorsement rules treat as a liability.
Provenance
Sources
- Kohavi, R., Tang, D., Xu, Y. (2020). Trustworthy Online Controlled Experiments: A Practical Guide to A/B Testing. Cambridge University Press (established)
- Gordon, B.R., Zettelmeyer, F., Bhargava, N., Chapsky, D. (2019). "A Comparison of Approaches to Advertising Measurement: Evidence from Big Field Experiments at Facebook." Marketing Science 38(2), 193-225 (established)
- Ehrenberg, A.S.C., Goodhardt, G.J., Barwise, T.P. (1990). "Double Jeopardy Revisited." Journal of Marketing 54(3), 82-91; McPhee, W. (1963). Formal Theories of Mass Behavior (established)
- AMEC (2020). Barcelona Principles 3.0. International Association for the Measurement and Evaluation of Communication (established)
- Federal Trade Commission. 16 CFR Part 255, "Guides Concerning the Use of Endorsements and Testimonials in Advertising" (established)ecfr.gov
- Analytics-Toolkit.com; Statsig, "Power Analysis for A/B Testing" (industry synthesis of standard power-analysis methodology) (established as statistical fact; traffic/time figures are industry-sourced illustrations)
- Aggarwal, P. et al. (2024). "GEO: Generative Engine Optimization." arXiv:2311.09735, KDD 2024 (emerging)arxiv.org
- Quote Investigator (2022). "One-Half the Money I Spend for Advertising Is Wasted, But I Have Never Been Able To Decide Which Half" (established that the Wanamaker attribution is unverified)
Every figure above is attributed to a real, dated source and tagged with its evidence tier. Where a claim could not be verified to a primary source, it is not stated as fact.