Measurement & Honesty · established evidence

The Statistical Power Problem Nobody Tells Small Businesses About

Last reviewed 2026-07-20. Written by Chandranshu Kumar, Founder, Raveneye Global. · 10 min read

A/B testing for small business runs into a wall most testing tools never mention: statistical power. To tell a real improvement from random noise, a test needs a minimum number of visitors, and that number is fixed by math, not effort. Standard settings, roughly 95% significance, 80% power, and a 5 to 10% minimum detectable effect, typically call for tens to hundreds of thousands of visitors per test. A single location seeing about a thousand visitors a month cannot reach that in any useful window. Detecting even a 20% relative lift can take around seven months, by which point the season, the offer, or the market has already moved. The problem is structural, not a failure of skill or discipline, and it explains why owners are so often sold traffic and ranking numbers they cannot act on. The way out is not a shortcut. It is aggregation: pooling many similar businesses to recover the statistical power no single one of them has.

The number a testing tool will not show you

An A/B test compares two versions of a page or an offer and asks whether the difference in outcomes is real or just noise. Answering that question is a statistics problem with four fixed inputs: the significance level (how often you accept a false positive), the statistical power (how often you catch a true effect when it exists), the minimum detectable effect (the smallest lift you want to be able to see), and the sample size. Set any three and the fourth is determined. You do not get to negotiate with it.

The conventional settings used across the industry are about 95% significance, 80% power, and a minimum detectable effect somewhere between 5 and 10%. Plug those into a standard sample-size calculation and the required audience is large. The commonly cited range is on the order of 50,000 to 500,000 users per test. This is not a controversial claim. The mathematics of power analysis is not in dispute; it is the same hypothesis-testing framework taught in every statistics course.

Large platforms run tests this way because they have the traffic to satisfy the math. Kohavi, Tang and Xu, documenting experimentation practice across more than 20,000 experiments a year at Google, LinkedIn and Microsoft, treat adequate sample size as a precondition, not an aspiration. The methods a small business is told to copy were written for organizations operating several orders of magnitude above its scale.

What that math means at a thousand visitors a month

Now impose the constraint a real local business lives under. A med-spa, a roofing contractor, or a solo law practice might see on the order of a thousand website visitors a month. Feed that traffic rate into the same power calculation and the timelines become the story.

On industry-standard illustrations consistent with the underlying power math, detecting even a 20% relative lift at roughly a thousand visitors a month takes about seven months of continuous testing, and a 10% lift takes about 31 months. Those horizons are longer than the useful life of the thing being tested. A seasonal offer, a new competitor, a website redesign, a change in ad mix: any one of them invalidates the experiment before it can reach significance. You would be measuring a business that no longer exists by the time the result arrives.

This is the quiet cruelty of the advice to "just run an A/B test." It is not wrong in principle. It is unreachable in practice at this scale, and the tools rarely say so. An owner sets up a test in good faith, watches it for a few weeks, sees no clear winner because there cannot be one yet, and concludes either that the change did not matter or that testing is a waste of time. Both conclusions are artifacts of insufficient power, not findings about the business.

The test does not get more honest if you wait

Faced with a test that will not resolve, the natural temptation is to peek at the dashboard and stop the moment it looks like a winner. This is where a hard problem becomes a misleading one.

Peeking, checking results repeatedly and stopping early on the first apparent significance, is a documented way to manufacture false positives, because each look is another chance for random noise to cross the threshold. Kohavi and colleagues catalogue it alongside sample ratio mismatch (the traffic split not matching what was intended, a signal the test is broken) and novelty and primacy effects (early behavior that does not represent the steady state) as standard failure modes of controlled experimentation. Each one is more likely, not less, when a small test is nursed along for months in the hope of an answer.

So the small business is caught twice. It cannot gather enough sample to run a valid test, and the workarounds that make an underpowered test appear to resolve are precisely the ones that make its results untrustworthy. Waiting longer does not fix an underpowered design; it just gives noise more opportunities to look like signal.

Why owners get sold vanity metrics instead

If rigorous individual testing is out of reach, something has to fill the reporting slot. What fills it is usually a number that is easy to compute and pleasant to read: sessions, impressions, keyword rankings, followers, an "audit score" with no method behind it. These are vanity metrics, figures that move without telling you whether anything that matters to the business moved.

It helps to see this as a structural outcome rather than a moral failing of the vendor. The buyer wants proof, real proof is unreachable at their scale, and a plausible-looking substitute is available, so the substitute gets sold. The pattern is old. The famous line that "half my advertising is wasted, I just do not know which half" has no verified original source, an apt detail: the industry's founding parable about measurement is itself unverified. And when the harder measurement does get run, the convenient numbers often do not survive it. Gordon and colleagues, comparing observational attribution against randomized ground truth across 15 large field experiments at Facebook, found the observational methods frequently pointed in the wrong direction or magnitude even with rich data behind them.

The reputable end of the industry has faced this openly. In 2020 the public relations profession ratified the Barcelona Principles 3.0, which formally reject Advertising Value Equivalency and other vanity metrics in favor of outcome-based measurement. That is what discipline looks like: an industry voting to stop reporting a number it could no longer defend. The lesson for a small business is not that all metrics are lies, but that a metric you cannot tie to a decision is decoration.

This is a scale floor, not a skill gap

It is worth being precise about what the small business lacks. It is not intelligence, effort, or good intentions. It is sample. Rigorous measurement has a scale floor, and owner-operated local businesses sit beneath it. That is a property of the arithmetic, not a verdict on the operator.

The point sharpens when you notice that even the marketing science owners are told to follow was built on pooled data, never on single small firms. The widely quoted 60:40 split between brand-building and activation spend was derived from an analysis of roughly 996 IPA case studies spanning 700 brands, 83 sectors, and more than 30 years. The double jeopardy law, that smaller brands have both fewer buyers and slightly lower loyalty, is one of the most replicated regularities in marketing, established across many categories at once. How Brands Grow rests on empirical generalizations the Ehrenberg-Bass Institute frames as replicated across categories over decades.

Every one of those findings earned its authority by aggregating across many firms. None of them could have been discovered inside a single small business, and applying their headline numbers directly to one small firm is an unproven leap, flagged as such: the 60:40 benchmark in particular is drawn from large branded advertisers, not single-location operators. The science that works does so precisely because it pools. That is the clue to the fix.

The fix is aggregation

If no single local business has the traffic to reach statistical power, the answer is not to pretend otherwise or to sell a number that hides the gap. It is to change the unit of analysis. Pool many similar businesses and the combined sample can recover the power that no one of them holds alone. A question a single med-spa can never answer with a thousand visitors a month becomes answerable across a hundred med-spas measured the same way.

This is the same move marketing science has always made. An empirical generalization, in the Ehrenberg-Bass sense, is a pattern that holds across many cases rather than a result from one. Aggregation turns a collection of underpowered individuals into a dataset with real statistical footing, and it lets you benchmark any single business against a comparable set rather than against noise. The right comparison for a roofing contractor is not its own month-over-month wobble; it is where it stands relative to many comparable roofing contractors read on the same method.

That is exactly the role of a pooled cross-client benchmark. RavenEye builds one, the Visibility Corpus, for a narrow and defensible purpose: to give a single business a comparison it could never generate on its own, and to keep any single-client claim skeptical until pooled evidence supports it. The credible unit of evidence for a small business is the aggregate, and a benchmark is only as trustworthy as the method behind it is disclosed.

What to measure when you cannot test

The practical takeaway for an owner without a data team is not to abandon measurement. It is to stop trying to force a platform-scale method onto a local-scale business, and to insist on two things instead: a metric tied to a decision, and a comparison that is not just your own noise.

  • Prefer decision metrics over vanity metrics. A number is worth reporting only if a plausible move by you would change it and you would act differently depending on where it lands.
  • Read direction and benchmark, not single-test significance. At your scale, standing relative to comparable businesses over time is more reliable than a p-value from an underpowered experiment.
  • Demand the method behind any score. If a vendor cannot tell you how a figure was sampled and computed, treat it as decoration, not evidence.
  • Trust pooled evidence over a single-client anecdote. A pattern that holds across many similar firms is stronger proof than one flattering case study.
  • State the uncertainty. A read that shows its limits is more useful than a confident number that cannot survive a real experiment.

The evidence

Key findings, with their sources

  • Standard sample-size settings, about 95% significance, 80% power, and a 5 to 10% minimum detectable effect, typically require on the order of 50,000 to 500,000 users per A/B test.

    established Industry synthesis of standard power-analysis methodology (Analytics-Toolkit.com; Statsig, "Power Analysis for A/B Testing"), consistent with the standard statistics of hypothesis testing.

  • At roughly 1,000 visitors a month, detecting even a 20% relative lift can take about 7 months, and a 10% lift about 31 months, past the point where the business or campaign has already changed.

    established Industry synthesis of power-analysis methodology (Analytics-Toolkit.com; Statsig), an illustration consistent with the undisputed math of power analysis, not a peer-reviewed study.

  • Peeking, sample ratio mismatch, and novelty or primacy effects are catalogued failure modes of controlled experimentation, documented across more than 20,000 experiments a year at Google, LinkedIn, and Microsoft.

    established Kohavi, Tang & Xu, Trustworthy Online Controlled Experiments: A Practical Guide to A/B Testing, Cambridge University Press, 2020.

  • Observational attribution methods frequently produced effect estimates in the wrong direction or magnitude versus randomized ground truth, across 15 large field experiments at Facebook (500M+ observations, 1.6B impressions).

    established Gordon, Zettelmeyer, Bhargava & Chapsky, "A Comparison of Approaches to Advertising Measurement: Evidence from Big Field Experiments at Facebook," Marketing Science 38(2), 2019.

  • The 60:40 brand-to-activation benchmark was derived from about 996 IPA case studies across 700 brands, 83 sectors, and more than 30 years, a large multi-brand dataset, not single small firms.

    established Binet & Field, The Long and the Short of It, IPA, 2013.

  • The public relations industry formally rejected Advertising Value Equivalency and other vanity metrics in favor of outcome-based measurement in its ratified Barcelona Principles 3.0.

    established AMEC, Barcelona Principles 3.0, 2020.

  • The "half my advertising is wasted" line has no verified original source; the earliest documented match is a secondhand 1919 speech, and the same sentiment is attributed to several people.

    established Quote Investigator, 2022, "One-Half the Money I Spend for Advertising Is Wasted."

Calibration

What is proven, what is promising, what is unproven

Evidence tierTacticsWhat the evidence says
establishedPower-analysis math: individual A/B tests need tens to hundreds of thousands of users, so single-location testing is structurally infeasible; peeking and related traps make underpowered tests worse, not better.Standard power analysis (Analytics-Toolkit / Statsig illustrations); Kohavi, Tang & Xu (2020).
establishedAggregation and empirical generalization recover statistical power a single firm lacks; the marketing science owners follow was itself built by pooling across many firms.Sharp / Ehrenberg-Bass (2010); Ehrenberg et al. double jeopardy (1990); Binet & Field (2013).
contestedApplying big-brand pooled benchmarks (for example the 60:40 split) directly to one single-location business; treating any single-client A/B result as proof.Binet & Field source data skews to large branded advertisers; flagged unproven at MSME scale in the pillar research.

Reference

Glossary

Statistical power
The probability that a test detects a real effect when one truly exists. Convention targets 80%. Reaching it requires a minimum sample size that small-traffic sites rarely meet.
Minimum detectable effect (MDE)
The smallest improvement a test is designed to be able to see. Asking to detect a smaller lift sharply increases the sample size required.
Significance level
The tolerated rate of false positives, usually set so results are called real at about 95% confidence. Tightening it also raises the sample needed.
Peeking
Repeatedly checking a running test and stopping the moment it looks significant. It inflates false positives because every look is another chance for noise to cross the line.
Sample ratio mismatch
When the actual traffic split between variants differs from the intended split, a red flag that the experiment is broken and its result cannot be trusted.
Empirical generalization
A pattern that holds across many cases rather than a finding from a single one. Aggregating across firms is how marketing science reaches conclusions no single business could.
Vanity metric
A number that is easy to report and pleasant to read but not tied to any decision, so it moves without telling you whether the business improved.

Straight answers

Frequently asked questions

How much traffic do you need to run an A/B test?

It depends on the effect you want to detect, but standard settings (about 95% significance, 80% power, a 5 to 10% minimum detectable effect) typically call for on the order of 50,000 to 500,000 users per test. That is set by the math of power analysis, not by how the test is built.

Can a small business with about 1,000 visitors a month run A/B tests?

Not usefully as individual tests. At that traffic, detecting even a 20% relative lift can take around seven months, and a 10% lift around 31 months, longer than the useful life of the thing being tested. The business changes before the experiment can resolve, so the result never arrives in time to act on.

What is statistical power in plain terms?

It is the chance that a test actually catches a real improvement instead of missing it. Low power means real effects go undetected and results look inconclusive even when a change genuinely helped. Reaching the conventional 80% power requires a minimum sample size that small-traffic sites usually cannot gather.

Why do agencies show me traffic and ranking numbers instead of proof of impact?

Often because rigorous individual proof is unreachable at your scale, so a number that is easy to compute fills the gap. That is a structural outcome, not necessarily bad faith, but it means those figures may be vanity metrics. The test is simple: if a number is not tied to a decision you would make, treat it as decoration.

If a small business cannot A/B test, what is the alternative?

Aggregation. Pooling many similar businesses measured the same way recovers the statistical power no single one of them has, and lets you benchmark a business against a comparable set rather than against its own noise. The credible unit of evidence at small scale is the aggregate, and any benchmark is only as trustworthy as its disclosed method.

Provenance

Sources

  1. Industry synthesis of standard power-analysis methodology (Analytics-Toolkit.com; Statsig, "Power Analysis for A/B Testing"), consistent with the standard statistics of hypothesis testing (established as statistical fact; specific traffic/time figures are industry illustrations) (established)
  2. Kohavi, R., Tang, D., Xu, Y., Trustworthy Online Controlled Experiments: A Practical Guide to A/B Testing, Cambridge University Press, 2020 (established)
  3. Gordon, B.R., Zettelmeyer, F., Bhargava, N., Chapsky, D., "A Comparison of Approaches to Advertising Measurement: Evidence from Big Field Experiments at Facebook," Marketing Science 38(2), 193-225, 2019 (established)
  4. Sharp, B., How Brands Grow: What Marketers Don't Know, Oxford University Press, 2010; Ehrenberg-Bass Institute for Marketing Science research program (established as a replicated empirical pattern; its application to AI-answer discovery is new and untested)
  5. Ehrenberg, A.S.C., Goodhardt, G.J., Barwise, T.P., "Double Jeopardy Revisited," Journal of Marketing 54(3), 82-91, 1990 (established)
  6. Binet, L., Field, P., The Long and the Short of It, IPA, 2013 (established for its large multi-brand dataset; contested as a direct prescription for a single-location MSME)
  7. AMEC, Barcelona Principles 3.0, 2020 (established)amecorg.com
  8. Quote Investigator, "One-Half the Money I Spend for Advertising Is Wasted, But I Have Never Been Able To Decide Which Half," 2022 (established that the attribution is unverified)

Every figure above is attributed to a real, dated source and tagged with its evidence tier. Where a claim could not be verified to a primary source, it is not stated as fact.

What this means for your business

If you have ever set up a test on your own site and watched it sit for weeks without a clear winner, that was not your failure. It was the math. At a thousand visitors a month, a rigorous individual test cannot resolve before your business changes, which is exactly why so much reporting quietly becomes traffic and ranking numbers you cannot act on. The credible read at your scale comes from comparison, where you stand against many businesses like yours, measured the same way. A Surface Intelligence Audit gives you that pooled read across all four gates buyers now use, benchmarked against comparable businesses, with its method shown so the number means something.

diagnostic Surface Intelligence Audit A measured read of where you stand across classic search, the local map pack, AI answers, and reputation, benchmarked against comparable businesses using the pooled Visibility Corpus, with the corrections that move you first. See how it works

Start free with a Machine-Readiness Score, a specialist-reviewed read of where you stand across search and AI answers, benchmarked against businesses like yours. No guaranteed number, and no obligation.