Primary Audits
Machine-Readiness Audit: US MSMEs, July 2026
The first edition of a quarterly benchmark. We measured 43,362 independent US small businesses; one in four was invisible to a machine reader, a majority had no type an engine can read as a local business, and readiness rose steadily with how established the business was. We will repeat it every quarter, using the same method, to track how the machine-readable web evolves.
Part of Discovery Science in the Insights library.
Abstract
Search and AI answers can only name a business a machine can actually read. We drew 51,001 independent small businesses from their Google Business Profile map listings, roughly 4,300 each across twelve trades in 88 US metros, and read what every homepage exposes to a machine; 43,362 were reachable. The human-facing basics were nearly universal: 96.8 percent were mobile-ready, 99.7 percent secure, 92.0 percent showed a phone number. The machine-facing signals were not. 26.6 percent carried no structured data at all, and 55.3 percent, a majority, had no LocalBusiness-family type that names a business as a local entity an engine can place. Only 33.1 percent were fully readable. The gap was highly structured: it differed sharply by trade (chi-square 819.8, p below 0.001, professional services lowest), it rose monotonically with review volume from 30.3 to 59.6 percent, and the individual weaknesses clustered together. Where our rates can be checked against independent web-wide measurements they are consistent, and they show these local businesses carrying local-entity markup far more often than the web at large (44.7 percent versus about 4 percent web-wide) while a majority still lack it. Because even this sample is drawn from businesses with an active profile, the true gap across all small businesses is very likely wider still. This is the first edition of a benchmark we will repeat each quarter, holding the method fixed, to track how the machine-readable web evolves. Every figure is measured, not modeled.
The way a small business gets found has split into two jobs that used to be one. For a human visitor, a clean homepage with a phone number and a booking button has always been enough, and on that measure the 43,362 sites we could read are in good shape: 96.8 percent were mobile-ready, 99.7 percent secure, 92.0 percent showed a phone number. The second job is newer and quieter. Search engines and AI answer engines increasingly read a page through its structured data, the machine-readable code that states plainly what kind of business this is and where. That layer is where the sample came apart. A quarter of these sites carried no structured data at all, and a clear majority never used a LocalBusiness-family type that names the business as a local entity an engine can place on a map. A site can look finished to every human who visits and still be half-legible to the machines now deciding who gets named.
For the buyer, none of this is visible, and that is the problem. A person landing on one of the machine-invisible sites has a perfectly good experience: the page loads, the phone number is there. What they never see is the moment upstream, when an engine assembling a shortlist skims the machine-readable signals and, finding little, has less to work with than it would for a competitor whose site says the same things in a format the engine can parse. And the gap is not random. It was deepest in the professional-services trades, where a buyer most often verifies credentials before committing, and it was deepest among the newest, quietest businesses, the ones with the fewest reviews and the most to gain from being found. The businesses least able to afford invisibility were the ones most likely to be invisible.
This is the ground-level view of the terrain the Visibility Corpus maps from above. If attention is moving toward answers that engines assemble, then whether an engine can read your site at all becomes a precondition for appearing on the surfaces where that attention now sits. The encouraging finding is that the expensive, hard things were mostly fine: nearly every site was mobile-ready and secure. The gap was in the cheap, high-impact things: naming yourself in structured data with a type an engine can place. That is the shape of the opportunity. For most of these 43,362 businesses the fix is not a rebuild, it is an afternoon of structured-data work, and it moves a site from half-legible to fully readable, which is the precondition for being named at all on the surfaces where attention is heading.
The data, in one read
A business now has two jobs, and only one of them is finished
For as long as small businesses have had websites, the job of a site was to satisfy a human: load quickly, look credible, show a phone number, make it easy to book. That job is now largely done. Across the sites in this study the human basics were nearly universal, and on that measure owner-run small businesses are not behind at all.
A second job has appeared underneath the first, and it is mostly unfinished. Search engines and, increasingly, AI answer engines do not read a page the way a person does. They read its structured data, the machine-readable code that states in plain terms what a page is: this is a dental practice, here is its address, here are its hours. A page can be beautiful to a human and nearly silent to a machine. This study set out to measure, at scale and first-hand, how silent small-business websites actually are.
We ran it because the question, are small businesses built to be found by a machine, had no large, transparent, first-party answer we could find. So we produced one.
A page can be beautiful to a human and nearly silent to a machine.
How we measured this
We assembled the sample from the map listings for twelve trades, dentistry, auto repair, med-spa, small law, plumbing, HVAC, chiropractic, veterinary, salon, accounting, roofing, and physical therapy, across 88 US metros, keeping the independent-business websites and filtering out national chains, franchises, directories, review sites, and aggregators. That produced 51,001 unique independent businesses, each carrying its own star rating and review count from its profile, roughly 4,300 per trade. Sampling from map listings rather than organic search results was a deliberate choice: it draws the businesses a person actually sees when they look for a local service, and it attaches each one to a real-world rating and review count we can analyze against.
We then fetched each homepage and read, directly from its HTML, whether it carried JSON-LD structured data and which types, whether any type was in the LocalBusiness family, and a set of on-page signals: a title, a meta description, an H1, a mobile viewport, HTTPS, Open Graph tags, and a visible phone number. 43,362 sites (85.0 percent) responded to an automated read. To summarize each site we assigned a Machine-Readiness Score from 0 to 100, weighting the local-entity type most heavily. For a deeper read of performance and accessibility, which need a real browser, we ran a full audit on a 12-site subsample. Every figure in this study is a value a tool actually returned in late July 2026, not a model estimate. We report Wilson confidence intervals on the headline rates, a chi-square test for the difference across trades, and the correlation between review volume and readiness, so a reader can weigh the precision of each claim rather than take it on faith.
The headline: only a third of sites are fully readable to a machine
Sorting every reachable site into one of four tiers makes the shape of the problem plain. Just 33.1 percent were fully readable, carrying a LocalBusiness-family type together with the basic on-page signals an engine expects. 28.7 percent were named but not placed: they carried some structured data, but nothing that told an engine they were a local business at a specific address. 26.6 percent were machine-invisible, exposing no structured data at all. The remaining 11.6 percent were placed but unfinished, carrying a local type but missing other basics.
The Machine-Readiness Score tells the same story in one number: a median of 75.0 out of 100, with a wide spread from the first quartile at 60 to a third quartile at 100. Half of these businesses scored at or below 75.0, which is to say half are, at best, only partly legible to the systems now deciding who gets named in an answer.
One in four is invisible; a majority has no local-entity type
The two most consequential gaps are worth stating precisely, because with a sample this size the margins are narrow. 26.6 percent of reachable sites, 11,543 businesses, carried no structured data whatsoever (95 percent confidence interval 26.2 to 27.0 percent). And presence alone was not enough to be placed: 55.3 percent, a clear majority of 23,985 sites, never used a LocalBusiness-family type (confidence interval 54.8 to 55.8 percent).
That distinction is the crux. A site can carry a generic WebSite or Organization tag, as many did, and still never tell an engine it is a dentist in Austin or a plumber in Tampa. Among the sites that did carry structured data, the median site described itself with about 7 typed entities, so the issue is rarely one of effort in general, it is the absence of the one type that lets an engine place a business locally. It is close to free to add, and fully within the owner's control, which is what makes it the highest-impact fix in this whole study.
The gap by trade: professional services lag most
The gap was not spread evenly. Ranking the twelve trades by how often they used a LocalBusiness-family type, the professional-services categories sat clearly at the bottom: Accounting at 25.5 percent and Physical therapy at 35.2 percent were the two lowest, while the hands-on local trades led, with Chiropractic highest at 50.8 percent. This difference is not noise. A chi-square test of trade against local-entity presence returns a statistic of 819.8 on 11 degrees of freedom, far past the threshold for significance at the 0.001 level.
There is a plausible reason, and it has a name in economics. An accountant, a physical therapist, or a lawyer sells what the literature calls a credence good: a service whose quality the buyer struggles to judge even after it is delivered, let alone before (Darby and Karni, 1973; Dulleck and Kerschbamer, 2006). Firms selling credence services compete on signals of trust, credentials, testimonials, the reassurance of the prose, and tend to pour their effort into the human-facing layer while leaving the machine-readable one thin. That is exactly the pattern here. The trades whose buyers most need to verify before committing, and which therefore stand to gain most from being surfaced and cited by a trusted engine, were the hardest for an engine to read as local businesses in the first place. We offer this as an interpretation the data is consistent with, not as something the data proves.
Machine-readiness by trade, all 43,362 reachable sites. Score is the mean Machine-Readiness Score (0 to 100).
| Trade | Sites | Any schema | LocalBusiness type | Mean score |
|---|---|---|---|---|
| Chiropractic | 3901 | 76.1% | 50.8% | 78.1 |
| Dental | 3836 | 76.3% | 50.7% | 77.8 |
| Med-spa | 4074 | 78.6% | 49.2% | 77.7 |
| HVAC | 3856 | 76.5% | 48.4% | 77.6 |
| Plumbing | 3822 | 75.6% | 47.3% | 77.2 |
| Roofing | 3598 | 77.8% | 47.1% | 77.7 |
| Veterinary | 3670 | 70.3% | 46.6% | 75.2 |
| Law firm | 3761 | 77.9% | 43.9% | 76.5 |
| Salon | 3782 | 68.4% | 43.7% | 71.0 |
| Auto repair | 2826 | 68.4% | 42.7% | 72.7 |
| Physical therapy | 2728 | 73.6% | 35.2% | 72.2 |
| Accounting | 3508 | 58.2% | 25.5% | 65.2 |
The investment effect: the busier the business, the more readable
Because every business came from its Google Business Profile, we could line machine-readiness up against how established it is, using review count as a proxy. The relationship was clear and monotonic: readiness rose at every step. Businesses with ten or fewer reviews used a LocalBusiness-family type just 30.3 percent of the time. That rate climbed steadily through each band, reaching 59.6 percent for businesses with more than five hundred reviews. The overall correlation between review volume and the readiness score was positive though modest (0.178), which is what a careful reading expects: reviews are a rough proxy for how long and how seriously a business has invested online, not a direct cause of good schema.
The implication is uncomfortable. The businesses that are already busy and well-known are also the ones an engine can most easily read and recommend, while the newer, quieter businesses, the ones with the most to gain from being surfaced, are the least legible. Machine-readiness, left to itself, compounds an advantage the established already hold.
Rating and geography
Two secondary cuts round out the picture. Star rating tracked with readiness in the same direction as reviews but more weakly: sites rated under 4.0 used a local-entity type 27.6 percent of the time, rising to 48.0 percent for those rated 4.8 to 5.0. Since the median business in this sample is highly rated, this is less a story about service quality than another reflection of the same investment effect.
Geography mattered less than trade or tenure, but it was real. Across the 73 metros with at least 200 reachable sites, a floor we set so every ranked metro is compared on solid ground, the strongest average readiness scores clustered around Denver, Austin, and Nashville (near 79.3), and the weakest around Buffalo, Miami, and Dayton (near 70.2). The spread across metros was narrower than the spread across trades, which is the more useful finding: what a business does matters more to its machine-readiness than where it is.
The gaps travel together
Machine-neglect turned out to be systemic rather than isolated. A site that was missing structured data was far more likely to be missing other basics too: among sites with no structured data, only 50.9 percent carried a meta description, against 85.0 percent of sites that did have structured data. The weaknesses cluster. A site is rarely strong on everything except schema; more often, a thin machine layer is one symptom of a site that was built once, for a human, and never revisited for the machines.
Reachability told a quieter version of the same story. Across the full 51,001-business sample, 15.0 percent did not respond to an automated read at all: dead domains, timeouts, and a handful behind security challenges that block any machine. Those sites are excluded from every rate above. A site a machine cannot even reach is, for the purpose of being crawled and cited, in the hardest position of all, and there were more of them among the businesses with the fewest reviews.
How these numbers compare to the published record
A first-party number is worth more when it can be checked against independent measurement, so we set ours beside the published record. The broadest yardstick is HTTP Archive's Web Almanac, which reads structured data across the whole web. Its 2024 measurement finds JSON-LD on 41 percent of pages. Our local-business homepages carried it at 73.4 percent, well above that, which is the direction to expect: homepages concentrate structured data more than interior pages, and businesses that maintain an active map profile are the more digitally-active tier of small business to begin with.
The second comparison reframes the headline finding, and it is important enough to state carefully. The LocalBusiness type is rare across the web at large; the same Web Almanac measurement puts it on under 4 percent of pages. Our businesses used it at 44.7 percent, more than ten times the web-wide rate. Read one way that is a quiet success: these are exactly the businesses the LocalBusiness vocabulary was built for, and they adopt it far more than the average site does. Read the way that matters for being found, a majority still do not, for a signal that is close to free and squarely within the owner's control. Both readings are true. Holding both at once, ahead of the web but short of their own ceiling, captures the finding accurately, and it is why the gap reads as an opportunity rather than an indictment.
Our figure also sits comfortably among the handful of other first-party audits in this space. An independent audit of 5,000 sites reported roughly 29 percent carrying no schema, within about two points of our 26.6 percent; a web-wide technology index puts any-structured-data adoption near 79 percent across all sites, a little above our small-business rate, which is exactly what you would expect if small businesses trail the web average. When independent measurements taken different ways converge on the same order of magnitude, that convergence is the strongest evidence a single study can give that its number is real rather than an artifact of how it was collected.
A closer look: Core Web Vitals and accessibility
The census reads what a homepage exposes in its HTML, which is what lets it scale to thousands of sites cheaply. Two things it cannot see that way are how a page performs and how accessible it is, both of which need a real browser to measure. To read those, we deep-audited a 12-site subsample with a full performance-and-accessibility pass.
The pattern rhymed with the census. Load speed was a solved problem: all twelve passed the largest-contentful-paint threshold. The vital that failed was layout stability, with 3 of twelve shifting more than the 0.1 good threshold, one as high as 0.49, content sliding under a visitor's thumb as the page settles. Accessibility was the widest spread of anything measured, from 64 to 97 on the automated scale, median 89. This subsample is far too small to generalize, and is offered as a deeper read alongside the census, not as a second census.
What it means
Read together, these findings describe a specific and fixable gap. The costly foundations of a modern small-business site, speed, security, mobile layout, are mostly in place. What is thin is the cheap, machine-facing finish that lets an engine read a business well enough to name it, and it is thinnest exactly where it would help most: among credence-service trades whose buyers rely on being recommended, and among the newer, quieter businesses still trying to be found.
It matters to be precise about what that finish does and does not buy, because it is easy to oversell. Structured data is not a ranking lever. Google is explicit that structured data can make a page eligible for richer search features and helps an engine understand what a page is about, but does not on its own make a page rank better; local ranking, in Google's own account, turns on relevance, distance, and prominence, and structured data is not among them. So the claim here is narrower and firmer than add schema and you will win. It is that machine-readability is a precondition of eligibility. A business a machine cannot parse as a local entity is harder to place, harder to feature, and harder to cite, and a quarter of these businesses have handed an engine nothing to read at all.
Whether crossing that threshold lifts how often an AI answer actually names you is a genuinely open question. Research on generative-engine visibility finds that structure and citability can raise how often a source is used, but a recent critical survey concludes no technique yet shows a stable, cross-platform causal effect, and a separate ten-thousand-business analysis found structured data only weakly correlated with AI visibility, with attention concentrating instead on review volume and confirmed local presence. Our own data points the same way: readiness rose with reviews, so the businesses a machine can read best are largely the ones customers already vouch for. The synthesis is that being machine-readable is necessary, not sufficient. It is the floor you have to be standing on before reputation and prominence can lift you into an answer at all.
It is also worth being careful about cause inside our own numbers. This study measures what sites expose, not why, and the relationships here are associations, not proven causes. More reviews do not add schema to a site; both are downstream of a business investing in its presence over time. And the sample, drawn from businesses that already maintain an active profile and appear in map results, is the visible tier of small business, not the whole of it. The businesses with no profile at all are not here, and there is every reason to think they fare worse, so these figures set a floor on the gap, not a ceiling.
That is why the gap is worth closing even under uncertainty. The foundations are expensive and mostly built; the local-entity signal is nearly free and mostly missing. On the surfaces where attention is heading, the difference between a business a machine can read and one it cannot is the difference between being an option a system can consider and one it never sees. For most of these businesses that is an afternoon of work, not a rebuild, and it is the rare fix that costs little and removes a ceiling the owner did not know was there.
A benchmark we intend to keep
This is the first edition of a measurement we plan to repeat every quarter. A single snapshot tells you where the machine-readable small-business web stands today; a series tells you which way it is moving, and how fast. That second question is the more valuable one. If structured-data adoption is climbing quarter over quarter, the businesses that wait are falling behind a rising bar. If it is flat, the gap this edition found is durable, and the advantage of closing it lasts longer.
To make the editions comparable, every wave will hold the method fixed: the same twelve trades, a fresh draw of independent businesses from Google Business Profile map listings across the same kind of US metro spread, the same homepage read, the same Machine-Readiness Score, the same tiering rules. Each wave is a fresh random-within-method sample rather than the same 43,362 sites re-checked, so the series measures the population moving, not a fixed panel aging. After four quarters we will publish an evolution study that sets the waves side by side and reports the trend, by trade, by tenure, and by metro, with the same care for what the numbers do and do not prove.
That caveat carries forward too. Because each wave is drawn from businesses that already maintain an active profile, the series tracks the visible tier of small business over time, not the whole of it. Read as a floor that moves, not a full census of every small business in the country.
Limits, stated plainly
Every number here is measured, and every number here has bounds. The sample is drawn from businesses with an active Google Business Profile appearing in map results, a visibility-selected group; businesses without a profile are absent and almost certainly fare worse, so these rates understate the population gap. We read the homepage only, so a site carrying structured data on inner pages we did not visit is counted by its homepage; homepage rates are a floor. Structured-data detection parses JSON-LD, the dominant format, so the small number of sites using older microdata or RDFa are undercounted.
The performance and accessibility figures come from a 12-site subsample and are a deeper read, not a population estimate. The 15.0 percent of sites that did not respond are excluded from every on-page rate. All measurements are a single snapshot from late July 2026, and the chain-and-directory filter is heuristic, so a small number of non-independent sites may remain. None of these limits changes the central finding, which is large, statistically clear, and consistent across every cut we made: the machine-readable layer of the small-business web is thin, and it is thinnest where it would matter most.
The evidence, in numbers
Key findings, dated and sourced
-
26.6% of reachable independent small-business homepages (11,543 of 43,362) carried no structured data at all (95% CI 26.2-27.0%)
emerging Raveneye Global machine-readiness census (n=43,362 reachable of 51,001 independent US small businesses sampled from Google Business Profile map listings, 12 trades x 88 metros), captured 2026-07-29 to 2026-07-30
-
55.3% (23,985 of 43,362), a majority, had no LocalBusiness-family type an engine can read as a local entity
emerging Raveneye Global machine-readiness census (n=43,362 reachable of 51,001 independent US small businesses sampled from Google Business Profile map listings, 12 trades x 88 metros), captured 2026-07-29 to 2026-07-30
-
Only 33.1% were fully readable; 26.6% were machine-invisible; median Machine-Readiness Score 75.0/100 (IQR 60-100)
emerging Raveneye Global machine-readiness census (n=43,362 reachable of 51,001 independent US small businesses sampled from Google Business Profile map listings, 12 trades x 88 metros), captured 2026-07-29 to 2026-07-30
-
Machine-readiness differed significantly by trade (chi-square 819.8, df 11, p<0.001): Accounting lowest at 25.5%, Chiropractic highest at 50.8%
emerging Raveneye Global machine-readiness census (n=43,362 reachable of 51,001 independent US small businesses sampled from Google Business Profile map listings, 12 trades x 88 metros), captured 2026-07-29 to 2026-07-30
-
The investment effect was monotonic: LocalBusiness-schema rate rose from 30.3% (under 10 reviews) to 59.6% (500+ reviews), correlation 0.178 with the readiness score
emerging Raveneye Global machine-readiness census (n=43,362 reachable of 51,001 independent US small businesses sampled from Google Business Profile map listings, 12 trades x 88 metros), captured 2026-07-29 to 2026-07-30
-
The gaps cluster: sites with no structured data carried a meta description 50.9% of the time versus 85.0% for sites with structured data
emerging Raveneye Global machine-readiness census (n=43,362 reachable of 51,001 independent US small businesses sampled from Google Business Profile map listings, 12 trades x 88 metros), captured 2026-07-29 to 2026-07-30
-
Human basics were near-universal: 96.8% mobile-ready, 99.7% HTTPS, 92.0% with a visible phone; 15.0% of sampled sites did not respond to an automated read
emerging Raveneye Global machine-readiness census (n=43,362 reachable of 51,001 independent US small businesses sampled from Google Business Profile map listings, 12 trades x 88 metros), captured 2026-07-29 to 2026-07-30
-
In a 12-site deep subsample, all passed largest-contentful paint, but 3 of 12 exceeded the 0.1 CLS threshold; accessibility ranged 64 to 97 (median 89)
emerging Raveneye Global deep audit (12-site subsample, automated accessibility audit + Core Web Vitals from performance traces), 2026-07-29
-
The sample is drawn from businesses with an active Google Business Profile, so the true rate of missing structured data across all small businesses is likely higher, not lower
emerging Raveneye Global machine-readiness census (n=43,362 reachable of 51,001 independent US small businesses sampled from Google Business Profile map listings, 12 trades x 88 metros), captured 2026-07-29 to 2026-07-30
-
These local businesses use the LocalBusiness type at 44.7%, more than ten times the ~4% web-wide rate, yet a majority still lack it; the census figure also converges with an independent 5,000-site audit (~29% no schema vs our 26.6%)
established Raveneye Global machine-readiness census (n=43,362 reachable of 51,001 independent US small businesses sampled from Google Business Profile map listings, 12 trades x 88 metros), captured 2026-07-29 to 2026-07-30; Structured Data, Web Almanac 2024, HTTP Archive
-
Structured data is a precondition of eligibility, not a proven ranking lever: Google states it enables rich-result features and page understanding but not generic ranking, and the causal effect on AI-answer citation is not yet established across platforms
contested Google Search Central; Aggarwal et al. 2024 (arXiv:2311.09735); GEO critical survey 2026
Learning outcomes
What this study teaches
- Check your own homepage for JSON-LD structured data today; a quarter of these businesses had none, leaving an engine to guess about them or skip them.
- Do not settle for a generic Organization or WebSite tag; use a LocalBusiness-family type so an engine can place you as a local entity. A majority of these sites had not.
- If you are newer or have fewer reviews, act now; readiness rose monotonically with review volume, and the businesses that most need to be found were the least readable.
- If you are in professional services, accounting, therapy, law, assume you are behind on this signal; your trades measured the lowest, and it is the cheapest gap to close.
- Treat a thin machine layer as a symptom, not an isolated fault; the sites missing schema were missing other basics too, so audit the whole homepage, not just one tag.
Honest limits
What this does not yet settle
- The sample is drawn from businesses with an active Google Business Profile appearing in map results, a visibility-selected group. Businesses without a profile are absent and almost certainly fare worse, so these figures set a floor on the population gap, not a ceiling.
- We read the homepage only; a site can carry structured data on inner pages we did not visit, so homepage rates are a floor.
- Structured-data detection parses JSON-LD, the dominant format; the small number of sites using microdata or RDFa are undercounted.
- Core Web Vitals and accessibility come from a 12-site deep subsample, a deeper read rather than a population estimate.
- The 15.0% of sites that did not respond are excluded from every on-page rate; a fuller crawl with retries could reach some of them.
- All measurements are a single snapshot from late July 2026; the chain-and-directory filter is heuristic, so a small number of non-independent sites may remain. Relationships reported are associations, not proven causes.
This is a synthesis of dated, attributed evidence, not a census. The AI-answer layer in particular has no independent, Nielsen-grade measurement yet, so readings of it are directional and named as a frontier, never presented as settled.
Straight answers
Frequently asked questions
What share of small-business websites are missing structured data?
In our census of 43,362 reachable independent small-business homepages, drawn from active Google Business Profiles, 26.6 percent carried no structured data at all and 55.3 percent, a majority, had no LocalBusiness-family type an engine can read as a local entity. Because the sample already excludes businesses without a profile, the true rate across all small businesses is likely higher, not lower.
Does having more reviews mean a business is more machine-ready?
In our data, yes, and the relationship was monotonic. Businesses with more than five hundred reviews used a LocalBusiness-family type 59.6 percent of the time, against 30.3 percent for those with ten or fewer. Reviews do not cause good schema; both reflect a business investing in its online presence over time, but the pattern is clear: the busier the business, the more readable it was.
Which types of business are least machine-ready?
The professional-services trades. In our census, Accounting sites used a LocalBusiness-family type only 25.5 percent of the time and Physical therapy 35.2 percent, versus 50.8 percent for Chiropractic. The difference across trades is statistically significant. The trades whose buyers most verify credentials before booking invested least in the machine-readable layer.
Is a fast, mobile-friendly site enough to be machine-ready?
No. In our census the human basics were near-universal, 96.8 percent mobile-ready and 99.7 percent secure, yet a quarter had no structured data and a majority no local-entity type. Only 33.1 percent were fully readable. Being machine-ready is a separate job from being fast and mobile-friendly, and it is the one most often left undone.
How does this compare to structured-data use across the web?
These local businesses are ahead of the web at large. HTTP Archive's 2024 Web Almanac finds JSON-LD structured data on about 41 percent of web pages and the LocalBusiness type on under 4 percent. Our businesses carried structured data at 73.4 percent and a LocalBusiness type at 44.7 percent, more than ten times the web-wide rate for the local-entity type. That is the tension in the finding: these are exactly the businesses the LocalBusiness vocabulary was built for and they adopt it far more than average, yet a majority still do not, for a signal that is nearly free.
Will adding schema make my business rank higher or get cited by AI?
Not by itself. Google is explicit that structured data enables richer search features and helps an engine understand a page, but does not on its own improve ranking, and the evidence that it directly increases AI-answer citations is not yet settled across platforms. The claim is narrower: machine-readability is a precondition. A business a machine cannot parse as a local entity is harder to place, feature, or cite, and a quarter of the businesses we read had given an engine nothing to read at all. It is necessary, not sufficient.
How was this measured, and how reliable is it?
We took 51,001 independent small businesses from their Google Business Profile map listings across twelve trades and 88 US metros, filtered out chains and directories, and read each homepage's structured data and on-page signals directly; 43,362 were reachable. Every figure is a value a tool returned in late July 2026, and with this sample size the margins are narrow, near half a percentage point on the headline rates. Our figures also converge with independent audits and web-wide indices, which is the best sign a single study can give that its numbers are real. The main limit is that the sample is visibility-selected, so it understates the gap rather than overstating it.
Provenance
References
- Raveneye Global machine-readiness census, 51,001 independent US small businesses drawn from Google Business Profile map listings across 12 trades and 88 metros (43,362 reachable), structured data and on-page signals read from each homepage, captured 2026-07-29 to 2026-07-30.
- Raveneye Global deep audit, 12-site subsample, accessibility on a 0-to-100 automated scale and Core Web Vitals from performance traces, captured 2026-07-29.
- Structured Data, Web Almanac 2024 (JSON-LD present on 41% of mobile pages; LocalBusiness on 3.97%; WebSite 12.73%, Organization 7.16%), HTTP Archive https://almanac.httparchive.org/en/2024/structured-data
- Schema.org adoption statement (over 45 million domains, over 450 billion Schema.org objects, 2024), Schema.org https://schema.org/
- General structured data guidelines, and Introduction to structured data markup (structured data enables rich-result features and page understanding, not generic ranking), Google Search Central https://developers.google.com/search/docs/appearance/structured-data/sd-policies
- Local business (LocalBusiness) structured data, Google Search Central https://developers.google.com/search/docs/appearance/structured-data/local-business
- Improve your local ranking on Google (local results based primarily on relevance, distance, and prominence), Google Business Profile Help https://support.google.com/business/answer/7091
- LocalBusiness type and its subtypes, Schema.org https://schema.org/LocalBusiness
- Darby, M. R. and Karni, E. (1973), Free Competition and the Optimal Amount of Fraud (the credence-goods concept), Journal of Law and Economics 16(1), 67-88.
- Nelson, P. (1970), Information and Consumer Behavior (search vs experience goods), Journal of Political Economy 78(2), 311-329.
- Dulleck, U. and Kerschbamer, R. (2006), On Doctors, Mechanics, and Computer Specialists: The Economics of Credence Goods, Journal of Economic Literature 44(1), 5-42.
- Aggarwal, P. et al. (2024), GEO: Generative Engine Optimization, ACM SIGKDD (KDD '24), arXiv:2311.09735. https://arxiv.org/abs/2311.09735
- Liu, N. F., Zhang, T. and Liang, P. (2023), Evaluating Verifiability in Generative Search Engines (only 51.5% of generated sentences fully supported by citations), EMNLP Findings, arXiv:2304.09848. https://arxiv.org/abs/2304.09848
- Clutch (2025), No-Code Tools Fuel Website Growth, Yet 17% of Small Businesses Are Still Offline (83% of US small businesses have a website).
- Core Web Vitals thresholds (LCP under 2.5s, CLS 0.1 or less at the 75th percentile), Google web.dev https://web.dev/articles/vitals
- Companion first-party analyses discussed above: The MSME Visibility Gap (an independent 5,000-site audit near 29% carrying no schema); The Digital Divide 2.0 (a web-wide index near 79% any structured data); Reputation Physics (a 10,000-business analysis on what AI answers weight), Raveneye Global, 2026.
Every measured figure is dated to its capture and tagged with an evidence tier. Every cited work is real and locatable. Where an engine could not be captured this round, it is named as uncaptured, not estimated. Small-sample readings are labelled as directional.