Discovery Science · emerging evidence
The State of AI Visibility: A Baseline Reading (Volume 1 of the Surface Report)
AI visibility is how often a business is named and cited when a buyer asks a generative engine, ChatGPT, Perplexity, Gemini, or a Google AI Overview, instead of typing a query and scanning a list of links. This article opens the Surface Report, a planned quarterly index that will publish a real cross-client baseline for AI visibility once the underlying data holds enough businesses to support a reliable reading. Volume 1 is deliberately a methodology, not a headline number. The evidence for why such an index matters is already established and independent: a 2025 Pew Research Center study of 68,879 Google searches, and the peer-reviewed founding paper on generative engine optimization. What does not yet exist is a trustworthy cross-business reading of who is actually surfaced inside AI answers. This piece states exactly what the Surface Report will measure, how it will sample, and the rules it holds to, so the first published reading can be judged against a standard set in advance.
Why AI visibility needs a baseline no one has yet published
There is no shortage of commentary about AI search. There is a shortage of trustworthy measurement. No generative engine publishes a full account of which businesses it names, how often, or why, and the numbers that circulate are largely vendor-produced, methodologically opaque, or drawn from a single account. For an individual owner the practical question is narrower and unanswered by any of it: when a buyer asks an engine who does this kind of work, how often does my business appear in the answer, and how does that compare with the businesses that appear instead of me?
A baseline is the precondition for answering that with evidence. Without a reference distribution built from real businesses, any single reading is a number without a scale, and any claim of improvement is an assertion rather than a measurement. The Surface Report exists to build that reference distribution and to publish it on a fixed cadence, so that a later reading can be compared against an earlier one and against a peer set, rather than against a marketing figure.
The strongest independent evidence available today establishes the problem the index addresses, not its solution. It shows that the behavior of AI answers is measurably different from the behavior of the classic results page, which is precisely why the old yardsticks do not carry over and a new baseline has to be constructed rather than borrowed.
What the established evidence already shows, and what it does not
The Surface Report is not a leap of faith. It rests on a body of primary evidence that is independent, methodologically transparent, and in several cases peer-reviewed. That evidence is strong enough to justify measuring AI visibility as its own discipline. It is not, by itself, a baseline of who wins, which is the gap the index is built to close.
The click behavior is measured
Pew Research Center tracked the real browsing of 900 US adults across 68,879 Google searches in March 2025, of which 12,593 contained an AI summary. Users clicked a traditional organic result in 8 percent of searches when an AI Overview was present, against 15 percent when it was not, and clicked a link inside the summary itself only about 1 percent of the time. This is the single strongest data point in the field for one reason: it was independent of any search or optimization vendor, and it reports what people actually did rather than what a survey said they intended.
The operational consequence is that organic traffic is now the wrong success metric for the AI-answer surface. A business can be read by the model, contribute to the answer, and receive almost no click for it. Being the named entity in the answer, not the buried link in a citation list, is the outcome that carries the buyer forward.
The optimization levers are measured
The founding study of the field, published at ACM SIGKDD in 2024, ran a controlled benchmark of roughly 10,000 queries across nine datasets to test which content-level changes alter whether a source is cited inside a generated answer. Adding citations to credible sources, including direct quotations, and replacing vague claims with specific statistics produced a 30 to 40 percent relative lift on the study's visibility metric against unoptimized baselines, and citing authoritative sources was consistently the single strongest lever. That such a discipline can be measured at all establishes the object of the index: the thing being optimized is no longer a rank, it is a citation.
The citation itself is not fully trustworthy
A sober counterweight sits beside the optimization result. Research on citation faithfulness in retrieval-augmented systems, the architecture underneath these engines, found that in the dominant production patterns a citation is attached to an answer independently of the evidence that actually produced it, so a cited source is not reliably the source the model reasoned from. A related attribution survey reports that over 95 percent of answers from the open-source models it tested contained at least one unattributed sentence. Winning a citation is therefore necessary but not sufficient, and any index that treats a citation as proof of influence would overclaim. The Surface Report measures presence in the answer, and states plainly what presence does and does not prove.
AI search statistics are abundant and mostly unreliable
The reason a careful index is needed rather than another dashboard is that the existing supply of AI search statistics is noisy and, in places, actively misleading. Two documented examples set the standard the Surface Report has to clear.
The first is llms.txt, the proposed root-level file promoted as the way to hand AI systems a clean summary of a site. An analysis of 137,210 domains found that 97 percent of valid llms.txt files received zero requests in May 2026, and Google has stated publicly that the file is not used for search. A tactic marketed as a control lever is, on the evidence, read by almost nothing. The second is E-E-A-T. Google's own documentation describes Experience, Expertise, Authoritativeness, and Trust as a framework its human raters use to evaluate the algorithm, and Google has repeatedly clarified that it is not itself a scored ranking factor a page can optimize into a document. A large share of commercial advice conflates the rater heuristic with a machine-readable target.
The pattern is consistent: the tooling layer sold as how you control AI visibility is younger, leakier, and less standardized than the marketing suggests. An index that wants to be trusted has to be visibly stricter than the field it reports on, which is why the method is published before any number is.
What the Surface Report will measure: the four pillars of the Machine-Readiness Score
The unit of the index is the Machine-Readiness Score, a 0 to 100 reading composed of four pillars, held separately rather than blended into a single rank. The pillars are classic search, the local map pack, AI answers, and reputation. They are kept distinct because the evidence says they behave differently and move for different reasons.
The AI-answer pillar is the one the Surface Report is built around, and it is defined against the evidence above. It measures citation and mention presence across the major generative engines, explicitly not organic traffic, because the Pew data shows organic traffic understates this surface. A business is scored on how often it is named when a real buyer question is asked, and on the quality and consistency of the sources the engine reaches for when it answers.
The remaining three pillars anchor the reading in the surfaces that still carry volume and still decide local trade. Classic search and the local map pack remain where a large share of intent lands, and reputation, the corroborating record of reviews and third-party mention, feeds both the classic vote and the AI citation. Separating them matters: a single blended number would hide exactly the decoupling that makes this era worth measuring.
Generative engine optimization and the decoupling of rank from citation
The intellectual case for measuring classic search and AI answers as separate pillars is that ranking and citation are now measurably different jobs. The generative engine optimization literature identifies content levers that lift citation, and those levers are not identical to the ones that lift a classic rank. A 2025 large-scale controlled follow-up went further and found that AI answer engines are systematically biased toward earned, third-party media over brand-owned and social content, a sharp contrast with classic Google, which the same study found sources more evenly. That finding is a single study and not yet replicated, so the Surface Report treats it as emerging rather than settled, but it points the same direction as the founding paper: the sources an engine trusts to answer a question are not the same set that a classic ranking rewards.
The direct implication for the index is that a business can rank on page one and still be absent from the answer written above it, and can be cited from a page that sits far outside the top organic results. A number that averaged the two together would report a comfortable middle that describes neither. Measuring them apart is the only way to tell an owner which of the two races they are actually losing.
The rules the Surface Report holds itself to
An index is only as credible as the rules it refuses to break under pressure. The Surface Report commits to a short list, in writing, ahead of the first release.
- No fabricated or naked number. Every figure carries its source and its method, and no metric is reported without the sampling behind it.
- Method before numbers. The protocol is published first, so a later reading can be checked against a standard set in advance rather than one written to fit the result.
- Evidence is tiered by strength. Established, emerging, and contested findings are labeled as such, and a single unreplicated study is never dressed up as settled fact.
- No guaranteed outcome. The index reports where a business stands and how it moves. It does not promise a ranking, a citation, or a position.
- The corpus must earn the baseline. No cross-business figure is published until the Visibility Corpus holds enough entities to report a reading rather than an anecdote.
What Volume 1 does and does not claim
To be unambiguous: this installment publishes no cross-client Machine-Readiness Score baseline, because a trustworthy one does not exist yet. The Visibility Corpus is still accruing, and a number reported before it reaches usable scale would violate the first rule on the list above. What Volume 1 delivers is the standard the first real reading will be held to, grounded in independent, primary evidence for why AI visibility has to be measured as its own thing.
The reading itself follows when the data supports it. In the meantime the starting point is not a published industry average, it is a measured read of a single business, the same read that, aggregated under one protocol, becomes the corpus this report will one day publish from.
The evidence
Key findings, with their sources
-
Users clicked a traditional organic result in 8% of searches with an AI Overview present, versus 15% without, and clicked a link inside the summary only about 1% of the time, across 68,879 real Google searches (12,593 with an AI summary) from 900 US adults.
established Pew Research Center, "Do people click on links in Google AI summaries?", July 22, 2025 (browsing-panel study).
-
Adding citations to credible sources, direct quotations, and specific statistics produced a 30 to 40 percent relative lift on a generated-answer visibility metric versus unoptimized baselines, with citing authoritative sources the single strongest lever, across ~10,000 queries and nine datasets.
established Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan, Deshpande, "GEO: Generative Engine Optimization", arXiv:2311.09735, ACM SIGKDD 2024 (peer-reviewed).
-
Over 95 percent of answers from tested open-source LLMs contained at least one unattributed sentence, and in dominant production pipelines a citation is attached independently of the evidence that produced the answer.
established RAG faithfulness study, arXiv:2409.11242, 2024; attribution survey arXiv:2601.19927, 2025-26.
-
AI answer engines are systematically biased toward earned, third-party media over brand-owned and social content, a contrast with classic Google, which sources more evenly.
emerging Chen, Wang, Chen, Koudas, "Generative Engine Optimization: How to Dominate AI Search", arXiv:2509.08919, 2025 (single large-scale study, not yet replicated).
-
97 percent of valid llms.txt files across 137,210 domains received zero requests in May 2026, and Google states the file is not used for search.
established Ahrefs, "We Analyzed 137K Sites: 97% of llms.txt Files Never Get Read", June 2026; Google Search Central statements.
-
E-E-A-T is a framework Google's human raters use to evaluate the algorithm, not a scored ranking factor a page can directly optimize into a document.
established Google Search Central, "E-A-T gets an extra E for Experience", December 2022; Google Search Quality Rater Guidelines.
Calibration
What is proven, what is promising, what is unproven
| Evidence tier | Tactics | What the evidence says |
|---|---|---|
| established | AI-answer click behavior and the organic-traffic reframe | Pew Research Center, 68,879 real Google searches, 2025. |
| established | Which content levers lift citation inside a generated answer | Aggarwal et al., GEO, ACM SIGKDD 2024, arXiv:2311.09735. |
| established | Citation faithfulness is an unresolved gap, not a solved feature | arXiv:2409.11242, 2024; attribution survey arXiv:2601.19927. |
| emerging | Earned, third-party media is favored over owned content in AI citation | Chen et al., arXiv:2509.08919, 2025 (single study). |
Reference
Glossary
- AI visibility
- How often, and how prominently, a business is named and cited when a buyer asks a generative engine a question, rather than typing a query into a classic results page.
- Surface Report
- A planned quarterly, owned-data index that will publish a cross-business baseline of AI visibility once the Visibility Corpus holds enough entities to support a reliable reading.
- Machine-Readiness Score
- A 0 to 100 reading of a single business across four pillars held separately: classic search, the local map pack, AI answers, and reputation.
- Visibility Corpus
- The reference distribution built by sampling many businesses under one consistent protocol; the data the Surface Report publishes from once it reaches usable scale.
- How often a business is named in AI answers to a sampled panel of real buyer questions, measured across the major engines and over time, instead of by organic clicks.
- Zero-click search
- A search that ends without a click to any external website, because the answer is satisfied on the results page itself.
Straight answers
Frequently asked questions
What is AI visibility?
AI visibility is how often and how prominently a business is named and cited when a buyer asks a generative engine such as ChatGPT, Perplexity, Gemini, or a Google AI Overview, rather than scanning a list of links. It is measured by presence in the answer, not by organic clicks, because the click behavior of AI answers is measurably different from a classic results page.
Is the Surface Report baseline available now?
No. Volume 1 is the published method, not a number. A cross-business baseline is only trustworthy once the Visibility Corpus holds enough entities across enough verticals to report a reading rather than an anecdote, and until it does, no baseline figure is published. This piece states the standard the first real reading will be held to.
How will AI visibility be measured?
By sampling a panel of the real questions a business's buyers ask, phrased the way buyers phrase them, and putting each to the major engines on a fixed cadence. The method records how often the business is named, alongside which competitors, and from which sources, producing a share-of-answer figure and its movement over time rather than a one-off screenshot.
Why not just use the AI search statistics already circulating?
Because most of them are vendor-produced, methodologically opaque, or drawn from a single account, and some are actively misleading. Documented examples include llms.txt, promoted as a control lever but read by almost nothing, and E-E-A-T, widely misread as a ranking factor when Google describes it as a rater framework. An index worth trusting has to be visibly stricter than the field it reports on.
Is AI visibility the same as an SEO ranking?
No. The evidence shows ranking and citation are now different jobs: a business can rank on page one and be absent from the AI answer above it, and can be cited from a page far outside the top organic results. That is why the Machine-Readiness Score holds classic search and AI answers as separate pillars instead of blending them into one number.
Provenance
Sources
- Pew Research Center, "Do people click on links in Google AI summaries?", July 22, 2025 (established)pewresearch.org
- Aggarwal, P., Murahari, V., Rajpurohit, T., Kalyan, A., Narasimhan, K., Deshpande, A., "GEO: Generative Engine Optimization", arXiv:2311.09735, ACM SIGKDD 2024 (peer-reviewed, established)arxiv.org
- Chen, M., Wang, X., Chen, K., Koudas, N., "Generative Engine Optimization: How to Dominate AI Search", arXiv:2509.08919, 2025 (emerging, single large-scale study)arxiv.org
- Lewis, P. et al., "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks", arXiv:2005.11401, NeurIPS 2020 (established)arxiv.org
- "Measuring and Enhancing Trustworthiness of LLMs in RAG through Grounded Attributions", arXiv:2409.11242, 2024; "Attribution Techniques for Mitigating Hallucinated Information in RAG Systems: A Survey", arXiv:2601.19927, 2025-26 (established gap, emerging on fixes)arxiv.org
- Singhal, A., "Introducing the Knowledge Graph: things, not strings", Official Google Blog, May 16, 2012 (established)blog.google
- Google Search Central, "E-A-T gets an extra E for Experience", December 2022, and the Search Quality Rater Guidelines (established)
- Ahrefs, "We Analyzed 137K Sites: 97% of llms.txt Files Never Get Read", June 2026 (established)ahrefs.com
Every figure above is attributed to a real, dated source and tagged with its evidence tier. Where a claim could not be verified to a primary source, it is not stated as fact.