Discovery Science · mixed evidence

The Honesty Audit: Which AI-Visibility Claims Are Actually Backed by Evidence

Last reviewed 2026-07-20. Written by Chandranshu Kumar, Founder, Raveneye Global. · 11 min read

Every week brings a new confident claim about AI visibility: that a text file at the root of your site controls whether ChatGPT names you, that schema markup delivers a measured accuracy jump, that the right vendor can guarantee an AI ranking. This is an honesty audit of those AI visibility claims, each one tested against the primary sources rather than the marketing around it. Some hold up cleanly. A controlled, peer-reviewed study does show that specific content changes raise how often a source is cited inside a generated answer. Others collapse on contact with the data: an analysis of 137,210 domains found that almost no one reads the llms.txt file marketed as a lever for AI visibility. And a few, guaranteed rankings chief among them, rest on no published evidence at all and are contradicted by how answer engines actually work. The pattern matters more than any single verdict: in this field, the confidence of a claim and the evidence behind it are rarely the same size.

What it takes to audit a visibility claim

The problem is not that AI-visibility advice is all wrong. It is that a peer-reviewed result, a single unreplicated study, and an unsupported sales line are all sold in the same confident register, so a buyer has no way to tell one from another. A rigorous audit restores that distinction by putting the same three questions to any claim before believing it.

These are not academic niceties. They are the difference between spending a budget on something the evidence supports and spending it on a tactic that a primary source has already quietly retired.

  • What primary source supports it? Not a blog citing a blog, but a study, a standards document, or a platform's own statement that can be read directly.
  • What evidence tier is it? Established (replicated or standards-backed), emerging (a real but single, unreplicated finding), or contested (an industry narrative with no primary source behind it).
  • What was it measured against? A number with no baseline, no control, and no place of publication is an assertion wearing the costume of a finding.

Backed by evidence: content changes can raise how often you are cited

Start with the claim that survives audit, because it sets the bar. The idea that you can influence generative-engine visibility through content is not vendor folklore; it is the founding, peer-reviewed result of the field. In the paper that coined generative engine optimization, Aggarwal and colleagues ran a controlled benchmark of roughly 10,000 queries across nine datasets and measured which content-level interventions changed whether a source was surfaced inside a generated answer.

The strongest levers were concrete: adding citations to credible sources, including direct quotations, and replacing vague statements with specific statistics. Together these produced a 30 to 40 percent relative lift on the study's visibility metric against an unoptimized baseline, and citing authoritative sources was consistently the single strongest move.

Read that verdict precisely, because precision is the point of an audit. This is a relative lift on one visibility measure, in the engines the researchers tested, versus content that had done nothing. It is not an accuracy figure, and it is not a promise that any single query will cite you. What it is, is real: published, controlled, and replicable, which is more than almost any other claim in this space can say.

Not backed by evidence: the llms.txt AI-visibility claim

Now the claim that collapses. A well-marketed tactic holds that placing an llms.txt file at your site root, a curated summary meant to hand AI systems a clean version of your content, improves how you appear in AI answers. The primary data does not support it.

Ahrefs analyzed 137,210 domains and found that 97 percent of valid llms.txt files received zero requests in the month studied. Of the small remainder that were fetched at all, 96 percent of fetches were generic bots, and only about a fifth of that residual traffic came from named AI tools. A second study, from SE Ranking across 300,000 domains, put adoption at roughly 10 percent, with no evidence of a visibility effect. Google's John Mueller has stated publicly that llms.txt is "not done for search," and Google's own guidance on optimizing for its AI features is explicit that no special file or markup is required to appear in them.

The verdict: llms.txt is not read at meaningful scale and is not treated as a ranking or citation signal by any major AI platform. It may still be a convenience for AI coding tools reading developer documentation, which is a narrow and different job. As a lever on whether a buyer's question names your business, the evidence is that it does nothing.

Half true, oversold: the structured-data accuracy jump

Some claims are harder because a real thing sits underneath an inflated number. Structured data is the clearest case. Schema.org is a genuine, standards-body vocabulary, founded jointly by Google, Bing, Yahoo, and Yandex in 2011 so that machine-readable markup means the same thing across every engine that consumes it. Marking up your pages has documented value: it lets engines parse your facts unambiguously and makes a page eligible for rich results.

The oversell is the statistic bolted on top. Claims of a specific "accuracy jump" or a percentage "lift in AI citations" credited to adding schema almost never arrive with a baseline, a control group, or a place of publication. The 30 to 40 percent figure that does exist in a controlled, peer-reviewed study is the one from the generative-engine-optimization paper above, and it is attributed to content-level changes, not to markup. Google, for its part, states plainly that no special markup is required to appear in its AI features, which run on the core index.

The audit's verdict is a split decision. Structured data is worth engineering for its parsing and eligibility value, and it is one of the few parts of the discovery stack you can inspect end to end through open validators. The quantified promise usually attached to it is a separate object, and that object generally fails the "what was it measured against" test.

Architecturally unsupported: getting cited proves the model read you

A subtler claim treats an AI citation as evidence that the model reasoned from your page. The retrieval-augmented-generation literature says that inference is unsafe. In the generate-then-retrieve and retrieve-then-generate pipelines that dominate production systems, a citation is attached to an answer largely independently of the evidence that actually produced it, so the source credited is not reliably the source the model used. Researchers argue that only inline, generation-time citation can guarantee faithfulness, and most deployed systems do not work that way.

The scale of the gap is documented. An attribution survey reports that more than 95 percent of answers from the open-source language models it tested contained at least one unattributed sentence. This is an established finding about the faithfulness problem; which mitigation actually closes it remains an open, emerging question.

The verdict: winning a citation is a real and useful signal, but it is not proof that your content shaped the answer, and it is certainly not a mechanism you can reverse-engineer with confidence. Any claim that treats "we got you cited" as a controllable, faithful outcome is overstating what the architecture currently delivers.

Unsupportable: the guaranteed AI ranking

This is the claim that fails hardest, and it fails on two counts at once. First, no primary source establishes that any intervention guarantees a citation or a fixed position inside a generated answer. Second, two documented properties of these systems make such a guarantee unsupportable rather than merely unproven.

The first property is non-determinism. Answer engines sample their output, so the same question, asked twice, can return a different set of named sources; a single check tells you almost nothing, and a guarantee about the next check would claim more certainty than the mechanism allows. The second is the decoupling and faithfulness gap above: rank and citation are not the same job, and the link between what a model reads and what it credits is unreliable. Stack these together and a promised, repeatable AI ranking is not a bold claim, it is a claim the evidence contradicts.

This is why measuring the AI-answer surface means reporting a presence rate with a confidence band, sampled many times per engine and stamped with the engine, locale, and date it was run against. That range reflects how these systems actually behave: sampled, non-deterministic, never fixed. A single guaranteed number would claim a certainty the field does not have.

Mislabeled: raise your E-E-A-T score

Some claims are not false so much as category errors, and "optimize your E-E-A-T score" is the common one. Experience, Expertise, Authoritativeness, and Trust are real concepts, but Google's own documentation is explicit about what they are: a framework that thousands of human raters use to evaluate the quality of the search algorithm's output, a feedback signal on the system rather than a score computed for any individual page.

Google has repeatedly clarified that E-E-A-T is not itself a direct ranking factor a page can optimize into a document. A large share of commercial GEO and AEO advice nonetheless treats a rater heuristic as if it were a machine-readable target with a dial you can turn. The underlying qualities genuinely matter for how content is judged over time; the "score" you are being sold the ability to raise does not exist as a thing an algorithm reads off your page.

Handle with care: AEO vs GEO as interchangeable terms

The last item is a terminology claim, and it earns a contested tier rather than a clean pass or fail. Industry accounts describe answer engine optimization as the older label, inherited from the featured-snippet and voice-search era of the mid-2010s, when the job was structuring content to be lifted as a direct, extractable answer. Generative engine optimization is the newer, academically coined term for being retrieved and cited by a large language model that synthesizes an answer.

In everyday usage the two labels have collapsed into near-synonyms. That is a terminology-hygiene problem, not a settled fact, and this narrative is an industry consensus rather than a peer-reviewed or standards-body claim. The more accurate handling is to keep the distinction the mechanisms deserve: AEO vs GEO share many tactics, but a pre-LLM featured snippet and an LLM-synthesized citation are produced by different machinery, and any content that flattens them into one thing should be read as convenient shorthand, not as established science.

A four-question test for any AI-visibility claim

The value of an audit is not the individual verdicts, which will age as the systems change. It is the method, which does not. The next confident claim you are handed, from us or from anyone, can be run through the same four questions before it earns your budget.

  • Where is the primary source? If the trail ends at a blog citing another blog, treat the claim as marketing until a study, a standard, or a platform statement is produced.
  • What is the evidence tier? Established, emerging, or contested are not the same and should never be priced the same.
  • What was the number measured against? No baseline, no control, and no publication means no finding, however precise the figure looks.
  • Does it survive how the systems actually work? A claim that ignores non-determinism, the ranking-citation split, or the faithfulness gap is describing a machine that does not exist.

The evidence

Key findings, with their sources

  • Content-level interventions (adding citations to credible sources, direct quotations, and specific statistics) produced a 30 to 40 percent relative lift on a generative-answer visibility metric against an unoptimized baseline, with citing authoritative sources the single strongest lever.

    established Aggarwal et al., "GEO: Generative Engine Optimization", arXiv:2311.09735, ACM SIGKDD 2024 (controlled benchmark, ~10,000 queries across nine datasets).

  • Of 137,210 domains analyzed, 97% of valid llms.txt files received zero requests in the month studied; of the fetched remainder, 96% of fetches were generic bots and only about 19.5% of that residual traffic came from named AI tools.

    established Ahrefs, "We Analyzed 137K Sites: 97% of llms.txt Files Never Get Read", June 2026.

  • A second large study put llms.txt adoption at roughly 10% across 300,000 domains, with no evidence of a visibility effect; Google states llms.txt is "not done for search" and no special file or markup is required to appear in its AI features.

    established SE Ranking analysis of ~300,000 domains, 2026; John Mueller / Google Search Central guidance on AI features, 2026.

  • More than 95% of answers from tested open-source large language models contained at least one unattributed sentence, and in dominant RAG pipelines a citation is attached largely independently of the evidence that produced the answer.

    established Attribution survey, arXiv, 2025-26; "Measuring and Enhancing Trustworthiness of LLMs in RAG through Grounded Attributions and Learning to Refuse", arXiv:2409.11242, 2024.

  • Schema.org is a standards-body vocabulary founded jointly by Google, Bing, Yahoo, and Yandex in 2011; Google states no special markup is required to appear in its AI features, which run on the core index.

    established Schema.org founding documentation, 2011; Google Search Central, guidance on optimizing for AI features, accessed 2026.

  • E-E-A-T is a human-rater evaluation heuristic used to assess the quality of the search algorithm, not a scored ranking factor a page can optimize into a document.

    established Google Search Central, "E-A-T gets an extra E for Experience", December 2022, and the public Search Quality Rater Guidelines.

Calibration

What is proven, what is promising, what is unproven

Evidence tierTacticsWhat the evidence says
Established (holds)Content-level GEO changes (citations, quotations, statistics) can raise generative-answer visibility.Aggarwal et al., GEO, arXiv:2311.09735, KDD 2024, controlled benchmark.
Established (holds)Structured data has real parsing and rich-result value as an open, inspectable standard.Schema.org founding, 2011; Google structured-data guidelines.
Established (refuted)llms.txt controls or improves your AI-answer visibility.Ahrefs 137K-domain study, 2026; SE Ranking ~300K domains; Google "not done for search".
Established (refuted)A specific "accuracy jump" percentage is delivered by adding schema markup.No primary source; the real 30-40% figure is content-level in the GEO paper, not markup.
Established gapAn AI citation proves the model reasoned from your page.RAG faithfulness literature; >95% of tested answers carry an unattributed sentence.
UnsupportableA vendor can guarantee an AI ranking or citation.No primary evidence; contradicted by non-determinism and the ranking-citation decoupling.
ContestedAEO and GEO are interchangeable terms describing one thing.Industry-consensus narrative, not primary-sourced; the mechanisms differ.

Reference

Glossary

AI visibility claim
Any assertion that a specific tactic changes whether or how a business appears in the answers AI engines generate. The subject of this audit, tested against primary sources rather than accepted on confidence.
Evidence tier
A verdict on how well-supported a claim is: established (replicated or standards-backed), emerging (a real but single, unreplicated finding), or contested (an industry narrative with no primary source).
llms.txt
A proposed root-level file meant to hand AI systems a curated summary of a site. Primary data shows it is largely unread and is not treated as a ranking or citation signal by major AI platforms.
Citation faithfulness
Whether the source an AI answer cites is genuinely the source the model reasoned from. In dominant retrieval-augmented pipelines it often is not, which is a documented, unresolved gap.
Generative engine optimization (GEO)
The academically coined practice of influencing whether a large language model retrieves and cites a source inside a synthesized answer. Its founding study is peer-reviewed and controlled.
E-E-A-T
Experience, Expertise, Authoritativeness, Trust: a framework human raters use to evaluate Google's output quality, per Google's own documentation, not a score computed for an individual page.

Straight answers

Frequently asked questions

Does llms.txt actually help my business show up in AI answers?

The primary data says no. An analysis of 137,210 domains found 97 percent of valid llms.txt files received zero requests in the month studied, a second study put adoption at roughly 10 percent with no visibility effect, and Google has stated the file is not used for search. It may be a convenience for AI coding tools reading developer docs, but as a lever on whether a buyer's question names you, there is no evidence it works.

Can anyone guarantee an AI ranking or a citation in ChatGPT or Google AI Overviews?

No, and the guarantee is not just unproven, it is contradicted by how the systems work. Answer engines sample their output, so the same question can return different named sources on repeat runs, and the link between what a model reads and what it cites is documented as unreliable. A read of this surface reports a presence rate with a confidence band, sampled many times, never a guaranteed number.

Is the "30 percent accuracy jump from schema markup" claim real?

Treat it as unverified until it arrives with its working. The 30 to 40 percent figure that exists in a controlled, peer-reviewed study is a relative visibility lift attributed to content-level changes, citations, quotations, and specific statistics, not to schema. Percentage jumps credited specifically to markup usually have no baseline, control, or publication. Structured data is worth doing for its real parsing value; the number bolted onto it generally is not supported.

What is the difference between AEO and GEO?

Industry accounts describe answer engine optimization as the older label from the featured-snippet and voice-search era, when the job was making content extractable as a direct answer, and generative engine optimization as the newer term for being retrieved and cited by a synthesizing language model. In practice the labels have blurred into synonyms. That blurring is an industry convention, not a settled fact, and the underlying machinery differs, so it is worth keeping the distinction rather than treating the two as one.

How can I tell if an AI-visibility claim is trustworthy?

Run it through four questions: where is the primary source, what evidence tier is it (established, emerging, or contested), what was any number measured against, and does the claim survive how the systems actually work. A claim that fails on the source, the baseline, or the mechanics is marketing, however confident it sounds.

Provenance

Sources

  1. Aggarwal, P., Murahari, V., Rajpurohit, T., Kalyan, A., Narasimhan, K., Deshpande, A., "GEO: Generative Engine Optimization", arXiv:2311.09735, ACM SIGKDD 2024 (established, peer-reviewed)arxiv.org
  2. Ahrefs, "We Analyzed 137K Sites: 97% of llms.txt Files Never Get Read", June 2026 (established)ahrefs.com
  3. SE Ranking, llms.txt adoption analysis across ~300,000 domains, 2026 (established)
  4. Google Search Central / John Mueller, statements that llms.txt is "not done for search" and that no special markup is required to appear in AI features, 2026 (established)
  5. Schema.org, structured-data vocabulary jointly founded by Google, Bing, Yahoo, and Yandex, 2011-present (established)
  6. Google Search Central, structured-data general guidelines and guidance on optimizing for AI features, accessed 2026 (established)
  7. "Measuring and Enhancing Trustworthiness of LLMs in RAG through Grounded Attributions and Learning to Refuse", arXiv:2409.11242, 2024 (established finding of the faithfulness gap)arxiv.org
  8. "Attribution Techniques for Mitigating Hallucinated Information in RAG Systems: A Survey", arXiv, 2025-26 (established gap; emerging on mitigation)
  9. Google Search Central, "Our latest update to the quality rater guidelines: E-A-T gets an extra E for Experience", December 2022, and the public Search Quality Rater Guidelines (established)
  10. Industry accounts of the AEO (featured-snippet/voice era) versus GEO (LLM era) lineage (contested, industry-consensus narrative, not primary-sourced)

Every figure above is attributed to a real, dated source and tagged with its evidence tier. Where a claim could not be verified to a primary source, it is not stated as fact.

Where this leaves your business

If most AI-visibility claims cannot survive an audit, the place to start is not another claim, it is a measurement of where you actually stand. A Surface Intelligence Audit reads your presence across classic search, AI answers, reputation, and your technical foundation, samples the AI-answer surface many times per engine, and reports it as a rate with a confidence band, dated and stamped with the engines and questions it was run against. No guaranteed number, because the evidence does not support one. Just a defensible starting position before anyone asks you to spend.

diagnostic Surface Intelligence Audit A measured, specialist-read diagnostic of where you stand across all four surfaces, with your named competitors scored on the same buyer-question panel and a ranked list of the corrections that move your number most. Every figure is stamped with how and when it was measured. See how it works

Start free with a Machine-Readiness Score, a specialist-reviewed read of where you stand across search and AI answers. No obligation.