Discovery Science · established evidence
The llms.txt Autopsy: What Happens When 137,000 Sites Actually Get Checked
llms.txt, a proposed root-level text file meant to hand AI systems a clean, curated summary of a website, has been promoted across the discovery-marketing industry as a way to control how a business appears in AI answers. In 2026 that promise was finally measured at scale, and the data is unkind. Ahrefs checked 137,210 domains and found that 97 percent of valid llms.txt files received zero requests in a single month, and of the small share fetched at all, almost all of the traffic was ordinary bots rather than named AI tools. Google's John Mueller has said plainly that llms.txt is not done for search. Read together, the large-scale read-rate data and the platform's own statements support one conclusion: llms.txt is, at present, essentially never read and is not treated as a ranking or citation signal by any major AI platform. This piece takes that evidence apart, then asks what the same body of research says actually moves AI visibility.
The promise, and the test that finally checked it
The pitch for llms.txt is intuitive enough to sell itself. A website is a tangle of navigation, scripts, and boilerplate; large language models have finite context windows; therefore, the argument goes, a single clean file at the root of a domain, listing a curated summary and the pages that matter, should help AI systems read a site correctly and represent it accurately in generated answers. Framed that way, publishing one sounds like basic hygiene for the AI era, and a large number of vendors began recommending it as a lever for AI-answer visibility.
Intuition, however, is not evidence, and a proposed convention is not an adopted standard. The relevant empirical question is narrow and answerable: when a site publishes llms.txt, does anything actually read it, and does any major AI platform treat it as an input to what it cites? For most of the file's short life that question went unmeasured, and the marketing filled the vacuum. In mid-2026 it stopped being unmeasured.
Ahrefs ran a server-log analysis across 137,210 domains and reported the result in a study titled, without much ceremony, We Analyzed 137K Sites: 97% of llms.txt Files Never Get Read. The value of the work is not rhetorical. It replaces an assumption about how AI systems behave with a direct observation of what they did, at a scale large enough that the finding is not an artifact of one unusual site or one unusual week.
Does llms.txt work? What 137,000 sites reveal
The headline number is the one that ended the debate on its own terms. Across the domains that carried a valid llms.txt file, 97 percent of those files received zero requests during the measurement month of May 2026. Not a low read rate. Zero. For the overwhelming majority of sites that had done exactly what the vendors recommended, nothing came to read the file at all.
The residual traffic was mostly not AI
A defender of the tactic might reach for the small minority of files that were fetched and argue that is where the value lives. The composition of that residual traffic closes off the argument. Of the fetches that did occur, roughly 96 percent came from ordinary bots, and only about 19.5 percent of that already-small remainder came from named AI tools. In other words, the sliver of llms.txt reads that happened at all were dominated by generic crawlers, and the fraction attributable to the AI assistants the file is meant to influence was a fraction of a fraction.
This matters because the entire theory of llms.txt rests on the AI answer engines reading it. A file that is retrieved almost exclusively by non-AI bots is not feeding the systems it was designed to feed, no matter how carefully it is written.
Adoption is thin, and independently so
The read-rate finding is reinforced by a separate adoption finding from a second source. SE Ranking examined roughly 300,000 domains and found llms.txt present on only around 10 percent of them. Two independent teams, different methods, different samples: one measuring whether the file is read, the other whether it is even published, and both landing on a picture of a convention that is neither widely adopted nor meaningfully consumed. Independent replication is what separates a durable finding from a single provocative chart, and here the replication points the same way.
Is llms.txt a ranking signal? Google's own answer
Behavioral data tells you what happened; it cannot by itself tell you what a platform intends. For that, the most authoritative source is the platform speaking about its own systems, and on llms.txt the platform has spoken. Google's John Mueller has stated publicly that llms.txt is not done for search, characterizing it, at most, as a token-saving convenience for AI coding tools that read developer documentation, not as a signal that shapes how a business surfaces in AI features.
This is consistent with Google's own generative-AI optimization guidance, which states explicitly that no AI-specific text file or special markup is required to appear in AI features. The instruction from the platform is not "publish this file to be included." It is closer to the opposite: there is no dedicated file to publish for this purpose.
The same pattern recurs across this field. Google has been equally explicit that E-E-A-T, the Experience, Expertise, Authoritativeness, and Trust framework, is a set of criteria for the human raters who evaluate search quality, not a machine-computed score a page can optimize into a ranking factor. In both cases the industry took a genuine concept and reframed it as a controllable lever, and in both cases the primary source says the lever is not what it is being sold as. Reading the platform's own documentation, rather than the secondary content built on top of it, is the single most reliable defense against this class of claim.
Why a text file was never going to be the lever
The empirical result is clear, but the mechanism behind it explains why the result was predictable. Generative answer engines are retrieval systems. When a query arrives, the engine assembles candidate passages from an index, ranks them, and generates an answer grounded in what it retrieved. The unit of work is the passage and the entity, drawn from the live index of the open web, not a hand-written manifest a site owner left at its root and hoped would be consulted.
A useful contrast is schema.org. Where llms.txt is a proposal, structured data via schema.org is a vocabulary founded jointly by Google, Bing, Yahoo, and Yandex in 2011, governed by a standards body, and actually consumed by the systems that were built to consume it. The difference is not popularity; it is provenance and adoption. A convention becomes an input to a machine when the machine's owners commit to reading it, and there is no comparable commitment behind llms.txt.
There is also a structural reason a static summary file is a weak instrument for AI visibility even in principle. Research on how AI Overviews assemble answers indicates that the passages a system reads and the pages it credits are not guaranteed to be the same set, and a page well outside the top organic results can still be cited. That reading of the underlying mechanics is emerging rather than settled, and rests on technical analysis of platform disclosures rather than a single peer-reviewed source, so it should be held loosely. But the direction is instructive: in a system where citation is decoupled from a site's own declarations, a self-authored file describing what a site would like to be cited for is exactly the kind of input such a system is least likely to weight.
What actually works instead: the generative engine optimization evidence
The most useful thing the llms.txt data does is clear space to ask what the evidence supports in its place. Here the field is not empty. The founding, peer-reviewed study of the category, published at ACM SIGKDD in 2024, ran a controlled benchmark of roughly 10,000 queries and measured which content interventions changed whether a source was surfaced inside a generated answer. The strongest levers were not files or markup. They were adding citations to credible sources, including direct quotations, and replacing vague claims with specific statistics, with citing authoritative sources the single most consistent driver, producing a meaningful relative lift on the study's visibility metric.
A large-scale 2025 follow-up points in the same direction from a different angle. It found that AI search engines are systematically biased toward earned and third-party media over brand-owned and social content, a sharp contrast with classic Google, which sources more evenly. That result is emerging rather than established, one methodologically large study not yet widely replicated, and it should be cited with that caveat. But taken with the founding benchmark, the shape of the advice is stable: what moves AI visibility is being a credible, well-cited, corroborated entity that other trustworthy sources reference, not a manifest a business writes about itself.
The gap between those two pictures is the whole point. The tactic with the strongest marketing behind it, llms.txt, is the one the read-rate data undercuts. The tactics with the strongest evidence behind them, credible citations and earned third-party corroboration, are slower, harder, and less packageable as a checkbox, which is precisely why they are undersold relative to their support.
The infrastructure honesty gap
llms.txt is not an isolated case. It is one instance of a broader pattern in which the tooling layer promoted as "how you control AI visibility" turns out to be younger, leakier, and less standardized than the marketing implies. The same year the llms.txt data landed, Cloudflare documented a case in which a retrieval crawler continued fetching pages after being blocked, using an undeclared browser-spoofing agent and rotating network identity to evade the block, and verified that answers still reflected content from domains that had explicitly disallowed it. That specific incident is documented and independently verifiable; its characterization as ongoing practice is contested by the vendor named in it, and it is a single company's report rather than a third-party audit, so it belongs in the "established as an incident, contested as a general claim" category.
Stack these findings together and a defensible, evidence-first position emerges. The public no-crawl directive is not always honored. The self-authored summary file is not read. The entity-identity mechanisms underneath "be the same business everywhere" are looser in practice than their specifications suggest. None of this means AI visibility is unmanageable. It means the levers being sold as clean, controllable switches are, on the evidence, neither as clean nor as controllable as advertised, and that any responsible read of the space has to tier its tactics by how much primary evidence actually stands behind each one.
How to read any AI-visibility claim honestly
The practical residue of this autopsy is a small set of questions that separate a substantiated tactic from a sales story, and they generalize well beyond llms.txt. First, is the mechanism a governed standard the platforms have committed to reading, like schema.org, or a proposal one party authored and others are hoping gets adopted? Second, is there primary data on whether the thing is actually consumed at scale, or only an intuitive argument that it should be? Third, does the platform itself describe the input as a signal it uses, or has it said the opposite? And fourth, is the supporting evidence established, emerging, or contested, and is whoever is selling the tactic honest about which?
Applied to llms.txt, all four questions resolve against it: a proposal, not a standard; measured and found unread; disavowed by the platform as not done for search; and defended, when it is defended, by intuition rather than data. Applied to earned citations and credible corroboration, they resolve the other way. That is not a reason for cynicism about the field. It is a method for spending effort where the evidence is, and it is the same method that should govern how a business decides which of its own visibility investments are worth making.
The evidence
Key findings, with their sources
-
97% of valid llms.txt files received zero requests in a single month, across a sample of 137,210 domains.
established Ahrefs, "We Analyzed 137K Sites: 97% of llms.txt Files Never Get Read", June 2026 (server-log analysis).
-
Of the small share of llms.txt files that were fetched at all, about 96% of fetches came from ordinary bots, and only ~19.5% of that residual traffic came from named AI tools.
established Ahrefs, "We Analyzed 137K Sites: 97% of llms.txt Files Never Get Read", June 2026.
-
llms.txt was present on only about 10% of roughly 300,000 domains examined, an independent replication of the low-adoption picture.
established SE Ranking analysis of ~300,000 domains, 2026 (reported via Search Engine Journal).
-
Google states llms.txt is "not done for search", at most a token-saving convenience for AI coding tools, and that no AI-specific text file or special markup is required to appear in AI features.
established John Mueller (Google) public statement, 2026; Google Search Central generative-AI optimization guidance.
-
In a controlled benchmark, adding citations to credible sources, direct quotations, and specific statistics were the strongest content levers for being surfaced in a generated answer; citing authoritative sources was the single most consistent driver.
established Aggarwal et al., "GEO: Generative Engine Optimization", arXiv:2311.09735, ACM SIGKDD 2024 (peer-reviewed).
-
AI search engines are systematically biased toward earned and third-party media over brand-owned content, unlike classic Google which sources more evenly.
emerging Chen, M. et al., "Generative Engine Optimization: How to Dominate AI Search", arXiv:2509.08919, 2025 (large-scale single study).
-
A retrieval crawler was documented continuing to fetch pages after being blocked, using an undeclared browser-spoofing agent and rotating network identity, with answers still reflecting content from domains that had disallowed it.
contested Cloudflare Blog, "Perplexity is using stealth, undeclared crawlers to evade website no-crawl directives", August 4, 2025 (vendor report; characterization disputed by Perplexity).
Calibration
What is proven, what is promising, what is unproven
| Evidence tier | Tactics | What the evidence says |
|---|---|---|
| established | Earning citations from credible third-party sources; adding statistics, direct quotations, and citations to your own content; structured data via schema.org (a standards-body vocabulary the platforms committed to reading). | GEO-bench controlled study (Aggarwal et al., SIGKDD 2024); schema.org standards history since 2011. |
| emerging | Prioritizing earned and third-party media over brand-owned pages for AI-answer surfaces; treating classic-search rank and AI citation as decoupled targets. | Chen et al. 2025 earned-media-bias study (large-scale, not yet widely replicated); AI Overviews retrieve-then-cite mechanics (synthesis of platform disclosures). |
| contested | Publishing llms.txt to influence AI answers; relying on a robots.txt no-crawl directive as a hard control over AI crawlers. | Ahrefs 137K read-rate study and Mueller statement (llms.txt); Cloudflare stealth-crawler incident, disputed by the vendor named (crawler compliance). |
Reference
Glossary
- llms.txt
- A proposed root-level text file meant to give AI systems a clean, curated summary of a website. As of 2026, large-scale data shows it is essentially never read and is not treated as a ranking or citation signal by major AI platforms.
- Ranking or citation signal
- An input a search or AI system actually uses to decide what to rank or which sources to name in an answer. A file a site publishes is only a signal if the platform commits to reading and weighting it.
- Retrieval crawler
- An AI bot that fetches live pages at query time to build an in-session answer, distinct from a training crawler that bulk-ingests for model pretraining. The two behave differently, which is why one crawl rule cannot govern both.
- Generative Engine Optimization (GEO)
- The academically-coined discipline of getting content retrieved and cited inside an LLM-generated answer. The founding benchmark found credible citations, quotations, and specific statistics to be the strongest content levers.
- Structured data (schema.org)
- A machine-readable markup vocabulary founded jointly by Google, Bing, Yahoo, and Yandex in 2011 and governed by a standards body. Unlike llms.txt, it is an adopted standard the search platforms actively consume.
Straight answers
Frequently asked questions
Does llms.txt work?
On the current evidence, no, at least not for its stated purpose of influencing AI answers. An Ahrefs analysis of 137,210 domains found 97 percent of valid llms.txt files received zero requests in a single month, and the small remainder that were fetched came overwhelmingly from ordinary bots rather than named AI tools. There is no primary data showing it changes whether a business is cited.
Is llms.txt a ranking signal?
No major AI platform treats it as one, and Google has said so. John Mueller stated publicly that llms.txt is not done for search, and Google's generative-AI guidance states no AI-specific text file or special markup is required to appear in AI features. Treating it as a ranking or citation lever is not supported by the platforms or the data.
Does llms.txt help SEO?
There is no evidence that it does. llms.txt is not consumed by classic search ranking, and the read-rate data shows AI systems largely are not reading it either. If a business has already published one, it is unlikely to cause harm, but it should not be counted as an SEO or AI-visibility investment with any measurable return behind it.
Is llms.txt worth adding to my site?
It is low-cost to add and low-risk to keep, but on the evidence it is not worth prioritizing, and it is certainly not worth paying for as a headline AI-visibility service. The tactics with real supporting research, earning credible third-party citations and structuring content to be extractable, deserve the effort that oversold files do not.
If llms.txt does not move AI visibility, what does?
The peer-reviewed evidence points to being a credible, well-corroborated entity: content that cites authoritative sources, includes specific statistics and direct quotations, and is referenced by trustworthy third-party media. A 2025 study found AI engines lean on earned and third-party sources more heavily than classic Google does. Those levers are slower and harder to package, which is exactly why they are undersold.
Provenance
Sources
- Ahrefs (2026). "We Analyzed 137K Sites: 97% of llms.txt Files Never Get Read." June 2026 (established)
- SE Ranking (2026). llms.txt adoption analysis across ~300,000 domains, reported via Search Engine Journal, June 2026 (established)
- Mueller, J. (Google) (2026). Public statement that llms.txt is "not done for search"; Google Search Central generative-AI optimization guidance (established)
- Aggarwal, P., Murahari, V., Rajpurohit, T., Kalyan, A., Narasimhan, K., Deshpande, A. (2023/2024). "GEO: Generative Engine Optimization." arXiv:2311.09735, ACM SIGKDD 2024 (peer-reviewed, established)arxiv.org
- Chen, M., Wang, X., Chen, K., Koudas, N. (2025). "Generative Engine Optimization: How to Dominate AI Search." arXiv:2509.08919 (emerging, single large-scale study)arxiv.org
- Google Search Central Blog (2022). "Our latest update to the quality rater guidelines: E-A-T gets an extra E for Experience"; Google Search Quality Rater Guidelines (established)
- Schema.org (2011-present). Structured-data vocabulary jointly maintained by Google, Bing, Yahoo, and Yandex (established)
- Cloudflare Blog (2025). "Perplexity is using stealth, undeclared crawlers to evade website no-crawl directives." August 4, 2025 (established as incident, contested as ongoing characterization)
Every figure above is attributed to a real, dated source and tagged with its evidence tier. Where a claim could not be verified to a primary source, it is not stated as fact.