Discovery Science · established evidence

Is Your Business Getting Recommended, or Just Indexed? A Framework for Telling the Difference

Last reviewed 2026-07-20. Written by Chandranshu Kumar, Founder, Raveneye Global. · 10 min read

The difference between getting recommended and just getting indexed is the difference between three separate facts most owners treat as one. Being indexed means a search or AI system holds a copy of your page in the corpus it can draw from. Being surfaced means that page is actually retrieved and cited when a real buyer asks a real question. Being recommended means the answer names your business as the one to choose. These three states are decoupling: a page can sit in the index and never surface, and a page can be cited yet never send a click or shape the buyer. The framework below separates the three, shows the primary evidence for each split, and gives a lay owner a clear way to tell, for their own business, which state they are actually in.

Three questions hiding inside one

Owners usually ask a single question, "am I showing up in AI search," and expect a single answer. The evidence says that question contains three distinct ones, and the answer to each can differ. First, is your page in the index or corpus at all. Second, when a buyer asks a question you should own, does the engine retrieve and cite your page in the live answer. Third, of the businesses the answer names, are you the one it recommends. A framework that collapses these into one number will mislead, because the mechanisms that govern each are different.

This is not a semantic distinction. It maps directly onto how modern answer engines are built. Generative answers are produced by retrieval-augmented generation, an architecture that pairs a language model with a separate index queried at the moment of the question, so that the answer is assembled from passages pulled at inference time rather than recited from memory. Presence in that index is a precondition, not an outcome. The steps that run after retrieval, ranking, generation, and citation, decide whether presence becomes a recommendation.

Being indexed was never the finish line

The idea that the index is a substrate rather than a destination predates generative AI by more than a decade. When Google introduced its Knowledge Graph in 2012, it framed the shift as indexing "things, not strings," building a store of entities and facts that search could reason over rather than a flat list of matched documents. The visible ranked list has always been a thin interface over a deeper representation of who and what exists.

Retrieval-augmented generation extended that logic. The founding work showed that pairing a parametric model with a non-parametric index produced more specific and factual output and could update its knowledge without retraining, precisely because the index is consulted fresh each time. What follows is a plain consequence: your page being in the corpus tells you almost nothing about whether it will be chosen from the corpus. Indexation is table stakes. The interesting question begins after it.

Ranking and citation have come apart

The most consequential split for an owner is between classic rank and AI citation. Industry analysis of Google's own technical disclosures describes AI Overviews assembling a query-specific custom corpus of candidate passages, ranking them, and then running a separate citation pass after the answer text has been written. If that description holds, two things follow that break the old mental model. The passages used to compose the answer and the pages credited beneath it are not guaranteed to be the same set, and a page ranking well outside the top organic results can still be cited while a page ranking first is left out.

We tier this claim as emerging rather than established, because it rests on secondary technical analysis of Google's patents and documentation rather than a single peer-reviewed primary source, and it should be read as directionally supported, not settled. But the direction matters. It means "we rank number one" is not evidence that you are cited, and "we are not cited" is not evidence that you rank poorly. Ranking and citation are now two different races, which is the structural reason a serious diagnostic measures classic search and AI answers as separate pillars rather than one blended score.

A citation is not a click, and often not a faithful read

Suppose you clear both bars: you are indexed, and you are cited. Two independent bodies of evidence show that a citation is still weaker than owners assume it to be.

The citation rarely forwards a visitor

Pew Research Center tracked the real browsing of 900 US adults across 68,879 Google searches in March 2025, of which 12,593 contained an AI summary. In searches where a summary appeared, users clicked a traditional search result about 8 percent of the time, against 15 percent when no summary was present, and clicked a link inside the summary itself only about 1 percent of the time. They also ended their browsing session more often after a page with a summary, 26 percent, than without one, 16 percent. Being the cited source, in other words, seldom becomes a visit. If your reporting counts clicks, the surface that increasingly decides who gets recommended is nearly invisible to it.

The citation may not be the source the model used

There is a second, more technical gap. In the dominant production pattern for RAG systems, citations are attached to an answer independently of the evidence that actually generated it, which means a cited source is not reliably the source the model reasoned from. A 2025 attribution survey reports that over 95 percent of answers from the open-source language models it tested contained at least one sentence with no supporting attribution at all. The practical lesson for a diagnostic is calibration, not cynicism: winning a citation is necessary but not sufficient, and "we got cited" should be treated as a signal to verify, not a result to celebrate.

How to tell which state you are in

No engine publishes whether it recommends you, so direct observation is the only way to find out. The procedure is simple to state and disciplined to run. Assemble a panel of the real questions your buyers ask, phrased the way they phrase them, not the keywords you wish they used. Run that panel across each engine that matters to your market, for example ChatGPT, Google AI Overviews, Perplexity, Gemini, and Copilot. For each question, record which of the four states you are in: absent, indexed but not surfaced, surfaced but not recommended, or recommended. Note who is named instead of you when you are not.

Two cautions matter here. Generative answers are not deterministic, so the same question can return different sources on different runs, which means a single check is anecdote and a repeated sample is evidence. And the answer surface barely forwards clicks, as the Pew data shows, so you cannot infer your AI-answer state from your analytics traffic. It has to be observed on the answer surface itself. The output of this procedure is not a rank. It is an appearance rate with a date and an engine attached, reported as a range rather than a single confident number, precisely because the surface is volatile and undocumented.

What does not move it: two common misconceptions

A diagnostic is only as good as the false leads it rules out. Two of the most common in this space fail on the primary evidence, and a framework that respects the reader has to say so.

The first is the llms.txt file, promoted as a way to hand AI systems a clean summary and thereby control visibility. An analysis of 137,210 domains found that 97 percent of valid llms.txt files received zero requests in a single month of 2026, and Google has stated publicly that the file is not used for search. Placing a file at your root is not a path to being recommended. The second is the belief that E-E-A-T, experience, expertise, authoritativeness, and trust, is a score you optimize a page into. Google's own documentation is explicit that E-E-A-T is a human-rater evaluation framework used to assess the algorithm's output quality, not a machine-computed ranking factor a page can directly target. Much commercial AI-visibility advice conflates a rater heuristic with an optimization dial. The diagnostic value of naming these is simple: it keeps budget aimed at the levers with evidence, and away from the ones with only marketing.

The evidence

Key findings, with their sources

  • In searches where a Google AI Overview appeared, users clicked a traditional search result about 8% of the time, versus 15% with no summary present, and clicked a link inside the summary only about 1% of the time.

    established Pew Research Center, "Do people click on links in Google AI summaries?", July 22, 2025 (browsing panel: 900 US adults, 68,879 searches, 12,593 with an AI summary).

  • Users ended their browsing session more often after visiting a page with an AI summary (26%) than after one without (16%).

    established Pew Research Center, "Do people click on links in Google AI summaries?", July 22, 2025.

  • AI Overviews is described as assembling a query-specific custom corpus and running a separate citation pass after the answer is written, so a page cited need not be one the model drew from, and a page outside the top organic results can still be cited.

    emerging Synthesis of Google technical disclosures and patent filings via secondary industry analysis, 2025 (needs primary-source verification).

  • Adding citations to credible sources, direct quotations, and specific statistics produced a 30 to 40 percent relative lift on a position-adjusted visibility metric in a controlled benchmark of roughly 10,000 queries; citing authoritative sources was the single strongest lever.

    established Aggarwal, P. et al., "GEO: Generative Engine Optimization", arXiv:2311.09735, ACM SIGKDD 2024 (peer-reviewed).

  • In dominant production RAG pipelines, citations are attached independently of the evidence that generated the answer; an attribution survey reports over 95% of answers from tested open-source language models contained at least one unattributed sentence.

    established arXiv:2409.11242 (2024); "Attribution Techniques for Mitigating Hallucinated Information in RAG Systems: A Survey", arXiv:2601.19927 (2025-26).

  • Across 137,210 domains, 97% of valid llms.txt files received zero requests in a single month of 2026, and Google has stated the file is not used for search.

    established Ahrefs, "We Analyzed 137K Sites: 97% of llms.txt Files Never Get Read", June 2026.

Calibration

What is proven, what is promising, what is unproven

Evidence tierTacticsWhat the evidence says
EstablishedSplits and levers the peer-reviewed or independent evidence supportsPew's click-suppression measurements (2025); the GEO-bench content levers of citations, quotations, and statistics (Aggarwal et al., SIGKDD 2024); the citation-faithfulness gap and unattributed-sentence rate; Ahrefs's llms.txt non-adoption (2026); Google's own statement that E-E-A-T is a rater framework, not a ranking factor.
EmergingDirectionally supported but from single or secondary sourcesThe two-stage retrieve-then-cite architecture of Google AI Overviews (synthesis of Google disclosures, needs primary verification); the finding that AI engines favor earned and third-party media over owned content (Chen et al., 2025, a single large-scale study not yet replicated).
ContestedIndustry narrative, not primary-sourcedThe exact retrieval or ranking weightings individual engines apply, and the precise AEO-versus-GEO historical lineage, are practitioner accounts rather than disclosed or peer-reviewed facts, and are excluded from the claims above.

Reference

Glossary

Indexed
Present in the corpus a search or AI system can draw from. A precondition for being cited, and nothing more.
Surfaced
Retrieved and cited in a live answer to a specific question, as distinct from merely sitting in the index.
Named by the answer as the business to choose, the only state that reliably moves a buyer.
Retrieval-augmented generation (RAG)
The architecture behind generative answers: a language model paired with a separate index queried at the moment of the question, so answers are assembled from passages pulled at inference time.
Custom corpus
The query-specific set of candidate passages an engine assembles for one question before ranking, generating, and citing.
Appearance rate
How often a business is named or cited across repeated runs of a fixed question panel, reported with an engine, a locale, and a date because the surface is non-deterministic.

Straight answers

Frequently asked questions

What is the difference between being indexed and being recommended?

Being indexed means a search or AI system holds your page in the corpus it can draw from. Being recommended means the answer to a real buyer question names your business as the one to choose. Between them sits a third state, being surfaced, where you are cited but not chosen. They are governed by different mechanisms, so they have to be measured separately.

Can I rank number one in Google and still not be recommended by AI?

Yes. Industry analysis of Google's disclosures describes AI Overviews building a separate custom corpus and running a distinct citation pass, so classic rank and AI citation are decoupled. A page ranking first can be left out of the answer, and a page ranking well below the top can still be cited. We tier the specific architecture as emerging, but the practical decoupling is why ranking and citation are treated as two different races.

How do I check whether AI search recommends my business?

Observe it directly, because no engine publishes it. Assemble a panel of the real questions your buyers ask, run them across each engine that matters, and record for each whether you are absent, indexed but not surfaced, surfaced but not recommended, or recommended. Repeat the runs, because generative answers are non-deterministic, and report an appearance rate with a date and engine rather than a single confident number.

Does submitting an llms.txt file get my business recommended?

No. An analysis of 137,210 domains found 97 percent of valid llms.txt files received zero requests in a single month of 2026, and Google has said the file is not used for search. Placing a file at your root is not a path to being cited or recommended.

Is being cited the same as getting traffic?

No. The Pew Research Center browsing-panel study found users clicked a link inside an AI summary only about 1 percent of the time, and clicked a traditional result about 8 percent of the time with a summary present versus 15 percent without. A citation rarely forwards a visitor, which is why AI-answer visibility cannot be read from your analytics traffic and has to be measured on the answer surface itself.

Provenance

Sources

  1. Pew Research Center, "Do people click on links in Google AI summaries?", July 22, 2025 (established)pewresearch.org
  2. Lewis, P. et al., "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks", arXiv:2005.11401, NeurIPS 2020 (established)arxiv.org
  3. Aggarwal, P., Murahari, V., Rajpurohit, T., Kalyan, A., Narasimhan, K., Deshpande, A., "GEO: Generative Engine Optimization", arXiv:2311.09735, ACM SIGKDD 2024 (established)arxiv.org
  4. Chen, M., Wang, X., Chen, K., Koudas, N., "Generative Engine Optimization: How to Dominate AI Search", arXiv:2509.08919, 2025 (emerging, single study)arxiv.org
  5. "Measuring and Enhancing Trustworthiness of LLMs in RAG through Grounded Attributions and Learning to Refuse", arXiv:2409.11242, 2024; "Attribution Techniques for Mitigating Hallucinated Information in RAG Systems: A Survey", arXiv:2601.19927, 2025-26 (established gap)arxiv.org
  6. Singhal, A., "Introducing the Knowledge Graph: things, not strings", Official Google Blog, May 16, 2012 (established)blog.google
  7. Ahrefs, "We Analyzed 137K Sites: 97% of llms.txt Files Never Get Read", June 2026 (established)ahrefs.com
  8. Google Search Central, "Our latest update to the quality rater guidelines: E-A-T gets an extra E for Experience", December 2022; Search Quality Rater Guidelines (established)
  9. Google AI Overviews two-stage retrieval and citation architecture, synthesis of Google technical disclosures and patent filings via secondary industry analysis, 2025 (emerging, needs primary verification)

Every figure above is attributed to a real, dated source and tagged with its evidence tier. Where a claim could not be verified to a primary source, it is not stated as fact.

Which state is your business in?

The framework above only becomes useful when it is pointed at your business, on your buyer questions, across the engines your market actually uses. Most owners cannot see the answer, because their reporting counts clicks and the answer surface barely forwards any. A Machine-Readiness Score does the observation for you: a specialist-reviewed read of where you actually stand across classic search, the local map pack, AI answers, and reputation, so you know whether you are absent, indexed, surfaced, or recommended before any work is scoped. It is the first step, and it is free.

diagnostic Machine-Readiness Score A specialist-reviewed read of where you actually stand across classic search, the local map pack, AI answers, and reputation, scored to one honest number so your AI-answer visibility is measured rather than assumed. See how it works

Start free with a Machine-Readiness Score, a measured read of where you stand across search and AI answers. No guaranteed number, and no obligation.