Discovery Science · emerging evidence
Two Separate Races: How Google AI Overviews Decouples Ranking from Citation
Understanding how Google AI Overviews works matters because it does not behave like the ranked list it sits above. The available technical accounts describe two separate operations. First the engine retrieves a query-specific set of candidate passages, ranks them against each other, and generates a synthesized answer from that set. Then, as a distinct step, it runs a citation pass that decides which sources to name and attach. Those two operations do not have to converge on the same pages. A page can help write the answer and never be credited, and a page well outside the top organic results can still be cited. Read carefully, this means classic ranking and AI-answer citation are two different races with two different finish lines. The direct evidence for the exact mechanics is still emerging and rests on secondary analysis of Google disclosures, so this piece tiers every claim and reports what is measured, not asserted.
A ranked list held one contest. An answer holds two.
The classic results page ran a single competition: pages were scored against a query and placed in order, and the searcher chose from the order. There was one race and one visible finish line, the position in the list. An AI Overview replaces that with something structurally different. The engine reads a set of sources, forms one synthesized answer, and attaches a short list of citations beside it. The reader sees a verdict and a few names, not a ranked menu.
The important detail is buried in that sequence. Writing the answer and choosing the citations are not the same act. The set of passages that inform the generated text and the set of pages the engine decides to credit are produced by different steps, and there is no rule that forces them to match. That is why "we rank number one" and "we are cited in the answer" have become two separate claims, each needing its own measurement.
How Google AI Overviews works: retrieval first, a separate citation pass second
The available technical accounts of how Google AI Overviews works describe a retrieve-then-cite pipeline rather than a single ranked lookup. Reporting drawing on Google technical disclosures and patent filings describes the engine assembling a query-specific custom corpus of candidate passages through embedding-based retrieval and pairwise model ranking, generating the overview text from that corpus, and then running a distinct citation-extraction step that scores and attaches sources after the text exists.
Two consequences follow directly from that ordering. The passages used to compose the answer and the pages credited in it need not be identical, because citation is decided by a separate scorer operating after generation. And a page ranking far outside the top organic results can still be selected for citation, because the citation pass is not reading from the organic rank order. In other words the answer and its citations are outputs of different subsystems that happen to appear together in one box.
Why this is tiered as emerging, not settled
This mechanism is the spine of the article, so its evidence status has to be stated plainly. The two-stage description is a synthesis of Google public technical documentation, patent filings, and secondary industry technical analysis. It is directionally consistent across multiple independent write-ups, but it is not drawn from a single peer-reviewed primary paper, and Google does not publish the live production pipeline in full. We therefore treat it as emerging: strong enough to reason from and to design measurement around, not strong enough to quote exact internal latencies or weightings as fact. Where a specific number would require primary confirmation Google has not provided, this piece declines to state one.
The architecture underneath: retrieval augmented generation
The retrieve-then-generate shape is not unique to Google. It is the standard architecture of nearly every generative answer engine, and it has a documented origin. The founding retrieval augmented generation paper paired a pretrained language model with a non-parametric document index queried at inference time, and showed the combination produced more specific and factual output than a parametric model alone while being updatable without retraining.
The research literature has since matured into a shared taxonomy, commonly split into naive, advanced, and modular patterns, that is now the common vocabulary for describing how AI Overviews, ChatGPT Search, Perplexity, and Gemini construct answers. The value of that vocabulary here is that it separates the steps most buyer-facing content collapses. Retrieval, generation, and citation are distinct stages, and the ranking-versus-citation split is exactly what you get when you take those stages seriously instead of treating "AI search" as one undifferentiated event.
Ranking vs citations: why the cited set and the ranked set diverge
Put the two subsystems side by side and the divergence is easy to see. Classic ranking answers the question "which page best matches this query, in order." The citation pass answers a different question: "given an answer that already exists, which sources should be named as support." A page can be a strong contributor to the retrieved corpus that shapes the wording and still not be the page the citation scorer prefers to display. And a niche, highly specific page can be an excellent citation for one sentence of the answer while ranking nowhere near the top of the organic results.
This is why a business can win a citation it did not expect and lose one it thought it had earned by ranking. Neither outcome is a glitch. They are the predictable result of two scorers optimizing for two different objectives on two different inputs. Treating them as one metric, a single blended rank number, hides exactly the gap that decides whether you are named in the answer.
The citation you cannot fully trust
There is a further, sobering layer. Even when a source is cited, the citation is not a guarantee that the model reasoned from it. Research on citation faithfulness in retrieval augmented generation found that in the dominant production patterns, citations are attached to an answer somewhat independently of the evidence that actually produced the text, so a displayed source is not reliably the source the model used. A related attribution survey reported that more than 95 percent of answers from the tested open-source models contained at least one sentence with no supporting attribution at all.
For anyone optimizing, this reframes the goal. Winning a citation is necessary but not sufficient, and a citation is not proof of influence. It also means vendor claims to guarantee a citation should be read skeptically, because the faithfulness gap is an unresolved technical problem in the research, not a solved one. The real target is presence and consistency across many real answers over time, measured, rather than a single citation treated as a trophy.
How AI search picks sources, and the earned-media tilt
If ranking does not determine citation, what does the citation pass appear to favor. The strongest controlled evidence in this area is the founding generative engine optimization study, which tested content-level interventions across roughly ten thousand queries and nine datasets and found that adding citations to credible sources, including direct quotations, and replacing vague claims with specific statistics produced a meaningful relative lift in a source visibility metric, with citing authoritative sources the single most consistent lever.
A later large-scale study pushed the point further, reporting that AI search engines skew toward earned and third-party media over brand-owned and social content more sharply than classic Google does. That finding is emerging rather than settled, a single large study not yet widely replicated, so it is offered as a direction, not a law. Both results point the same way: the citation pass rewards corroboration the business does not fully control, which is precisely why AI-answer visibility behaves unlike on-page ranking.
What actually happens when the summary appears
The two-race structure would matter less if the answer box were a minor feature. The click data says it is not. Pew Research Center behavioral tracking of nearly 69,000 real Google searches found that when an AI Overview was present, users clicked a traditional organic result in about 8 percent of searches, versus about 15 percent without a summary, and clicked a link inside the summary itself in only about 1 percent of cases.
That is the operational reason organic traffic is the wrong yardstick for this channel. If most sessions with a summary never click through, then the value of being in the answer is being named in it, not the trickle of clicks it forwards. The buyer-facing outcome is being the entity the answer credits and recommends, which the citation pass governs, and which classic rank does not measure.
The evidence, tiered: established, emerging, and contested
This topic sits on a mix of evidence, and pretending otherwise would undercut the whole point. The RAG architecture and the click-suppression behavior are established, drawn from a foundational paper and an independent behavioral study. The specific two-stage AI Overviews mechanism is emerging, synthesized from Google disclosures and patents and consistent across independent analyses but not confirmed by a single primary source. The earned-media tilt is emerging from one large study. The precise internal weightings and timings of the citation pass are, as of this writing, not disclosed and should not be quoted as fact.
The correct response to that mix is not to overclaim and not to dismiss. It is to design around what the strongest evidence supports, that ranking and citation are separable, and to measure the citation surface directly rather than infer it from rank. Where a number would require access Google has not granted, we measure our own reads and report them, and we say plainly which claims are still awaiting primary confirmation.
What this means for measurement
If ranking and citation are two races, one score cannot referee both. A page can hold a strong classic position and be absent from the answer written above it, and a page can be named in the answer while ranking modestly. Collapsing the two into a single blended metric averages away the exact gap a business needs to see. The implication is architectural: classic search visibility and AI-answer visibility deserve to be measured as separate pillars, each with its own method, before any work is scoped.
This is the evidentiary basis for treating the AI-answer surface as its own discipline rather than a footnote to SEO. It is measured by sampling real buyer questions across the engines and recording how often the business is named and cited, explicitly not by counting organic clicks, because the click data shows organic traffic no longer describes this channel. The first step is a reading of where you actually stand on that surface today.
The evidence
Key findings, with their sources
-
Available technical accounts describe AI Overviews assembling a query-specific custom corpus via embedding retrieval and model ranking, generating the answer, then running a separate citation pass, so the pages that write the answer and the pages credited are not guaranteed to be the same set.
emerging Synthesis of Google public technical disclosures and patent filings via secondary industry technical analysis, 2024-2026 (needs primary-source verification).
-
When an AI Overview was present, users clicked a traditional organic result in about 8% of searches versus about 15% without one, and clicked a link inside the summary itself in only about 1% of cases.
established Pew Research Center, "Do people click on links in Google AI summaries?", July 22, 2025 (900 U.S. adults, 68,879 searches, 12,593 with an AI summary).
-
Adding citations to credible sources, direct quotations, and specific statistics produced a meaningful relative lift in a source-visibility metric across roughly 10,000 queries and nine datasets, with citing authoritative sources the single strongest lever.
established Aggarwal, P. et al., "GEO: Generative Engine Optimization", arXiv:2311.09735, ACM SIGKDD 2024.
-
A cited source is not reliably the source the model reasoned from, and more than 95% of answers from tested open-source LLMs contained at least one sentence with no supporting attribution.
established arXiv:2409.11242, "Measuring and Enhancing Trustworthiness of LLMs in RAG", 2024; RAG attribution survey, arXiv 2601.19927, 2025-26.
-
Retrieval augmented generation pairs a pretrained language model with a document index queried at inference time and produces more specific and factual output than a parametric model alone.
established Lewis, P. et al., "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks", arXiv:2005.11401, NeurIPS 2020.
-
AI search engines skew toward earned and third-party media over brand-owned and social content more sharply than classic Google does.
emerging Chen, M. et al., "Generative Engine Optimization: How to Dominate AI Search", arXiv:2509.08919, 2025 (single large-scale study, not yet widely replicated).
Calibration
What is proven, what is promising, what is unproven
| Evidence tier | Tactics | What the evidence says |
|---|---|---|
| established | Treat classic ranking and AI-answer citation as separable; measure the citation surface directly; use RAG-stage vocabulary (retrieval, generation, citation) to reason about what moves each. | Lewis et al. 2020 (RAG); Pew Research 2025 (click behavior); Aggarwal et al. 2024 (GEO-bench levers). |
| emerging | Design around the two-stage retrieve-then-cite model of AI Overviews and the earned-media tilt; favor corroboration and authoritative citation over on-page tweaks alone. | Google disclosures/patents via secondary analysis (AI Overviews mechanics); Chen et al. 2025 (earned-media bias). |
| contested | Do not quote exact internal weightings, latencies, or a guaranteed-citation promise; the precise citation-pass scoring and RAG faithfulness fixes are unresolved. | Undisclosed production pipeline internals; unresolved citation-faithfulness literature (arXiv:2409.11242). |
Reference
Glossary
- Retrieval augmented generation (RAG)
- The architecture behind most answer engines: a language model paired with a document index that is queried at answer time, so the model composes text from freshly retrieved passages rather than from memory alone.
- Custom corpus
- The query-specific set of candidate passages an engine assembles for a single query, from which the answer is generated. It is built per query, not read from a fixed ranked list.
- Citation pass
- A distinct step that runs after the answer text is generated and decides which sources to name and attach. Because it is separate from generation, its chosen sources need not match the passages that shaped the text.
- Reranking
- Scoring retrieved candidates against each other to order them before generation. It is one of several stages, and it is not the same as the citation decision.
- Citation faithfulness
- Whether a displayed citation is actually the evidence the model used. The research shows this is often only loosely guaranteed, so a citation is not proof the source influenced the answer.
- How often a business is named or cited across a panel of real buyer questions on a given engine. It is the AI-answer analogue of visibility, measured by presence in answers rather than by organic clicks.
Straight answers
Frequently asked questions
Does ranking number one in Google mean I will be cited in the AI Overview?
No, and that is the core point. The available accounts of how Google AI Overviews works describe a citation step that runs separately from organic ranking, on a query-specific corpus rather than the rank order. A top-ranked page can be absent from the answer, and a lower-ranked page can be cited. They are two different contests.
How does Google AI Overviews decide which sources to cite?
The reported mechanism is that the engine retrieves and ranks candidate passages, generates the answer, then runs a distinct citation pass that scores and attaches sources afterward. The exact internal scoring is not publicly disclosed, so it should be reasoned about as an emerging model and measured directly, not quoted as a fixed formula.
My page is cited in an answer but ranks low. Is that a mistake?
Not necessarily. Because citation is decided by a separate step that is not reading from the organic rank order, a specific, highly relevant page can be an excellent citation for one part of an answer while ranking modestly. The two outcomes come from two different subsystems optimizing for two different goals.
Is being cited the same as getting traffic?
No. Pew Research found that when an AI summary is present, only about 1 percent of searches result in a click on a link inside the summary. The value of a citation is being named and credited in the answer the buyer reads, not the clicks it forwards, which is why organic traffic is the wrong success metric for this channel.
Why measure AI answers separately from classic search?
Because ranking and citation are separable, a single blended score averages away the exact gap that decides whether you are in the answer. Measuring the AI-answer surface on its own, by sampling real buyer questions and recording how often you are named, is the only way to see that gap clearly.
Provenance
Sources
- Lewis, P. et al., "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks", arXiv:2005.11401, NeurIPS 2020 (established)arxiv.org
- Gao, Y. et al., "Retrieval-Augmented Generation for Large Language Models: A Survey", arXiv:2312.10997, 2023-24 (established)arxiv.org
- Aggarwal, P. et al., "GEO: Generative Engine Optimization", arXiv:2311.09735, ACM SIGKDD 2024 (established)arxiv.org
- Chen, M., Wang, X., Chen, K., Koudas, N., "Generative Engine Optimization: How to Dominate AI Search", arXiv:2509.08919, 2025 (emerging)arxiv.org
- Various authors, "Measuring and Enhancing Trustworthiness of LLMs in RAG through Grounded Attributions and Learning to Refuse", arXiv:2409.11242, 2024 (established finding of the faithfulness gap)arxiv.org
- "Attribution Techniques for Mitigating Hallucinated Information in RAG Systems: A Survey", arXiv:2601.19927, 2025-26 (emerging on mitigations)arxiv.org
- Pew Research Center, "Do people click on links in Google AI summaries?", July 22, 2025 (established)pewresearch.org
- Google AI Overviews retrieve-then-cite mechanics: synthesis of Google public technical disclosures and patent filings via secondary industry technical analysis, 2024-2026 (emerging, needs primary-source verification)
Every figure above is attributed to a real, dated source and tagged with its evidence tier. Where a claim could not be verified to a primary source, it is not stated as fact.