Discovery Science · established evidence

What "GEO" Actually Means: Reading the Founding Paper So You Don't Have To

Last reviewed 2026-07-20. Written by Chandranshu Kumar, Founder, Raveneye Global. · 10 min read

Generative engine optimization, or GEO, is not a marketing slogan; it began as a specific, peer-reviewed experiment. In 2024 a team of researchers published "GEO: Generative Engine Optimization" at ACM SIGKDD, built a benchmark of roughly ten thousand queries across nine datasets called GEO-bench, and measured which changes to a web page actually made a generative engine more likely to feature that page inside a synthesized answer. Three content moves stood out: citing credible sources, adding direct quotations, and replacing vague claims with specific statistics. Citing an authoritative source was the single strongest lever. That is the real result, and it is genuinely useful. What the paper did not do is prove that any of this books you a customer, survives every future engine change, or means the model actually reasoned from the page it credited. Reading the founding paper closely is how you separate what the evidence supports from what the market has since oversold.

What the GEO paper actually is

Most of what circulates about generative engine optimization is second-hand: a checklist copied from a blog that copied it from another blog. The primary source is a single document. Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan, and Deshpande published "GEO: Generative Engine Optimization" as an arXiv preprint (2311.09735) and then at ACM SIGKDD 2024, one of the established venues for data-mining research. The paper coined the term and, more importantly, ran a controlled test of it.

The test rested on a purpose-built benchmark the authors named GEO-bench: roughly ten thousand queries drawn from nine datasets spanning different domains and question types. For each query, the researchers took the source pages a generative engine might draw on, applied a specific content change to some of them, and then measured whether that change made the engine more likely to surface and prominently feature that source inside its generated answer.

The measurement is the part worth understanding, because it is where careful reading pays off. The authors did not measure clicks, rankings, or revenue. They designed a metric called Position-Adjusted Word Count, which approximates how much of the generated answer a given source occupies and how prominently it appears. Everything the paper claims is a claim about that visibility proxy, not about a buyer picking up the phone.

What actually moved the needle

The headline finding is refreshingly concrete. Across the benchmark, a handful of content-level interventions produced a relative lift of roughly 30 to 40 percent on the Position-Adjusted Word Count metric compared with unoptimized baselines. Not every change helped, and the ones that did were not the tricks the SEO industry might expect.

Three levers stood out consistently, and one of them was the strongest of all.

Citing authoritative sources

Adding citations to credible, authoritative sources was the single most reliable lever in the study. A page that grounds its claims in reputable references was measurably more likely to be featured inside the generated answer than the same page making the same claims unsupported. This is the closest thing the paper offers to a durable principle: the machine rewards content that behaves like it has done its homework.

Including direct quotations

Adding relevant direct quotations, the kind of specific, attributable language a human editor would treat as evidence, also raised a source's visibility inside generated answers across the tested engines. Quotation is a form of specificity the retrieval-and-synthesis pipeline can latch onto.

Replacing vague claims with statistics

Swapping soft, general assertions for concrete statistics lifted visibility as well. "Many customers prefer" is weaker, in this measurement, than a specific figure with a source attached. The pattern across all three levers is the same: the interventions that worked added verifiable substance, not persuasion or keyword density.

What is GEO, precisely, and what sits underneath it

It helps to be exact about what the term names. Generative engine optimization is the practice of shaping content so that a generative answer engine, the layer that reads sources and writes a synthesized response, is more likely to retrieve, feature, and cite it. The object being optimized is a citation or a mention inside an answer, not a position in a list of links.

GEO does not float in a vacuum. The engines it targets are built on retrieval-augmented generation, the architecture introduced by Lewis and colleagues in 2020, which pairs a language model with a searchable index queried at the moment a question is asked. Understanding that substrate matters, because it explains why the paper's levers work: a system that retrieves passages and synthesizes from them will naturally favor passages that are specific, sourced, and quotable. GEO is, in effect, content design for a reader that is itself a retrieval system.

GEO vs SEO: the same page, a different target

The most common misreading is that generative engine optimization is a rebrand of search engine optimization. It shares tactics with SEO, and a well-structured, credible page tends to do well in both, but the target is different. SEO optimizes for a rank in a list of links. GEO optimizes for being featured and cited inside a single synthesized answer.

That difference is not cosmetic, and it is worth stating plainly: classic ranking and AI-answer citation are measurably decoupled. A page can rank outside the top organic results and still be cited in an AI answer, and a page can rank well and go unmentioned. The founding paper does not prove this decoupling on its own, but its choice to invent a new visibility metric rather than reuse rank position is an early acknowledgment that the old yardstick does not fit the new surface. This is the intellectual basis for treating classic search and AI answers as separate things to measure, rather than one blended score.

AEO vs GEO: a terminology caveat worth keeping

You will see "answer engine optimization" (AEO) and "generative engine optimization" (GEO) used interchangeably. They are not quite the same idea, and the difference is worth holding onto even though the market has largely collapsed them together.

Industry accounts describe AEO as the older label, inherited from the mid-2010s work of optimizing content to win featured snippets ("position zero") and voice-search answers. GEO is the newer, academically coined term specific to content retrieved and cited by a generative model. They share tactics, structuring content to be extractable as a direct answer, but the underlying machinery differs: a featured snippet is an extraction from one page, while a generated answer is a synthesis across many. We flag this as an industry-consensus narrative rather than a settled, primary-sourced fact, which is exactly the kind of distinction the founding paper trains you to make.

What the paper never proved

A founding result is valuable precisely because its limits are legible. The GEO paper earns its authority by measuring one thing carefully, which means it is silent on several things people now claim it settled. Reading it carefully means holding both.

It measured visibility, not customers

The 30-to-40-percent lift is a lift on Position-Adjusted Word Count, a visibility proxy inside the answer. The paper does not measure clicks, bookings, or revenue, and it makes no claim to. Being more visible inside an answer is a plausible input to being chosen, but the study does not, and cannot, promise a business outcome. Anyone citing this result as proof of leads is extending it past what it tested.

Being cited is not proof you were read

A separate strand of research complicates the whole idea of a citation as a trophy. In the dominant production pattern, citations are attached to a generated answer independently of the evidence that actually produced it, so a cited source is not reliably the source the model reasoned from. A 2024-25 attribution survey found that over 95 percent of answers from tested open-source language models contained at least one unattributed sentence. Winning a citation, in other words, is necessary but not sufficient, and it is not the same as the model faithfully using your content. The founding GEO paper does not address this faithfulness gap.

A citation is not a click

Even a genuine, faithful citation may go nowhere. Pew Research Center's 2025 behavioral study of nearly 69,000 real Google searches found that users clicked a link inside an AI summary only about 1 percent of the time, and clicked any traditional result in just 8 percent of searches where an AI summary appeared, versus 15 percent without one. The GEO paper optimizes for presence inside the answer, which is the right target for this surface, but it is a reminder that presence, not pass-through traffic, is what this channel actually delivers.

One controlled study is a starting point, not a standing law

The result was produced against specific engines and datasets at a specific moment. Generative engines change their retrieval and ranking behavior often, and a later large-scale study (Chen and colleagues, 2025) found AI search engines systematically biased toward earned, third-party media over brand-owned content, a pattern the founding paper did not test for. Treat the GEO paper as the field's origin data point, not as a fixed specification you can optimize against forever.

How to read a founding result

The discipline the paper models is the one worth keeping. It defined a metric before it made a claim, it reported relative lift rather than a single naked number, and it scoped its conclusions to what it measured. That is the opposite of most generative engine optimization advice in circulation, which borrows the paper's authority while dropping its caveats.

For a business, the practical takeaway is narrow and specific. The levers the study validated, credible citations, real quotations, specific statistics, are good content practice regardless of engine, because they add verifiable substance a retrieval system can use. What no study, including this one, can hand you is a guaranteed citation or a guaranteed customer. The only responsible way to know whether any of this is working for your business is to measure your actual presence in real answers to your real buyer questions, and to keep measuring as the engines change.

The evidence

Key findings, with their sources

  • The founding GEO study tested content interventions across a purpose-built benchmark (GEO-bench) of roughly 10,000 queries drawn from nine datasets.

    established Aggarwal, P. et al., "GEO: Generative Engine Optimization", arXiv:2311.09735, ACM SIGKDD 2024 (peer-reviewed).

  • Adding citations to credible sources, direct quotations, and specific statistics produced a relative lift of roughly 30 to 40 percent on the study's Position-Adjusted Word Count visibility metric; citing authoritative sources was consistently the single strongest lever.

    established Aggarwal, P. et al., "GEO: Generative Engine Optimization", arXiv:2311.09735, ACM SIGKDD 2024.

  • Over 95% of answers from tested open-source language models contained at least one unattributed sentence, and in dominant RAG pipelines a cited source is not reliably the source the model reasoned from.

    established Attribution/faithfulness research, arXiv:2409.11242 (2024) and attribution survey arXiv:2601.19927 (2025-26).

  • Users clicked a link inside a Google AI summary only about 1% of the time, and clicked a traditional result in about 8% of searches with an AI summary present versus 15% without.

    established Pew Research Center, "Do people click on links in Google AI summaries?", 2025 (browsing panel, 68,879 real searches).

  • A 2025 large-scale follow-up found AI search engines are systematically biased toward earned, third-party media over brand-owned content, a factor the founding paper did not test.

    emerging Chen, M., Wang, X., Chen, K., Koudas, N., "Generative Engine Optimization: How to Dominate AI Search", arXiv:2509.08919, 2025.

  • The generative answer engines GEO targets are built on retrieval-augmented generation, which pairs a language model with an index queried at inference time.

    established Lewis, P. et al., "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks", arXiv:2005.11401, NeurIPS 2020.

Calibration

What is proven, what is promising, what is unproven

Evidence tierTacticsWhat the evidence says
establishedCiting authoritative sources, adding direct quotations, replacing vague claims with specific statisticsControlled GEO-bench experiment, ~30-40% relative visibility lift; authoritative-source citation the strongest lever (Aggarwal et al., ACM SIGKDD 2024)
emergingPrioritizing earned, third-party media over brand-owned pages for AI-answer visibilitySingle large-scale 2025 follow-up study, not yet independently replicated (Chen et al., arXiv:2509.08919)
contestedTreating AEO and GEO as one identical playbookIndustry-consensus historical narrative, not primary-sourced; shared tactics but different underlying mechanisms (dossier terminology note)

Reference

Glossary

Generative engine optimization (GEO)
Shaping content so a generative answer engine is more likely to retrieve, feature, and cite it inside a synthesized answer. The optimization target is a citation or mention, not a rank in a list of links.
GEO-bench
The benchmark built for the founding GEO paper: roughly 10,000 queries across nine datasets, used to test which content changes raise a source's visibility inside generated answers.
Position-Adjusted Word Count
The visibility metric the GEO paper designed to approximate how much of a generated answer a given source occupies and how prominently it appears. It measures presence in the answer, not clicks or revenue.
Retrieval-augmented generation (RAG)
The architecture underneath most generative answer engines: a language model paired with a searchable index that is queried at the moment a question is asked, so the model answers from retrieved passages.
Answer engine optimization (AEO)
The older label, inherited from featured-snippet and voice-search optimization, for structuring content to be extracted as a direct answer. It shares tactics with GEO but predates the LLM-synthesis mechanics GEO describes.
Citation faithfulness
Whether a source credited in a generated answer is actually the source the model reasoned from. Research shows the two are frequently decoupled, so a citation is not proof of faithful use.

Straight answers

Frequently asked questions

What is generative engine optimization?

Generative engine optimization (GEO) is the practice of shaping content so a generative answer engine, like ChatGPT, Perplexity, Gemini, or Google AI Overviews, is more likely to retrieve, feature, and cite it inside a synthesized answer. The term was coined and empirically tested in a 2024 peer-reviewed paper, which found that citing credible sources, adding quotations, and using specific statistics measurably raised a source's visibility inside generated answers.

What did the founding GEO paper actually test?

It built a benchmark called GEO-bench of roughly 10,000 queries across nine datasets, applied specific content changes to candidate source pages, and measured whether those changes made a generative engine more likely to feature the page inside its answer. The metric was Position-Adjusted Word Count, a proxy for visibility inside the answer, not clicks, rankings, or revenue.

Does GEO replace SEO?

No. They share tactics, and a credible, well-structured page tends to do well in both, but the targets differ. SEO optimizes for a position in a list of links; GEO optimizes for being featured and cited inside a single synthesized answer. Classic ranking and AI-answer citation are measurably decoupled, which is why it is more accurate to measure them separately than to blend them into one score.

What is the difference between AEO and GEO?

Industry accounts describe answer engine optimization (AEO) as the older label from featured-snippet and voice-search optimization, and generative engine optimization (GEO) as the newer, academically coined term for content retrieved and cited by a generative model. They overlap in tactics but differ in mechanism: a snippet is an extraction from one page, while a generated answer is a synthesis across many. This is an industry-consensus distinction rather than a settled, primary-sourced one.

Did the GEO paper prove it gets you more customers?

No, and this is the most important thing to read correctly. The paper measured visibility inside the generated answer, not clicks, bookings, or revenue. Separate research also shows a citation is not proof the model actually used your content, and that most AI-summary sessions never click through at all. The levers the study validated are good content practice, but no study, including this one, can promise a guaranteed citation or a customer.

Provenance

Sources

  1. Aggarwal, P., Murahari, V., Rajpurohit, T., Kalyan, A., Narasimhan, K. & Deshpande, A., "GEO: Generative Engine Optimization", arXiv:2311.09735, ACM SIGKDD 2024 (established)arxiv.org
  2. Lewis, P. et al., "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks", arXiv:2005.11401, NeurIPS 2020 (established)arxiv.org
  3. Pew Research Center, "Do people click on links in Google AI summaries?", 2025 (established)pewresearch.org
  4. Attribution and faithfulness research, arXiv:2409.11242, 2024; attribution survey arXiv:2601.19927, 2025-26 (established finding of the faithfulness gap; emerging on mitigations)arxiv.org
  5. Chen, M., Wang, X., Chen, K. & Koudas, N., "Generative Engine Optimization: How to Dominate AI Search", arXiv:2509.08919, 2025 (emerging, single large-scale study)arxiv.org
  6. Gao, Y. et al., "Retrieval-Augmented Generation for Large Language Models: A Survey", arXiv:2312.10997, 2023-24 (established)arxiv.org
  7. Industry-consensus AEO/GEO terminology narrative (contested, flagged as not primary-sourced)

Every figure above is attributed to a real, dated source and tagged with its evidence tier. Where a claim could not be verified to a primary source, it is not stated as fact.

What this means for your business

The founding paper proves a narrow, specific thing: certain content moves raise how visible you are inside an AI answer, but visibility inside the answer is the only thing it measured, and no study can hand you a guaranteed citation or a customer. That leaves every owner with one question the research cannot answer for them: right now, when your buyers ask an engine about what you do, are you named in the answer at all? The only responsible way to know is to measure your real presence in real answers, which is exactly where a free Machine-Readiness Score starts.

diagnostic Free Machine-Readiness Score A specialist-reviewed read of where you actually stand across classic search, the local map pack, AI answers, and reputation, scored to one number so you can see the gap before any work is scoped. See how it works

Start free with a Machine-Readiness Score, a measured read of where you stand across search and AI answers. No obligation.