The Macro Shift · established evidence
The Generative Engine Optimization Paper: What Aggarwal et al. (2024) Actually Proved
Generative engine optimization began as an academic paper, not a vendor pitch. In 2024, researchers from IIT Delhi, Princeton, and Georgia Tech published "GEO: Generative Engine Optimization" at KDD, one of the top data-mining conferences, giving the field its first peer-reviewed footing. The paper did three concrete things: it defined the generative engine as a distinct response category that synthesizes an answer rather than returning a ranked list, it built a benchmark to measure how visible a source is inside that synthesized answer, and it tested specific content interventions such as citing sources, adding statistics, and quoting authorities to see which ones raised a source's visibility. This matters because most claims about how to be cited by AI are trade assertions with no measurement behind them. The GEO paper is the rare exception, and reading what it actually proved, and what it did not, is the starting point for any work on AI-answer visibility.
What the generative engine optimization paper set out to do
The paper, formally "GEO: Generative Engine Optimization" by Pranjal Aggarwal, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, Karthik Narasimhan, and Ameet Deshpande, appeared in the Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining in 2024. KDD is one of the most selective venues in data mining, which is the first thing that separates this work from a marketing blog: it survived peer review at a top conference.
Its starting observation is structural. A classic search engine returns an ordered list of links and leaves the selection to the reader. A generative engine reads across sources and returns a single synthesized answer with a handful of sources named inside it. The authors argue this is a genuinely new response category, not a cosmetic feature, because the unit a business competes for is no longer a rank in a list but a place inside the composed answer.
From that observation the paper poses a researchable question: if the object being targeted has changed, can we measure a source's presence inside a generated answer, and can we identify content changes that reliably increase it? That framing is what makes the work testable rather than anecdotal.
What is generative engine optimization, in the paper's own terms
In common usage, generative engine optimization now names a broad set of tactics for being cited by AI answers. In the paper it means something narrower and more precise: a method for adjusting content so that it is more visible inside the answer a generative engine composes, measured against a defined benchmark rather than asserted.
The distinction is worth holding onto. A vendor uses the phrase to describe a service. The authors use it to describe a measurable optimization problem with an input, the content, and an output, the source's visibility in the synthesized response. When this article says the GEO paper proved something, it means proved in that second sense: shown to move a defined metric on a shared benchmark, under stated conditions.
The benchmark: how the researchers measured visibility inside an answer
The paper's most durable contribution is an instrument, not a single tactic. The authors built a benchmark to evaluate generative engine responses across a range of real user queries, and defined visibility metrics that capture how prominently and how substantively a given source appears inside the composed answer, rather than simply whether a link was returned.
This is the part that raises GEO above the level of opinion. A benchmark lets a claim be checked. If an intervention is said to help, the benchmark makes it possible to run the query with and without the change and observe the difference in the source's measured visibility. That is the ordinary machinery of empirical science, applied for the first time to the question of being named inside an AI answer.
It is also why the paper is a better foundation than the volume of "GEO tips" that followed it. Most of that trade material offers tactics with no measurement attached. The academic work offers a way to measure, which is the thing that lets any later claim be tested rather than believed.
The tested interventions: what actually moved a source's visibility
Against that benchmark the authors tested a set of content interventions to see which ones changed a source's visibility inside generated answers. The interventions that measurably helped in the systems they tested were credibility signals in the content itself: adding cited statistics, including relevant quotations, and referencing authoritative sources.
The direction of that result is intuitive once stated, but the value is that it was measured rather than assumed. It points at a specific mechanism. Generative engines appear to favor sources whose content carries the markers of substantiation, numbers with provenance, quotable expert language, and citations, when composing an answer. That is a different profile from the link-and-keyword signals that classic ranking rewarded.
What the finding does and does not license
The reading is bounded. The paper shows that certain content changes raised measured visibility in the engines it studied, at the time it studied them. It does not show that any single tactic works on every engine, that effects are stable as models are retrained, or that adding a statistic to a page will get that page cited. The contribution is a demonstrated, measurable relationship, not a formula with a promised outcome.
That boundary is exactly why substantiation is a discipline and not a checkbox. A number without a real source, or a quotation invented to fit, fails the honesty test the moment anyone checks it, and fabricated credibility signals are a liability rather than a lever.
GEO vs SEO: what changes when the object is a citation, not a rank
The clearest way to see what the paper implies is to set it against the model it succeeds. In 1998, Sergey Brin and Lawrence Page published the PageRank paper, which reframed relevance as a link-based vote of confidence and became the foundation of algorithmic search. In that world the object of pursuit is a position in an ordered list, and the currency is links.
The GEO paper describes a different object. When an engine synthesizes an answer, there is no list to hold a position in. A source is either drawn into the composed answer or it is not, and the currency shifts toward the substantiation signals the study measured. This is the substance behind the common shorthand of GEO vs SEO: a change in what is being pursued, from being ranked among links to being cited inside an answer, not a rebrand of the same work.
Neither model replaces the other outright. Classic ranking still governs the link results that persist alongside AI answers, and the two share underlying signals such as authority and clarity. The paper's point is narrower and more useful: the generative case has its own measurable dynamics, and treating it as identical to ranking leaves those dynamics unaddressed.
Answer engine optimization and GEO: same shift, different labels
The market uses several names for the discipline this paper helped launch. Answer engine optimization, often written AEO, and generative engine optimization are frequently used interchangeably, and the difference between AEO and GEO is more a matter of emphasis than of substance. Both describe working to be the source an engine names when it answers a question rather than the link a user clicks.
The terminology is still settling, and that is itself worth stating plainly rather than papering over with a confident definition. What the academic literature contributes is a fixed method, not a fixed vocabulary: a benchmark and a set of measured interventions that any of these labels can be held to. When the label matters less than the measurement, the field has matured a step.
The real limits: what the GEO paper does not prove
An evidence-first reading has to be as clear about the ceiling as the floor. The GEO paper is a single study, however rigorous, and generative engines change underneath it as models are updated. Effects measured in 2024 are not guaranteed to hold on a 2026 system, and the paper never claimed otherwise.
The surrounding market data reinforces the caution. Claims about how large the AI-search shift already is remain genuinely contested: reported figures for how much of search behavior AI now commands vary by roughly five to ten times depending on the vendor and, crucially, on which denominator is being counted. Any single "AI search share" statistic should be read as provisional until the denominator is named.
The forecasting record is a further reason for humility. In February 2024, Gartner predicted that traditional search-engine volume would drop 25 percent by 2026 as AI chatbots absorbed queries. As of this writing that collapse has not materialized as stated, and Google still holds the large majority of the search market. The GEO paper is strong precisely because it measures a specific effect rather than forecasting a revolution, and it should be used for what it demonstrates, not stretched to underwrite predictions it never made.
Why this grounds an AI-answers measurement in evidence, not assertion
The practical value of the paper is that it moves the AI-answers question from belief to measurement. Being present inside a synthesized answer is not a vanity concern. Pew Research Center tracked the browsing of 900 consenting US adults across 68,879 Google searches in March 2025 and found that users clicked through to a traditional result in about 8 percent of searches that showed an AI summary, against 15 percent without one. When the answer is composed, the click that used to reward a good ranking is often never made.
That is the same pressure the long rise of zero-click search has applied for years, with independent clickstream analysis putting the no-click share of US Google searches above two thirds by early 2026. In an environment where the answer increasingly resolves on the surface itself, whether a business is named inside that answer is a real commercial fact, and it deserves to be measured with the same rigor the GEO paper brought to visibility in the lab.
This is where the academic literature does its quiet work. A measurement of AI-answer visibility that rests on a peer-reviewed benchmark and a set of tested interventions is standing on evidence, not on a vendor's claim about what the engines want. That is the difference between a number you can defend and a number you have to trust.
The evidence
Key findings, with their sources
-
The first peer-reviewed method for optimizing content to be cited inside a generative engine's synthesized answers, "Generative Engine Optimization" (GEO), was published at KDD 2024, a top data-mining venue, by researchers from IIT Delhi, Princeton, and Georgia Tech.
established Aggarwal, P., Murahari, V., Rajpurohit, T., Kalyan, A., Narasimhan, K. & Deshpande, A., "GEO: Generative Engine Optimization", Proceedings of KDD '24, arXiv:2311.09735.
-
Adding cited statistics, quotations, and authoritative sources measurably raised a source's visibility inside generated answers in the engines the study tested.
established Aggarwal et al., "GEO: Generative Engine Optimization", KDD 2024, arXiv:2311.09735 (peer-reviewed).
-
Users clicked through to a traditional search result in about 8% of searches that showed an AI summary, versus 15% of searches without one.
established Pew Research Center, "Do people click on links in Google AI summaries?", 2025 (browsing panel, 900 US adults, 68,879 searches).
-
The no-click share of US Google searches reached about 68% in early 2026, up from roughly 58.5% in 2024.
established SparkToro, "2024 Zero-Click Search Study" and "In 2026, Less than One Third of Google Searches Still Send a Click" (clickstream panel data).
-
A February 2024 forecast that traditional search volume would fall 25% by 2026 has not materialized as stated; Google still holds the large majority of search.
contested Gartner press release, 2024; reality-check reporting, 2026.
-
Reported figures for how large the AI-search shift already is vary by roughly five to ten times across vendors, depending on the denominator counted.
contested Cross-vendor comparison of AI-search share figures (Semrush, Similarweb, and trade aggregators), 2026.
Calibration
What is proven, what is promising, what is unproven
| Evidence tier | Tactics | What the evidence says |
|---|---|---|
| established | A peer-reviewed benchmark exists; content credibility signals (cited statistics, quotations, authoritative sources) measurably raised source visibility in tested engines. AI summaries measurably reduce click-through (Pew). Zero-click is the majority of Google searches. | Aggarwal et al. KDD 2024; Pew 2025; SparkToro 2024 and 2026. |
| emerging | Which specific interventions generalize across engines, locales, and successive model versions; how stable measured effects are as generative engines are retrained. | Single-study or short-window results; direction is credible but replication across systems and time is still thin. |
| contested | The overall size of the AI-search shift and the pace of any search-volume decline. Every single-vendor "AI search share" figure until the denominator is named. | Gartner 2024 forecast not borne out as stated; AI-share figures vary five to ten times by source and denominator. |
Reference
Glossary
- Generative engine
- A search interface that reads across sources and returns a single synthesized answer, naming a few sources inside it, rather than returning a ranked list of links. ChatGPT, Perplexity, Gemini, Copilot, and Google AI Overviews are examples.
- Generative engine optimization (GEO)
- In the academic sense, a measurable method for adjusting content so it is more visible inside the answer a generative engine composes, evaluated against a defined benchmark rather than asserted.
- Benchmark
- A shared, defined test set and metric that lets a claim be checked. In the GEO paper, it measures how prominently and substantively a source appears inside a generated answer.
- Citation (in an answer)
- A source named or drawn upon inside a synthesized answer. In the generative model, being cited replaces holding a rank as the thing a source competes for.
- Answer engine optimization (AEO)
- A near-synonym for GEO, working to be the source an engine names when it answers a question. The distinction between the two is mostly one of emphasis, not substance.
Straight answers
Frequently asked questions
What did the generative engine optimization paper actually prove?
It proved three concrete things under peer review: that the generative engine is a distinct response category from ranked-list search, that a source's visibility inside a synthesized answer can be measured against a benchmark, and that specific content interventions, among them adding cited statistics, quotations, and authoritative sources, measurably raised that visibility in the engines it tested. It did not prove that any single tactic works everywhere or that effects stay fixed as models change.
Is generative engine optimization the same as SEO?
No. SEO targets a position in a ranked list of links, with links as the core currency inherited from the PageRank model. GEO targets being cited inside a synthesized answer where there is no list to rank in, and the paper measured a shift toward content substantiation signals as what moves that visibility. They share some underlying signals but they target different objects.
Does adding statistics and citations guarantee my content gets cited by AI?
No. The paper showed that these credibility signals measurably raised visibility in the engines it tested, not that they guarantee a citation. Effects vary by engine and change as models are updated, and any statistic or quotation used has to be real and verifiable, because fabricated credibility signals are a liability rather than a lever.
What is the difference between GEO and AEO?
Generative engine optimization and answer engine optimization are used largely interchangeably. Both describe working to be the source an engine names when it answers a question. The difference is more a matter of emphasis and evolving vocabulary than of substance; what matters is the shared underlying method of measuring visibility and testing interventions.
Who wrote the GEO paper and where was it published?
It was written by Pranjal Aggarwal, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, Karthik Narasimhan, and Ameet Deshpande, researchers affiliated with IIT Delhi, Princeton, and Georgia Tech, and published in the Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD '24). The preprint is arXiv:2311.09735.
Provenance
Sources
- Aggarwal, P., Murahari, V., Rajpurohit, T., Kalyan, A., Narasimhan, K. & Deshpande, A., "GEO: Generative Engine Optimization", Proceedings of the 30th ACM SIGKDD Conference (KDD '24), 5-16, arXiv:2311.09735 (peer-reviewed, established)arxiv.org
- Brin, S. & Page, L., "The Anatomy of a Large-Scale Hypertextual Web Search Engine", Computer Networks and ISDN Systems, 30(1-7), 1998 (established)doi.org
- Pew Research Center, "Do people click on links in Google AI summaries?", 2025 (established, primary panel data)pewresearch.org
- SparkToro, "2024 Zero-Click Search Study" (with Datos/Semrush data), 2024 (established)sparktoro.com
- SparkToro, "In 2026, Less than One Third of Google Searches Still Send a Click" (with Similarweb data), 2026 (established)sparktoro.com
- Gartner, press release forecasting a 25% decline in search volume by 2026, 2024 (contested, used as a forecast-accuracy check)gartner.com
- Semrush and Similarweb AI-search traffic analyses, 2026 (contested, cited to document the spread in reported figures, not a settled number)
Every figure above is attributed to a real, dated source and tagged with its evidence tier. Where a claim could not be verified to a primary source, it is not stated as fact.