Measurement & Honesty · emerging evidence

Measuring the Answer Engine: What Generative Engine Optimization (GEO) Research Says So Far

Last reviewed 2026-07-20. Written by Chandranshu Kumar, Founder, Raveneye Global. · 10 min read

Generative engine optimization, or GEO, is the practice of making a source more likely to be named and cited inside a generated answer rather than ranked in a list of links. It has an academic starting point. In 2024 a peer-reviewed paper, "GEO: Generative Engine Optimization" by Aggarwal and colleagues, introduced both a set of optimization methods and an evaluation benchmark for measuring visibility inside generated answers. That work establishes that the object being optimized has changed, from a rank to a citation, and that the change can be studied rather than merely asserted. It is also, by the standards of a measurement discipline, very new. The founding paper is barely two years old, there is no agreed methodology for measuring share of answer across engines, and the systems being measured are non-deterministic by construction. What follows is what the research supports today, tiered by evidence, and where the limits still sit.

What the first generative engine optimization study established

For most of the search era, the academic and industry literature on visibility studied one object: rank, meaning a position in an ordered list of links. Generative engines changed the object. When ChatGPT, Perplexity, Google AI Overviews, Gemini, or Copilot answer a question, they synthesize a response and name a small number of sources inside it. The relevant outcome is no longer where a page sits in a list but whether it is drawn into the answer at all.

The paper "GEO: Generative Engine Optimization" (Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan, and Deshpande), accepted to KDD 2024, is the first academic framework and benchmark built for that new object. Its contribution is two-sided. It proposes optimization methods, content-level changes a source can make to improve its chances of being cited, and it introduces an evaluation methodology, a way to measure that visibility systematically rather than anecdotally. Before this work, "optimizing for AI answers" was a claim without a shared yardstick. After it, there was at least one.

The significance is less any single number and more the existence of the discipline. That a peer-reviewed framework and benchmark exist at all is the signal worth reading: the field now has a founding reference point, and a way to test assertions against evidence instead of accepting them on confidence.

The levers the benchmark tested

The study did not treat all content changes as equal. It measured which content-level interventions actually moved whether a source was cited inside a generated answer across the engines it tested. The interventions that raised a source's visibility were the ones that made a passage more verifiable and more quotable: adding cited statistics, incorporating direct quotations, and citing authoritative sources.

That result matters because it inverts a common instinct. Writing tuned to please a ranking algorithm is not the same as writing structured to be extracted and trusted by a system synthesizing an answer. The research points toward the second: concrete, attributable, self-contained passages that a model can lift and stand behind.

Read the effect sizes with care

The framing is qualitative, not a headline percentage. The methods measurably raised visibility inside the specific systems the paper evaluated, on the datasets it used, at the time it ran. Generalizing a single benchmark's effect size to every engine, every query type, and every locale would overstate what one study can carry. The direction is evidenced. The magnitude is context-bound.

Generative engine optimization versus classic SEO

The clearest way to read the research is as a change of target. Classic search engine optimization improves a page's position in a ranked list. Generative engine optimization improves a source's odds of being named and cited inside a synthesized answer. The two overlap, but they are not the same outcome, and the GEO framework exists precisely because the second could not be measured with the tools built for the first.

This is why the seo versus geo comparison is more than terminology. A page can rank well and still be absent from the answer written above the links, and a source can be cited by an engine that retrieves and evaluates content independently of Google's ranked results. The measurement that tells you about the first tells you little about the second. The academic contribution was to define and benchmark the second on its own terms.

Answer engine optimization and the vocabulary problem

The field carries several labels: generative engine optimization, answer engine optimization, AI-search optimization. In practice they describe the same job, being found and cited when a buyer asks an engine a question, and the research literature is thin enough that the vocabulary is still settling. That immaturity is not a reason to dismiss the work; it is a reason to be precise about what is evidenced and what is merely named.

The useful discipline is to separate the settled part from the marketed part. The settled part is the shift in object, from rank to citation, and the existence of a peer-reviewed framework and benchmark to study it. The marketed part is the growing catalog of confident tactics and proprietary "AI visibility scores" that outrun what any published study has yet shown. A reader is well served by treating the first as established groundwork and the second as claims to be checked.

Why measuring AI visibility has no standard yet

The largest gap is measurement itself. There is no standardized, agreed methodology for measuring share of answer or AI-search visibility. The reason is structural, not a matter of effort. Generative engines are non-deterministic: the same query can return different answers across sessions and runs. They personalize responses, and they are not fully observable from outside. No engine publishes query volume, impressions, or citation rates.

The consequence is unavoidable and worth stating plainly. Any claimed "AI visibility tracking" number is a sample-based estimate, not a census, and its reliability depends entirely on the disclosed sampling methodology behind it: which prompts, how many engines, how many repeated runs, which locales, on what dates. A number offered without that method attached is not a measurement; it is an assertion wearing a measurement's clothes. This is a genuinely open problem, not a solved one, and reading GEO research well means saying so.

How to measure GEO responsibly today

The practical answer that follows from the research is transparency, not false precision. A defensible read samples a fixed panel of realistic buyer prompts, runs it repeatedly across each engine, records whether and how a source appears, and stamps every reading with engine, locale, and date. It reports a range with the method shown rather than one tidy figure that hides its own uncertainty. That posture is the correct one for a first-generation measurement field, and it is the posture the founding paper's own methodology models.

What a mature measurement field looks like, and why GEO is not there yet

It helps to compare GEO with fields that have already grown up. Online controlled experimentation has a canonical methods text, Kohavi, Tang, and Xu's "Trustworthy Online Controlled Experiments" (2020), drawn from running tens of thousands of experiments a year at major platforms, cataloguing the traps and the discipline in one place. GEO has no equivalent. Its founding paper is barely two years old, and the field has not yet produced a shared, canonical methodology on that scale.

The communications industry offers a second comparison. In 2020 AMEC ratified the Barcelona Principles 3.0, an industry-wide standard that formally rejected Advertising Value Equivalency in favor of transparent, outcome-based measurement. That is what a field disciplining its own metrics looks like: a ratified standard, arrived at collectively, that rules a vanity metric out. AI-answer measurement has no such standard yet. The lesson is not that GEO is unserious; it is that the field is at the beginning of the maturation that experimentation and communications measurement have already been through.

How to read GEO research

Tier the evidence, and hold the tiers apart. The shift in object, from rank to citation, and the existence of a peer-reviewed framework and benchmark are established groundwork. The specific optimization levers are emerging, evidenced in the systems one study tested and not yet replicated at scale across engines. And the measurement of share of answer is open by nature, constrained by non-determinism and the absence of an agreed methodology.

The correct posture for anyone selling or buying work in this space is to disclose the field's immaturity rather than paper over it with confident numbers. The research supports a real, useful practice. It does not yet support the certainty that the noisier end of the market projects. Holding that line, evidenced where the evidence exists and plain about where it does not, is the difference between measurement and marketing.

The evidence

Key findings, with their sources

  • The first academic framework and benchmark for measuring and improving a website's visibility inside generated answers, introducing both optimization methods and an evaluation methodology.

    emerging Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan & Deshpande, "GEO: Generative Engine Optimization", arXiv:2311.09735, KDD 2024 (peer-reviewed).

  • Adding cited statistics, quotations, and authoritative sources measurably raised a source's visibility inside the generated answers of the engines the study tested.

    emerging Aggarwal et al., "GEO: Generative Engine Optimization", KDD 2024, arXiv:2311.09735.

  • There is no standardized, agreed methodology for measuring share of answer; because generative engines are non-deterministic, personalized, and not fully observable, any AI-visibility number is a sample-based estimate whose reliability depends on the disclosed sampling method.

    contested Inference from the GEO paper's own stated methodology (Aggarwal et al., 2024) plus the documented non-determinism of LLM outputs.

  • Online controlled experimentation has a canonical methods text drawn from tens of thousands of experiments a year at major platforms; GEO has no equivalent standard text yet.

    established Kohavi, Tang & Xu, "Trustworthy Online Controlled Experiments: A Practical Guide to A/B Testing", Cambridge University Press, 2020.

  • A communications-industry standard formally rejected Advertising Value Equivalency in favor of transparent, outcome-based measurement; AI-answer measurement has no comparable ratified standard yet.

    established AMEC, Barcelona Principles 3.0, 2020.

Calibration

What is proven, what is promising, what is unproven

Evidence tierTacticsWhat the evidence says
establishedThe disciplines GEO is measured against: canonical experimentation methodology, and a ratified standard renouncing a vanity metric. Benchmarks of maturity the field has not yet reached.Kohavi, Tang & Xu (2020); AMEC Barcelona Principles 3.0 (2020).
emergingThe GEO framework itself: the shift from rank to citation, the content-level levers (cited statistics, quotations, authoritative sources), and a benchmark to evaluate them.Aggarwal et al., "GEO: Generative Engine Optimization", KDD 2024, arXiv:2311.09735.
contestedMeasuring share of answer across engines. No agreed methodology exists; every number is a sample-based estimate governed by non-determinism and disclosed sampling.Inference from Aggarwal et al. (2024) plus documented LLM non-determinism.

Reference

Glossary

Generative engine optimization (GEO)
The practice of improving whether a source is named and cited inside a generated answer, as distinct from ranking it in a list of links. Named in the founding 2024 academic paper by Aggarwal and colleagues.
Answer engine optimization (AEO)
A near-synonym for GEO in common use: optimizing to be the cited answer an engine returns rather than a position in a ranked list.
Benchmark
A shared, repeatable evaluation setup that lets different methods be compared on the same footing. The GEO paper introduced one for AI-answer visibility.
Share of answer
The share of a defined panel of buyer prompts in which a source is named or cited by an engine. There is no standardized way to measure it yet, so any figure is a sample-based estimate.
Non-determinism
The property, built into generative engines, that the same query can return different answers across runs, which is why AI-visibility measurement is sampling, not a census.

Straight answers

Frequently asked questions

What did the GEO paper actually establish?

That the object of optimization has changed from a rank to a citation, and that this can be studied with a shared framework. Aggarwal and colleagues (KDD 2024) introduced both optimization methods and an evaluation benchmark for visibility inside generated answers. It is the first peer-reviewed reference point for the field, which is its main significance.

Is generative engine optimization the same as SEO?

No. Classic SEO optimizes for a position in a list of links. Generative engine optimization optimizes for being named and cited inside a synthesized answer, which depends more on entity consistency, extractable and attributable content, and third-party corroboration than on ranking position alone. The GEO framework exists because the second outcome could not be measured with the tools built for the first.

Can anyone reliably measure my AI-search visibility?

Only as a disclosed estimate, not as a fixed number. There is no standardized methodology for share of answer, and generative engines are non-deterministic, personalized, and not fully observable. A responsible read samples a fixed prompt panel across engines, repeats the runs, stamps each reading with engine, locale, and date, and reports a range with the method shown. Be skeptical of any single "AI visibility score" offered without its sampling method attached.

How mature is GEO research?

Young. The founding academic paper is barely two years old, and the field has no canonical methods text on the scale of Kohavi's experimentation guide, nor a ratified measurement standard like the communications industry's Barcelona Principles. The practice is real and evidenced in parts; the certainty projected by the noisier end of the market runs ahead of what any published study has shown.

How would I know if my business is being cited in AI answers?

You have to measure it directly, because no engine publishes the data. The method is a repeated sample of your real buyer questions across each engine, recorded by engine, locale, and date, with mentions counted separately from citations, and reported as a range. That measured starting point is what any credible plan is built on.

Provenance

Sources

  1. Aggarwal, P., Murahari, V., Rajpurohit, T., Kalyan, A., Narasimhan, K., Deshpande, A., "GEO: Generative Engine Optimization", arXiv:2311.09735, KDD 2024 (peer-reviewed, emerging)arxiv.org
  2. Kohavi, R., Tang, D., Xu, Y., "Trustworthy Online Controlled Experiments: A Practical Guide to A/B Testing", Cambridge University Press, 2020 (established)doi.org
  3. AMEC (International Association for the Measurement and Evaluation of Communication), Barcelona Principles 3.0, 2020 (established)amecorg.com
  4. State of AI-answer measurement: inference from the GEO paper's stated methodology (Aggarwal et al., 2024) plus the documented non-determinism of LLM outputs (emerging/contested by nature)

Every figure above is attributed to a real, dated source and tagged with its evidence tier. Where a claim could not be verified to a primary source, it is not stated as fact.

What this means for your business

The research supports a real practice and a disclosed posture: being cited in AI answers can be worked on and measured, but only as a disclosed estimate, never as a guaranteed number. That is exactly how the AI-answers pillar of your Machine-Readiness Score is built. It samples a panel of your real buyer questions across ChatGPT, Google AI Overviews, Perplexity, Gemini, and Copilot, stamps every reading with engine, locale, and date, and reports where you stand as a range with the method shown, so you can see the gap before anyone scopes a fix.

service AI-Answer & GEO Visibility The directed program that gets your business named and cited inside AI answers, built on the peer-reviewed method, entity consistency, extractable content, and earned corroboration, and tracked as a disclosed share of answer. See how it works

Start free with a Machine-Readiness Score, a specialist-reviewed read of where you stand across search and AI answers. No guaranteed number, and no obligation.