Trust, Ethics & Regulation · established (benchmark) / emerging (live-engine generalization) evidence
Optimizing to Be Quoted: What the Princeton Generative Engine Optimization Study Proved (and What It Didn't)
Generative engine optimization is the practice of shaping content so AI answer engines are more likely to quote it, and the study that named the field proved the effect is real. In controlled tests, Aggarwal and colleagues, working out of Princeton, showed that restructuring a source to add direct quotations and cited statistics raised its visibility inside generated answers by 22 to 41 percent, while keyword density did nothing comparable. That is a genuine, peer-reviewed result published at ACM SIGKDD 2024. But it was measured on GEO-bench, a benchmark the authors built, not on the live commercial engines your buyers actually use. ChatGPT, Perplexity, and Google AI Overviews run undisclosed, frequently changing selection logic that no external benchmark reproduces. So the precise reading is this: the lever exists and its direction is proven, but the exact number is a laboratory finding, not a claim about what happens in production.
What the generative engine optimization study set out to test
For twenty-five years the object being optimized was a rank: a position in a list of blue links. AI answer engines changed the object. A model now reads a set of retrieved passages, forms a synthesized answer, and names a few sources inside it. The question stopped being "where do I rank" and became "am I the source that gets cited." That is the gap answer engine optimization was coined to close.
In 2024, Pranjal Aggarwal, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, Karthik Narasimhan, and Ameet Deshpande published the first controlled study of that question, "GEO: Generative Engine Optimization," at ACM SIGKDD 2024. Their aim was empirical, not rhetorical: hold the underlying claims of a page constant, vary only how the content is structured and worded, and measure whether specific, repeatable choices change how often a generative engine quotes the source. It is the first controlled academic evidence that AI-answer visibility responds to measurable content design distinct from classic search ranking.
What the study proved: quotations and statistics beat keyword density
The headline result is clean. Across the methods the authors tested, restructuring a source to add direct quotations and cited statistics was the strongest lever, raising the source's visibility inside generated answers by 22 to 41 percent in controlled tests. Adding authoritative, citation-backed material moved the needle; padding the page with keyword density did not.
Read that comparison carefully, because it is the part with the most practical weight. The technique that worked is the technique that also happens to make a passage more trustworthy to a human: a real quotation from a named source, a real statistic with a citation attached. The technique that failed is the one classic SEO spent two decades abusing. In other words, the measured lift rewards the shape of genuine evidence, not the shape of manipulation. That alignment is the single most useful thing the study established, and it is why the finding survives outside the lab as a design principle even where the exact number does not.
What "measured on GEO-bench" actually means
Here is the caveat the confident headlines drop. The 22 to 41 percent figure was measured on GEO-bench, a benchmark the authors constructed to evaluate the methods under controlled conditions. It is a rigorously built test environment made of real search queries. It is not a panel of observed traffic from live ChatGPT, Perplexity, or Google AI Overviews sessions as buyers experience them.
This distinction is not pedantry, it is the whole reliability question. A benchmark holds variables still so an effect can be isolated and reproduced. A live commercial engine does the opposite: it changes its retrieval, ranking, and citation logic continuously, keeps that logic undisclosed, personalizes results, and blends in signals a benchmark cannot see. A lift that is stable in the first setting can be larger, smaller, or absent in the second. The correct way to state the result is that the study proved a controlled effect and identified which lever produces it, not that it measured what any given business will see inside a production engine next quarter.
Why the number should not be quoted as a promise
RavenEye's rule against a naked metric, a number stated without its source and its conditions, exists for exactly this case. The 22 to 41 percent figure is a textbook example. Cited with its benchmark context, it is strong evidence for a method. Cited alone, as "add quotes and get 40 percent more AI visibility," it becomes a claim the underlying study never made and no one can stand behind. The figure belongs in the first sentence of the finding and never in a guarantee.
Generative engine optimization vs SEO: why this is a different object
The reason the study matters at all is that it measures something classic SEO metrics cannot. Search engine optimization scores a position in a ranked list. Generative engine optimization scores whether a passage gets lifted and quoted into a synthesized answer. A page can rank well and never be quoted, and a page can be quoted while ranking modestly, because the engine selects passages, not whole pages, and scores each passage on whether it can stand alone when extracted.
That is why the winning technique was passage shape and evidence density rather than link volume or keyword targeting. The unit of competition moved from the page to the passage. For a small business, the operational translation is direct: the answer to a real buyer question needs to sit in a self-contained chunk, lead with the answer, and carry its own proof, so an engine can quote it without dragging three other paragraphs along. The GEO study is the first controlled reason to believe that shape is what gets rewarded.
Why live engines do not simply inherit the benchmark number
Two structural facts about generative systems explain why a benchmark lift does not transfer cleanly to production, and both are established outside the GEO paper itself.
First, generative models hallucinate as a structural property of how they are trained and decoded, not as a bug that more data removes. The canonical survey, Ji et al. in ACM Computing Surveys, finds no natural-language generation task is immune. An engine that can confidently state something false about a business can also cite unpredictably, which means citation behavior on live traffic carries noise a clean benchmark deliberately excludes.
Second, the commercial engines are moving targets by design. Their selection and citation logic is undisclosed and revised frequently. A technique validated against a fixed benchmark in 2024 is being applied, in 2026, against systems that have since changed. The direction of the GEO finding, evidence-shaped passages get quoted more, is durable because it aligns with what these systems are built to retrieve. The precise magnitude is not portable, and anyone who tells you it is has left the evidence behind.
How to rank in AI answers and get cited in AI search without overclaiming
The accurate application of the GEO study is narrower and more useful than the hype around it. It does not license a citation guarantee, and it does not endorse fabricating quotes or statistics to trip a benchmark. It supports a specific, defensible craft: take the true claims a page already makes, attach real quotations and real cited data where they belong, and restructure each key passage so it leads with the answer and stands on its own.
Every step of that is simultaneously what the study rewarded and what a human reader trusts more. That is the test to apply to any tactic sold as generative engine optimization: does it make the content more genuinely evidenced, or does it only game a signal? The first is craft and it compounds. The second is the manipulation the next section takes apart.
Where optimization ends and manipulation begins
The GEO study proved the citation signal is gameable, and that is the uncomfortable half of the finding. If adding quotations and statistics lifts visibility, a dishonest operator can add fabricated quotations and invented statistics to be disproportionately quoted, at least until an engine learns to discount them. This is where "optimizing to be cited" risks becoming "gaming an unaudited trust signal," and it is why this article sits in the Trust, Ethics and Regulation pillar rather than a tactics column.
The boundary is not subjective. Google's own Search Quality Rater Guidelines treat trustworthiness as the terminal criterion of E-E-A-T: a page can show experience, expertise, and authority and still be rated low quality if its content is inaccurate or deceptive. And the FTC's fake-review rule, 16 CFR Part 465, effective October 21, 2024, already makes fabricated endorsements, including fake reviews created by AI systems and purchased engagement, rule violations carrying civil penalties. The legal and algorithmic systems are converging on the same line: evidence must be real. Legitimate generative engine optimization uses true quotes, true data, and true expertise, and structures them to be found. Everything past that line is not optimization, it is a liability waiting for an audit.
Reading the evidence
The GEO study is genuinely important and routinely misquoted, and both things are true at once. It gave the field its first controlled proof that AI-answer visibility responds to deliberate content design, and it identified the specific lever, evidence density, that produces the effect. That is a real contribution and a sound basis for method.
What it did not do is measure what your business will gain inside a live engine, and no responsible reading claims otherwise. The correct posture is the one the whole answer era demands: treat the direction as established, treat the magnitude as a benchmark result, and measure your own visibility on the engines your buyers actually use rather than borrowing a number from a lab. The study tells you which way to build. Only measurement tells you where you stand.
The evidence
Key findings, with their sources
-
Restructuring a source to add direct quotations and cited statistics raised its visibility inside generated answers by 22 to 41 percent in controlled tests, while keyword density showed no comparable effect.
established (benchmark) Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan & Deshpande, "GEO: Generative Engine Optimization", ACM SIGKDD 2024, arXiv:2311.09735 (peer-reviewed).
-
The 22 to 41 percent lift was measured on GEO-bench, a benchmark the authors constructed, not on live commercial engines whose selection and citation logic is undisclosed and frequently changing.
emerging (live-engine generalization) Aggarwal et al., "GEO: Generative Engine Optimization", ACM SIGKDD 2024, arXiv:2311.09735.
-
It is the first controlled academic evidence that AI-answer visibility responds to specific, measurable content design choices distinct from classic search ranking.
established (benchmark) Aggarwal et al., "GEO: Generative Engine Optimization", ACM SIGKDD 2024, arXiv:2311.09735.
-
Generative models produce fluent, confident, false content as a structural property of training and decoding, not a bug removable with more data, and no natural-language generation task is immune.
established Ji, Lee, Frieske, et al., "Survey of Hallucination in Natural Language Generation", ACM Computing Surveys, 55(12), 2023; preprint arXiv:2202.03629.
-
Since October 21, 2024, the FTC fake-review rule (16 CFR Part 465) makes fabricated endorsements, including fake reviews created by AI systems and purchased engagement, rule violations carrying civil penalties up to $51,744 per violation.
established FTC, 16 CFR Part 465, Federal Register 2024-18519, effective Oct 21, 2024.
-
Google treats trustworthiness as the terminal member of E-E-A-T: a page can show experience, expertise, and authority and still be rated low quality if its content is inaccurate or deceptive.
established Google, Search Quality Rater Guidelines; Search Central, "Creating Helpful, Reliable, People-First Content".
Calibration
What is proven, what is promising, what is unproven
| Evidence tier | Tactics | What the evidence says |
|---|---|---|
| Established (benchmark) | Adding real quotations and cited statistics raises quotation rate; keyword density does not; the effect is a controlled, reproducible finding on GEO-bench. | Aggarwal et al., GEO, ACM SIGKDD 2024, arXiv:2311.09735. |
| Established (general) | Evidence must be genuine: hallucination is structural, trust is the terminal search-quality criterion, and fabricated proof is already an FTC rule violation. | Ji et al. 2023; Google Rater Guidelines; FTC 16 CFR 465. |
| Emerging (live-engine generalization) | The exact 22 to 41 percent magnitude transferring to live ChatGPT, Perplexity, or Google AI Overviews; direction is credible, magnitude is unproven on production traffic. | Extrapolation beyond GEO-bench; live engines use undisclosed, changing logic (RavenEye research note, 2026). |
| Contested / unsupported | Any guaranteed AI citation, any promised percentage lift on a live engine, or fabricated quotes and statistics to game the signal. | No study supports it; fabricated evidence is an FTC 16 CFR 465 and Section 5 exposure. |
Reference
Glossary
- Generative engine optimization (GEO)
- Shaping content so AI answer engines are more likely to retrieve and quote it inside a synthesized answer. The term and its first controlled study come from Aggarwal et al., ACM SIGKDD 2024.
- GEO-bench
- The benchmark the GEO study authors built to test their methods under controlled conditions. A constructed test environment of real queries, not a panel of live commercial-engine traffic.
- Answer engine
- A search interface (ChatGPT, Perplexity, Gemini, Copilot, Google AI Overviews) that returns a synthesized answer naming a few sources, instead of only a ranked list of links.
- Citation (in an AI answer)
- The event of being named or quoted as a source inside a generated answer. The object GEO optimizes for, distinct from a ranking position.
- Passage-level retrieval
- The behavior of AI engines selecting and scoring individual passages rather than whole pages, which is why self-contained, answer-first chunks are more quotable.
- Naked metric
- A number stated without its source and the conditions it was measured under. This piece never states one on its own, which is why the 22 to 41 percent figure always travels with its benchmark caveat.
Straight answers
Frequently asked questions
What did the Princeton GEO study prove?
It proved, under controlled conditions, that restructuring a source to add real quotations and cited statistics raises how often AI answer engines quote it, by 22 to 41 percent in its tests, while keyword density does not help. It is the first controlled evidence that AI-answer visibility responds to specific content design distinct from classic search ranking.
Does the 22 to 41 percent lift apply to ChatGPT or Google AI Overviews?
Not directly. The figure was measured on GEO-bench, a benchmark the authors built, not on live ChatGPT, Perplexity, or Google AI Overviews traffic. Those engines use undisclosed, frequently changing selection logic that no external benchmark reproduces, so the direction of the finding is credible but the exact magnitude should not be treated as what a live engine will deliver.
Is generative engine optimization the same as SEO?
No. SEO optimizes for a position in a ranked list of links. Generative engine optimization optimizes for being quoted inside a synthesized answer, which the study found depends on passage shape and evidence density rather than link volume or keyword targeting. A page can rank well and still never be quoted.
What actually makes content more likely to be quoted by an AI answer?
On the study's evidence, real quotations and cited statistics, placed in self-contained passages that lead with the answer. The technique that worked is the one that also makes content more trustworthy to a human reader, and the technique that failed, keyword density, is the one classic SEO abused.
Can anyone guarantee an AI citation?
No, and any offer that does should be treated as a warning sign. Engines change their logic continuously and disclose none of it, the underlying study measured a benchmark effect rather than a live-traffic guarantee, and fabricating evidence to force a citation is already an FTC rule violation. Sound practice improves the odds and measures the result; it does not promise a number.
Provenance
Sources
- Aggarwal, P., Murahari, V., Rajpurohit, T., Kalyan, A., Narasimhan, K. & Deshpande, A., "GEO: Generative Engine Optimization", ACM SIGKDD 2024, arXiv:2311.09735 (established, benchmark; emerging on live-engine generalization)arxiv.org
- Ji, Z., Lee, N., Frieske, R., et al., "Survey of Hallucination in Natural Language Generation", ACM Computing Surveys, 55(12), Article 248, 2023; preprint arXiv:2202.03629 (established)arxiv.org
- FTC, 16 CFR Part 465, Trade Regulation Rule on the Use of Consumer Reviews and Testimonials, Federal Register 2024-18519, effective Oct 21, 2024 (established)ecfr.gov
- Google, Search Quality Rater Guidelines, and Search Central, "Creating Helpful, Reliable, People-First Content" (established)
- Google Search Central Blog, "E-A-T gets an extra E for Experience", Dec 2022 (established)
Every figure above is attributed to a real, dated source and tagged with its evidence tier. Where a claim could not be verified to a primary source, it is not stated as fact.