Discovery Science · emerging evidence
Earned Media Beats Owned Content in AI Search: What a 2025 Generative Engine Optimization Study Found
A 2025 controlled study of generative engine optimization reported a pattern that should reshape how businesses think about visibility: AI answer engines cite earned, third-party authoritative sources at a systematically higher rate than they cite brand-owned pages or social content, and they do so far more lopsidedly than classic Google search, which draws on owned and earned sources more evenly. The work, by Chen and colleagues, tested the effect across multiple verticals, languages, and reworded versions of the same query. The practical reading is uncomfortable for anyone who has invested only in their own website: in an AI answer, what other credible sources say about you can weigh more than what you say about yourself. This is a single, not-yet-replicated study, so we treat it as an emerging signal rather than settled law. Read against the wider evidence, though, it points in a consistent direction.
What the 2025 study actually found
In 2025 a research team led by Chen and colleagues published Generative Engine Optimization: How to Dominate AI Search (arXiv:2509.08919), a large-scale controlled follow-up to the earlier academic work that coined the term. Its central reported result is a source-selection bias: across the systems it examined, generative answer engines drew disproportionately on earned and third-party media, the independent publications, industry resources, and outside authorities that talk about a business, rather than on the brand-owned pages a business publishes about itself or the social content it posts.
What makes the finding notable is the comparison it draws. The same study contrasts that lopsided behavior with classic Google search, which the authors report sources more evenly across owned and earned material. In other words, the tilt toward third-party authority is not simply how all search has always worked; it is reported to be sharper in the generative-answer layer than in the ranked-link layer that preceded it.
The authors tested the pattern across multiple verticals, multiple languages, and paraphrased versions of the same underlying query, which is a meaningful design choice: it guards against the finding being an artifact of one industry, one phrasing, or one language. That breadth is why the result is worth taking seriously even while it remains a single study.
Earned media and owned content are two different levers
The distinction at the heart of this piece is old in communications, even if it is newly consequential in search. Owned content is everything a business controls and publishes itself: its website, its landing pages, its blog, its own social accounts. Earned media is what independent third parties say about it without being paid to say it: a journalist's article, an industry directory listing, a review corpus, an expert citation, a reference in a resource other people maintain.
The two levers behave differently because their credibility comes from different places. Owned content is a first-party claim; its author and its subject are the same party. Earned media is a second-party endorsement; the source has no obligation to be flattering, which is precisely what gives it evidential weight. A system trying to answer "who is good at this" has a structural reason to lean on the second kind of signal, because it is harder to manufacture.
This is not the same as the paid-media lane. Buying links, renting placements, or manufacturing coverage is neither earned nor durable; it violates search platforms' link-spam policies and is increasingly filtered out of AI citations. The lever the evidence points to is genuine earned authority, mention won on merit, not purchased corroboration dressed up to look like it.
The founding GEO experiment pointed the same way
The 2025 result does not sit in isolation. The peer-reviewed study that established generative engine optimization as a field, Aggarwal and colleagues at ACM SIGKDD 2024, ran a controlled benchmark of roughly 10,000 queries across nine datasets and measured which content-level changes moved whether a source was surfaced inside a generated answer.
Its strongest levers are directly relevant here. Adding citations to credible sources, including direct quotations, and replacing vague claims with concrete statistics produced a relative lift on the order of 30 to 40 percent on the study's visibility metric versus unoptimized baselines, and citing authoritative sources was consistently the single most effective intervention it tested. That earlier experiment measured the tactic (borrow authority through citation) while the 2025 study measures the pattern (engines prefer authority they can borrow). They are two readings of the same underlying pull toward corroborated, attributable material.
When an independent 2024 result and a large-scale 2025 result point in the same direction from different angles, the convergence is worth more than either finding alone, even though neither, on its own, closes the question.
Why retrieval would reward corroboration over self-assertion
A plausible mechanism sits underneath both findings, and it is worth stating carefully as reasoned inference rather than proven fact. Nearly every generative answer engine is built on retrieval-augmented generation: the founding architecture, described by Lewis and colleagues in 2020, pairs a language model with a searchable index of outside documents that is queried at answer time, which the authors showed produces more specific and factual output than a model relying on its parameters alone.
In that design, the answer is assembled from whatever passages the retrieval step surfaces as most relevant and trustworthy for the query. A page that asserts its own excellence offers the retriever a first-party claim. A cluster of independent pages that corroborate the same fact offers something a ranking or relevance system has long treated as stronger: agreement from sources that do not share the subject's incentive. It is reasonable to expect a retrieval-first system to over-index on that corroboration, and the two studies above are consistent with exactly that expectation. The mechanism is inferred from the architecture and the results; it has not been isolated experimentally, and we flag it as such.
How this departs from classic search
For two decades, a business could win a great deal of search visibility on the strength of its own site. Publish enough well-structured, useful pages, earn a reasonable link profile, and you could rank. Owned content was the primary lever, and earned media was an accelerant. The 2025 study's reported contrast, that classic Google sources more evenly while generative engines tilt harder toward third parties, suggests that habit transfers poorly to the answer layer.
The consequence is concrete. A business can hold a strong classic ranking and still be missing from the AI answer written above or instead of those links, because the two surfaces are drawing on different mixes of sources. This is one more piece of evidence that classic search visibility and AI-answer visibility are decoupling into related but separate disciplines, which is why measuring them as one blended number obscures more than it reveals.
Ranking is not the same as being cited
Industry practitioners have documented cases of pages cited inside AI answers that sit well outside the top classic results, and pages that rank highly yet go uncited. The source-selection bias in the 2025 study is a structural reason to expect that mismatch rather than a surprise: if the answer layer weights third-party authority more heavily, the pages it names need not be the pages that rank, and the sites it draws from need not be sites the business owns.
The limits of this finding
Rigor here means naming the limits as plainly as the finding. First, the 2025 study is a single, not-yet-independently-replicated result. Its breadth across verticals and languages strengthens it, but one paper, however large, is an emerging signal, not an established law. Anyone selling the earned-media tilt as a settled certainty is ahead of the evidence.
Second, being cited is not the same as being read faithfully, nor the same as being visited. The RAG faithfulness literature shows that in the dominant production pipelines, a citation can be attached to an answer independently of the passages the model actually reasoned from; one 2024 attribution survey reported that more than 95 percent of answers from the open-source systems it tested contained at least one unattributed sentence. A citation is a claim of provenance, not a guarantee of it.
Third, even a citation that is faithful may never turn into a click. Pew Research Center's 2025 behavioral study of 68,879 real Google searches found users clicked a traditional result in about 8 percent of searches where an AI summary appeared, versus about 15 percent without one, and clicked a link inside the summary itself only around 1 percent of the time. Winning earned citations is necessary for AI-answer presence; it is not, by itself, a traffic strategy. The right yardstick for this surface is presence in the answer, measured directly, not organic sessions.
What the study does not claim
It is as important to state the boundaries of the result as its substance. The study does not claim that owned content no longer matters; a business still needs a coherent, well-structured, machine-readable site for any engine to have a first-party entity to attach earned mentions to. It does not provide a public per-vertical citation rate that a specific business can benchmark itself against. It does not prove that buying coverage works; the durable lever is earned, not purchased, and manufactured corroboration is both against platform policy and increasingly filtered out.
And it does not promise that a given amount of earned media produces a given amount of AI visibility. There is no guaranteed dose-response curve in the evidence. What the study supports is a directional reweighting of strategy, not a formula.
What this means for a visibility strategy
Read together, the evidence argues for a shift in where effort goes. A visibility program built almost entirely on owned pages is optimized for the surface that is losing share of the decision. Reallocating some of that effort toward earning genuine third-party authority, the digital-PR discipline of building citable assets and placing them on sources engines already trust, aligns the strategy with how the answer layer appears to select what it cites.
The reallocation is a rebalance, not a reversal. Owned foundations still carry classic search and give earned mentions an entity to point at; earned authority is what the generative layer disproportionately reaches for. The pairing is the point. And because none of this can be assumed, the discipline that ties it together is measurement: sampling real buyer questions across each engine and recording how often a business is actually named, then watching whether earned placements move that number over time, reported with variance rather than asserted.
The evidence
Key findings, with their sources
-
A 2025 large-scale controlled study reported that generative answer engines cite earned, third-party authoritative sources at a systematically higher rate than brand-owned pages or social content, and more lopsidedly than classic Google, which sources more evenly; tested across multiple verticals, languages, and query paraphrases.
emerging Chen, M., Wang, X., Chen, K., Koudas, N., "Generative Engine Optimization: How to Dominate AI Search", arXiv:2509.08919, 2025.
-
In a controlled benchmark of ~10,000 queries across nine datasets, adding citations to credible sources, direct quotations, and concrete statistics produced roughly a 30 to 40 percent relative lift on the study's visibility metric, with citing authoritative sources the single strongest lever.
established Aggarwal, P. et al., "GEO: Generative Engine Optimization", ACM SIGKDD 2024, arXiv:2311.09735 (peer-reviewed).
-
Retrieval-augmented generation pairs a language model with an outside document index queried at answer time, producing more specific and factual output than parametric generation alone; it is the architecture underneath most generative answer engines.
established Lewis, P. et al., "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks", arXiv:2005.11401, NeurIPS 2020.
-
More than 95 percent of answers from the tested open-source LLM systems contained at least one unattributed sentence, and citations can be attached independently of the passages actually used, so a citation is not proof the source was reasoned from.
established Attribution/faithfulness literature, arXiv:2409.11242, 2024 (and related 2025 attribution survey).
-
Across 68,879 real Google searches, users clicked a traditional result in about 8 percent of searches with an AI summary present versus about 15 percent without, and clicked a link inside the summary only around 1 percent of the time.
established Pew Research Center, "Do people click on links in Google AI summaries?", July 2025 (browsing panel).
Calibration
What is proven, what is promising, what is unproven
| Evidence tier | Tactics | What the evidence says |
|---|---|---|
| established | Citing credible third-party sources and concrete statistics measurably lifts a source's presence in generated answers; RAG is the shared architecture; citation is not the same as faithful use or a click. | Aggarwal et al. KDD 2024; Lewis et al. 2020; arXiv:2409.11242; Pew 2025. |
| emerging | Generative engines tilt toward earned/third-party media over owned and social content more sharply than classic Google does, across verticals and languages. | Chen et al., arXiv:2509.08919, 2025 (single large-scale study, not yet replicated). |
| contested | The exact mechanism (retrieval rewarding corroboration over self-assertion) and any fixed dose-response between earned placements and AI visibility. | Inferred from architecture and results; not isolated experimentally; no published dose-response curve. |
Reference
Glossary
- Earned media
- Coverage or mention won on merit from independent third parties, a journalist, a directory, a reviewer, an outside authority, without paying for the placement.
- Owned content
- Everything a business publishes and controls itself: its website, landing pages, blog, and its own social accounts. A first-party claim.
- Generative engine optimization (GEO)
- The practice, and the academic field, of shaping how a business is retrieved and cited inside the synthesized answers AI engines return rather than ranked in a list of links.
- Retrieval-augmented generation (RAG)
- The architecture behind most answer engines: a language model paired with a searchable index of outside documents queried at answer time, so answers are assembled from retrieved sources.
- A measure of how often a business is actually named or cited across a panel of real buyer questions in each AI engine, as distinct from where it ranks in classic search.
Straight answers
Frequently asked questions
Does earned media really matter more than my own website for AI search?
A 2025 large-scale study reports that generative engines cite third-party authoritative sources at a systematically higher rate than brand-owned pages, more so than classic Google. That is an emerging finding from a single study, not settled law, but it converges with the peer-reviewed 2024 GEO experiment, which found citing authoritative sources was the strongest lever for AI visibility. The read is a rebalance toward earned authority, not abandoning your site, which still anchors the entity everything else points at.
What is the difference between earned media and owned content in AI search?
Owned content is what you publish about yourself and control: your site, pages, and social accounts. Earned media is what independent parties say about you without being paid to. Answer engines appear to weight the second kind more heavily, because a source with no incentive to flatter you is harder to manufacture and carries more evidential weight.
Is generative engine optimization the same as SEO?
No. Classic SEO optimizes for a position in a ranked list of links, a surface a business can move largely on the strength of its own pages. Generative engine optimization concerns being retrieved and cited inside a synthesized answer, which the current evidence suggests depends more on third-party corroboration and entity authority than on owned pages alone. The two are related but decoupling into separate disciplines.
If I get cited by an AI engine, does that bring traffic?
Not reliably. Pew Research Center found that when an AI summary appears, people click a traditional result in roughly 8 percent of searches versus about 15 percent without one, and click inside the summary only around 1 percent of the time. A citation wins presence in the answer, which is the outcome that now matters most on that surface, but it is not the same as a click, and organic sessions are the wrong yardstick for it.
How would I know whether AI engines cite third-party sources about my business?
You have to measure it directly, because no engine publishes this. A structured read samples a panel of your real buyer questions across each engine and records how often you are named and which sources the answer draws on. That reading is the starting point before any earned-media work is scoped.
Provenance
Sources
- Chen, M., Wang, X., Chen, K., Koudas, N., "Generative Engine Optimization: How to Dominate AI Search", arXiv:2509.08919, 2025 (emerging, single large-scale study)arxiv.org
- Aggarwal, P., Murahari, V., Rajpurohit, T., Kalyan, A., Narasimhan, K., Deshpande, A., "GEO: Generative Engine Optimization", ACM SIGKDD 2024, arXiv:2311.09735 (established, peer-reviewed)arxiv.org
- Lewis, P. et al., "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks", arXiv:2005.11401, NeurIPS 2020 (established)arxiv.org
- Gao, Y. et al., "Retrieval-Augmented Generation for Large Language Models: A Survey", arXiv:2312.10997, 2023-24 (established)arxiv.org
- Attribution and citation-faithfulness literature, arXiv:2409.11242, 2024, and related 2025 attribution survey (established as the faithfulness gap; emerging on mitigations)arxiv.org
- Pew Research Center, "Do people click on links in Google AI summaries?", July 2025 (established)pewresearch.org
- Google Search Central, Search Quality Rater Guidelines and "E-A-T gets an extra E for Experience", December 2022 (established; on authoritativeness as a human-rater evaluation, not a direct ranking factor)developers.google.com
Every figure above is attributed to a real, dated source and tagged with its evidence tier. Where a claim could not be verified to a primary source, it is not stated as fact.