The Macro Shift · established evidence
Who Gets Paid When the Machine Answers: The New Economics of AI Licensing
For most of the web's history, a publisher's business model rested on a simple exchange: a search engine sent a reader to a page, and the page carried an advertisement that paid for the visit. An AI answer engine breaks that exchange. It reads the page, folds the sentence into an answer, and the reader often never clicks through. Out of that collapse a second economy has formed: direct payment for the right to train on or quote a publisher's work. The New York Times chose to sue, filing against OpenAI and Microsoft in December 2023 over millions of articles it says were used without consent or payment. News Corp chose to license, signing a deal with OpenAI in 2024 worth more than $250 million over five years. Reddit signed two separate deals worth roughly $130 million a year combined. The pattern is not random. A small cluster of American AI labs, OpenAI, Google, Microsoft, and Anthropic, has become the effective sole buyer of the world's written content, and buyer power now decides who gets paid and who gets nothing.
The lawsuit that named the terms
For most of the web's history, a publisher's business model rested on a simple exchange. A search engine sent a reader to a page, and the page carried an advertisement that paid for the visit. An AI answer engine changes that exchange at the root: it reads the page, folds a sentence from it into a generated answer, and the reader who receives that answer frequently never visits the original page at all. The click that used to fund the article is no longer guaranteed to happen.
The dispute over what that change owes the publisher reached a court on December 27, 2023, when The New York Times filed suit against OpenAI and Microsoft in the Southern District of New York. The complaint alleged that millions of the paper's copyrighted articles had been used to train ChatGPT without consent and without payment, and it named Microsoft as a defendant alongside OpenAI because of the financial and technical partnership between the two companies. It was among the first major publisher lawsuits of its kind, and it set the terms the rest of the industry would argue over for the next two years: was training a language model on a copyrighted article a research use, protected the way earlier courts protected search engines that indexed the same page, or was it a use that required a license and a payment, the way republishing the article itself would?
OpenAI answered in public. In a blog post on January 8, 2024, the company called the Times' claims 'without merit,' arguing that training a model on text already available on the open internet fell within the fair use doctrine that has long let search engines, archives, and researchers work with copyrighted material without paying for each use. That argument, that mass ingestion of public text for machine learning is a research use rather than a competing republication, remains the central legal position of every AI lab now facing a similar suit. Whether it holds is not yet settled, and the uncertainty is exactly what pushed a second track to open alongside litigation: instead of arguing the point in court, a publisher could simply get paid.
The price of being licensed
OpenAI answered that opening with money as often as with argument. In 2024 the company signed a licensing deal with News Corp, Rupert Murdoch's publishing group, worth more than $250 million over five years, one of the largest agreements disclosed in the new market. The deal covers content from The Wall Street Journal, the New York Post, and the Daily Telegraph in the UK, and it lets OpenAI train on and cite that reporting inside ChatGPT under contract rather than under contest. For a publishing group that owns some of the highest-authority mastheads in English-language business and political reporting, the deal converted a legal exposure into a revenue line.
Reddit found a similar opening from the other direction, licensing the same underlying asset, its archive of user discussion, to two buyers at once. In early 2024 the company signed a deal with Google reported at roughly $60 million a year to train Gemini, and separately reached an agreement with OpenAI reported at roughly $70 million a year, a combined figure close to $130 million a year flowing to one company from two AI labs. Reddit's position was structural: its forum threads capture the kind of first-person, conversational text that formal news writing rarely produces, and that a chat-style model needs in order to sound like a person rather than a press release.
A third deal, reported in 2026, showed the same publisher can keep collecting. Meta signed its own AI content-licensing agreement with News Corp, valued at roughly $50 million, a figure still emerging in the reporting rather than independently confirmed at the level of detail attached to the OpenAI deal. Stacked against the News Corp-OpenAI number, it suggests a publisher with a large enough archive and enough legal standing can sell training rights to more than one buyer without exhausting the asset, the way a single piece of music can be licensed to several films.
Across the market as a whole, disclosed deal values have ranged from roughly $5 million to $250 million per agreement between 2023 and 2026, according to a market tracker published by Quartz. That is not a price for a unit of content; it is closer to a price for a relationship, set by a publisher's archive depth, its legal standing, and how badly a given AI lab wants its brand associated with a credible newsroom rather than an anonymous scrape. No visible market rate has formed. Each deal is negotiated from a blank page.
The publishers who chose to sue instead
Not every publisher got a call from an AI lab's business-development team, and not every publisher that did took the deal. News Corp's own Dow Jones and New York Post sued Perplexity AI on October 21, 2024, alleging what the complaint called 'massive' and 'illegal' copying of copyrighted articles by the AI search company. The suit was amended on December 11, 2024, to add claims of trademark dilution, arguing Perplexity's habit of attributing fabricated or altered content to Dow Jones and Post bylines damaged the mastheads' names directly, not just their revenue.
Reddit, one of the best-paid publishers in the new licensing market, sued a different AI lab entirely in June 2025. The complaint against Anthropic alleged the company scraped and used Reddit's user data to train its models without ever signing a license, the same conduct Reddit had already negotiated a paid agreement to permit for Google and OpenAI. The suit is a clean natural experiment: the same content, the same publisher, one buyer that paid and another the publisher says did not, and only one of the two ended up in court.
What links the Perplexity and Anthropic suits is not the publisher; it is the absence of a deal. A company that scrapes without an agreement inherits exactly the legal exposure OpenAI tried to defend against in the New York Times case, and inherits it from every publisher with the resources to notice the scraping and file. The publishers without those resources, the great majority of the web's writers, sit in a third category this reporting does not directly measure: neither paid nor suing, because most have neither the archive depth nor the legal budget to be worth negotiating with or worth defending against in court.
The same masthead, two strategies
News Corp's own record makes the cleanest case that this is not a divide between publishers who cooperate and publishers who resist. The same company signed licensing deals with OpenAI and Meta while its Dow Jones and New York Post units sued Perplexity, all inside the same eighteen-month window. The company did not choose a single posture toward AI companies; it negotiated a separate answer for every buyer, license where a buyer offered a fair price, sue where a buyer did not offer one at all.
An industry tracker following the licensing market through 2026 describes the wider pattern the same way: publishers have split into a license camp and a sue camp, but the split runs by relationship rather than by company. A publisher licenses the buyer willing to pay and sues, or waits to sue, the buyer that is not. The strategy variable is not the publisher's philosophy about AI; it is whether the specific company on the other side of the negotiation chose to make an offer.
The same publisher takes both positions inside a single year, which rules out reading this as a referendum on whether training AI on news content is acceptable. What is actually happening is a negotiation, buyer by buyer, over whether a company with the scale to train a frontier model would rather pay a licensing fee or risk a multi-year lawsuit and an uncertain fair-use ruling. So far, the largest labs have often chosen to pay at least some publishers, which is itself a signal about how seriously they read their own legal exposure.
A four-company buyer's market
Every deal recorded here sits on one side of the table with a strikingly small set of names. OpenAI, Google, Microsoft, as OpenAI's primary financial partner and co-defendant, and Anthropic account for essentially the entire disclosed market for AI content licensing and for the lawsuits that stand in for it. All four are headquartered in the United States. The sellers on the other side of these deals include a British-Australian-American media conglomerate, an American internet forum, and, in the lawsuits, the same set of publishers plus whichever outlet a given AI company scraped without asking.
That asymmetry is the geopolitical shape of the market, not an incidental fact about it. A newspaper in India, a broadcaster in Brazil, or a wire service in continental Europe has no domestic AI lab of comparable scale to negotiate with instead. Its practical choices narrow to the same three this reporting has traced inside the US market: license to one of four American buyers, sue one of the same four American buyers in an American court, or absorb the scraping with no forum and no payment at all. The News Corp deal already reaches across that boundary once, since it folds in the Daily Telegraph, a British masthead, inside a contract negotiated and priced in the United States.
Every dominant medium in this history has done both things at once: it has opened a new channel for a shortlist of participants and concentrated authority over everyone else. The licensing deals genuinely open a new channel for the publishers who land one, handing News Corp and Reddit a real, disclosed, non-advertising revenue line that did not exist three years ago. The same deals concentrate the terms of a global content trade, price, scope, and payment schedule, inside four boardrooms that set each contract unilaterally, publisher by publisher, with no visible common rate and no publisher-side coalition large enough to negotiate back.
What the courts have not settled
The legal question underneath every one of these deals is still open. A federal judge in the Southern District of New York ruled on April 4, 2025, allowing the core copyright claims in the Times' suit against OpenAI to proceed, rejecting OpenAI's motion to dismiss on several counts. That ruling is procedural. It means the case survived an early attempt to end it, not that a court has decided whether training a model on copyrighted text without a license is fair use. The Reddit suit against Anthropic, filed two months later, raises the identical unresolved question from a different publisher's chair.
Until one of those cases, or another like it, reaches a final verdict, the licensing deals already signed function as a private answer to a question the public court system has not yet given. News Corp, Reddit, and the other publishers with disclosed deals in the $5 million to $250 million range are effectively setting a market price for AI training rights ahead of any settled law, one contract at a time, while every publisher without a comparable deal waits on a ruling that could still go either way.
The dispute traced here sits one step upstream of a second question this publication tracks directly: when an answer engine names a source inside a generated answer, the reader sees the name, not whether the source behind it was ever paid to be there. The lawsuits and licensing deals in this record settle, or attempt to settle, who gets compensated for supplying the material a model trains on. They do not settle who a model chooses to cite once it is answering a live question, a related decision the industry has so far negotiated in public far less than it has negotiated the first.
The evidence
Key findings, with their sources
-
The New York Times filed suit against OpenAI and Microsoft in the Southern District of New York on December 27, 2023, alleging millions of its copyrighted articles were used without consent or payment to train ChatGPT.
established NPR, drawing on court filings (2023 to 2025).
-
OpenAI's licensing deal with News Corp, announced in 2024, is worth more than $250 million over five years, covering The Wall Street Journal, the New York Post, and the Daily Telegraph.
established AI Business reporting (2024).
-
Reddit signed a content-licensing deal reported at roughly $60 million a year with Google to train Gemini, and a separate deal with OpenAI reported at roughly $70 million a year, close to $130 million a year combined.
established Columbia Journalism Review reporting (2024).
-
News Corp's Dow Jones and New York Post sued Perplexity AI on October 21, 2024, alleging 'massive' and 'illegal' copying, with an amended complaint filed December 11, 2024 adding trademark dilution claims.
established CNBC reporting (2024).
-
Reddit sued Anthropic in June 2025, alleging the AI company scraped and used Reddit user data to train its models without a license, the same conduct Reddit had separately licensed to Google and OpenAI.
established Bloomberg reporting (2025).
-
Meta signed an AI content-licensing agreement with News Corp reported at roughly $50 million.
emerging MediaCopilot reporting (2026).
-
Disclosed AI content-licensing deal values have ranged from roughly $5 million to $250 million per agreement across the 2023 to 2026 market, with no visible common price forming.
established Quartz market tracking (2026).
-
OpenAI's public response to the Times suit, posted January 8, 2024, argued the case was 'without merit' and that fair-use doctrine protects training on publicly available text.
established Harvard Law Review, citing OpenAI's blog and Euronews coverage (2024).
-
A federal judge allowed the core New York Times copyright claims against OpenAI to proceed in a ruling issued April 4, 2025, rejecting OpenAI's motion to dismiss on several counts, a procedural result that leaves the fair-use question undecided.
established The Hollywood Reporter, citing the SDNY court opinion (2025).
Calibration
What is proven, what is promising, what is unproven
| Evidence tier | Tactics | What the evidence says |
|---|---|---|
| established | The direct-licensing payments confirmed in public reporting (News Corp's OpenAI deal, Reddit's combined Google and OpenAI deals) and the lawsuits publishers have filed against AI companies that did not offer a deal. | Corroborated across court filings and multiple independent outlets, including NPR, AI Business, Columbia Journalism Review, CNBC, Bloomberg, and The Hollywood Reporter; not dependent on any single source. |
| emerging | The Meta-News Corp deal figure of roughly $50 million, and the broader read that a two-tier economy of licensed-few and scraped-many is forming industry-wide. | The Meta figure comes from a single reporting source not yet corroborated at the level of detail attached to the OpenAI-News Corp deal; the two-tier framing is a pattern read across a market where most deal terms remain undisclosed. |
| contested | Whether training a model on copyrighted text without a license constitutes fair use under US copyright law. | OpenAI argues it does; The New York Times and Reddit argue it does not. A federal judge allowed the Times' core claims to proceed on April 4, 2025, but that is a ruling on a motion to dismiss, not a decision on the merits, so the underlying legal question remains open. |
Reference
Glossary
- Direct licensing deal
- A commercial agreement in which an AI company pays a publisher directly for the right to train on, or cite, its content, instead of relying on unlicensed scraping.
- Fair use
- A doctrine in US copyright law that permits limited use of copyrighted material without permission in some circumstances, the central legal defense AI labs have raised against publisher lawsuits over training data.
- Motion to dismiss
- An early request asking a court to end a lawsuit before trial; a judge denying the motion lets the case proceed but does not decide who wins.
- Trademark dilution
- A legal claim that a company's use of another's brand name or masthead, for example attributing fabricated content to a real publication's byline, damages that brand's distinctiveness or reputation.
- Two-tier publishing economy
- A market structure, visible in AI licensing since 2023, in which a small number of publishers with archive scale and legal standing negotiate paid deals while most publishers are scraped without payment or a clear legal remedy.
Straight answers
Frequently asked questions
Why did The New York Times sue OpenAI instead of signing a licensing deal?
The Times filed suit on December 27, 2023, alleging millions of its articles were used to train ChatGPT without consent or payment. Suing, rather than negotiating, let the paper press the underlying legal question, whether training on copyrighted text without a license is fair use, directly in court instead of settling it privately for an undisclosed fee.
How much are AI companies paying publishers for content?
Disclosed deal values have ranged from roughly $5 million to $250 million per agreement between 2023 and 2026. OpenAI's deal with News Corp is worth more than $250 million over five years, Reddit's combined deals with Google and OpenAI are worth roughly $130 million a year, and a separate Meta-News Corp deal is reported at roughly $50 million.
Has a court decided whether training AI on copyrighted content is fair use?
No. A federal judge allowed The New York Times' core copyright claims against OpenAI to proceed on April 4, 2025, but that ruling only rejected OpenAI's motion to dismiss; it did not rule on the merits. The fair-use question at the center of the dispute remains legally unresolved.
Why did News Corp both license content to AI companies and sue one of them?
News Corp licensed to OpenAI and Meta, buyers that offered payment, while its Dow Jones and New York Post units sued Perplexity, a company it says copied its content without ever making an offer. The company's strategy tracks each buyer's willingness to pay, not a single company-wide position on AI.
Why does AI content licensing matter beyond the publishing industry?
Because the buyers on the other side of nearly every disclosed deal, OpenAI, Google, Microsoft, and Anthropic, are American companies, publishers on every continent effectively negotiate the terms of the global written-content trade with the same small cluster of US-headquartered firms, whether or not they ever sign a deal.
Provenance
Sources
- NPR, on The New York Times' lawsuit against OpenAI and Microsoft, drawing on court filings (2023 to 2025)npr.org
- AI Business, "OpenAI Inks Licensing Deal With News Corp for ChatGPT Training Data" (2024)aibusiness.com
- Columbia Journalism Review, on Reddit's AI licensing deals with OpenAI and Google (2024)cjr.org
- CNBC, "Murdoch firms Dow Jones and New York Post sue Perplexity AI" (2024)cnbc.com
- LLM Pulse, industry tracker of AI content licensing deals (2024 to 2026)llmpulse.ai
- Bloomberg, "Reddit Sues Anthropic, Says AI Company Exploited User Data" (2025)bloomberg.com
- MediaCopilot, on the Meta-News Corp AI content licensing deal (2026)mediacopilot.ai
- Quartz, market tracking of AI training data pricing and licensing deals (2026)qz.com
- Harvard Law Review, "NYT v. OpenAI: The Times's About-Face," citing OpenAI's public response (2024)harvardlawreview.org
- The Hollywood Reporter, on the SDNY ruling advancing The New York Times' lawsuit against OpenAI (2025)hollywoodreporter.com
Every figure above is attributed to a real, dated source and tagged with its evidence tier. Where a claim could not be verified to a primary source, it is not stated as fact.