← All dispatches
Dispatches · #intelligence · ScrapeOps

AI Training Data Licensing Hit $49M for Wiley in FY26

August 5, 2026 · Abhishek Gupta
Infographic showing Wiley's AI training data licensing revenue growth: $49 million in fiscal 2026, up 23 percent, versus $250 million News Corp OpenAI deal

Wiley made $49 million licensing its books and journals to AI companies in fiscal 2026 — up 23% from the year before, while the rest of the 218-year-old publisher's revenue barely moved. AI training data licensing just became the fastest-growing line on the income statement of a company that mostly sells academic textbooks.

That's not a scraping story. It's the other half of the same fight: while Cloudflare and publishers write AI bots out of robots.txt, a smaller group of content owners are getting paid directly, in dollars, for the same underlying asset — clean, structured, comprehension-ready text.

The short version

  • Wiley's AI and Data Analytics revenue hit $49 million in fiscal 2026 (reported June 16, 2026), up 23% year-over-year, while total company revenue held flat at $1.677 billion.
  • Wiley's cumulative lifetime AI licensing revenue crossed $110 million during fiscal 2026 — most of it earned in the last two fiscal years.
  • News Corp's licensing deal with OpenAI, signed in May 2024 and covering the Wall Street Journal, Barron's, MarketWatch, and a dozen other titles, is worth up to $250 million over five years — roughly $50 million a year.
  • Wiley signed a separate fiscal 2026 agreement with Anthropic to put peer-reviewed content directly into Claude, alongside new partnerships with IQVIA and OpenEvidence in healthcare AI.
  • Both deals route around the open-web scraping fight entirely: the content is delivered under contract, never crawled, blocked, or argued about in a robots.txt file.

Why Is a Publisher's Fastest-Growing Line Item AI Licensing?

Because it's nearly pure margin on content Wiley already owns and has already digitized. There's no new printing, no new distribution, no new sales team — just a contract that lets an AI company ingest a catalog Wiley built over two centuries.

That's why a segment worth roughly 3% of Wiley's $1.677 billion in FY26 revenue accounted for effectively all of the company's growth. Total revenue was flat. AI licensing grew 23%. Every dollar of expansion the company reported came from one place.

Is $49 Million Actually a Big Number?

Relative to Wiley's own book, yes — it's the difference between a flat year and a growth year. Relative to the wider market, it's still early: News Corp's OpenAI deal alone runs at roughly the same $50 million-a-year pace, and that's a single deal at a single publisher.

Scale that comparison out and the pattern holds across the industry. Wiley, an academic and professional publisher, is pulling in AI revenue at nearly the same annual run-rate as News Corp, a general news publisher with a much larger audience. That suggests the ceiling on per-publisher AI licensing revenue isn't audience size — it's how structured, current, and legally clean the underlying dataset is. Wiley's peer-reviewed journals and reference works are exactly that: verified, versioned, and already organized by discipline.

How Does the Wiley-Anthropic Deal Change the Scraping Calculus?

It removes the incentive to scrape Wiley at all. Anthropic gets Claude's answers grounded in peer-reviewed sources through a direct feed instead of a crawler, and Wiley gets paid per the terms of the contract instead of showing up as an unpaid line in someone's training run.

That's the quiet argument licensing deals are making to every other AI company still running crawlers against paywalled or rate-limited sites: the legally clean version of this transaction is available, it has a price, and increasingly, publishers would rather sell it than fight about it in court or in robots.txt.

What Does This Mean If You're Not Wiley or News Corp?

Almost nobody building an AI product gets to sign a $49 million licensing deal. Most of what a RAG pipeline or a training run actually needs isn't locked inside a handful of publisher catalogs — it's scattered across thousands of smaller sources, some licensable, most just publicly available and messy: duplicate coverage of the same story, five versions of the same regulatory filing, stale pages nobody's re-crawled in months.

Licensing solves the legal and reputational risk for maybe a dozen large content owners a year. It doesn't solve the acquisition problem for everyone else, which is still: find the right sources, deduplicate them, and turn them into something a model can actually use. We built ScrapeOps for exactly that gap — the part of the data supply chain that stays hard even after the licensing headlines move on. More on how we think about the wider data-acquisition problem is in our research.

Where This Goes Next

Wiley's fiscal 2027 guidance calls for "another strong year in AI and data analytics" without a specific number attached, and its recurring AI revenue — $8 million of the $49 million — is expected to double or triple next year on its own. If that holds, licensing stops being a one-off headline deal and starts looking like a real, repeating revenue category for content owners with the right kind of catalog.

The publishers who get there first will be the ones whose content was already structured enough to license in a week instead of a year. Everyone else is still choosing between getting scraped for free, blocking crawlers outright, or building the deduplication and structuring pipeline no one else will build for them. We cover more of this in our regular dispatches.

Frequently Asked Questions

How much did Wiley make from AI licensing in fiscal 2026? Wiley's AI and Data Analytics segment generated $49 million in fiscal 2026, up 23% year-over-year, according to the company's June 16, 2026 earnings release. Cumulative lifetime AI revenue crossed $110 million during the same fiscal year.

How much is the News Corp-OpenAI licensing deal worth? News Corp's May 2024 deal with OpenAI is worth up to $250 million over five years — roughly $50 million a year — and covers titles including the Wall Street Journal, Barron's, MarketWatch, and the New York Post.

What does Wiley's deal with Anthropic cover? Wiley's fiscal 2026 agreement with Anthropic puts the publisher's peer-reviewed academic and professional content directly into Claude, giving Anthropic a licensed, structured source instead of relying on crawling Wiley's paywalled journals.

Does AI content licensing replace the need for web scraping? No. Licensing works for a small number of large publishers with clean, structured catalogs. Most of the data AI teams need is scattered across smaller, unlicensed sources that still have to be found, deduplicated, and structured before a model can use them.

Abhishek Gupta is Co-Founder at Dekrypt Labs, building ScrapeOps — the data acquisition engine that turns any question into clean, deduplicated, comprehension-ready sources. dekryptlabs.com