Market evidence

What data actually
sells for.

"Would anyone really pay for our data?" is the first question every owner asks. Here is the public record. These are disclosed, reported transactions between data owners and AI companies. None of them are ours, and that is the point: this market exists with or without us.

News Corp
$250M+ / 5yrs
Shutterstock
$104M in 2023
Reddit
~$60M / yr
Wiley
$44M
Springer Nature
~$23M
Taylor & Francis
$10M+

Reported deal values, publicly disclosed 2023 to 2025

SellerDealBuyerReported value
Reddit Licensing of user discussion data for AI training, announced ahead of its IPO in early 2024 Google ~US$60M per year
News Corp Multi-year licence covering news content across its mastheads, May 2024 OpenAI US$250M+ over 5 years
Shutterstock Image, video and metadata licensing to multiple AI developers; anchor customers reported at ~$10M a year each Meta, Alphabet, Amazon, Apple, OpenAI US$104M in 2023 alone
Wiley Academic content rights for AI model training, disclosed across FY2024 and FY2025 Undisclosed tech companies US$44M across two deals
Taylor & Francis (Informa) Data access agreement covering academic content, May 2024 Microsoft US$10M+ initial
Springer Nature One-time licence over a defined corpus of published papers, 2024 Google ~US$23M
Axel Springer News content licensing and product integration, December 2023 OpenAI Tens of millions over 3 years

Sources: deal values as publicly reported by Reuters, Bloomberg, the Wall Street Journal, The Bookseller and company disclosures, 2023 to 2025. Figures are the reported values at announcement; some parties have not confirmed exact terms. This page reports the market; it is not financial advice.

What it means for you

Three things this table tells you.

1. The buyers are the largest companies on earth, and they are not stopping. Every frontier lab now runs a formal buying program for licensed data. OpenAI operates a public data partnerships program; Google negotiates content and data licences directly. Estimates put total lab spending on training data above US$10B a year and rising, because quality data, not compute, has become the binding constraint on model progress.

2. Nobody in this table is a data company. A newspaper group, a forum, a stock photo library, three academic publishers. Each was sitting on years of accumulated content and operational history that suddenly acquired a market. The same shift is now reaching business operational data: how companies actually communicate, decide and execute is exactly what the next generation of AI needs and cannot find on the public internet.

3. Clean rights command the premium. Every deal above exists because scraping is no longer defensible. Buyers pay for licensed, warranted, consented data because it removes legal risk from their models. That is precisely what Australian supply offers by default, and why we operate here.

Your business is smaller than Reddit. Your data is also something Reddit does not have: the complete private record of how a real company operates. Different asset, same market.

Where would yours sit on this table?

The assessment gives you a defensible range, grounded in a market we have transacted in ourselves.

Get a value assessment