The world's most valuable resource used to be oil — now it's whatever's sitting in your old email server. AI labs are racing to feed their models with real-world corporate data, and the appetite is insatiable. The result is a fast-emerging market where bankruptcies, bookshelves, and social media threads are all suddenly worth serious money.
Feeding the beast: The AI data licensing race kicked into view when Google won a bankruptcy auction to buy Spirit Airlines’ software code, internal messages, and financial and operational data for $10M. The deal points to where AI companies are looking next for the massive volumes of human-generated information their models need to mimic real-world processes. With scraping the open internet becoming more legally fraught, corporate archives are emerging as the next frontier.
- Reddit brings in ~$60M annually from each of its OpenAI and Google licensing deals, with a larger Google agreement now in talks.
- John Wiley & Sons has generated over $110M in AI licensing revenue since 2024 through deals with leading AI developers.
Jumping In on the Data Hunt
It's not just corporate email threads on the menu. Amazon is buying rare books at scale, scanning them, and destroying the originals in Las Vegas warehouses, joining Anthropic and Meta in the race for high-quality training data. Books published before 2022 are particularly valuable because they predate the explosion of AI-generated content. Training models on too much synthetic material can degrade their outputs over time, a phenomenon known as “model collapse.”
- Anthropic’s “Project Panama” was a secretive effort to mass-scan books for AI training data without drawing public attention.
- Micro1 committed over $20M to license operational data in 11 days, while Cloudflare acquired data marketplace Human Native.
Data meets privacy: A bankruptcy court delayed Google's Spirit purchase after former flight attendants argued that deidentifying records doesn’t fully protect sensitive payroll files, tax forms, and time cards. That privacy tension could become a major hurdle for this market. Consumer data often has explicit protections, while employee data can fall through the cracks. Companies eyeing the licensing payday may find the legal and reputational bill comes due first.
