Microsoft Exec Calls AI Scraping 'Largest Theft of Labor in Human History'
Unsealed filings in the NYT lawsuit reveal Microsoft and OpenAI knew their models were an existential threat built on scraped data.
Policy · Source: Hacker News
What happened
The New York Times copyright lawsuit against OpenAI and Microsoft just got ugly. Unredacted filings reveal explosive internal communications from both companies. A top Microsoft executive privately called their AI training practices the largest theft of labor in human history. OpenAI leadership admitted their models pose an existential threat to the publishers whose work trained them.
The documents detail exactly how the tech giants built their datasets. OpenAI employees allegedly bypassed paywalls undetected to scrape content. They pulled millions of articles from Common Crawl and built custom datasets heavily reliant on news. Researchers even deliberately stripped copyright notices from the training data. They did this so the models would not output those notices to users.
Microsoft knew the damage it was causing. Internal data showed their Copilot answer engine caused click-through rates to the NYT domain to plummet by up to 93 percent compared to traditional Bing search. A Microsoft director called this decline a doom loop. He warned it would eventually hurt the performance of their own models by destroying the web. Satya Nadella even testified that paywalled content should be licensed. He claimed he would have forced OpenAI to retrain models if he knew they scraped paywalled data.
Key facts
- 93% — Maximum drop in click-through rates to NYT from Microsoft Copilot compared to Bing search
- 91,692 — Copies of works published by NYT, Daily News, and CIR in OpenAI mid-training datasets
- 2 million — Documents from nytimes.com included in a Common Crawl-derived dataset
- 160,903 — Unique works from news publishers in the Project Mango training dataset
Why it matters
The fair use defense is the bedrock of the current generative AI boom. These unsealed quotes directly undermine that defense. Fair use requires that a new product does not substitute or harm the market for the original work. When the head of ChatGPT internally calls the product largely substitutive, the legal shield cracks. Microsoft data proving a 93 percent drop in publisher traffic makes the substitution argument undeniable. If courts rule against OpenAI, every startup training models on scraped data faces immediate legal and financial ruin. The precedent would force a massive restructuring of how AI companies acquire data.
The content supply chain is breaking. Microsoft internally recognized that destroying the economic foundation of their data suppliers is a massive risk. If publishers die, the flow of high-quality human data stops. Builders relying on LLMs will eventually face model degradation. The web will fill with synthetic garbage instead of original reporting. You cannot train the next generation of reasoning models on AI-generated slop. The companies building these foundation models are cannibalizing the very creators they need to survive. This doom loop threatens the entire AI ecosystem.
For builders
Audit your training data pipelines immediately
If you are scraping paywalled content, stop right now. Satya Nadella testified that paywalled data must be licensed for grounding or training. Startups caught bypassing paywalls will face indefensible copyright claims and massive fines.
Prepare for mandatory licensing fees
The era of free data is ending. If courts reject the fair use defense for AI training, foundation model providers will pass massive licensing costs down to API users. Factor much higher intelligence costs into your future margins.
Build synthetic data generation engines
Relying on human web data is becoming a legal minefield. Founders who figure out how to train highly capable models on synthetic or fully licensed proprietary data will win the next decade. Start investing in data partnerships today.
My take
I always knew the AI data land grab was shady, but seeing it in writing is wild. You cannot build a sustainable business by bankrupting your supply chain. We need to stop pretending mass scraping is fair use and start paying for the data that makes our products work. If we kill the creators, we kill the models.
Original reporting: Hacker News. This is my rewrite and opinion.