Microsoft exec called AI scraping the largest theft in history
Unsealed court filings reveal Microsoft and OpenAI executives privately admitted their scraping practices were theft and bypassed paywalls.
Policy · Source: Hacker News
What happened
Unredacted court filings in The New York Times copyright lawsuit against OpenAI and Microsoft just dropped. They reveal damning internal communications that shatter the public narrative around fair use. Executives at both companies privately acknowledged the severe economic impact of their data scraping practices. Brent Hecht, a Microsoft director of applied science, called the mass ingestion of data an astonishing theft of unprecedented proportions and the largest theft of labor in human history.
The documents show a highly coordinated effort to ingest copyrighted data at scale. OpenAI employees reportedly shared hacks to bypass paywalls undetected. When told about a hack to get around the Times paywall, OpenAI President Greg Brockman reportedly replied with approval. Microsoft and OpenAI also swapped massive datasets through internal initiatives named Project Taxi and Project Mango. Furthermore, they deliberately stripped copyright notices from the training data so their models would not output those notices to end users.
Microsoft CEO Satya Nadella testified he would have demanded OpenAI retrain its models if he knew they scraped paywalled content. Yet internal Microsoft data showed their Copilot engine caused click-through rates to the Times to plummet by up to 93 percent compared to standard Bing search. OpenAI leadership, including head of ChatGPT Nick Turley, admitted their chatbots act as a direct, existential threat to publishers. They knew the product was largely substitutive and would only get more substitutive as the models improved.
Key facts
- 93% — Drop in NYT click-through rates via Copilot compared to Bing search
- 91,692 — Copies of works from NYT and others in OpenAI mid-training datasets
- 160,903 — Unique news works in the shared Project Mango training dataset
- 2 million — Documents from nytimes.com included in a Common Crawl-derived dataset
Why it matters
The fair use defense is the legal bedrock of the current generative AI boom. AI companies rely heavily on the argument that training models on copyrighted web data is transformative and does not harm the original creators. These unsealed quotes directly undermine that defense by proving the companies knew they were building a substitutive product. If courts decide that scraping paywalled data and replacing original sources violates copyright law, the cost of building foundational models will skyrocket overnight. Startups will face massive licensing fees just to acquire the training data they currently take for free.
We are looking at a potential collapse of the open web data pipeline. Microsoft executives internally warned of a doom loop where starving content creators hurts model performance and the web simultaneously. If publishers aggressively lock down their data or win massive settlements, the era of free scraping ends for good. Only tech giants with deep pockets will afford the licensing deals required to train state-of-the-art models. This creates a massive moat for incumbents while pulling the ladder up on open source builders and early stage founders.
For builders
Prepare for expensive data licensing
The days of scraping Common Crawl without consequence are rapidly ending. If courts reject the fair use defense based on these admissions, founders building proprietary models will need to pay for high-quality data. Start budgeting for licensing agreements now or risk devastating copyright lawsuits.
Paywalls are no longer safe havens
OpenAI reportedly used specific hacks to bypass publisher paywalls during their training runs. If you build content platforms or proprietary datasets, expect aggressive scraping attempts from AI agents. You must invest in advanced bot detection and legal terms of service to protect your proprietary data.
Search traffic is permanently altered
Microsoft internal data showed a staggering 93 percent drop in click-through rates for publishers when users queried Copilot instead of Bing. If your startup relies on inbound SEO traffic for customer acquisition, you need a completely new distribution strategy. AI answers are replacing traditional search clicks entirely.
My take
I am entirely tired of the blatant hypocrisy from big tech. You cannot build a trillion-dollar product by hacking paywalls and stripping copyright notices, only to turn around and claim you are just doing fair use research. If early stage founders pulled this exact same stunt, we would be sued into oblivion by corporate lawyers. We desperately need clear, enforceable rules on data ownership before the entire web turns into a barren wasteland of AI-generated slop.
Original reporting: Hacker News. This is my rewrite and opinion.