Unsealed Briefs Show OpenAI and Microsoft Knew They Used Pirated Books

New court filings reveal OpenAI executives knowingly trained models on pirated books and actively tried to cover their tracks in 2022.

Drama ยท Source: Hacker News

What happened

Unsealed court documents in the Authors Guild lawsuit against OpenAI and Microsoft just dropped. The filings reveal internal communications from top executives at both companies. They show leadership knew they were training models on pirated books from LibGen as early as 2019. Sam Altman and Dario Amodei reportedly presented an early version of GPT-3 to Bill Gates and Microsoft Chief Technology Officer Kevin Scott. During this meeting, they explicitly disclosed the use of LibGen.

The internal culture at OpenAI was blunt about the consequences for human creators. In May 2020, OpenAI Policy Director Jack Clark stated their work would substitute human labor and make people unemployed. He noted they would likely ignore artist concerns and release products anyway. Another researcher, Tarun Gogineni, was hired in 2022 to improve writing quality. He called author complaints acceptable economic disruption. Gogineni even joked about autocompleting George R.R. Martin unfinished A Song of Ice and Fire series with GPT-5. He predicted the death of the reader as machines created slop for other machines.

When public scrutiny increased, OpenAI actively tried to hide the evidence. In the summer of 2022, executives launched a cover-up called Project Clear to delete LibGen files from their systems. Internal Slack messages show researchers were worried about the bad optics of using a sketchy Russian website. Sam McCandlish specifically worried about the news showing up on Hacker News. Vice President of Research Bob McGrew pushed to excise LibGen from their storage because OpenAI was in the news too much.

Key facts

Why it matters

This shatters the standard defense that AI companies accidentally ingested copyrighted data while scraping the open web. For founders building AI products, the legal shield of fair use is getting heavily tested by these revelations. If courts rule that OpenAI intentionally engaged in mass piracy and actively covered it up, the foundation models we all rely on could face massive legal consequences. We might see forced data deletion or astronomical licensing fees paid out to publishers and authors. The era of scraping first and asking for forgiveness later is officially closing.

The second-order effect here is all about enterprise trust and regulatory backlash. The optics of Project Clear and executives laughing about acceptable economic disruption will invite brutal regulatory scrutiny from lawmakers. Enterprise customers care deeply about compliance and legal liability. If foundation models are deemed legally toxic due to stolen training data, large corporations will hesitate to integrate them into their workflows. This creates a massive wedge for open-source models trained on provably clean, licensed data. Startups that can guarantee copyright-free AI will easily steal enterprise market share from OpenAI.

For builders

Audit your training data immediately

If you train custom models or fine-tune existing ones, you must document your data sources right now. Ignorance is no longer a valid legal defense in court. Enterprise buyers will demand ironclad proof that your datasets are clean before signing any contracts. Founders who ignore compliance will lose lucrative enterprise deals to competitors who prioritized data provenance.

Prepare for API cost spikes

If OpenAI is forced to settle this lawsuit and pay massive licensing fees to the Authors Guild, they will inevitably pass those costs down to developers. Builders relying entirely on cheap GPT-4 or GPT-5 APIs should model their unit economics for sudden, steep price hikes. You need a fallback plan with open-source models to protect your margins.

My take

Founders love to move fast and break things, but deleting evidence in a company Slack channel is just amateur hour. You cannot build a trillion-dollar enterprise business while acting like a pirate startup hiding from Hacker News. Clean data is about to become the ultimate competitive moat in the AI industry.

Original reporting: Hacker News. This is my rewrite and opinion.

More AI news for builders