Microsoft Staff Asked If AI Scraping Was 'Largest Theft of Labor in Human History'
Court filings unsealed Thursday in the New York Times' copyright lawsuit against OpenAI and Microsoft reveal internal warnings at both companies about the impact of using news content to train large language models. Microsoft memos and OpenAI employee messages described the practice as tantamount to mass appropriation of journalists' work and flagged risks that the models could degrade by ingesting the very output they were meant to synthesize.

Why It Matters
The documents illuminate internal concerns at two of the industry's largest AI developers about legal, ethical and technical consequences of training on copyrighted news content, material directly at issue in ongoing litigation that has drawn in multiple publishers and remains before a federal judge.
Key Facts
- Source of documents: Unsealed filings from the New York Times' copyright suit against OpenAI and Microsoft
- Date of suit: Filed in late 2023
- Other plaintiffs: Case later joined by eleven other publishers
- Judge: Judge Sidney Stein, Southern District of New York
- Preserved data: OpenAI has been ordered to preserve 20 million ChatGPT conversation logs
Documents unsealed as part of the New York Times' copyright case against OpenAI and Microsoft reveal internal debate and alarm over the companies' use of news content to train large AI models. A 2023 Microsoft memo, attributed to Brent Hecht, warned that the public would view models that "hoover up" creators' work as an unprecedented theft and argued that large models could "destroy" their own information supply chain. Microsoft told the court the memos reflected one employee's perspective and not official company policy.
The filings also record candid exchanges among OpenAI staff. An internal message reported by the Times quoted company employees asking whether using news articles amounted to "the largest theft of labor in human history," and expressing concern that feeding models on published work could create a "doom loop" that eroded output quality. Other notes from OpenAI personnel warned that AI systems were becoming substitutes for cultural labor and could harm the industries they relied on.
Microsoft CEO Satya Nadella is reported to have testified that paywalled content should be licensed by whoever wants to use it, and said he would have sought retraining of OpenAI models had he known paywalled material was used. The filings also include references to attempts by some OpenAI staff to circumvent paywalls and to internal observations that chatbots do not reliably send traffic back to publishers — a point relevant to the companies' arguments about substitution versus complementarity.
Both Microsoft and OpenAI have maintained that their training practices fall within fair use, arguing that models transform source material rather than merely replicating it. Plaintiffs and their lawyers say the newly unsealed documents reveal what the companies themselves thought about the legality and propriety of their behavior. The case is advancing through summary judgment motions, and additional materials have been unsealed as the court considers those filings.
Keep Reading

This cartridge-playing Game Boy clone is smaller and cheaper than Analogue’s Pocket

Lenovo’s Yoga Slim 7X is the most laptop that $1,000 can currently buy
