OpenAI and Microsoft knew they were starting a ‘doom loop’ for the web
Newly unsealed court documents in the New York Times’ lawsuit against OpenAI and Microsoft show internal warnings that large language model development could harm the web and publishing ecosystem. Company employees described extensive web scraping as tantamount to massive unpaid labor, warned the technology could supplant referral traffic to news sites, and used terms like a looming “doom loop.”

Why It Matters
The filings suggest that Microsoft and OpenAI were aware their data-collection and product designs risked reducing traffic and revenue for publishers while potentially degrading the content supply that models rely on. Those admissions are central to legal and policy debates over copyright, fair use, and how AI systems should be trained and deployed.
Key Facts
- Source: Recently unsealed court documents from the New York Times’ case against OpenAI and Microsoft (92-page filing excerpted by The Verge).
- Internal term used: An internal Microsoft document warned the companies’ AI content strategy had started a “doom loop.”
- Quote on labor: Microsoft’s Brent Hecht described the scraping as the “largest theft of labor in human history.”
- Fair use critique: Hecht said Microsoft’s defense made a “complete mockery of the idea of ‘fair use.’”
- Microsoft distancing: Spokesperson Alex Haurek said those comments reflect one employee’s individual perspective and do not represent the company’s views.
Court filings unsealed in the New York Times’ suit against OpenAI and Microsoft contain candid internal assessments suggesting the companies understood the risks their AI training practices posed to publishers and the broader web. Multiple employees and documents cited in the filing warn that large-scale scraping of internet content for model training could undermine the economic foundations of content creators and reduce direct traffic to source sites. Among the most striking remarks attributed to Microsoft staff, Director of Applied Science Brent Hecht is reported to have called the data harvesting the “largest theft of labor in human history” and argued that the company’s legal rationale turned the idea of fair use into a “complete mockery.” An internal Microsoft memo quoted in the filing warns that the firms’ AI content strategy had entered a “doom loop,” a situation in which the product both depends on and damages its content supply chain. OpenAI employees expressed related concerns about memorization and reproduction of copyrighted material: the filing quotes staff acknowledging that GPT-4 had “memorized a ton of data” and could reproduce source text verbatim, and it cites instances where the model returned long passages from outlets including the Times, Mercury News, The Denver Post, LifeHacker, and Eurogamer. The documents also attribute observations to OpenAI and Microsoft teams that AI-driven summaries and chat interfaces can displace clicks to original reporting, with some internal estimates suggesting search referrals may have fallen by as much as 60 percent for affected sites. Microsoft has sought to limit the reach of individual comments cited in the filings. Spokesperson Alex Haurek told media outlets that the quoted statements reflect a single employee’s perspective and are not a legal analysis or company position, and another Microsoft executive described the employee as expressing a divergent, academic viewpoint rather than speaking for the company. The filings nonetheless place those internal concerns alongside other remarks from senior figures at both firms and add new detail to ongoing legal and policy questions about training data practices, copyright, and how AI products interact with the online information ecosystem.
Keep Reading
Most of what you know about data centers is wrong
Amazon, Palantir and 12 more top tech stock picks from UBS analysts

Researchers used Claude to hack OpenAI
