Microsoft exec called AI scraping ‘the largest theft of labor in human history,’ new unredacted filings reveal
Newly unredacted material from The New York Times’ copyright lawsuit against OpenAI and Microsoft reveals internal company documents and communications characterizing the mass scraping of news content for AI training as tantamount to theft and warning that AI products threaten publishers. The filings allege the firms bypassed paywalls, stripped copyright notices from training data, and assembled datasets containing large numbers of Times and other publishers’ articles used to train large language models.

Why It Matters
The disclosures bear directly on central legal and economic questions in the suit—whether large-scale scraping of copyrighted journalism for model training constitutes fair use and whether AI products materially substitute for publishers’ original work, harming their markets and employment.
Key Facts
- Lawsuit: The New York Times filed a copyright lawsuit against OpenAI and Microsoft three years ago; new unredacted material was submitted in that case.
- Internal description of scraping: A Microsoft executive described the mass copying of publisher content as comparable to theft and a separate Microsoft document warned generative AI could disrupt employment of those who produced the training data.
- Impact on traffic: Microsoft data cited in the filings showed Copilot reduced click-through rates to nytimes.com by as much as 93% compared with traditional Bing search.
- Scale of copied works: OpenAI’s mid-training datasets reportedly contained over 91,692 copies of works from the NYT, Daily News, and Center for Investigative Reporting; Common Crawl-derived data reportedly included more than 2 million nytimes.com documents.
- Specific programs: The filings allege data-sharing initiatives named Project Taxi and Project Mango exchanged training datasets between OpenAI and Microsoft; Project Mango was said to include at least 160,903 unique works from the plaintiffs.
Unsealed portions of court filings in The New York Times’ copyright action against OpenAI and Microsoft present internal communications and documents that portray the companies’ content-acquisition practices for model training as deeply problematic. According to the new material, executives and researchers within Microsoft and OpenAI used language underscoring the moral and commercial stakes of large-scale scraping of news publishers’ work, with at least one Microsoft researcher characterizing the scope of copying as unprecedented theft. The filings allege concrete steps by the companies to acquire publisher material, including scraping content via the Bing index, creating datasets such as WebText and WebText2 heavily weighted toward news articles, and harvesting millions of pages from Common Crawl. They also describe purported efforts to bypass paywalls and to remove copyright notices from material before it was fed into training datasets, and they claim the two firms exchanged substantial training data through initiatives called Project Taxi and Project Mango. The documents cited in the filing point to measurable commercial effects: internal Microsoft analysis attributed dramatic drops in click-through rates to nytimes.com when users received Copilot answers instead of traditional search results. Company presenters warned that such substitution could damage both publishers’ revenues and the quality of the web content supply that underpins foundation-model businesses. OpenAI and Microsoft executives are also quoted as acknowledging that models are “substitutive” for publishers’ work and could pose an existential risk to news organizations. The newly disclosed excerpts challenge key elements of the companies’ fair-use defense in the litigation by documenting both the scale of copyrighted content used and internal recognition that models can supplant users’ visits to original sources. The filings state specific counts of allegedly copied works in training sets and describe internal awareness of paywall circumvention methods. OpenAI and Microsoft did not provide comment on the newly unsealed material in the reporting; many underlying exhibits remain sealed and some quoted lines in the filings appear without their original context.
Keep Reading

Crusoe raises $3.9B to build massive data centers and small modular ‘AI factories’

Scott, Baldwin ask FTC to investigate Amazon, Walmart AI over ‘Made in USA’ fraud detection

Khosla-backed Mazama Energy just raised $135M to drill deeper into super-hot-rock geothermal

Moore says he would ‘absolutely sign’ statewide data center moratorium
Original source: TechCrunch
Also reported by Ars Technica AI.