The New York Times and a coalition of news publishers are suing OpenAI and Microsoft for the unauthorized use of millions of articles to train AI. Court documents reveal internal warnings from a Microsoft executive who described the process as a massive theft of labor.
The 10 million articles and Brent Hect's 'astonishing theft' warning
According to a court filing reported by Reuters, OpenAI scraped content from more than 10 million articles to fuel its artificial intelligence models, with nearly one-third of that content originating from the New York Times. The legal battle has taken a sharp turn with the revelation of internal Microsoft communications. Brent Hect, the Director of Applied Science at Microsoft, allegedly described the practice as an "astonishing theft of unprecedented proportions" and suggested it could be the "largest theft of labor in human history."
The filings further suggest that OpenAI may have engaged in an "accidental cover-up" while attempting to identify and remove content from the New York Times and other plaintiffs from its systems. While Microsoft has dismissed Brent Hect's comments as the individual perspective of a single employee that does not reflect the company's official stance, the internal admission provides significant ammunition for the publishers' claims of willful infringement.
Satya Nadella's admission on chatbot substitution
The core of the legal dispute rests on whether the use of copyrighted news is "transformative" or merely a replacement for the original product. Microsoft and OpenAI argue that their AI models fall under "fair use" laws.. However, Microsoft CEO Satya Nadella admitted under oath that interacting with chatbots has effectively substituted the need for users to visit the underlying source websites to find information, as reported by Reuters.
This admission undermines the tech giants' claim that AI serves as a complementary tool rather than a competitor.. The erosion of traffic is a primary concern for publishers, and the court documents reveal a candid admission from one of OpenAI's own engineers, who noted that regardless of how prominently links are displayed, users simply will not click through to the original news sites.
The DOJ's 'national security' defense for AI training
The conflict has expanded beyond a private civil dispute to include the interests of the United States government. In early September, the US Department of Justice filed a brief supporting OpenAI and Microsoft.. The Department of Justice invoked the necessity of "scientific progress," economic growth, and "national security" as justifications for the current trajectory of AI development, suggesting that overly restrictive copyright rulings could hinder American competitiveness in the global AI race.
This intervention highlights a growing tension between the protection of intellectual property and the strategic drive for technological dominance. While OpenAI has attempted to mitigate these tensions by signing content licensing deals with various global publishers since the launch of ChatGPT in late 2022, these agreements have not satisfied the broader group of plaintiffs who believe the initial training was fundamentally illegal.
The unknown damages sought by Ziff Davis and other publishers
The scope of the lawsuit extends far beyond the New York Times. The legal action includes Ziff Davis—the owner of CNET—as well as the parent company of Mother Jones, the investigative site The Intercept, and various local American newspapers. These entities are seeking damages for every individual article allegedly stolen and utilized by OpenAI's models.
Despite the scale of the claims, the total financial penalty remains unknown, as the specific damages for each article have not been finalized. The legal process is expected to be protracted; if US District Judge Sidney Stein allows the case to proceed, a final ruling is not anticipated until 2027. Until then, the industry remains in a state of precarious uncertainty regarding the cost of training large language models on proprietary data.
Comments 0