The New York Times and several major publishers are taking OpenAI and Microsoft to court over the unauthorized use of copyrighted material. The lawsuit alleges that millions of news articles were scraped to train the ChatGPT artificial intelligence model.

Advertisement

The alleged theft of 10 million news articles

The legal dispute centers on the massive scale of data ingestion used to build modern AI. According to the report, the San Francisco-based OpenAI allegedly scraped content from more than 10 million articles to train its systems. The New York Times claims that nearly one-third of that scraped content originated from its own archives alone. In a court document, Microsoft Director of Applied Science Brent Hect described the practice as an "astonishing theft of unprecedented proportions," suggesting it could be one of the largest thefts of human labor in history.

A coalition including Ziff Davis and The Intercept

This legal action is not a solo effort by a single media outlet. As reported by the source, the lawsuit includes a diverse group of plaintiffs, such as:

  • Ziff Davis, the owner of CNET and other tech publications
  • The parent company of Mother Jones
  • The investigative site The Intercept
  • Multiple local newspapers across the United States
  • These organizations are collectively seeking financial damages for every individual article they claim was stolen and utilized by OpenAI's flagship models .

    Why the DOJ is backing Microsoft and OpenAI's fair use claim

    Microsoft and OpenAI have mounted a defesne based on the concept of "transformative" use. They argue that their utilization of news content falls under existing fair use laws. This position has gained significant momentum from the federal government; in early September, the U.S. Department of Justice filed a brief in support of the tech companies. The DOJ's filing argued that protecting these AI developments is vital for national security, economic growth, and scientific progress.

    The egnineer's admission that users won't click links

    While AI companies have attempted to mitigate concerns by providing citations and links, the economic impact on publishers remains a primary concern. The source notes that OpenAI has already signed various content licensing deals, but critics argue these efforts are insufficient. A significant point of contention is an admission from one of OpenAI's own engineers, who stated in a court document that users are unlikely to click on provided links regardless of how prominently they are displayed.

    Who is responsible for the alleged 'accidental cover-up'?

    Several critical details regarding the alleged misconduct remain unverified. while the court document mentions a warning from Brent Hect about a potential "accidental cover up" by OpenAI to hide the origins of certain content, the specific nature of this attempt has not been fully detailed. Additionally, while the plaintiffs are seeking damages for each stolen article, the total financial penalty that could be imposed remains unknown. If U.S . District Judge Sidney Stein grants certain motions, the case may avoid a trial, though a final ruling is not expected until 2027.