
AI companies used vast troves of published text to train their large language models, and executives at those firms were aware that doing so constituted a theft of copyrighted materials, according to comments cited by news publishers in a court filing that was unredacted this week.
The brief was initially filed earlier this month in a court proceeding that brings together several related copyright infringement cases against ChatGPT maker OpenAI and its partner Microsoft, including lawsuits brought by The New York Times and by Ziff Davis, which owns CNET.
Microsoft’s director of applied science, Brent Hecht, called it “an astonishing theft of unprecedented proportions” and maybe the “largest theft of labor in human history,” according to the publishers’ filing. The full exhibits showing the context of the quotes in the filing remain under seal.

The comments cited by publishers appear to undercut the argument by AI companies that the use of copyrighted material was legally protected as fair use. Under this legal doctrine, use of copyrighted material is protected depending on how it’s used, the nature of the work, how much of it was used and the effect of the use on the market for that work.
Microsoft CEO Satya Nadella said under oath that conversations with chatbots gave information “right there on the website on the AI platform versus needing to go to the underlying source,” like the website of a publisher that reported the information, according to the filing. Similarly, an OpenAI executive wrote that publishers faced an “existential threat” from products like the company’s chatbot.
The publishers’ brief also describes steps AI developers took to evade paywalls, like that for The New York Times. It describes how an OpenAI employee told the company’s president, Greg Brockman, about “a hack to get around nytimes paywall,” to which Brockman replied, “ah nice.”
At the same time, the filing says Nadella testified that anything behind a paywall “should be licensed by anyone who wants to use it” for AI development and that if he had been aware OpenAI had trained on paywalled content, he would have required OpenAI to retrain its models.
In a statement provided to CNET, a Microsoft spokesperson said that Hecht’s comments “reflect one employee’s perspectives” and that the company’s position is in its court filings, which “explain why these transformative uses are consistent with copyright law and why Copilot is not a substitute for publishers’ journalism.”
Nadella’s comments addressed changes in how people consume information and are “perfectly consistent” with the company’s legal position, the spokesperson said. “Those observations should not be confused with conclusions about copyright questions before the Court, which Microsoft addresses in its filings.”
In its own brief in the case, Microsoft said the use of published content to train LLMs was significantly transformative. “Copyright law does not permit rightsholders to block transformative technologies like LLMs; it encourages such uses on the expectation that rightsholders will adapt and the public will be better off for it,” the company argued.
Representatives for OpenAI and Ziff Davis did not immediately respond to requests for comment.
LLMs are trained on as much published content as developers can get their hands on, and the litigation over copyright — including related cases from book authors — could have a significant impact on both AI companies and media outlets. Publishers argue AI companies should have to pay to license the content they use, while the tech companies say doing so is unnecessary and burdensome to the development of more capable AI. The Trump administration weighed in on the case earlier this month with a statement arguing a licensing requirement would impede American AI development in a race with China.
OpenAI and Microsoft employees were also aware that their chatbots’ ability to find, copy, summarize and sometimes regurgitate content from publishers was keeping readers from visiting the sites of news organizations, the publishers argue.
“Defendants’ own experts acknowledge that grounded LLMs exploit their sources rather than promote them like search uses that have been deemed fair use,” the publishers’ filing states.







