AI companies used massive amounts of published text to train their large language models, and officials at these companies were aware this could amount to an appropriation of copyrighted material, according to statements cited by news publishers in a court filing unsealed this week.
اضافة اعلان
The filing had been submitted earlier this month as part of legal proceedings consolidating a number of related copyright infringement lawsuits filed against OpenAI, developer of ChatGPT, and its partner Microsoft.
The cases include lawsuits filed by The New York Times and Ziff Davis, owner of CNET.
According to the publishers' filing, Brent Hecht, Microsoft's director of applied sciences, described the matter as "an astonishing theft of unprecedented scale," and possibly "the biggest theft of work in human history."
However, the full attachments providing context for these statements remain sealed and unavailable to the public.
The publishers say the statements contained in the court filing could weaken AI companies' argument that using copyrighted material falls within the scope of "fair use," which provides legal protection under certain circumstances.
Fair use assessment depends on several factors, including how the material is used, the nature of the original work, the amount of the portion used, and the effect this use has on the market for the original content.
According to the filing, Microsoft CEO Satya Nadella said under oath that conversations with chatbots provide information "directly on the AI platform's site, rather than requiring a trip to the original source," such as the publisher's website that produced the information.
Similarly, an OpenAI official wrote that publishers face an "existential threat" from products like the company's chatbots.
The publishers' filing also describes steps it says AI developers took to bypass paywalls that block content from non-subscribers, including The New York Times' paywall.
The filing notes that an OpenAI employee told company president Greg Brockman about "a trick to bypass the New York Times paywall," to which Brockman replied, according to the documents: "Oh nice."
At the same time, the filing states that Nadella testified that any content behind a paywall "should be licensed by anyone who wants to use it" in developing AI.
He added that had he known OpenAI trained its models on paid content, he would have demanded the company retrain those models.
In a statement provided by Microsoft to CNET, a company spokesperson said Hecht's comments "reflect the view of a single employee," and that the company's official position is set out in its court filings, which, according to the spokesperson, explain why these transformative uses comply with copyright law, and why Copilot is not a substitute for the journalism publishers produce.
The spokesperson added that Nadella's statements addressed changes in how people consume information, and that they are "entirely consistent" with the company's legal position.
He said: "These remarks shouldn't be conflated with conclusions on the copyright issues before the court, which Microsoft addresses in its court filings."
In a separate filing submitted by Microsoft in the case, the company argued that using published content to train large language models is a highly transformative use.
The company said: "Copyright law does not allow rights holders to block transformative technologies like large language models; rather, it encourages such uses, based on the expectation that rights holders will adapt to them and that the public will benefit."
Representatives for OpenAI and Ziff Davis did not immediately respond to requests for comment.
Large language models are typically trained on as much published content as developers can access, and copyright-related lawsuits, including similar cases filed by book authors, could have major implications for both AI companies and media outlets alike.
Publishers argue that AI companies should pay to license the content they use, while tech companies say such a requirement is unnecessary and would burden the development of more capable AI models.
The administration of US President Donald Trump had intervened in the case earlier this month, submitting a position arguing that imposing licensing requirements could hinder the development of American AI amid competition with China.
Publishers also say that employees at OpenAI and Microsoft were aware that chatbots' ability to find, copy, and summarize, and in some cases reproduce, publishers' content, reduces readers' need to visit the original news organizations' websites.
The publishers' filing states: "The defendants' own experts acknowledge that source-backed large language models exploit their sources rather than promoting them, unlike search engine uses, which have been deemed fair use."
Resource: Al Ghad.