Is it legal to train AI models on copyrighted books? It’s complicated
Most published authors have, without their knowledge or consent, contributed to the development of the same AI tools that threaten to undermine their livelihoods. That seems illegal, right?
The training of AI models on copyrighted materials, such as books, is a complex issue with uncertain legal implications. In one notable ruling, Judge William Alsup found that Anthropic's use of books to train its AI models was legal, despite the fact that the books were pirated from illegal online sources. Judge Alsup ruled that Anthropic was not infringing on copyright by training its AI on works, but instead by using them to create something new and different, similar to how a writer might study literature.
Cathy Gellis, an attorney with expertise in intellectual property, copyright, and technology, believes that this ruling is favorable for AI companies, as the $1.5 billion fine ordered by the judge is a small price to pay compared to the projected $200 billion in annual revenue for AI companies by 2028. She argues that copyright law is concerned with copying, not with using or experiencing the work, and thus, using copyrighted materials for training AI models does not violate copyright law.
However, copyright law has not been updated since 1976, and judges are struggling to interpret guidelines from 50 years ago to address legal questions that could shape the future of the AI industry. The issue often hinges on fair use law, which allows for the use of copyrighted materials without permission for purposes such as criticism, parody, education, and other means. Fair use is determined by considering factors such as the purpose and nature of the work, the amount used, and its impact on the market.
In cases where AI models are trained on copyrighted materials to create new, synthetic content, the argument that the AI is competing with the original works has not yet been successful in court. However, when it comes to training AI models on copyrighted materials, the relationship between AI and copyright is different from that of AI-generated content.
In one case, Thaler v. Perlmutter, the court ruled that if a work is 100% AI-generated, it is not copyrightable, raising questions about how to prove whether a work was generated using AI and, if so, what percentage of it was created or assisted with AI.
As AI companies face pending litigation over these issues, it remains unclear when a definitive solution will be reached. The decisions made in these cases are shaping the landscape of AI and copyright, and it would be unwise for AI companies to ignore them.
Written by urgent.news from TechCrunch's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.