Gen AI outputs are unattributable, study finds
Researchers discovered a phenomenon they call “attribution decay,” where the more data a generative model is trained on, the harder it becomes to trace a generated image to a single image from the training data.
A new study from MIT's Computer Science and Artificial Intelligence Laboratory has found that AI-generated images trained on large datasets cannot be traced back to the original data used for training, a phenomenon called "attribution decay." This can create issues in cases of intellectual property theft. As the amount of data a generative model is trained on increases, it becomes increasingly difficult to link a generated image to a specific image from the training data, even if the AI-generated image resembles a well-known artist, like Picasso.
Artificial intelligence companies have increasingly used various works, often without seeking permission from authors, to train their models, leading to numerous copyright infringement lawsuits. In 2023, Disney, NBCUniversal, and DreamWorks filed an intellectual property lawsuit against AI image-generator MidJourney. The same year, The New York Times filed a similar lawsuit against OpenAI and Microsoft.
This research raises legal questions regarding the use of training data and the concept of fair use policy, as noted by the authors. Zheng Dai, a former MIT researcher and the lead author of the work, expressed concerns about the future of intellectual property. According to Dai, "We might have to rethink what intellectual property means. You can't just assume it, and the attribution link sort of vanishes."
Written by urgent.news from Semafor's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.