Does generative AI actually copy artists? Researchers say it’s up for debate
Generative AI has long been accused of copying artists’ work outright (see the numerous copyright lawsuits winding their way through court). But a new study out of MIT makes a very different argument—one that could complicate how those cases hold up. In the study, published in Nature , the MIT researchers Zheng Dai and David Gifford set out to to test whether a generated image can be traced back…
Generative AI systems, particularly diffusion models, have been accused of copying artists' work, leading to numerous copyright lawsuits. However, a new study from MIT suggests that attributing generated images to specific training data may be more complex than previously thought. Researchers Zheng Dai and David Gifford discovered that as training data sets grow larger, it becomes increasingly difficult to trace the origin of a generated image back to a single piece of training data.
This phenomenon, termed "attribution decay," occurs because larger datasets contain a high degree of visual redundancy. Many images within the dataset share similar features, making it challenging to identify which specific image contributed to the generated image. To test their hypothesis, the researchers employed a technique called "ablation," where they removed specific pieces of training data from the model and observed the impact on the generated output.
When the removed data had little to no effect on the output, it indicated that the model could generate an image resembling a particular artist's work without a provable causal link to that artist's contribution to the training data. The researchers caution that their findings should be viewed as a best explanation rather than a definitive proof.
The phenomenon of attribution decay is likely due to the distributive and redundant encoding of important features throughout the training set. As datasets grow larger, the influence of individual images becomes intertwined with the rest of the data, making it nearly impossible to isolate and attribute specific contributions to particular pieces of training data.
This effect becomes more pronounced at scales of 10,000 to 100,000 images, which is the range of data often used in commercially deployed AI image generators. The study provides an example using the artwork of Andy Warhol. If a model were trained on a dataset containing 50,000 artworks, including Warhol's silkscreens, removing his specific images would likely have little to no impact on the generated output.
This is not because the model "understands" Warhol's style, but rather because other artists within the dataset were using similar visual features, such as bold flat colors, repeated grids, and pop culture subjects. Thus, Warhol's paintings were not a unique ingredient but one of many sources contributing to the same visual patterns.
In conclusion, the MIT researchers' findings complicate the ability to attribute generative AI images to specific artists, challenging existing arguments that AI has directly copied an artist's work. Artists may find it increasingly difficult to establish a direct causal link between their artwork and the output generated by AI systems.
Written by urgent.news from Fast Company's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.