AI may be learning from billions of images without copying any one of them
An MIT study finds that AI-generated images often cannot be traced to individual training images, as larger datasets cause the influence of specific examples to fade.
A recent study from MIT's Computer Science and Artificial Intelligence Laboratory explores the concept of "attribution decay" in large AI models. This phenomenon occurs when the influence of any individual piece of training data becomes increasingly difficult to identify as the dataset grows in size. The researchers tested this by removing specific images from the training data and observing the impact on the generated output.
Their findings suggest that, as models and datasets scale, removing a single image can make little or no difference to the final result. This means that it becomes challenging to attribute specific AI-generated outputs to particular training images. While this does not imply that training data is irrelevant, nor does it settle the broader debate surrounding AI companies' use of copyrighted material without permission, the study does highlight that a model may learn broad visual patterns from an extensive pool of data without any single image being directly responsible for a particular output.
Written by urgent.news from Digital Trends's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.