AI’s attribution problem gets worse as models scale
Diffusion models are becoming sophisticated enough that they can reproduce an image even when they don’t have access to the original. In a series of ‘what if’ scenarios, researchers associated with MIT’s Computer Science & Artificial Intelligence Laboratory (CSAIL) swapped out different training datasets to test the impact on image outputs when original image data was completely removed. It turns…
Diffusion models, capable of creating images without access to original data, are becoming increasingly sophisticated. Researchers from MIT's Computer Science & Artificial Intelligence Laboratory (CSAIL) conducted experiments to examine the impact of removing training datasets on the models' outputs. They found that as models scale up, individual inputs become less influential, a phenomenon referred to as "attribution decay."
This discovery has significant implications for intellectual property (IP) and copyright infringement concerns. The researchers suggest that understanding attribution decay could help address legal issues, model interoperability, fairness, privacy, and ethical concerns related to AI models. The MIT researchers trained various ensembles on datasets containing up to 160,000 images and observed that removing certain pieces of data didn't affect the model's output.
This finding highlights the challenge of tracing model-generated content back to its original sources, potentially complicating legal questions around derivative works, fair use, and copyrightability.
Written by urgent.news from Computerworld's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.