{
  "id": 1679632,
  "title": "AI models get convenient amnesia about source material as they grow, MIT boffins find",
  "url": "https://urgent.news/2026/08/18/ai-models-get-convenient-amnesia-about-source-material-as-they-grow",
  "topic": "ai",
  "section": "AI",
  "published": "2026-08-18T09:00:00.000Z",
  "source": {
    "name": "The Register",
    "slug": "the-register",
    "url": "https://www.theregister.com/ai-and-ml/2026/08/18/ai-models-get-convenient-amnesia-about-source-material-as-they-grow-mit-boffins-find/5288846"
  },
  "original_language": "en",
  "account": "A paradox emerges as AI models grow in size - the more memory they possess, the less they can attribute the source of that memory back to the original training data. Researchers from MIT's Computer Science & Artificial Intelligence Laboratory (CSAIL) set out to identify a method for linking AI model outputs to specific training data, aiming to facilitate AI regulation. However, their findings, detailed in a paper titled \"Outputs of Generative Diffusion Models are Often Unattributable,\" may complicate the regulatory landscape.\n\nThe paper, to be published in the scientific journal Nature Communications, reveals a phenomenon dubbed \"attribute decay.\" As diffusion models like Midjourney and Stable Diffusion, used to create images, videos, and audio, expand in size, their ability to attribute output back to the source data diminishes. This decay occurs regardless of whether specific training data, such as images of the Mona Lisa or works by Leonardo Da Vinci, is removed from the model's learning process.\n\nMIT's Zheng Dai and David K Gifford, the paper's authors, explain that attributing the origin of AI-generated content becomes increasingly elusive as models grow in size. This attribute decay poses significant challenges for various applications, including machine unlearning, data poisoning, interpretability, fairness, and privacy. Moreover, as these models are increasingly adopted for creative and commercial purposes, attributability carries substantial ethical, policy, financial, and legal implications.\n\nThe researchers argue that the findings suggest AI models exhibit a level of creativity, as they do not merely copy their training data. If the outputs are unrelated to any individual piece of training data, it raises questions about fair use, copyrightability, and compensation for authors whose work is reproduced without clear attribution. Furthermore, the ability to test whether an output is derivative could compel companies to demonstrate their models cannot be traced back to a specific source.\n\nWhile the research suggests a potential strategy for liability avoidance - building larger models that render output attribution virtually impossible - legal experts caution that existing copyright claims against AI companies have not yet focused on the extent to which model outputs can be linked to specific artists' work. Current cases primarily revolve around whether training constitutes fair use or not. Nonetheless, the MIT study highlights the need for alternative methods to assess copying in the context of increasingly complex AI models.",
  "summary": "Attributing diffusion model output to a specific input becomes more difficult with more training data",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 2,
    "also_reported_by": [
      {
        "outlet": "The Register Science",
        "title": "AI models get convenient amnesia about source material as they grow, MIT boffins find",
        "url": "https://urgent.news/2026/08/18/ai-models-get-convenient-amnesia-about-source-material-as-they-grow-1681859",
        "published": "2026-08-18T09:00:00.000Z"
      }
    ]
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}