{
  "id": 1681859,
  "title": "AI models get convenient amnesia about source material as they grow, MIT boffins find",
  "url": "https://urgent.news/2026/08/18/ai-models-get-convenient-amnesia-about-source-material-as-they-grow-1681859",
  "topic": "ai",
  "section": "AI",
  "published": "2026-08-18T09:00:00.000Z",
  "source": {
    "name": "The Register Science",
    "slug": "the-register-science",
    "url": "https://www.theregister.com/ai-and-ml/2026/08/18/ai-models-get-convenient-amnesia-about-source-material-as-they-grow-mit-boffins-find/5288846"
  },
  "original_language": "en",
  "account": "MIT researchers have discovered a perplexing paradox in the realm of AI model training – as these models grow larger, they seemingly lose the ability to attribute their outputs to the specific source material from which they learned. The MIT Computer Science & Artificial Intelligence Laboratory (CSAIL) researchers, Zheng Dai and David K Gifford, published their findings in a paper titled \"Outputs of Generative Diffusion Models are Often Unattributable\" in the scientific journal Nature Communications.\n\nDiffusion models, such as those used by Midjourney and Stable Diffusion, have gained popularity for generating various types of media, including images, videos, and audio. However, this paper suggests that attributing model output to the training data becomes increasingly challenging as the models expand in size. This phenomenon, referred to as \"attribution decay,\" means that even if specific training data is removed through a process called ablation, large models can still reproduce the same image or style.\n\nThe authors argue that this discovery complicates efforts to regulate AI models, as it becomes difficult to establish clear attribution between generated content and its training sources. This lack of attributability raises ethical, policy, financial, and legal implications, particularly in the context of creative and commercial uses of AI models. Gifford explains that if outputs are not tied to specific inputs, questions arise regarding fair use, the novelty of the outputs, and the potential compensation for authors.\n\nMIT's press release quotes Gifford as stating that the findings suggest AI models exhibit creativity beyond mere copying of training data. He also highlights the potential obligation for companies to demonstrate the inability of their models to attribute outputs to specific sources. At the same time, the research proposes a liability avoidance strategy, which involves training models large enough to render outputs unattributable to any single input.\n\nHowever, the research also presents a challenge for legal proceedings involving AI models. James Grimmelmann, a law professor at Cornell Law School and Cornell Tech, suggests that the inability to attribute outputs may create difficulties in establishing whether similarities between a model's output and a copyrighted work are due to copying or coincidence. While current copyright claims against AI companies have not specifically focused on the extent to which outputs resemble real-world artwork, Grimmelmann believes that output similarity for image models remains largely untested in court.",
  "summary": "Attributing diffusion model output to a specific input becomes more difficult with more training data",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 2,
    "also_reported_by": [
      {
        "outlet": "The Register",
        "title": "AI models get convenient amnesia about source material as they grow, MIT boffins find",
        "url": "https://urgent.news/2026/08/18/ai-models-get-convenient-amnesia-about-source-material-as-they-grow",
        "published": "2026-08-18T09:00:00.000Z"
      }
    ]
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}