Urgent.News

What's breaking now, across thousands of outlets.

AI

Microsoft exec called AI scraping ‘the largest theft of labor in human history,’ new unredacted filings reveal

Newly unsealed court filings show Microsoft privately called OpenAI's data practices "theft" while both companies scraped paywalled Times content, built datasets from it, and warned internally it would gut publishers.

New unredacted documents in the copyright lawsuit brought by The New York Times against OpenAI and Microsoft reveal that AI scraping was deemed equivalent to theft and a major danger to publishers. A high-ranking Microsoft official described the company's AI training methods as "theft." OpenAI's leadership asserted that its AI models pose an "existential threat" to the journalists and publishers whose work trained them.

The unsealed material also details how the companies covertly obtained and utilized that content by bypassing paywalls undetected, amassing training datasets through mass scraping, and deliberately removing copyright notices from the training data. Much of the new information originates from The Times' own brief, not the unsealed exhibits.

The unredacted filing is the latest development in the lawsuit that initially accused the firms of copyright infringement by training generative AI models on its content. The question of whether AI companies can legally employ copyrighted material to train AI lacks a clear answer, but courts have generally sided with AI companies' arguments that training qualifies as "fair use."

Certain admissions, however, contradict OpenAI's fair use defense, especially the requirement that use doesn't replace or damage the market for the original work. For instance, Microsoft's internal data indicates that its Copilot "answer engine" caused The New York Times' click-through rates to plummet by up to 93% compared to traditional Bing search.

An internal Microsoft presentation by Director of Applied Science Brent Hecht in January 2024 describes this decline as a "doom loop" that would "hurt the performance of our models and the entire web at the same time." Microsoft's CEO, Satya Nadella, testified that "anything that is paywalled should be licensed by anyone who wants to use it...for grounding or training," and emphasized that if he had been aware of OpenAI scraping paywalled information, he would have required OpenAI to retrain its models.

Other admissions undermine different aspects of the fair-use test: OpenAI's Head of ChatGPT, Nick Turley, wrote in internal communications that publishers face an "existential threat" from AI products like ChatGPT, which are "largely substitutive" and "will get more and more substitutive as they get better." OpenAI President Greg Brockman described the models as "excellent at news."

Nadella confirmed under oath during a deposition that conversing with chatbots "has substituted...giving you the information right there on the website on the AI platform versus needing to go to the underlying source." This suggests the technology could directly compete with, rather than enhance, the original work. Microsoft documents state that there is a "real risk" that generative AI could "significantly disrupt the employment of the very people who generated the data on which the foundation model was trained."

The amount of copying is staggering. The documents disclose for the first time that OpenAI's mid-training datasets alone contain over 91,692 copies of works published by The New York Times, The Daily News, and the Center for Investigative Reporting. A Common Crawl-derived dataset included over 2 million documents from nytimes.com alone.

In a January 2023 internal memo, Hecht called it "an astonishing theft of unprecedented proportions" and "the largest theft of labor in human history." The filing provides additional details on how OpenAI and Microsoft acquired the plaintiffs' content, including scraping it from the Bing Index. OpenAI researcher Nick Ryder reportedly shared a "hack to get around nytimes paywall" with Brockman, who responded "ah nice."

OpenAI employees allegedly created training datasets like WebText and WebText2 that heavily relied on scraped news content. They also allegedly gathered millions of articles from Common Crawl, a free, open repository of web crawl data. The findings also reveal deliberate efforts to remove copyright notices from training data before it reached the model, as researchers didn't want the model to output "copyright notices" to users. OpenAI and Microsoft did not respond to requests for comment.

Written by urgent.news from TechCrunch's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Also reported by 1 other outlet

Read the original at techcrunch.com →

More in AI

More from Thursday 17 September →