Big AI's content problem: Take the work, keep the money
The more we learn about how AI does business, the more unfair it looks
The AI industry faces a significant content problem, as companies must consume other people's work to train their large language models (LLMs). This includes books, news stories, photographs, code, and websites. However, AI companies repeatedly disregard copyright laws and license restrictions, arguing that the use of public content for training helps promote public knowledge.
They also claim this practice does not act as an economic substitute for the original creators. The New York Times has filed a lawsuit against OpenAI and Microsoft, alleging a massive theft of labor, with Microsoft's Director of Applied Science comparing it to "the largest theft of labor in human history." Microsoft's internal documents acknowledge that generative AI could disrupt employment in the very sector that created the training data.
Judge Sidney H. Stein is yet to rule on the case, while OpenAI and Microsoft dispute the allegations under a fair use defense. The AI industry continues to downplay the issue, focusing on revenue generation rather than addressing the ethical concerns surrounding their practices.
Written by urgent.news from The Register Science's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.
Also reported by 1 other outlet
- Big AI's content problem: Take the work, keep the money theregister.com