Urgent.News

What's breaking now, across thousands of outlets.

AI

The AI Data Satiation Point: Why 'Wild' AI Text is Breaking Chinchilla Scaling Laws

The AI Data Satiation Point: Why "Wild" AI Text is Breaking Chinchilla Scaling Laws As we approach the end of 2026, the internet looks fundamentally different than it did just two years ago. According to new research from Pangram Labs and the University of Massachusetts Amherst, nearly 31% of all web tokens collected in August 2026 are AI-generated. This isn't just about synthetic data generated…

A new study from Pangram Labs and the University of Massachusetts Amherst suggests that the internet is increasingly populated with AI-generated text, with nearly 31% of web tokens collected in August 2026 being AI-generated. This "wild" AI text, which includes sentences, articles, and reviews written by language models (LLMs) intended for human consumption, presents a significant challenge for researchers developing next-generation frontier models.

The study, titled "How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text," provides data-backed insights into the impact of AI-generated text on model pretraining. Contrary to earlier concerns about "model collapse" when training models on their own output, this research reveals that the value of AI tokens—measured in terms of their contribution to model performance—changes based on the amount of human text a model has already been trained on.

In data-starved models, AI-generated tokens can initially improve performance by providing a scaffold for learning basic grammar and reasoning. However, once a model has been trained on a substantial amount of high-quality human text, adding more AI tokens can actually harm the model's ability to understand and generate human-like content.

This new understanding of scaling laws for AI-generated text challenges existing theoretical frameworks, such as the Chinchilla scaling laws, and underscores the need for updated models to accurately predict the effects of AI-generated data on model performance.

Brief written by urgent.news from Dev.to's own syndicated text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

More from Thursday 1 October →