Urgent.News

What's breaking now, across thousands of outlets.

Business

Pretraining progress is mostly coming from data

Breaking down 6 years of pretraining progress into data vs model improvements

Pretraining progress is mostly coming from data

From 2019 to 2025, AI progress has primarily been driven by advancements in data engineering, rather than improvements in model architecture. Researchers trained combinations of model recipes and data corpora, evaluating the models based on end capabilities measured by the OLMES eval, which aggregates 10 different benchmarks. They found that 12 times more compute efficiency gains came from data improvements compared to model improvements at a 1e19 FLOPs budget.

The gains from both data and model improvements were largely independent and additive, accounting for 88% of the variance in the OLMES score. The main contribution of model improvements was enabling larger amounts of compute to be usable, rather than significantly increasing performance on their own. Data improvements had a more noticeable impact on smaller models, while larger models could handle larger datasets regardless of their quality.

As models continue to grow in size, data engineering will likely remain a crucial factor in driving AI progress.

Written by urgent.news from Dwarkesh Patel's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dwarkesh.com →

More in Business

More from Tuesday 8 September →