Urgent.News

What's breaking now, across thousands of outlets.

AI

Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs

Tech Report | Code | Interactive demo

OpenAI has launched Olmo-core 3, an enhanced framework for creating large language models with a new, open mixture-of-experts (MoE) training system. This upgrade enables MoE training to reach the trillion-parameter range while maintaining computational efficiency. The framework is a key component behind the upcoming generation of Olmo models and aligns with OpenAI's goal of making the tools and training infrastructure behind each new model more accessible to researchers and smaller labs.

Training large AI models consumes significant computational resources, leading to increased costs and energy usage. MoE models provide a more efficient solution by allowing for a larger number of learned components (parameters) without each input needing to use all of them. However, as MoEs grow in size, the computational advantages diminish due to the costs associated with storing the full model across GPU memory, updating it during training, and directing inputs to the appropriate experts within a cluster.

Olmo-core 3 addresses these challenges by optimizing the distribution of the model and its training state across GPU clusters. Three key techniques contribute to this optimization: rowwise expert parallelism, GPU-resident routing, and grouped GEMM. Rowwise expert parallelism places routed data directly into expert input buffers, minimizing the need for additional arrangement.

GPU-resident routing keeps routing metadata on the GPUs, allowing the CPU to queue work without waiting for the information to be copied back. Grouped GEMM (matrix multiplication) combines multiple small expert computations so GPUs can execute them more efficiently.

To further enhance performance, Olmo-core 3 supports MXFP8, a lower-precision number format that represents certain values with fewer bits. This approach can reduce computation and data movement between GPUs, provided the benefits outweigh the costs associated with converting between number formats. In tests using NVIDIA B300 GPUs, Olmo-core 3 demonstrated a 21% increase in training throughput compared to a higher-precision format (BF16) while also reducing peak active memory from 103 GiB to 95 GiB.

These gains primarily resulted from improved feed-forward computation and data movement between experts.

Olmo-core 3 incorporates various techniques for scaling MoE training, from a single GPU to a large cluster of GPUs. Benchmarks have been conducted across a range of configurations, including a 1.2-trillion-parameter model with 58.36 billion parameters active per token across 512 GPUs. The highest observed throughput reached 858 TFLOP/s/GPU.

While these tests utilized random routing to measure system performance rather than the quality of trained models, the framework demonstrates the scalability of Olmo-core 3 for sustained training runs.

Furthermore, the technical report includes experiments that informed the development of Olmo-core 3, providing insights into the trade-offs and optimizations necessary for efficient large-scale training of MoE models.

Written by urgent.news from Hugging Face's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at huggingface.co →

More in AI

Unstable connections

Sociologist Sherry Turkle explores the impacts of developing emotional relationships with artificial intelligence.

More from Thursday 1 October →