Urgent.News

What's breaking now, across thousands of outlets.

AI

Scaling MoE reinforcement learning on Amazon EKS with EFA and DeepEP with 40% more throughput

Learn how to scale Mixture-of-Experts (MoE) reinforcement learning on Amazon EKS using Elastic Fabric Adapter (EFA) and DeepEP. This post presents an architecture that combines Amazon EKS, EFA, and Amazon S3 and increased aggregate reinforcement learning rollout throughput by 40% for large-scale RLHF and GRPO training.

In order to scale Mixture-of-Experts (MoE) reinforcement learning on Amazon Elastic Kubernetes Service (Amazon EKS), AWS leverages Elastic Fabric Adapter (EFA) and DeepEP. These technologies address the challenges of heterogeneous compute, high-throughput communication, and dynamically orchestrating subsystems to maintain balance.

Large-scale Reinforcement Learning (RL) training of MoE models, such as those using Reward Models, Verifiers, and Checkpoint Updates, places significant demands on infrastructure due to the combination of elastic inference work and tightly coupled model training. With newer sparse MoE architectures, communication becomes the primary constraint.

DeepEP optimizes expert-parallel communication over EFA, thereby accelerating MoE training on Amazon EKS.

Written by urgent.news from AWS Machine Learning's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at aws.amazon.com →

More in AI

More from Friday 25 September →