Impactful scheduling for GPU clusters
The AI Infrastructure team at Ai2 is responsible for providing the institute's GPU compute capacity, specifically targeting large, distributed training workloads. They use a pyramid of four metrics that build on each other: availability, occupancy, impact, and utilization. The team replaced a priority-based scheduler with a system that includes GPU time budgets, hierarchical fair-share allocation, and a time-slicing contract.
This change transformed the debate about GPU time allocation from a case-by-case operational task to a transparent administrative budgeting process. The team manages thousands of NVIDIA H100, B200, and B300 GPUs in clusters ranging from 88 to 1024 GPUs. These clusters are used for large-scale distributed training of AI models, serving about 150 internal researchers from various AI domains.
Despite having demand for GPUs that exceeds supply, the team faced issues like GPU "squatting," priority inflation, and challenges in identifying root causes of problems. The team initially tried to prioritize research workloads by assigning them GPUs, but this led to idle GPUs due to seasonality and manual solving of a knapsack problem.
To address these issues, the team introduced a hierarchical system where managers could proportionally allocate GPU time to research efforts based on their judgment of likely impact. This enabled leadership to think like investors, deciding how to fund each research effort with GPU time before the workloads existed. The new system improved GPU utilization and reduced issues related to priority inflation and preemptability.
Written by urgent.news from Hugging Face's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.