A $22,000 GPU bill with a very quiet night shift
The one I think about most is the account where the infrastructure was correct and the cost was still wrong ML startup. Eight engineers. Production inference API with real traffic Their AWS bill was $22,000 a month. Mostly GPU compute for inference I went in expecting to find over-provisioned instances. I found the opposite. The instance types were appropriate for the workload. Utilization was…
The story centers around a machine learning startup that was being billed $22,000 a month for GPU compute. While the infrastructure and instance types were well-suited for the workload, the cost was high due to inefficient scheduling. The company's inference API received a predictable traffic pattern, with 90% of requests arriving between 9am and 11pm US Eastern time. The remaining hours of the day saw minimal traffic, primarily consisting of health checks and background jobs.
The issue lay in the fact that the GPU instances ran continuously, maintaining full capacity even during these low-usage periods. This resulted in paying full price for eight hours of near-zero utilization every night. To address this, the team implemented a scaling schedule. At 11pm Eastern, the inference cluster was scaled down to a minimal warm state, sufficient to manage background jobs and health checks. At 8am Eastern, the cluster was scaled back up to accommodate incoming traffic.
The process of implementing and testing this new schedule took two days. However, the change proved fruitful, leading to monthly savings of $4,800. The improvement came not from altering the architecture, the model, or the product's functionality, but rather from aligning the cost with the actual demand pattern. The original infrastructure remained unchanged, and the schedule was the only aspect that needed adjustment.
Sometimes, the solution lies in fine-tuning the scheduling rather than making significant changes to the system.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.