Urgent.News

What's breaking now, across thousands of outlets.

Tech

Kubernetes can run AI inference. But can it count the real cost?

Welcome to another edition of Road to KubeCon, where we’re tracking the Kubernetes and cloud-native ecosystem on the way into The post Kubernetes can run AI inference. But can it count the real cost? appeared first on The New Stack .

Kubernetes can run AI inference. But can it count the real cost?

In the cloud-native ecosystem, Kubernetes has been making strides in supporting AI inference workloads. HPE, a presenting sponsor of Road to KubeCon, has been recognized for its virtualization and cloud operations portfolio, highlighting the need for unified governance and provisioning of workloads, including AI inference. At the same time, Red Hat's Kubernetes project has introduced two new Alpha features for container storage security, offering native controls to harden volume mounts and improve overall storage management.

Meanwhile, KubeCon has added an AI Inference + Agentic track, reflecting the growing role of Kubernetes in production AI workloads. China Merchants Bank has demonstrated a successful implementation of a Kubernetes-based AI infrastructure, unifying management of nearly 10,000 accelerator cards and achieving a 60% increase in average utilization.

However, as the demand for AI inference grows, concerns have been raised about the cost implications. WEKA's chief AI officer, Val Bercovici, argues that Kubernetes' resource model is not designed to handle the unique costs associated with large-scale AI inference, such as request mix and accelerator memory usage. He suggests that a new scheduling and memory layer may need to emerge alongside Kubernetes to better manage these costs, rather than seeing Kubernetes as a hindrance.

Written by urgent.news from The New Stack's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at thenewstack.io →

More in Tech

More from Friday 18 September →