Red Hat AI 3.5 tackles the GPU queue that can stall AI pilots
Red Hat released Red Hat AI 3.5 this week, a move designed to let software engineering teams run AI with The post Red Hat AI 3.5 tackles the GPU queue that can stall AI pilots appeared first on The New Stack .
Red Hat has released Red Hat AI 3.5, a platform designed to operationalize AI with the same rigor as enterprise applications on mission-critical infrastructure. This release focuses on enhanced multi-tenancy for AI service providers, allowing for complete hardware-to-software isolation and priority-aware service requests on shared GPU infrastructure.
Tushar Katarki, Senior Director of Product for Red Hat AI, emphasizes that running enterprise AI without safety controls is "like driving a supercar blindfolded." Red Hat AI 3.5 addresses this by delivering operational guardrails, verifiable trust, and multi-tenant controls needed to run AI as a mission-critical service.
Joshua Estrin, an applied mathematician and data scientist, highlights that every GPU request now becomes a priority decision. He believes that efficiency without isolation could lead to security and reliability crises, and that the winners in this market will be those who can share capacity while maintaining visibility into live production workloads.
Red Hat AI 3.5 introduces several new features, including EvalHub for verifying models before deployment through risk-focused safety benchmarking and regulatory compliance certifications. New observability dashboards provide platform teams with real-time metrics on inference health, GPU utilization, and AI model performance. Non-admin users can access dashboard features such as per-user token consumption showback and distributed inference workloads.
The platform also introduces shared GPU control for multi-tenant inference, managing resource allocation across tenants through fair-share GPU scheduling and providing priority-aware serving for admission control and priority-based request routing. Additionally, Red Hat's VP of product management, Anindo Sengupta, stresses the importance of secure tenant isolation for running multi-tenant AI at scale, with Kubernetes on bare metal being a potential solution.
Written by urgent.news from The New Stack's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.