AI safety efforts will require more compute, not less: experts
AI safety and alignment, according to top frontier model developers, are demanding substantial computational resources, which are expected to spur a rise in infrastructure spending, according to a Citi research report. This assertion came during the opening day of the AI Infra Summit held in Santa Clara, where experts emphasized that the industry's main bottleneck lies in acquiring massive compute, power, and facility capacities rather than merely advancing model architecture.
Companies are designing next-generation data centers with 30% to 40% more GPUs, and software orchestration platforms like Astra, as well as hardware upgrades such as Nvidia’s Vera Rubin platform, are being developed to enhance system throughput. Ian Buck, Nvidia's Vice President of Hyperscale and HPC, explained that transitioning from conversational chatbots to autonomous, agentic AI workloads demands hardware resources 100 times more than those needed in 2023.
These agentic workloads require larger input sequence lengths, KV cache sizing, multi-turn interactions, and sub-agent spawning. Intel CEO Lip-Bu Tan presented the chipmaker's strategic transition from a traditional CPU vendor to a comprehensive AI infrastructure provider. Tan highlighted that while competition is shifting from GPU horsepower to system-level optimization, CPUs will remain crucial for reinforcement learning, orchestration layers, and agentic workflows.
To enhance decision-making speed, Tan restructured Intel's management structure, reducing layers from 10-12 to 4-5. Intel's 18A process node is now in volume production, and the advanced 14A node is expected in the coming quarters, with 0.9 and 1.0 PDK releases also on the horizon. Meanwhile, Meta Platforms' Head of Infrastructure, Santosh Janardhan, discussed the company's scaling of its massive training infrastructure, from its 1-gigawatt-plus Prometheus cluster to the 5-gigawatt Hyperion footprint planned for the next few years.
The primary operational hurdle, according to Janardhan, is multi-region, loss-less networking, which will enable Meta to limit GPU cluster interruptions to just two per 8,000 GPUs per day, achieved through the deployment of proprietary Back-End Aggregation (BAG) networking architecture.
Written by urgent.news from Investing.com's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.