Wiring and powering GPUs differently can swing AI latency by orders of magnitude, says CoreWeave
The neocloud market is moving past its origins as a stopgap for scarce graphics processing units. AI-native startups now choose their infrastructure on latency, burst capacity and openness, not just chip availability. That shift is playing out at CoreWeave Inc., which is expanding beyond GPU compute into networking, storage and software as inference demand grows. […] The post Wiring and powering…
At Fully Connected, CoreWeave CEO Jerry Liu and CoreWeave SVP AI Initiatives Lukas Biewald discussed how differences in powering and wiring GPUs can significantly affect AI latency. The neocloud market has shifted from being a stopgap for scarce GPUs to prioritizing latency, burst capacity, and openness, with AI-native startups focusing on inference demand instead of just chip availability.
CoreWeave is expanding beyond GPU compute into networking, storage, and software as inference demand grows. LlamaIndex, an open-source framework for retrieval-augmented generation, has evolved into a model builder that rents its compute rather than owning it. Co-founder and CEO Liu emphasized the importance of tailoring performance, cost, and latency for customers in document parsing and extraction.
Liu and Biewald noted that CoreWeave's compute footprint has grown to 75% inference and 25% training, processing millions of document pages daily for finance, legal, and insurance customers with bursty workloads. CoreWeave does not own a GPU cluster, making guaranteed capacity crucial for serving customers without getting throttled.
Biewald highlighted the significant impact of networking and power distribution configurations on chip performance, stating that the difference can be orders of magnitude. CoreWeave's openness, using standard networking protocols recommended by Nvidia, sets it apart from the hyperscalers, which use proprietary APIs that limit workload portability.
The company's portfolio expansion mirrors early AWS days, but CoreWeave acknowledges that most workloads will host on AWS or GCP rather than CoreWeave. This shift indicates a larger change in who builds intelligence, with Liu believing that abundant neocloud capacity and improving coding agents could make post-training small open-weight models accessible to a broader audience.
Written by urgent.news from SiliconANGLE's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.