OpenAI and Cerebras Bring GPT-5.6 Sol Ultrafast to Enterprise Inference
OpenAI is expanding its inference infrastructure through a multi-year partnership with Cerebras, aiming to support faster responses for real-time AI workloads. The centerpiece is GPT-5.6 Sol Ultrafast , a Cerebras-backed deployment that OpenAI says can reach up to 750 tokens per second during a limited preview. For enterprises, the development is less about a minor model setting and more about…
OpenAI has announced a collaboration with Cerebras to enhance its inference infrastructure for real-time AI workloads. The main focus is on GPT-5.6 Sol Ultrafast, a deployment backed by Cerebras that can handle up to 750 tokens per second during a preview period. This development is significant for businesses as it addresses the importance of latency in workflows where delays can impact the user experience or business processes.
The multi-year partnership with Cerebras aims to provide OpenAI with 750 megawatts of ultra-low-latency AI inference capacity, which will be deployed in multiple phases through 2028. OpenAI will integrate Cerebras wafer-scale compute into its inference stack to deliver faster responses and enable real-time AI experiences across customer workloads.
However, it is important to note that the Ultrafast performance claim of 750 tokens per second is only applicable during the preview period and not a guaranteed service-level commitment for every customer or deployment. The actual enterprise access and rollout of this capability will depend on the staged availability and the capacity made available to customers over time.
For teams considering potential use cases, the most promising near-term applications are those where a faster response can significantly alter the workflow, rather than just making an existing chat interface feel quicker. Examples include interactive decision support, high-volume assistance, and AI systems requiring multiple model calls before a user can take action.
Nonetheless, pricing and governance details for this high-speed deployment have not been disclosed by OpenAI, and companies should assess which workflows warrant premium low-latency access and establish appropriate governance controls before implementing the Ultrafast capability.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.