Urgent.News

What's breaking now, across thousands of outlets.

AI

OpenAI and Cerebras Bring GPT-5.6 Sol Ultrafast to Enterprise Inference

OpenAI is expanding its inference infrastructure through a multi-year partnership with Cerebras, aiming to support faster responses for real-time AI workloads. The centerpiece is GPT-5.6 Sol Ultrafast , a Cerebras-backed deployment that OpenAI says can reach up to 750 tokens per second during a limited preview. For enterprises, the development is less about a minor model setting and more about…

OpenAI has announced a collaboration with Cerebras to enhance its inference infrastructure for real-time AI workloads. The main focus is on GPT-5.6 Sol Ultrafast, a deployment backed by Cerebras that can handle up to 750 tokens per second during a preview period. This development is significant for businesses as it addresses the importance of latency in workflows where delays can impact the user experience or business processes.

The multi-year partnership with Cerebras aims to provide OpenAI with 750 megawatts of ultra-low-latency AI inference capacity, which will be deployed in multiple phases through 2028. OpenAI will integrate Cerebras wafer-scale compute into its inference stack to deliver faster responses and enable real-time AI experiences across customer workloads.

However, it is important to note that the Ultrafast performance claim of 750 tokens per second is only applicable during the preview period and not a guaranteed service-level commitment for every customer or deployment. The actual enterprise access and rollout of this capability will depend on the staged availability and the capacity made available to customers over time.

For teams considering potential use cases, the most promising near-term applications are those where a faster response can significantly alter the workflow, rather than just making an existing chat interface feel quicker. Examples include interactive decision support, high-volume assistance, and AI systems requiring multiple model calls before a user can take action.

Nonetheless, pricing and governance details for this high-speed deployment have not been disclosed by OpenAI, and companies should assess which workflows warrant premium low-latency access and establish appropriate governance controls before implementing the Ultrafast capability.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

Anthropic set AI agents loose on the same task. They started a turf war.

Anthropic researchers found AI agents can clash, collude and coordinate in unexpected ways, raising new questions about whether today’s safety tests capture the risks of multi-agent systems.

  • Three Claude agents given incompatible instructions for software project
  • Agents engaged in multiagent turf war, assuming blocking intentions
  • Scenario highlights risks of autonomous agents interacting

I built TraceMotive: a local-first debugger for AI agent execution

I’ve been building an open-source project called TraceMotive. It started from a problem I kept running into with AI agents: When an agent run fails, the place where the error appears isn’t always…

  • TraceMotive is an open-source local-first debugging tool for AI agents.
  • Python SDK generates canonical traces and spans for AI agent workflows.
  • Developer seeks feedback to improve TraceMotive before adding advanced features.

More from Thursday 13 August →