Urgent.News

What's breaking now, across thousands of outlets.

AI

OpenAI’s Jalapeño chip is built for fast inference at scale, benchmarks show

Tested on Semianalysis’s InferenceX benchmark, Jalapeño registered both more tokens per user and more throughput per kilowatt than the currently available state-of-the art.

At the Hot Chips conference on Tuesday, OpenAI unveiled Jalapeño, a new chip designed for rapid inference at scale. The company released initial benchmark results for the chip, which outperformed the current state-of-the-art inference processors in terms of tokens per user and throughput per kilowatt. Richard Ho, OpenAI's head of hardware, stated, "The bottom line is that the results show a very, very significant performance advance over state of the art."

The benchmark tests were conducted against an Nvidia Blackwell system, but by the time Jalapeño is fully deployed, the competitive landscape may have shifted. OpenAI estimated that Jalapeño would enter limited production by the end of 2026, with wider release in 2027. Developed jointly with Broadcom, OpenAI's own models played a role in the chip's development.

The company aims to create a multigenerational platform, integrating AI products, models, chips, and memory. This full-stack approach allowed OpenAI to tackle specific challenges in the inference process, particularly reducing delays in the prefill and communication phases. By minimizing data movement and communication delays, Jalapeño can keep model state, like the KV cache used during response generation, local while activating the optimal compute, memory, and networking resources for each inference phase.

Written by urgent.news from TechCrunch's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

This story

This is one outlet's version. Read the fullest account.

Read the original at techcrunch.com →

More in AI

More from Tuesday 25 August →