OpenAI’s Jalapeño chip is built for fast inference at scale, benchmarks show
Tested on Semianalysis’s InferenceX benchmark, Jalapeño registered both more tokens per user and more throughput per kilowatt than the currently available state-of-the art.
At the Hot Chips conference on Tuesday, OpenAI unveiled Jalapeño, a new chip designed for rapid inference at scale. The company released initial benchmark results for the chip, which outperformed the current state-of-the-art inference processors in terms of tokens per user and throughput per kilowatt. Richard Ho, OpenAI's head of hardware, stated, "The bottom line is that the results show a very, very significant performance advance over state of the art."
The benchmark tests were conducted against an Nvidia Blackwell system, but by the time Jalapeño is fully deployed, the competitive landscape may have shifted. OpenAI estimated that Jalapeño would enter limited production by the end of 2026, with wider release in 2027. Developed jointly with Broadcom, OpenAI's own models played a role in the chip's development.
The company aims to create a multigenerational platform, integrating AI products, models, chips, and memory. This full-stack approach allowed OpenAI to tackle specific challenges in the inference process, particularly reducing delays in the prefill and communication phases. By minimizing data movement and communication delays, Jalapeño can keep model state, like the KV cache used during response generation, local while activating the optimal compute, memory, and networking resources for each inference phase.
Written by urgent.news from TechCrunch's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.
This story
This is one outlet's version. Read the fullest account.
- OpenAI's upcoming Jalapeño chip looks like it'll be an inference beast theregister.com
- OpenAI says its Jalapeno AI chip delivers faster responses than rivals like Nvidia digitaltrends.com
- OpenAI says its Jalapeño chip delivered 1.5x-1.9x more AI work per watt and 1.7x-3.6x lower latency vs. Nvidia chips across GPT-OSS, DeepSeek R1, Kimi K2.5 1T (Emma Roth/The Verge) theverge.com
- A detailed look at Jalapeño, OpenAI's ASIC developed with Broadcom in 16 months, which beat Nvidia, AMD, and Google chips on multiple top open-source models (SemiAnalysis) newsletter.semianalysis.com