OpenAI built a chip in nine months. Then it let AI rewrite the code.
When OpenAI unveiled Jalapeño, its first custom inference chip, in June, the company made some big promises. The chip, developed The post OpenAI built a chip in nine months. Then it let AI rewrite the code. appeared first on The New Stack .
OpenAI unveiled Jalapeño, its first custom inference chip, in June, promising substantial performance improvements. The company, in collaboration with Broadcom, developed the chip specifically for large language model inference. However, OpenAI withheld detailed performance results at the time. Now, OpenAI has published its first results, showcasing the chip's capabilities with GPT-OSS 120B, DeepSeek R1, and Kimi K2.5.
The findings confirm OpenAI's objectives: higher throughput without longer response times. Agents completing multiple steps in sequence can experience compounded delays, but Jalapeño addresses this by reducing waiting times between cores and chips. The chip keeps model state local while optimizing data movement within the system, resulting in faster processing.
OpenAI's benchmarks reveal that Jalapeño achieves 1.5 to 1.9 times more work per watt and reduces end-to-end latency by 1.7 to 3.6 times across the three models. In highly interactive workloads, it outperforms competitors by 2.1 to 4.1 times. The power efficiency comparisons, based on each accelerator's published power rating, show Jalapeño delivering between 8.6 and 104.3 times more work per watt, depending on the model.
Notably, OpenAI's development of Jalapeño accelerated AI-generated code. By leveraging AI-generated implementations, the hardware team shortened design, measurement, and verification cycles, cutting the time from initial design to tapeout to nine months. OpenAI plans to integrate Jalapeño into its infrastructure by the end of the year and is already working on subsequent generations.
While OpenAI will continue using accelerators from Nvidia and other partners, building its own chips grants the company greater control over the hardware evolution alongside its models.
Written by urgent.news from The New Stack's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.