Urgent.News

What's breaking now, across thousands of outlets.

AI

OpenAI's first custom chip just benchmarked past NVIDIA. Jalapeño changes the inference equation.

OpenAI just published benchmark results for Jalapeño — its first custom AI inference chip — and the numbers are credible and significant. Compared to leading NVIDIA hardware (GB200 and GB300), Jalapeño delivers better throughput, better latency, and better power efficiency simultaneously. Not a tradeoff between them. All three at once. "Jalapeño delivers both higher throughput and lower latency…

OpenAI unveiled its first custom AI inference chip, Jalapeño, and the results are impressive. Compared to NVIDIA's leading hardware (GB200 and GB300), Jalapeño offers better throughput, lower latency, and better power efficiency all at once, without any tradeoffs. The chip delivers 1.5–1.9x more throughput per watt, 1.7–3.6x lower end-to-end latency, and 2.1–4.1x higher performance for interactive low-latency workloads.

The benchmark tests were conducted on the InferenceX public benchmark across three models: GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T. Notably, latency dropped significantly from 5.99s to 1.65s end-to-end for DeepSeek R1. With a rated power of 700W, Jalapeño ran at or below 550W during the tests, providing half the power with a faster output compared to the GB300 which runs at 1,400W.

The chip is designed to handle language model inference phases – prefill (compute-heavy) and decode (memory-bandwidth-heavy) – on the same balanced chip, minimizing data movement between phases. This architecture optimizes for the inference process of language models, making every 3x latency improvement on a single call more significant when chaining 10 calls together.

The chip was designed, programmed, and produced by OpenAI in just nine months, with AI-generated kernels running 1.5–1.8x faster than human-expert implementations for selected attention and MoE blocks. Currently, using the OpenAI API will not change, as Jalapeño will only deploy within OpenAI's infrastructure by the end of 2026.

However, running agents or multi-step workflows will see significant gains, as the faster per-call inference adds up quickly across a full agentic loop. This custom chip also proves that custom silicon for LLM inference is viable and competitive, posing a challenge to existing players like NVIDIA, Google TPU, Amazon Trainium, and Meta MTIA.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

This story

This is one outlet's version. Read the fullest account.

Read the original at dev.to →

More in AI

No signup. No card. Free GPT & Claude.

Just go to: https://Duck.ai Duck.ai (by DuckDuckGo ) lets you chat with several AI models for free. Your requests are anonymized, and chats aren’t stored on DuckDuckGo’s servers by default.

More from Wednesday 26 August →