Urgent.News

What's breaking now, across thousands of outlets.

AI

OpenAI's upcoming Jalapeño chip looks like it'll be an inference beast

128 chips, 1.7 exaFLOPS, and 27 TB of HBM give Altman and crew a leg up over Blackwell, and maybe even Rubin

OpenAI's upcoming Jalapeño chip looks like it'll be an inference beast

OpenAI has unveiled its innovative Jalapeño AI accelerator at the Hot Chips conference, a custom chip developed in partnership with Broadcom. This marks the first in a series of custom silicon chips being created by OpenAI, designed (in part) by AI for AI purposes. Compared to Nvidia's GPU systems, the Jalapeño chip is projected to offer superior throughput and lower latency when released later this year and enter volume production in 2027.

However, OpenAI will continue to utilize existing hardware partners like AMD and Nvidia for their training needs, gradually transitioning to their own in-house silicon. Memory bandwidth is a key factor in inference tasks, and the Jalapeño chip appears to excel in this area. Benchmarks indicate that the chip delivers between 1.5x and 1.9x more "AI work" at peak throughput, and 1.7x to 3.6x lower end-to-end latency compared to its competitors, including GPT-OSS-120B, DeepSeek R1, and Kimi K2.5.

When it comes to ultra-low-latency inference, OpenAI claims their chips are 2.1x to 4.1x faster. Each Jalapeño system boasts 1.7 exaFLOPS of 4-bit compute, 27.5 TB of HBM4 memory, and a memory bandwidth of nearly 2 petabytes per second. While power consumption details are not yet available, it is estimated that each rack will consume between 40 and 60 percent of the power of competing GPU systems.

Despite its impressive performance, the Jalapeño chip is not designed to replace OpenAI's existing hardware partners, but rather to complement them by excelling specifically in inference tasks.

Written by urgent.news from The Register's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

This story

This is one outlet's version. Read the fullest account.

Read the original at theregister.com →

More in AI

How I Built a Reliable LLM Pipeline for Ad Creative Evaluation (with Strict Pydantic Contracts)

Art directors routinely spend 20–40 minutes per creative just checking brand guidelines, mandatory elements, and forbidden techniques.

  • CreativeAudit pipeline automates ad creative evaluation, returning PASS/NEEDSREVISION/FAIL verdicts
  • Pydantic contracts ensure strict output format, preventing invalid responses from becoming scores
  • Binary compliance for critical brand rules, scoring system assigns 0 or 10 for violations

More from Tuesday 25 August →