Urgent.News

What's breaking now, across thousands of outlets.

AI

Hot Chips 2026: Intel dives deep on Crescent Island AI accelerator — larger caches and deeper XMX engines target maximum AI FLOPS per watt

At Hot Chips 2026, Intel detailed more about its Crescent Island AI accelerator, which uses the Xe3P architecture. The accelerator will use liquid-cooled chips and HBM4 memory to serve inference workloads in data centers.

Hot Chips 2026: Intel dives deep on Crescent Island AI accelerator — larger caches and deeper XMX engines target maximum AI FLOPS per watt

Intel unveiled further details of its Crescent Island AI accelerator during the Hot Chips symposium. Unlike Nvidia's power-hungry GPUs and AMD's MI455X, Crescent Island targets the lower-power, inference-first segment of the AI accelerator market. This 350W air-cooled PCIe card utilizes up to 480 GB of LPDDR5X memory, enabling deployment in standard servers without special power and cooling needs.

Crescent Island comprises four Xe3P slices, each housing eight Xe Cores, resulting in a total of 32 cores. Every Xe Core features eight Xe Vector Engines and eight XMX matrix accelerators, totaling 256 of each resource. The Xe3P architecture refines the GPU cache hierarchy, similar to that used in Intel's Panther Lake processors, aiming to boost utilization and reduce register spills.

Each Xe Core boasts 1MB of general-purpose register file space, a significant increase from the 512KB on Battlemage and Xe2. It also includes 512KB of L1 cache per Xe Core, up from 256KB on Battlemage, and 32MB of shared L2 cache. These expanded caches support the chip's emphasis on matrix accelerators for AI compute-focused tasks.

Xe3P's XMX systolic engines have a 16-deep design, allowing for processing of larger matrix chunks compared to the four-deep systolic design of Xe2 and Xe3. This capability is crucial for its inference-focused objectives. The chip supports various data types, from FP4 formats with microscaling support (MXFP4) to full-rate double-precision (64 FP64 FMA units per Xe Core), making it suitable for both AI and high-performance computing.

Intel also integrated sigmoid and tanh transcendental functions into Xe3P, essential for many AI inference operations. While the chip lacks graphics-specific features such as RT cores, it offers a comprehensive suite of reliability, availability, and serviceability features, including ECC and parity protection across the die and memory reliability features.

Intel positions Crescent Island as a high FLOPS per watt data-center-focused component, optimized for compute-bound workloads like prefill, which involves processing prompt data and key-value cache construction. The chip's design prioritizes compute over memory bandwidth, making it well-suited for speculative decoding methods that generate draft tokens for model serving. This approach can improve decode performance by utilizing compute resources that would otherwise remain idle.

Written by urgent.news from Tom's Hardware's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at tomshardware.com →

More in AI

More from Tuesday 25 August →