Urgent.News

the world's headlines, one feed

AI

AMD acquires AI chip startup Taalas to boost inference performance by etching models into silicon

Early tech demos show model-specific integrated circuits churning out up to 17,000 tokens a second

AMD acquires AI chip startup Taalas to boost inference performance by etching models into silicon

Amidst its ongoing competition with Nvidia in the AI hardware market, AMD has acquired AI chip startup Taalas. The primary objective of this acquisition seems to be to enhance the performance of inference services, which are crucial for AI applications such as code assistants. Despite AMD not disclosing the deal's terms, it is believed that this is a full acquisition rather than a simple acquihire.

Taalas, founded in 2023 and located in Toronto, employs a unique approach to inference that diverges from traditional GPUs or dataflow architectures used by companies like Groq and Cerebras. Instead of relying on high-bandwidth memory (HBM) to store model weights, Taalas integrates them directly into the silicon. This technology is termed as model-specific integrated circuits or MSICs.

The startup's first test chip, HC1, built on TSMC's 6nm process, demonstrated impressive capabilities. In February 2024, it could process Meta's Llama 3.1 8B model at an astounding 16,960 tokens per second, outperforming Nvidia's GPUs by 48 times and Cerebras' accelerators by 8.5 times. Although Taalas remains tight-lipped about its chip's functionality, it comprises two main regions: one for storing model weights (mask-ROM recall fabric) and another for storing key-value caches and fine-tuning adapters (SRAM recall fabric).

AMD intends to pair its Instinct-based Helios racks with Taalas' accelerators, suggesting a disaggregated architecture where compute-heavy tasks are handled by GPUs and token generation is processed by Taalas accelerators. This setup could potentially lead to a tick-tock cadence where models are initially deployed on Instinct accelerators and later transitioned to Taalas accelerators.

However, it's worth noting that this technology comes with a significant drawback: once deployed, models are locked into the chips, and any substantial model changes would require a re-spin of the chips, which can be both costly and time-consuming.

Written by urgent.news from The Register Science's reporting — not their text. Machine-written; read the original for the full account.

Also reported by 5 other outlets

Read the original at theregister.com →

More in AI

AI models design viruses not found in nature for first time

AI models design viruses not found in nature for first time

Researchers from Stanford University have synthesised brand-new, self-replicating viruses using genomes designed by artificial intelligence for the first time, raising the possibility of significant medical advancements but also prompting concerns the technology could be misused.