Urgent.News

What's breaking now, across thousands of outlets.

AI

Presentation: Producing the World's Cheapest Tokens: A How-to Guide

Meryem Arik discusses strategies for designing low-cost LLM inference architectures for high-volume, non-real-time workloads. She explains how software architects and engineering leaders can achieve order-of-magnitude cost reductions by making critical trade-offs across hardware, inference runtimes, speculative decoding, and smart queue reordering. By Meryem Arik

Meryem Arik, founder and CEO of Doubleword, a self-hosted AI inference platform, discusses strategies for producing the world's cheapest tokens in a low-cost LLM inference architecture. Arik, an Oxford University alumnus, frequently speaks at leading conferences such as TEDx, QCon, and Forbes 30 Under 30. In her presentation, she explains how software architects and engineering leaders can achieve significant cost reductions by making critical trade-offs across hardware, inference runtimes, speculative decoding, and smart queue reordering.

Arik highlights that most companies overpay for inference by 2x to 5x, sometimes even an order of magnitude, for their highest token use cases. She argues that companies often fail to optimize their inference stack according to their specific use case priorities. Arik aims to walk the audience through the process of designing inference systems from scratch, focusing on producing tokens at the lowest possible cost.

Arik begins by emphasizing the importance of striking a balance between latency, cost, and token quality. Depending on the use case, one might prioritize either low latency, low cost, or high token quality. For instance, coding assistants benefit from low latency and high-quality outputs, while cost can be traded off. Conversely, real-time use cases, such as guardrails or on-device models, may require a different trade-off.

The focus of Arik's talk is on trading off latency and cost, aiming for high token quality and very low costs. She explores various techniques such as batch-specific optimizations, queue reordering, smart scheduling, and model quantization. As models become more advanced, the number of use cases requiring independent operation of models without constant human supervision is expected to grow, making this trade-off an increasingly popular choice.

Arik concludes by comparing the pursuit of ultra-low-latency inference providers, like Cerebras and Groq, to traditional, chatbot-style inference frameworks. She emphasizes that the latter offers a cost-effective alternative for the specific use case of producing the world's cheapest tokens while maintaining high quality.

Written by urgent.news from InfoQ's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at infoq.com →

More in AI

More from Tuesday 11 August →