Urgent.News

600+ sources. One page. See who else covered it.

Editions

Tech

I'm (mostly) picking models on speed now, not intelligence

Abstract editorial illustration

In a recent shift in approach, the author is no longer prioritizing raw intelligence when selecting daily driver models, but rather focusing on their speed. This new strategy seems to be working well for most daily tasks, including code, research, and data analysis. The author recalls a time when models around the ~Opus 4.6 level were deemed intelligent enough, but the recent US government shutdown provided an opportunity to test out Opus again after being excited about Fable. However, Fable proved to be disappointingly slow, leading the author back to Opus.

The author emphasizes the importance of speed in software interaction, stating that even the most beautiful product becomes frustrating when it's slow. Conversely, a basic product can feel incredibly efficient if it's fast. This highlights the author's belief that speed is crucial for ensuring a positive user experience.

The author draws upon their career in software development, where they've learned that fast software feels much better to use. They contrast this with the feeling of working with an outdated or slow model, which can be reminiscent of using dial-up internet. The author expects that as models become faster, they will reach a point where a human can process the output in real-time, around 100-200 tokens per second.

Looking at the speed rankings of various models on platforms like OpenRouter, the author notes a wide range of serving speeds, from under 30 tokens per second to over 129 tokens per second. The open weights ecosystem offers a diverse selection of models with varying speeds, providing customers with more options and potentially better value.

The author also touches on the trade-off between model speed and the time spent on tool calls, such as local machine processing. While a 5x speedup in model processing may only result in a 2x speedup in overall turn processing due to the time spent on tool calls, this bottleneck could eventually limit the benefits of faster models.

As the author anticipates the release of new GPUs like Nvidia's Vera Rubin series and AMD's MI400s, with improved memory bandwidth, they believe that models will reach speeds of 500 tokens per second or higher in the future. This could make the current sweet spot of speed and affordability seem primitive. The author questions whether models of this caliber will truly outperform the current "good enough" models in terms of everyday tasks and the impact of more intelligence on the user experience.

Written by urgent.news from Lobsters's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at martinalderson.com →

More in Tech

More from Sunday 2 August →