{
  "id": 44132,
  "title": "I'm (mostly) picking models on speed now, not intelligence",
  "url": "https://urgent.news/2026/08/02/im-mostly-picking-models-on-speed-now-not-intelligence",
  "topic": "tech",
  "section": "Tech",
  "published": "2026-08-02T13:49:23.000Z",
  "source": {
    "name": "Lobsters",
    "slug": "lobsters",
    "url": "https://martinalderson.com/posts/speed-vs-intelligence/"
  },
  "original_language": "en",
  "account": "In a recent shift in approach, the author is no longer prioritizing raw intelligence when selecting daily driver models, but rather focusing on their speed. This new strategy seems to be working well for most daily tasks, including code, research, and data analysis. The author recalls a time when models around the ~Opus 4.6 level were deemed intelligent enough, but the recent US government shutdown provided an opportunity to test out Opus again after being excited about Fable. However, Fable proved to be disappointingly slow, leading the author back to Opus.\n\nThe author emphasizes the importance of speed in software interaction, stating that even the most beautiful product becomes frustrating when it's slow. Conversely, a basic product can feel incredibly efficient if it's fast. This highlights the author's belief that speed is crucial for ensuring a positive user experience.\n\nThe author draws upon their career in software development, where they've learned that fast software feels much better to use. They contrast this with the feeling of working with an outdated or slow model, which can be reminiscent of using dial-up internet. The author expects that as models become faster, they will reach a point where a human can process the output in real-time, around 100-200 tokens per second.\n\nLooking at the speed rankings of various models on platforms like OpenRouter, the author notes a wide range of serving speeds, from under 30 tokens per second to over 129 tokens per second. The open weights ecosystem offers a diverse selection of models with varying speeds, providing customers with more options and potentially better value.\n\nThe author also touches on the trade-off between model speed and the time spent on tool calls, such as local machine processing. While a 5x speedup in model processing may only result in a 2x speedup in overall turn processing due to the time spent on tool calls, this bottleneck could eventually limit the benefits of faster models.\n\nAs the author anticipates the release of new GPUs like Nvidia's Vera Rubin series and AMD's MI400s, with improved memory bandwidth, they believe that models will reach speeds of 500 tokens per second or higher in the future. This could make the current sweet spot of speed and affordability seem primitive. The author questions whether models of this caliber will truly outperform the current \"good enough\" models in terms of everyday tasks and the impact of more intelligence on the user experience.",
  "summary": null,
  "key_points": [
    "Author prioritizes speed over intelligence in model selection for daily tasks.",
    "Recent US government shutdown allowed testing of Opus after Fable disappointment.",
    "Fast models enable real-time human processing of output around 100-200 tokens per second."
  ],
  "editors_take": null,
  "illustration": "https://urgent.news/ill/44132.png",
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}