{
  "id": 9079008,
  "title": "Fastest LLM 2026: Mercury 2.5 Beats Luna and Haiku",
  "url": "https://urgent.news/2026/09/22/fastest-llm-2026-mercury-2-5-beats-luna-and-haiku",
  "topic": "ai",
  "section": "AI",
  "published": "2026-09-22T03:44:49.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/shaam_ai/fastest-llm-2026-mercury-25-beats-luna-and-haiku-2nef"
  },
  "original_language": "en",
  "account": "In September 2026, Inception Labs' Mercury 2.5 emerged as the fastest LLM you can call from an API, boasting an impressive 1,107 tokens per second on widely available NVIDIA GPUs. This competitive speed comes at a price of $0.20/$0.75 per million tokens, which is more affordable than other options in the market. However, for general-purpose chat with long context, GPT-5.6 Luna still leads the pack. Google's Gemini 3.5 Flash-Lite sits in between the two in terms of speed, but it comes with a higher price tag of $0.30/$2.50 per million tokens. Claude Haiku 4.5, on the other hand, is the slowest and most expensive option, with a rate of around 82 tok/s non-reasoning and a cost of $1.00/$5.00 per million tokens.",
  "summary": "Verdict: Inception Labs' Mercury 2.5 is the fastest LLM you can call from an API in September 2026 — 1,107 tokens per second on widely available NVIDIA GPUs, at a list price of $0.20/$0.75 per million tokens that undercuts every rival here on output. For latency-bound work (voice agents, search pipelines, coding subagents), Mercury wins. For cheap general-purpose chat with long context, GPT-5.6…",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}