{
  "id": 12563005,
  "title": "⚡️ 675 B LLM Shatters Expectations – One GPU, Sub‑Second Latency!",
  "url": "https://urgent.news/2026/10/07/675-b-llm-shatters-expectations-one-gpu-sub-second-latency",
  "topic": "ai",
  "section": "AI",
  "published": "2026-10-07T06:17:18.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/amrithesh_dev/675-b-llm-shatters-expectations-one-gpu-sub-second-latency-1ic7"
  },
  "original_language": "en",
  "account": "Mistral Large 4, a 675-billion-parameter open-weight LLM, was released by Mistral AI on September 30, 2024. This model is notable for running on a single A100 GPU with sub-second latency and comes under a permissive research license. The launch of ML-4 has set a new benchmark for open-source AI, as it outperforms many large models in terms of accuracy and performance.\n\nA fintech startup was able to use ML-4 to improve code completion in a risk-engine IDE. They were able to achieve a 4.3-point improvement in HumanEval benchmark scores compared to using a 70-billion-parameter Llama 3 model, with lower latency and cheaper token costs.\n\nKey technical innovations in ML-4 include Grouped-Query Attention (GQA) and Sliding-Window Attention (SWA). GQA reduces the attention matrix size by grouping queries, while SWA reduces the computational cost of self-attention on long sequences. These methods, combined with 4-bit quantization and zero-copy offload, enable the model to run efficiently on a single A100 GPU.\n\nDespite the impressive capabilities of ML-4, there are some open questions. The hype around the parameter count may not translate to superior performance, as diminishing returns are observed beyond ~700 billion parameters. Nonetheless, ML-4 demonstrates that large-scale open-source LLMs can be both powerful and practical for real-world applications.",
  "summary": "# Mistral Large 4: The First Open‑Weight Giant That Actually Works at Scale The Lead A 675‑billion‑parameter model that you can pull from Hugging Face today still turns heads in every AI‑focused Slack channel I monitor. On 30 September 2024 , Mistral AI turned that headline into reality with the general‑availability launch of Mistral Large 4 (ML‑4) —an open‑weight LLM that immediately eclipsed…",
  "key_points": [
    "Mistral AI releases 675-billion-parameter open-weight LLM, ML-4, on September 30, 2024.",
    "ML-4 runs on a single A100 GPU with sub-second latency, outperforming many large models.",
    "Fintech startup improves code completion by 4.3 points using ML-4 over 70-billion-parameter Llama 3."
  ],
  "editors_take": "The release of Mistral Large 4 sets a new benchmark for open-source AI, enabling large-scale models to be both powerful and practical for real-world applications with efficient hardware requirements.",
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}