⚡️ 675 B LLM Shatters Expectations – One GPU, Sub‑Second Latency!
# Mistral Large 4: The First Open‑Weight Giant That Actually Works at Scale The Lead A 675‑billion‑parameter model that you can pull from Hugging Face today still turns heads in every AI‑focused Slack channel I monitor. On 30 September 2024 , Mistral AI turned that headline into reality with the general‑availability launch of Mistral Large 4 (ML‑4) —an open‑weight LLM that immediately eclipsed…
Mistral Large 4, a 675-billion-parameter open-weight LLM, was released by Mistral AI on September 30, 2024. This model is notable for running on a single A100 GPU with sub-second latency and comes under a permissive research license. The launch of ML-4 has set a new benchmark for open-source AI, as it outperforms many large models in terms of accuracy and performance.
A fintech startup was able to use ML-4 to improve code completion in a risk-engine IDE. They were able to achieve a 4.3-point improvement in HumanEval benchmark scores compared to using a 70-billion-parameter Llama 3 model, with lower latency and cheaper token costs.
Key technical innovations in ML-4 include Grouped-Query Attention (GQA) and Sliding-Window Attention (SWA). GQA reduces the attention matrix size by grouping queries, while SWA reduces the computational cost of self-attention on long sequences. These methods, combined with 4-bit quantization and zero-copy offload, enable the model to run efficiently on a single A100 GPU.
Despite the impressive capabilities of ML-4, there are some open questions. The hype around the parameter count may not translate to superior performance, as diminishing returns are observed beyond ~700 billion parameters. Nonetheless, ML-4 demonstrates that large-scale open-source LLMs can be both powerful and practical for real-world applications.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.