Mistral unveils Large 4 AI model ahead of wider release
Mistral said it trained the model over two months on 4,000 Nvidia Grace Blackwell GPU.
Mistral AI SAS opened access to its most capable large language model, Mistral Large 4, on launch through its cloud platform. The model features a mixture of experts architecture with 1 trillion parameters, but only uses 49 billion at a time for computational efficiency. Mistral Large 4 can answer questions across over 160 languages and earned an AA Cyber Index score of 82%, outperforming open-source rivals.
Mistral trained the model on 3,800 Grace Blackwell chips, each combining two Nvidia Blackwell graphics cards with one CPU. Advanced LLMs are trained through trial and error, with the model attempting tasks without human guidance and receiving feedback from an AI model to refine its results. Mistral's software stack enabled parallel rollouts of the model, producing 33 billion tokens per day.
While Mistral Large 4 outperformed several leading open-source LLMs on automation benchmarks, it falls behind frontier models like Astra on coding benchmarks. However, Mistral AI plans to build more specialized models optimized for specific use cases using Mistral Large 4.
Written by urgent.news from SiliconANGLE's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.
This story
This is one outlet's version. Read the fullest account.
- Mistral launches open-source Mistral Large 4, details AI roadmap siliconangle.com