DeepSeek debuts multimodal language model competitive with Opus 4.8
DeepSeek today debuted a new addition to its flagship V4 series of large language models. On launch, V4 Flash Vision Exp is only available via the Chinese startup’s paid developer platform. The company may release a free version later on given that it has open-sourced many of its earlier models. Those models include V4 Flash, […] The post DeepSeek debuts multimodal language model competitive with…
DeepSeek has unveiled a new iteration of its V4 series of large language models, known as V4 Flash Vision Exp. The model was exclusively released to the company’s paid developer platform at its launch. However, DeepSeek may provide a free version in the future, as they have previously open-sourced several of their earlier models, including V4 Flash which serves as the foundation for V4 Flash Vision Exp.
Comparative testing across seven benchmarks revealed that V4 Flash Vision Exp outperformed its predecessor, V4 Flash, in all tests, except one – the Cybergym benchmark, which assesses an LLM’s capacity to identify software vulnerabilities. The model excelled significantly more in image analysis, with a notable 10% improvement on two out of four image analysis benchmarks.
Notably, V4 Flash Vision Exp also outperformed Anthropic’s Opus 4.8 in two visual benchmarks – ALE and ZeroBench. ALE features more than 1,000 multi-step tasks that require LLMs to interact with applications, write code, and interpret media files, while ZeroBench contains 100 highly challenging image analysis tasks for frontier LLMs.
Although DeepSeek has not disclosed the architecture of V4 Flash Vision Exp, its Hugging Face page provides detailed information on the V4 Flash model. V4 Flash is a mixture of experts model consisting of 284 billion parameters, comprising multiple neural networks, each with 13 billion parameters. Upon receiving a prompt, the LLM activates only the neural network best suited to generate an answer, thereby using considerably less hardware compared to running the entire LLM.
To optimize the KV cache, which holds information for answering prompts, DeepSeek utilized two techniques – HCA and CSA – resulting in a 73% reduction in the computing power required to process prompts with 1 million tokens. Trained on 32 trillion tokens, V4 Flash employs the Muon algorithm to expedite the training process by minimizing the time needed to calibrate LLM’s hidden layers.
V4 Flash is one of two LLMs released by DeepSeek in April, with the other being V4 Pro, which boasts over five times more parameters. Given that V4 Flash Vision Exp is derived from V4 Flash, it is possible that V4 Pro will eventually serve as the basis for specialized models optimized for tasks like image analysis.
Written by urgent.news from SiliconANGLE's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.