DeepSeek’s updated V4 Pro AI model struggles on benchmarks, shines in cybersecurity
Chinese artificial intelligence start-up DeepSeek has quietly released DeepSeek-V4-Pro-0813, an updated version of its latest flagship model, leaving some developers underwhelmed by its overall capabilities and disappointed in its pricing – but impressing researchers in niche areas like cybersecurity. The stealth update to April’s preview version came with a brief statement on DeepSeek’s official…
Chinese AI startup DeepSeek has recently unveiled DeepSeek-V4-Pro-0813, a refined iteration of its flagship model. While developers have expressed mixed reactions to the release, researchers in specialized fields like cybersecurity have been notably impressed. The update, announced on DeepSeek's official website, touts "significantly enhanced agent capabilities" but was later withdrawn.
Following the launch of DeepSeek's cheaper V4 Flash model, which garnered attention for its cost-efficiency, the new flagship model has not met expectations in benchmark tests. It scored 53 on the Artificial Analysis Intelligence Index, matching Zhipu AI's GLM-5.2 but lagging behind OpenAI's GPT-5.6 and Moonshot AI's Kimi K3. In Vals AI's evaluation, the model ranked 12th, trailing behind OpenAI's GPT-5.5 and frontier systems like Kimi K3 and Anthropic's Claude Opus 5.
DeepSeek-V4-Pro-0813 encountered difficulties in completing tasks within a sandboxed terminal and generating complex financial models in Excel, according to Vals AI. The model struggled with premature completion during lengthy coding tasks. Social media reactions were largely negative, with some developers labeling the model as "disappointing."
However, cybersecurity analysts noted DeepSeek-V4-Pro-0813's superior performance in discovering system vulnerabilities. Belgian firm Aikido Security reported that the model outperformed rivals in detecting vulnerabilities, even though it exhibited "poor precision." The model's pricing has also garnered attention, with DeepSeek charging 44 US cents per million input tokens, described as "somewhat expensive" compared to the median of 33 US cents.
Written by urgent.news from SCMP Tech's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.