DeepSeek's top-ranked V4 Flash stumbles on real agent tasks as its prices surge
DeepSeek's V4 Flash has topped model leaderboards and been hailed by developers as a "total monster" since its rollout. But in real-world testing, it completed just 53.8% of a batch of complex agent tasks. Composio ran the model through eight different agent harnesses , including Claude Code, Codex, and OpenCode, on 30 deliberately difficult, multi-step tasks spanning live tools like Gmail,…
DeepSeek's V4 Flash, once hailed as a top-performing model, is struggling with real-world agent tasks, completing only 53.8% of a batch of complex tasks. Multiple agent harnesses, including Claude Code, Codex, and OpenCode, were used to test the model on 30 difficult workflows involving tools like Gmail, GitHub, Slack, and Google Sheets.
The model's performance varied greatly depending on the harness, tool configuration, caching behavior, retries, and provider stack it ran on. DeepSeek is hiking the prices for V4 Flash and Pro, two models that have gained popularity among developers building coding assistants and agents. The new pricing structure increases costs significantly, ranging from 57% to 371% for Flash and 51% to 355% for Pro.
While this may undermine the appeal of the ultra-low-priced models, it moves the discussion beyond the "cheap Chinese model" narrative, as enterprises begin to explore the best use cases for different models in their tech stacks.
Written by urgent.news from VentureBeat's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.