Urgent.News

One page, thousands of outlets. See who else covered it.

Editions

AI

DeepSeek's top-ranked V4 Flash stumbles on real agent tasks as its prices surge

DeepSeek's V4 Flash has topped model leaderboards and been hailed by developers as a "total monster" since its rollout. But in real-world testing, it completed just 53.8% of a batch of complex agent tasks. Composio ran the model through eight different agent harnesses , including Claude Code, Codex, and OpenCode, on 30 deliberately difficult, multi-step tasks spanning live tools like Gmail,…

DeepSeek's top-ranked V4 Flash stumbles on real agent tasks as its prices surge

DeepSeek's V4 Flash, once hailed as a top-performing model, is struggling with real-world agent tasks, completing only 53.8% of a batch of complex tasks. Multiple agent harnesses, including Claude Code, Codex, and OpenCode, were used to test the model on 30 difficult workflows involving tools like Gmail, GitHub, Slack, and Google Sheets.

The model's performance varied greatly depending on the harness, tool configuration, caching behavior, retries, and provider stack it ran on. DeepSeek is hiking the prices for V4 Flash and Pro, two models that have gained popularity among developers building coding assistants and agents. The new pricing structure increases costs significantly, ranging from 57% to 371% for Flash and 51% to 355% for Pro.

While this may undermine the appeal of the ultra-low-priced models, it moves the discussion beyond the "cheap Chinese model" narrative, as enterprises begin to explore the best use cases for different models in their tech stacks.

Written by urgent.news from VentureBeat's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at venturebeat.com →

More in AI

A Litigant Hid White-Text Prompt Injection in a Court Filing. A Human Caught It, Not an AI.

A court employee in Connecticut noticed some odd whitespace in a legal filing. That's it. That's the entire detection mechanism that stood between a working prompt injection attack and whatever AI…

  • Litigant hid white-text prompt injection in court filing
  • Human detected unusual whitespace, flagged filing
  • Attack exploits low-tech methods, AI systems read hidden instructions

More from Sunday 16 August →