{
  "id": 1272272,
  "title": "DeepSeek's top-ranked V4 Flash stumbles on real agent tasks as its prices surge",
  "url": "https://urgent.news/2026/08/16/deepseeks-top-ranked-v4-flash-stumbles-on-real-agent-tasks-as-its",
  "topic": "ai",
  "section": "AI",
  "published": "2026-08-16T13:00:00.000Z",
  "source": {
    "name": "VentureBeat",
    "slug": "venturebeat",
    "url": "https://venturebeat.com/orchestration/deepseeks-top-ranked-v4-flash-stumbles-on-real-agent-tasks-as-its-prices-surge"
  },
  "original_language": "en",
  "account": "DeepSeek's V4 Flash, once hailed as a top-performing model, is struggling with real-world agent tasks, completing only 53.8% of a batch of complex tasks. Multiple agent harnesses, including Claude Code, Codex, and OpenCode, were used to test the model on 30 difficult workflows involving tools like Gmail, GitHub, Slack, and Google Sheets. The model's performance varied greatly depending on the harness, tool configuration, caching behavior, retries, and provider stack it ran on. DeepSeek is hiking the prices for V4 Flash and Pro, two models that have gained popularity among developers building coding assistants and agents. The new pricing structure increases costs significantly, ranging from 57% to 371% for Flash and 51% to 355% for Pro. While this may undermine the appeal of the ultra-low-priced models, it moves the discussion beyond the \"cheap Chinese model\" narrative, as enterprises begin to explore the best use cases for different models in their tech stacks.",
  "summary": "DeepSeek's V4 Flash has topped model leaderboards and been hailed by developers as a \"total monster\" since its rollout. But in real-world testing, it completed just 53.8% of a batch of complex agent tasks. Composio ran the model through eight different agent harnesses , including Claude Code, Codex, and OpenCode, on 30 deliberately difficult, multi-step tasks spanning live tools like Gmail,…",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}