Urgent.News

One page, thousands of outlets. See who else covered it.

Editions

AI

The Official DeepSWE Number Is Footnoted to DeepSeek Harness

On August 13, DeepSeek shipped four things in one day. V4-Pro left preview. Peak and off-peak API prices start at 16:00 UTC on August 16. An 88-page draft paper went up on spatiotemporal composability. And deepseek-harness dropped as a developer preview, MIT license, one command: npx @deepseek-ai/dsh web . The model headlines ate the English YouTube tab. I stayed on the harness. Official figure…

On August 13, DeepSeek released four significant updates in a single day. First, the V4-Pro model preview was made available. Next, peak and off-peak API pricing began at 16:00 UTC on August 16. A 88-page draft paper on spatiotemporal composability was also published. Finally, deepseek-harness was released as a developer preview under the MIT license, with a single command to run: npx @deepseek-ai/dsh web.

The official DeepSWE number for the DeepSeek-V4-Pro-0813 model was 62.7, compared to 12.8 for the preview version. This figure is footnoted to indicate that the testing of public code-agent sets was conducted within the DeepSeek Harness in minimal mode.

Additionally, Fable 5 maintained its score of 70.0 on the same grid. A loop plugin was introduced, and its implementation was named Cordis. The peak output for V4-Pro was reported at $3.96 per million tokens.

To illustrate these points, a four-minute video was created, covering the footnote details, the plugin topic, and the peak-hour billing information. The original August 13 write-up focused on the plumbing version, which framed DeepSeek's Harness as the price signal. This story presents the scoreboard version, emphasizing the key metrics and updates.

For those already using Claude Code with a DeepSeek key, the article raises the question of whether the plugin topic alters their experience.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

I published our agent-security benchmark, including the attacks we fail to catch

Every company building AI agent security publishes a detection rate. We were doing it. The problem is that none of these numbers are checkable.

  • Published agent-security benchmark with 497 attacks and 1,172 benign samples
  • Detection engine achieves 99.8% detection rate, 0.09% false positive rate
  • Failures include pretext opener attack and Stack Overflow benign sample

More from Sunday 16 August →