Urgent.News

What's breaking now, across thousands of outlets.

AI

What Independent Benchmarks Say About Opus 5.5

Two independent benchmark results for Opus 5.5 came out this week, and they don't really agree. One has it in first place. The other has it fast and cheap, but third on secure code once you take out the answers it memorized. A few days ago I read through Anthropic's Opus 5.5 docs to see what they say got better. That's Anthropic talking about their own model, though. I wanted to see what people…

Two independent benchmark results for Opus 5.5 were released this week, but the findings are inconsistent. One analysis places Opus in first place, while another considers it third in secure code performance. These results come from two separate evaluations: Artificial Analysis and Endor Labs.

Artificial Analysis' Intelligence Index ranks Opus 5.5 first with 58 points, followed by OpenAI's GPT-6 Astra and Anthropic's Fable 5.1, both at 53 points. SciCode (scientific programming) saw Opus 5.5 leading with an 11-point margin over Astra and 4 points ahead of Fable. However, the indexing was conducted at the highest settings, which may not reflect typical usage.

When considering cost, Opus 5.5 is cheaper than GPT-6 Astra in terms of token usage, but Astra completes tasks with fewer tokens, making it more cost-effective per task. However, Fable 5.1 incurs higher costs per token and per task. Average costs per task are $5.98 for Opus 5.5, $3.26 for Astra, and $7.63 for Fable 5.1.

Endor Labs focused on secure code, testing a coding agent within open-source projects to ensure it adheres to security best practices. Opus 5.5 finished third with 68.7% FuncPass and 33.5% SecPass, compared to Fable 5.1's 87.2% FuncPass and 37.4% SecPass. The memorization of known fixes was a significant factor, with Opus 5.5 scoring 51 out of 51 confirmed cases, versus 17 for Opus 5 and 38 for Claude Code.

In summary, while Opus 5.5 excels in speed and cost, its performance in secure code remains less impressive compared to Fable 5.1. Analysts suggest testing models on personal code to gauge their suitability better.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

What Anthropic Says Has Improved in Opus 5.5

I went back to Anthropic's docs to see what 5.5 changes about the verbosity, repeated checks, and unnecessary delegation I wrote about with Opus 5.

  • Opus 5.5 improves verbosity, repeated checks, and unnecessary delegation.
  • Model verifies work without prompting, may cause excessive verification.
  • Faster response generation, fewer tokens per task, and matches or exceeds Opus 5 performance.

I Linted 14 Public AI SDK Repos. 12 Ship a Call With No Token Ceiling.

I argued last week that an AI SDK call with no bounds is three CWEs in one missing config object . Fair question back: does anyone actually ship that? So I linted for it.

  • 12 AI SDK repositories ship calls with no token ceiling
  • 116 files contain text generation or prompt completion calls
  • 5 SDK vendor repositories include unbounded calls

More from Wednesday 7 October →