Urgent.News

What's breaking now, across thousands of outlets.

AI

Artificial Analysis' Intelligence Index: Sonnet 5.5 (max) ranks above GPT-6 Astra (max) and behind only Opus 5.5 (max) but has the highest token use of them all (Artificial Analysis)

With max effort, Sonnet 5.5 gains 18 points over Sonnet 5 and moves to #2 on the Intelligence Index, behind only Opus 5.5 (max).

Sonnet 5.5 has emerged as a top performer on the Artificial Analysis Intelligence Index, surpassing GPT-6 Astra and ranking just behind Opus 5.5. Developed by Anthropic, this artificial intelligence model has shown significant improvements in performance, with a notable increase in its output tokens per task compared to its predecessor, Sonnet 5.

Though its cost per task has risen by roughly 50% over Sonnet 5, Sonnet 5.5 (max) maintains a competitive edge by delivering a higher number of output tokens. In fact, it demonstrates the highest token usage among the models examined, clocking in at approximately 193,000 output tokens per task – a rate that is around 60% higher than Opus 5.5 and significantly more than GPT-6 Astra.

The pricing for Sonnet 5.5 remains consistent with the previous iteration, at $2/$10 per million tokens of input/output, setting it apart from GPT-6 Sol, which operates at a lower cost. Despite this, Sonnet 5.5 (max) holds a respectable position on the Intelligence vs. Cost per Task Pareto Frontier, particularly at higher effort levels, where it remains competitive with GPT-6 Sol.

However, it's worth noting that Sonnet 5.5 (max) lags behind Opus 5.5 when it comes to factual knowledge and scientific reasoning. It scores 54% in factual accuracy, compared to Opus 5.5's 66%, though it does exhibit a lower hallucination rate. Additionally, Sonnet 5.5 (max) underperforms in the Humanity's Last Exam and SciCode benchmarks when compared to Opus.

The context window for Sonnet 5.5 (max) remains unchanged at 1 million tokens, allowing for image and text input. The model comes with five effort settings, ranging from low to max. While the highest effort setting (max) yields the best performance, it also carries a higher token usage cost.

Sonnet 5.5 (max) has shown substantial improvements in terminal use and knowledge work benchmarks, scoring 64% in Terminal-Bench 4.0, a 50-point increase over Sonnet 5 (max), and slightly above the performances of Opus 5.5 and GPT-6 Astra. It also reached 53% in the Terminal-Bench-Science benchmark, placing it behind Opus 5.5 and GPT-6 Astra in the Artificial Analysis Intelligence Index.

To achieve its impressive performance, Sonnet 5.5 (max) leverages the highest token usage among the models examined, an aspect that underscores its strengths and areas for potential optimization.

Written by urgent.news from Techmeme's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at artificialanalysis.ai →

More in AI

Silicon Valley 101: Unpacking the Wild Growth of the AI‑Data Industry | 硅谷101:深度解析AI数据行业的野蛮生长

https://www.youtube.com/watch?v=I-rLxiIGf-4 谁在出题、卖题、判卷?深度解析AI数据行业的野蛮生长 说明:百分比按照文档token总长度12876做占比估算,代表内容在全文的位置占比。 第(一)部分 节目开篇预告:线下AI活动介绍 (0%‑7%) 1 主持人开场首先宣传《硅谷101》即将在湾区举办的黑客松活动,活动归属Alignment…

More from Monday 28 September →