Urgent.News

What's breaking now, across thousands of outlets.

AI

Anthropic releases Sonnet 5.5, saying it generates outputs 30%+ faster than Sonnet 5 and costs up to 30% less per task, and plans to release Haiku 5.5 soon (Anthropic)

Sonnet 5.5 is a faster, lower-cost complement to Claude Opus 5.5. Where Opus 5.5 is built for complex work requiring careful judgment …

Anthropic has unveiled Claude Sonnet 5.5, an enhanced iteration of their Claude Sonnet 5 model. This new model boasts a remarkable 30%+ speed increase and offers up to 30% cost savings across most tasks. Positioned as a faster, more affordable counterpart to Claude Opus 5.5, Sonnet 5.5 excels at well-defined daily tasks such as bug fixing and producing professional documents, slides, and spreadsheets.

Additionally, it boasts exceptional design capabilities. Meanwhile, Claude Haiku 5.5, tailored for high-volume and cost-sensitive applications, is set to join the Claude 5.5 lineup soon. In terms of performance, Sonnet 5.5 excels, scoring 70.6% on Terminal-Bench 4.0, a coding evaluation, compared to Sonnet 5's 10.3%. It also trails Opus 5.5 by just two points on GDPval-AA, a test of real-world work across various professions.

Sonnet 5.5 demonstrates strong performance in long-horizon tasks and image understanding, marking it as the first Sonnet model to surpass Pokémon Red solely from screenshots. Collaboratively, Sonnet 5.5 outperforms its predecessor, producing clearer and more collaborative writing. It also generates outputs 30%+ faster than Sonnet 5, making it ideal for rapid iteration on less complex tasks.

Regarding cost, Sonnet 5.5 is priced identically to Sonnet 5, at $2 per million input tokens, $10 per million output tokens, and $0.20 per million tokens for cache reads. However, it typically incurs 30% less cost per task. Speed-wise, Sonnet 5.5 is the fastest Sonnet model thus far. In terms of safety and alignment, Sonnet 5.5 matches or surpasses Sonnet 5 on most alignment measures, thanks to cybersecurity capabilities comparable to Opus 5.

It also shares the same biology safeguards as Sonnet 5. Across various domains, Sonnet 5.5's performance varies—on some evaluations, it performs comparably to Opus 5.5, while in others, Opus 5.5 still leads. However, benchmark scores alone do not encompass a model's full capabilities; real-world testing and external evaluations reveal Opus 5.5's superiority in complex, open-ended tasks demanding sustained judgment.

The charts provided compare each model's score to its cost per task at various effort levels. As effort increases, models typically require longer processing time, resulting in higher cost per task but also higher scores. Sonnet 5.5's performance advantage is most pronounced in coding; at High effort on FrontierCode, it scores 10 points higher than Sonnet 5 at the same setting, at roughly one fifteenth of the cost per task.

On CursorBench, which assesses models on tasks from genuine coding sessions, Sonnet 5.5's best score is within two points of Opus 5.5. Early users appreciated Sonnet 5.5's swift understanding of codebases and its efficiency, particularly in batching tool calls, leading to fewer steps and lower costs. The model also shows improvements in multiple domains of knowledge work.

On GDPval-AA, which evaluates models on real-world tasks across 44 occupations and nine major industries, Sonnet 5.5 scores nearly on par with Opus 5.5 and about 400 points above Sonnet 5. It matches Opus 5.5 in computer use and chart recognition, and clearly outperforms Sonnet 5 and GPT-6 Sol in long-horizon knowledge work. Users praised Sonnet 5.5's natural conversational abilities and design skills, noting its polished output and minimal editing requirements for slide decks.

In an internal test, a public company's quarterly earnings materials and call transcripts, coupled with a slide template, were used to generate a 10-slide operating review. Two experts deemed the first draft ready for immediate use. In terms of cost, Sonnet 5.5 is less expensive to run than Sonnet 5 due to its lower token usage.

Additionally, it generates output 30%+ faster. Claude Code and Claude Platform default to Medium and High effort settings, respectively. Lower settings result in quicker responses and fewer tokens, making it suitable for routine tasks, while higher settings allow Claude to reason more thoroughly and check its work more rigorously.

Sonnet 5.5 does not elevate the frontier of model capabilities, so its alignment assessment concentrated on targeted risks applicable to models of any capability level, such as acting against users' interests, misleading them, and cooperating with high-stakes misuse. On an automated behavioral audit testing Claude across approximately 1,850 scenarios, Sonnet 5.5 matched or exceeded Sonnet 5 on most alignment measures, resistance to misuse, and honesty.

In newer containment evaluations, Sonnet 5.5 closely aligns with Opus 5.5 in escaping its sandbox and is the least likely of all models to probe its container limits. Overall, Opus 5.5 remains slightly superior in the full audit.

Written by urgent.news from Techmeme's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Also reported by 5 other outlets

Read the original at anthropic.com →

More in AI

More from Monday 28 September →