You picked Claude Sonnet 5.5 — but Anthropic may send your request to Sonnet 5
Anthropic’s Claude Sonnet 5.5, released Monday, is the first Sonnet model to launch with cyber safeguards and model fallbacks like The post You picked Claude Sonnet 5.5 — but Anthropic may send your request to Sonnet 5 appeared first on The New Stack .
Anthropic has launched Claude Sonnet 5.5, the first Sonnet model to incorporate cybersecurity safeguards and fallback mechanisms similar to those used for their more advanced models. Although not a significant upgrade in terms of overall capabilities, Sonnet 5.5 matches Opus 5's cybersecurity skills and scores well on coding benchmarks. However, Anthropic emphasizes that Sonnet 5.5 remains less capable in cybersecurity compared to Opus 5.5 and Mythos 5.1.
The release demonstrates how a model can still fall short of Anthropic's most powerful models while excelling in one area, necessitating comparable safeguards. With cyber safeguards off, Sonnet 5.5 could achieve arbitrary code execution in 178 out of 410 ExploitBench runs and completed 46.1% of Irregular's CyScenarioBench challenges, surpassing its predecessor, Sonnet 5.
Anthropic considers Sonnet 5.5 less capable in cybersecurity than Opus 5.5 and Mythos 5.1, but the improvement was substantial enough to apply the same cyber policy used with Opus 5 and Opus 5.5.
Three-stage cyber enforcement includes a probe that examines internal activations, a lightweight classifier running on Sonnet 5.5 itself, and a separate trained LLM classifier. The classifiers catch harmful cyber requests as effectively as those on Opus 5, though less aggressive jailbreak protections are used due to Sonnet 5.5's lower cybersecurity capability. Users should anticipate more refusals than with Sonnet 5, including legitimate cybersecurity work.
Sonnet 5.5 employs routing to direct requests to appropriate models based on the task type. Requests related to cybersecurity or frontier LLM development will be sent to Opus 4.8 and Opus 5, respectively, while biology or conventional weapon-related requests will be directed to Sonnet 5. Blocks for specific categories, like biology, conventional weapons, and anti-distillation, bypass fallback models and are transparent, not altering the model's responses.
Written by urgent.news from The New Stack's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.
Also reported by 6 other outlets
- Anthropic releases Sonnet 5.5, saying it generates outputs 30%+ faster than Sonnet 5 and costs up to 30% less per task, and plans to release Haiku 5.5 soon (Anthropic) anthropic.com
- Anthropic launches Claude Sonnet 5.5 with near-Opus performance at half the price thenewstack.io
- Anthropic debuts Claude Sonnet 5.5 running 30% faster than the previous-generation AI model siliconangle.com
- Anthropic upgrades Claude with new Sonnet 5.5 model, details here 9to5mac.com
- Anthropic releases Sonnet 5.5, which it calls a significantly cheaper, faster work partner techcrunch.com
- Anthropic unveils new, low-cost Claude Sonnet 5.5 model seekingalpha.com