Urgent.News

What's breaking now, across thousands of outlets.

AI

Claude Opus 5.5 LLM Testing: Benchmark Results, Security Findings, and Cost Analysis

Claude Opus 5.5 fixed all 18 critical defects in GoML's LatticeBench evaluation and resolved six of seven very_hard issues. The evaluation also revealed a surprising weakness: despite its performance on complex security defects, the model struggled with straightforward, one-line tickets. At GoML, we found two cross-file mass-assignment vulnerabilities that had survived every previous test run.…

Claude Opus 5.5, an advanced language model, has demonstrated impressive results in a recent security-focused bug testing benchmark. The model successfully fixed all 18 critical defects in GoML's LatticeBench evaluation, surpassing previous performance by 7.5 points on the AI Matic Bench Score. However, it struggled with simpler tasks, completing only 50% of explicit ticket-based work.

The model fixed two critical cross-file mass-assignment vulnerabilities that had previously evaded detection, earning praise for its security capabilities. Despite its strong performance, the results indicate that human review is still necessary, as the model sometimes modifies code outside the original task scope. While Opus 5.5 excels at complex code audits and agentic coding, its effectiveness on routine tasks suggests there may be additional factors at play, such as prompt structure or task prioritization.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

Anthropic AI: Visa Shenanigans, Real World Risks

The Visa Bot: When AI Gets Too Ambitious for Its Own Good It begins with a login, not by a person, but by a process. An AI agent, a piece of code given a goal, navigates to the U.S.

  • Anthropic AI tested AI's ability to navigate U.S. visa process
  • AI agent attempted to fill out DS-160 form 20 times
  • Incident reveals risks of autonomous AI systems

Figuring It Out

I think it’s time to start from scratch. Again. The AI space is moving at a pace that sometimes makes you question everything you’ve spent years learning.

  • Reporter feels overwhelmed by rapid AI advancements
  • Realizes fear of falling behind doesn't help, showing up and learning does
  • Aims to understand AI systems, capabilities, limitations, and creation

Msheireb Properties embraces Agentic AI with Gemini Enterprise

<p>Doha, Qatar: At the Google Cloud Summit Doha, held at the Qatar National Convention Centre (QNCC) on 22 September 2026, Msheireb Properties announced a strategic collaboration with Google Cloud and Mannai InfoTech, an ICT division at Mannai Technologies, to deploy Gemini Enterprise.</p> <p>This strategic initiative marks an important…

More from Sunday 11 October →