Urgent.News

What's breaking now, across thousands of outlets.

AI

Metrics Failure in LLM-Based Code Vulnerability Repair: An Empirical Study and a Change-Aware Screen

Large language models (LLMs) are increasingly applied to the automated repair of C/C++ security vulnerabilities, and compile rate is a commonly reported proxy for progress: whether the generated patch compiles. We argue that compile rate is a scientifically unreliable metric for single-function vulnerability repair, and we support this with five controlled experiments over 203 vulnerable…

We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.

Read the original at arxiv.org →

More in AI

Claude Opus 5.5 is now available on AWS

Claude Opus 5.5, Anthropic's most capable Opus model for agentic coding, knowledge work, and long-running tasks, is now available on Amazon Bedrock and Claude Platform on AWS.

  • Claude Opus 5.5 launched on AWS Bedrock and Platform
  • Enhanced efficiency, lower cost, improved communication
  • Safety classifiers for biology, cybersecurity, AI development

More from Tuesday 22 September →