Urgent.News

What's breaking now, across thousands of outlets.

AI

AI failed to properly patch software flaws 74% of the time, 1Password's study warns

We might be eager to thread AI into every area of our cybersecurity defenses, but new research reveals why we should pull back.

A new study has found that artificial intelligence failed to properly patch software flaws 74% of the time, according to research by 1Password's security team, Off-By-1 Labs. The FLAWED study explored what happens when large language models (LLMs) are given free rein to generate fixes for complex vulnerabilities in open source software.

The researchers selected six recently disclosed vulnerabilities in open source software and tested the capabilities of AI models to produce patches. They generated 6,080 patch attempts, with LLMs successfully creating suitable patches only 26% of the time. In 21% of results, the patches fixed the bug but also altered the application's behavior. The AI models failed 53.9% of the time, either by adding new bugs or both failing to create a patch and introducing new vulnerabilities.

The main issue appears to be that regardless of environmental conditions or prompts, LLMs generated "Fix-Like Artifacts with Embedded Defects," which superficially appear to fix the vulnerability but lack robust security mechanisms and may introduce additional bugs. These "FLAWED" patches can change an application's typical behavior.

Off-By-1 Labs released its tooling, FLAWED, on GitHub for researchers to conduct their own studies. The researchers suggest that while LLMs excel at discovering vulnerabilities, they are currently only effective at patching a narrow subset of them. They emphasize the need for human oversight in the patch process, as having an informed understanding of which bugs are most impactful in a codebase is crucial for human defenders and AI tooling.

Written by urgent.news from ZDNet's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at zdnet.com →

More in AI

Google Shakes Up AI Leadership, Demis Hassabis Steps Down As DeepMind CEO

Google is making its biggest AI leadership change in years. The company shakes up AI leadership and still maintains upbeat memos.

  • Demis Hassabis steps down as DeepMind CEO, becomes chairman and chief scientist
  • Jeff Dean leaves Google to start new venture with senior AI colleagues
  • Koray Kavukcuoglu becomes DeepMind CTO, reports directly to Google CEO

More from Thursday 6 August →