AI-generated vulnerability patches require human review
Researchers have discovered that AI-generated vulnerability patches, created by large language models (LLMs), often contain defects (FLAWED) in 53.9% of cases when dealing with complex patches. The study, led by 1Password's Off-by-1 Labs, aimed to provide better tooling and methodology for vulnerability remediation at scale. The findings, published in a research paper titled "Frontier Models’ Vulnerability Patches are Often F.L.A.W.E.D.," reveal that only 26.0% of generated patches successfully resolved the vulnerability without altering the application's behavior.
In 20.1% of cases, patches fixed the issue but changed the application's behavior, such as reimplementing file-local parsers or modifying "allow list" logic to "deny list" logic. In 53.9% of instances, LLM-generated patches didn't resolve the vulnerability, introduced a new vulnerability, or both. The study targeted six recently disclosed, complex vulnerabilities across open-source software, including Linux privilege escalation, ActiveMQ Remote Code Execution, and Chrome's File System Access API.
Despite the open-source code potentially existing within the models' training datasets, the success rates were significantly lower than expected. The researchers emphasize that while LLMs can generate patches, they still require human review to ensure effectiveness and safety.
Written by urgent.news from Lobsters's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.
This story
This is one outlet's version. Read the fullest account.