Shock horror — AI-generated security patches fall short of actually solving all the problems they were meant to fix
AI without oversight creates patches that rarely fix the issue entirely and sometimes just create new problems.
Security experts have cautioned that AI-generated patches for vulnerabilities are often merely a trade-off, swapping one set of problems for another. Researchers from 1Password's Off-by-1 Labs conducted an experiment to test AI-generated patches on six recently disclosed CVEs using two advanced models. Out of 6,080 patches generated, only 26% fully resolved the issue, while 49.3% failed to fix at least one existing exploit path.
A fifth (20.1%) fixed the original issue but altered application behavior, and 2.3% introduced new security issues. Even among the more promising patches, more than a third were found to be fragile and only partially addressing the underlying problem.
The researchers coined the acronym FLAWED (Fix-Like Artifacts With Embedded Defects) to describe these automated LLM patches and urged against relying on AI without human oversight. They explained that when given proper guidance, AI's success rate improves to 65%, but without guidance, it drops to a dismal 15.2%. This suggests that while AI can be a useful tool, it should not replace human developers entirely.
The researchers addressed this concern by releasing a patch evaluation harness called FLAWED, allowing organizations to assess the effectiveness of their AI-generated fixes.
Written by urgent.news from TechRadar's reporting — not their text. Machine-written; read the original for the full account.
This story
This is one outlet's version. Read the fullest account.


