Urgent.News

What's breaking now, across thousands of outlets.

AI

AWS Benchmark Aims to Reduce Number of False Positives Found by AI Vulnerability Scanners

Amazon Web Services (AWS) has developed a benchmark that can be used to test whether a model can distinguish real vulnerabilities from code that looks risky but is actually safe. The Deception Benchmark was created following an evaluation of the capabilities of 12 models from five different providers. In all, the benchmark includes 14,822 samples […]

AWS Benchmark Aims to Reduce Number of False Positives Found by AI Vulnerability Scanners

Amazon Web Services (AWS) has unveiled a benchmark to assess AI vulnerability scanners, aiming to minimize false positives. Named the Deception Benchmark, it was crafted from testing 12 models across five providers, encompassing 14,822 code samples in 16 languages and 70 Common Weakness Enumeration (CWE) categories. The benchmark's core purpose is to differentiate between genuine vulnerabilities and code that appears risky but is safe.

As AI models flag up to 95% of real vulnerabilities, they also incorrectly identify 41 to 99% of safe code as potential threats. Adversarial loops are employed to refine the samples, with models facing a 17 to 74 percentage point reduction in false positives using proof-of-exploit prompting, though this approach may overlook 7 to 44% of genuine issues. The environment-gated challenges prove particularly problematic, as models often flag code while ignoring relevant Kubernetes Network Policy settings.

No tested configuration successfully keeps both false positives and false negatives under 10%, according to the benchmark. To address this, the Deception Benchmark employs code-level challenges that present vulnerable and safe variants differing by subtle fixes, and environment-gated challenges that test specific IT environments, such as Kubernetes network policies blocking server-side request forgery (SSRF) paths.

The benchmark is designed to treat labeling as a continuous audit loop, with multiple independent reviewers examining each label and escalating disagreements to direct adjudication. This ensures that unresolved cases are passed on for human review. AWS Director of Applied Science Neha Rungta emphasizes that the ultimate goal is to identify AI models generating the fewest false positives for vulnerability scanning, a task becoming increasingly challenging as AI is integrated into DevSecOps workflows.

The benchmark encourages AI models to consider the broader impact of remediation on the surrounding environment, enabling them to better understand what good looks like when fixing vulnerabilities. Futurum Group's Mitch Ashley stresses that false positives, or verification debt, undermine confidence in AI findings, as AI scanners that lack environmental context cannot distinguish exploitable paths from blocked ones.

Ashley highlights that false positives create verification debt, increasing the workload for DevSecOps teams and potentially diminishing confidence in AI results. While the extent of AI's impact on DevSecOps teams remains unclear, other reports suggest that complex environments hinder AI effectiveness. As AI continues to evolve, AWS hopes that its benchmark will not only help discover vulnerabilities but also automate their resolution, requiring minimal human intervention.

Written by urgent.news from DevOps.com's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at devops.com →

More in AI

More from Wednesday 30 September →