Humans in the loop miss a third of dangerous AI coding agent requests
You wouldn't let Claude Code cat your AWS credentials or Kubernetes config on request, would you?
A web-based game designed to assess humans' ability to safely approve AI coding agent requests has revealed that humans are struggling to spot dangerous commands, approving about one-third of malicious requests on average. The study, which analyzed over 40,000 runs and 409,000 approved and denied commands, found that 35% of scope violations, such as requests to access sensitive data, were missed.
The most commonly caught dangerous commands were destructive ones, like deleting files or granting full permissions to locations. The most frequently missed potentially malicious command, "npm run analyze," was approved 65% of the time despite its potential to run arbitrary code defined in a project's package.json file. Wauters, the game's creator, noted that manually approving all agent actions is a draining activity that can lead to mistakes due to fatigue and lack of context.
He emphasized the need for better permission models, tooling, and AI-assisted surveillance to address this issue.
Written by urgent.news from The Register Science's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.
This story
This is one outlet's version. Read the fullest account.