Urgent.News

What's breaking now, across thousands of outlets.

AI

How Loopjacking Hijacks Human Approval in AI Workflows

Loopjacking hijacks human approval in AI workflows. Explore documented cases and practical defenses that keep approval bound to the action executed.

How Loopjacking Hijacks Human Approval in AI Workflows

In a scenario that highlights the concept of Loopjacking, an AI agent requests approval for a small payment to a known vendor. After verifying the details, the payment is approved. However, a larger payment is sent to an entirely different destination. When the user returns to check the approval record, they find it still displays the initial request. This situation demonstrates the potential risks of Loopjacking, a failure where human approval for one operation can authorize or release a significantly different outcome.

The experiments conducted in this research involved harmless mock transfers and local recording tools, with no actual money being transferred. The key issue here is that humans can still make the correct decision, but the software fails to preserve what was authorized by that decision. Adithyan Arun Kumar, the researcher behind the study "Loopjacking: Hijacking Human-in-the-Loop Approval," explores this gap in the field.

There are two primary ways Loopjacking can occur. First, the operation can change after approval. Second, the approval view might omit essential details that were present all along. Both scenarios break the expectation that what you approve should be what executes. When you click "Approve," you should be clearly aware of what is being authorized.

When someone approves a payment, they are essentially approving an amount, a recipient, and potentially other details depending on the context. For a payment, this could mean approving a specific amount to a particular vendor. For a file upload, it means approving the destination and the files being uploaded. For a command, it means approving the arguments and the execution context of that command.

The agent workflow must carry this meaning through various components, from proposing an action to preparing the approval view and ultimately performing the operation.

A crucial security requirement in this workflow is that a record indicating "this task was approved" cannot simply establish that every later action within the task was approved. Even if the same tool-call identifier is used, changed arguments can result if the implementation permits such replacements. An attacker could exploit this by influencing the operation while retaining someone else's authorization.

To succeed, the attacker must first establish a genuine approval, introduce a materially different effect outside the scope of that approval, and find a way for the product to consume that approval. Importantly, the attacker should not have a direct route to the affected effect.

The study of Loopjacking reveals two main failure modes: state substitution and representation mismatch. In state substitution, an attacker-influenced workflow state replaces the original operation after approval. This was demonstrated in the Agno AgentOS experiment, where a mock transfer of 20 units was approved but later changed to a transfer of 2,000 units to a different recipient.

Despite the user's inability to approve the different request, the substituted operation was executed successfully. Similarly, in the approval view leaving out consequential details, the approval record might accurately represent the original operation, but the actual executed operation could be materially different, as seen in the OpenClaw issue where additional arguments were absent from the approval representation.

A complete defense against Loopjacking requires maintaining the approval boundary throughout the workflow. For example, in OpenAI Agents SDK versions 0.22.0 and 0.22.2, ordinary per-call approval remains bound to the invocation, even when workflow state is saved and restored. This means that unchanged approved operations continue to execute correctly, while any changes to the invocation arguments, retaining the same call identifier, are rejected.

This implementation successfully preserves the operational authorization and effectively mitigates Loopjacking risks.

Written by urgent.news from HackerNoon's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at hackernoon.com →

More in AI

More from Friday 9 October →