Idempotency Before Retries: Safer Tool Execution for AI Agents
A tool call can succeed even when an AI agent never receives the response. Imagine an agent that submits a refund request to a payment service. The service processes the refund, but the network connection times out before the agent receives confirmation. The agent sees an error and retries. If the second request creates another refund, a transient communication failure has become a duplicate…
Executing tool calls can yield results even without the receiving AI agent ever getting the confirmation. Picture an agent that submits a refund request to a payment service. The service processes the refund, but the network connection drops before the agent gets the confirmation. The agent sees an error and tries again. If the second attempt creates another refund, a brief communication issue has turned into a duplicated business operation.
The issue isn't necessarily with the model's reasoning. It's about the execution boundary between the agent and the systems it can modify. Retries help, but they don't substitute for idempotency. Before allowing an agent to retry a state-changing tool call, the application needs a dependable method to know if the requested operation has already been accepted or finished.
The difference between a failed request and a failed operation is tricky in distributed systems. They make it hard to tell whether an operation failed or succeeded without delivering its result. Suppose the sequence goes like this: The agent asks for an action through a tool gateway. The gateway forwards the request to a downstream service.
The service records the change. The response disappears or turns up after the client's deadline. The agent gets a timeout and thinks about retrying. At step four, the caller doesn't know whether the side effect occurred. Treating every timeout as a sign of failure is risky. The same confusion also appears when a worker crashes after committing a database transaction but before responding to a queue message.
A message might be sent again even though its original processing succeeded. That's why retry plans and operation rules must be planned together. AWS's advice to make changing operations idempotent so repeated requests have the same effect as one request is a good rule. The key point is that idempotency safeguards the effect of repeating, not that it guarantees a network request will only happen once.
Give each logical action a stable operation key. An agent might call the same tool many times for different reasons. The application should differentiate a retry of one logical action from a genuinely new action. For instance, consider a workflow that makes a support ticket. A good operation key might come from a permanent workflow-run identifier and a stable step identifier: workflow-8472:create-ticket.
This is just an example, not a rule. Real systems should use identifiers that are unique, fit the right scope, and keep tenant isolation. The key must stay the same when the same logical operation is tried again. Generating a new random key for every attempt defeats the point: the downstream service sees each attempt as a different request.
On the other hand, two intentional ticket creations must not share a key just because their inputs look similar. A solid operation record should link the key to: The tenant or principal that's authenticated. The tool and operation type. A hashed version of the standardized request parameters. The current execution state. The downstream operation or resource identifier, if available.
The final result or enough information to find it. The hashed parameters are important because an operation key should not automatically approve a different action. If the same key shows up with different parameters, reject the conflict instead of using the old result. Stripe explains a similar method for idempotent requests: the same key shows retries, while different parameters are turned away.
Its exact way of keeping records and reusing responses is Stripe-specific and shouldn't be taken as the rule for every tool gateway. Think of running an AI agent's actions like a state machine. A simple "completed" yes or no isn't enough to handle a tool operation safely. The system must tell apart actions that haven't started, are in progress, are finished, and have uncertain outcomes.
One way to do this is: PENDING: the operation has been noted but execution hasn't started. RUNNING: a worker is handling an execution attempt. SUCCEEDED: the side effect is confirmed and its result is saved. FAILED_FINAL: the operation failed in a way that shouldn't be tried again automatically. UNKNOWN: the result can't be figured out yet.
What the system needs is states that fit the downstream system. "UNKNOWN" is key when an external service might have done a task but the application can't confirm it. A simple flow might look like this: Receive tool request | Validate identity, rules, and parameters | Look up operation key | Complete the action | Recheck the result | Unknown The steps and checking the result must be safe from multiple things happening at once.
Two workers shouldn't both see a missing record and start the same action. A database rule that makes each operation key unique, a conditional write, or a similar exact time coordination method can help decide which worker gets to do a job. But keeping a key in a database doesn't mean an external action will stay true to that reservation.
If the worker crashes after the external action works but before the local record is updated, the application might end up with a unclear result. That gap needs a clear way to fix things. Don't keep retrying without thinking about the error and the side effect. A connection problem before the request is sent might be okay to try again.
A mistake in the request usually needs a corrected one. A "too many requests" reply might be okay to try again after a little wait. A timeout after a possible successful write is different: the application might need to ask the downstream system or sort out the operation before trying again. For a tool that changes the state, a good policy is: Keep the same operation key for the same logical action.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.