My refund handler checked the ledger before paying. It still paid twice.
A refund handler commits the refund, then loses its reply. Maybe the HTTP response timed out. Maybe the worker died before it acknowledged the queue message. Maybe it was an AI agent's tool call, and the agent saw a timeout and called the tool again with the same arguments. The sender did nothing wrong by retrying. The question is what your handler does with the second delivery. The usual fix is…
Refund handlers must first check the ledger before issuing a refund. However, they may still end up paying twice due to various reasons such as timeout, lost acknowledgement, or repeated tool calls. To test this issue, a pytest plugin was created that delivers the same event multiple ways and counts how many times the effect lands in a shared ledger.
Four different handlers were tested: naive, check_then_act, guarded_raises, and guarded. The naive handler paid twice in every scenario, while the others paid correctly only once out of 200 trials. The webhook flow and an async version of the refund tool also yielded the same verdicts. Although reading the ledger before writing handles most re-deliveries, it fails when overlapping deliveries occur.
The only foolproof solution is to use a unique constraint on the key, as it enforces the desired behavior. However, this alone is not enough, as the handler still needs a pending and a done state to handle crashes between the claim and the API call. The drill doesn't cover real-world scenarios like using a database transaction for the refund call, but it still provides valuable insights into the issue.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.