How I Tracked Down a Race Condition in Async Python With Claude Code
TL;DR A background worker in one of my Python services was sending duplicate webhooks about once every 3,000 jobs. It never happened locally, never in tests, and only sometimes in production. I used Claude Code to turn a "can't reproduce" ticket into a deterministic, 100%-failing test in one afternoon, and the actual fix ended up being 11 lines. Here's the exact process, plus the lessons I'd…
A background worker in a Python service was sending duplicate webhooks about once every 3,000 jobs. The issue only appeared in production and not locally or during tests. The worker pulled send webhook jobs off a queue with a concurrency limit of 20. Each job checked if the webhook was already delivered and if not, sent it and recorded the delivery.
Customers reported receiving the same order.paid event twice, but only rarely. The bug was hard to find because it never reproduced locally or in tests, and adding debug logging made it less frequent.
Claude Code was used to turn a "can't reproduce" ticket into a deterministic, 100%-failing test in one afternoon. The process involved giving the agent raw evidence of the symptom, listing every code path where the same delivery could be sent twice, and ranking the paths by likelihood. The agent found four candidate paths, and the most likely one involved two different jobs for the same delivery running concurrently in the same worker process.
The next step was to find the check-then-act gap in the handler code. Every await in the handler could potentially switch the event loop to another task, creating a race condition. The agent suggested using asyncio.Event to control the interleaving of tasks and create a red test that would fail before the fix and pass after the fix.
The final fix involved two layers. The first layer was to claim the delivery atomically in the database using a unique identifier for each worker. If the delivery was already claimed, the handler would return without sending the webhook. The second layer was to add a lock per delivery ID, but it was decided that this was unnecessary since the atomic claim in the database already protected against race conditions across multiple worker processes.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.