Urgent.News

What's breaking now, across thousands of outlets.

Tech

Retries don't fix eventual consistency

Retries often appear to be a logical solution, yet questioning their efficacy proves enlightening. When something fails, the instinctive reaction is to attempt again. If it fails again, a bit more time is granted before another attempt. If all efforts prove futile... the cycle continues. While this approach can prove effective at times, there are instances where retries are unnecessary.

Recently, I engaged in a discussion about processing events in a distributed system. Picture a scenario where one service broadcasts a 'user created' event, another sends a 'subscription created' event, and a third service requires both before proceeding to process the subscription payment. The proposed resolution was simple. If the subscription event arrived prior to the user's record, classify it as an error.

Transfer the message to a dead-letter queue, conduct a manual review if needed, and retry it later. Initially, this proposal seemed logical. After all, the necessary data was absent, but the underlying assumption glosses over a crucial point: the absence of an error. Distributed systems do not guarantee simultaneous or ordered information delivery.

They only ensure eventual consistency. This principle underscores that information will reach all points eventually. Understanding eventual consistency as a characteristic, rather than an error, alters the system's architecture. Instead of perceiving missing data as a failure necessitating human intervention, it becomes a mere possible state of the system.

Each incoming data piece should be recorded. Upon the arrival of another, verify if all required components are now present. If so, execute the task. If not, remain idle. No retries, no dead-letter queues, no manual replay of messages. As information becomes available, the system's internal state evolves naturally. One of the most gratifying aspects of resolving the correct issue is observing ancillary problems vanishing along with it.

Consider a service's temporary unavailability for a minute. Messages accumulate, with some entering a dead-letter queue while newer messages continue flowing through the system once restored. This scenario introduces another predicament: replaying the older messages in the correct sequence. Should they leap ahead of newer events?

Should processing halt until they are replayed? What if the correct order is essential? These problems weren't unavoidable; they were self-inflicted. If every message is merely stored until all prerequisites are satisfied, such orchestration vanishes. Events can become eligible for processing as the missing information arrives. A transient network timeout?

Retry it. A packet dropped? Retry it. Distributed systems are rife with short-lived failures, and retrying once often suffices to mitigate the inherent glitches of real infrastructure. However, it's worth asking a simple question as retries proliferate. If one retry doesn't resolve the issue, why would six? Is there evidence that the dependency will recover, or are we merely hoping it will?

The broader lesson extends beyond retries. It emphasizes the importance of accurately identifying the problem before deploying a solution. Availability issues warrant one set of tools. Eventual consistency requires another. Misidentifying one for the other typically results in additional infrastructure, increased operational burdens, and systems that are more challenging to comprehend.

At times, the most straightforward resolution isn't refining retry strategies. It entails recognizing that there was never anything to retry in the first place. I have produced a video version of this discussion for those who prefer visual learning: Retries and mental models.

Written by urgent.news from Lobsters's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at var0.xyz →

More in Tech

More from Monday 3 August →