Urgent.News

What's breaking now, across thousands of outlets.

Tech

Dynamic machine ID leases in Elixir

Distributed ID generators often look simple: combine a timestamp, a counter, and a machine ID. NoNoncense uses that idea for fast counter, sortable, and encrypted nonces. It is a wonderfully fast design, right up until two running nodes receive the same machine ID. The hard part is not incrementing a counter. It is deciding who may use each machine ID, especially when pods are replaced, nodes…

Distributed ID generators typically combine a timestamp, a counter, and a machine ID to create unique identifiers. The popular NoNoncense library takes this concept and builds fast, sortable, and encrypted nonces. However, trouble arises when two nodes receive the same machine ID. The challenge lies not in incrementing a counter, but in determining ownership of each machine ID, especially in dynamic environments such as replacing pods, autoscaling nodes, overlapping deployments, and handling network partitions.

Earlier versions of NoNoncense (1.x) left this decision to the startup process, having the application derive an ID, initialize the nonce factory, and ensure topology safety independently. Version 2.0 shifts this decision into the supervision tree, providing a suitable way for different deployments to establish exclusivity.

The exclusivity strategy is not an afterthought but a fundamental aspect of ensuring an ID's uniqueness within a specific deployment. It can be based on SQL leases using PostgreSQL or MySQL for a durable and inspectable source of truth, or on Valkey or Redis leases using per-field hash TTLs for a compact, server-managed lease registry.

Kubernetes StatefulSets can utilize the pod ordinal through an environment variable when this ordinal is globally unique among nonce-generating pods. Alternatively, host identifiers can maintain the fixed-node model for deployments with a known, stable node list.

When dynamic deployments occur, NoNoncense treats a machine ID as a lease, initiating a LeaseManager under supervision. The manager waits for the acquisition of an initial lease before application startup can complete, ensuring that only a proven ID is used. The manager then renews the lease in the background, distinguishing between a confirmed loss and an uncertain result.

Confirmed loss indicates that the coordinator has positively rejected ownership, while uncertainty suggests a timeout or a dropped connection. In such cases, the manager retries uncertainty for a bounded period, but upon confirmed loss or local expiry, it erases the affected factories and calls an optional :on_lease_lost callback.

For dynamic strategies, the manager performs backoff retries and reinitializes the factories only after acquiring a new lease, integrating lease recovery into the component's regular lifecycle rather than a bespoke application failure path. The manager also features an in-memory lease cache, which the manager attempts to renew upon a crash or forced kill instead of needlessly claiming another ID.

However, keep in mind that this cache is not durable, with normal lease expiry serving as the source of truth after a node crash or forced kill. The strategy's design is not confined to the provided backends; applications can bring their own coordinator with unique allocation rules while retaining the same startup, recovery, and safety policy.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in Tech

More from Wednesday 19 August →