Preventing Duplicate Password-Reset Notifications (Under SMS Timeout and Retry Pressure)
Treat an SMS timeout as an unknown outcome, not a failed send: accept each password-reset event once, persist its expiry and idempotency key before dispatch, and retry only through a worker that can reconcile the original attempt. For a short-lived e-commerce reset token, compliance evidence is the deciding constraint. The system must be able to show what it accepted, what it attempted, when it…
When managing password-reset notifications in an app, it's crucial to address the challenge of duplicate messages caused by SMS timeouts and retry pressure. To prevent this issue, implement a system that treats an SMS timeout as an unknown outcome rather than a failed send. Accept each password-reset event once, and persist its expiry and idempotency key before dispatching the notification. Use a worker to handle retries in a controlled manner, ensuring only one logical notification is sent per event.
To distinguish between business actions and logical notifications, utilize two identifiers: event_id for the business action (such as a password-reset request) and idempotency_key for the logical notification command. By enforcing a unique constraint on the idempotency_key, two concurrent HTTP requests will converge on a single stored record, eliminating duplicates.
In case of a timeout, the system should enter a dispatch_unknown state, keep the provider's attempt identifier if available, and move through reconciliation before allowing another send.
Status polling serves a distinct purpose from retry. Instead of creating a second message, polling should read the provider's view and update the local record. This separation ensures that no duplicate notifications are sent. Additionally, the expiry of the reset token should act as a dispatch boundary. Before each attempt, compare the current time with the expires_at timestamp. If the reset token is close to expiring, mark the notification as expired and stop further actions.
The delivery contract and failure states should be narrowly defined. Admit a password-reset notification exactly once for a stable event identifier, never dispatch it after expiry, retain evidence of each state transition, and allow callers to inspect the logical result without triggering additional work. The contract focuses on the admission of the logical command, not the external SMS network's participation in the database transaction.
Local state should be managed with durably stored evidence of each transition, including timestamps, old and new states, event ID, attempt number, and reason codes. Avoid storing credentials, reset URL, token, or full phone number in these records. Instead, use a redacted destination fingerprint for correlation purposes, subject to the same retention and authorization controls as the rest of the evidence.
When deciding between managed notification services and direct provider integration, consider the trade-offs. A managed service can handle provider reconciliation and channel routing, reducing on-call load but may not align with auditor expectations. Direct provider integration offers more detail and fewer translation layers but requires the team to own leases, retry classification, retention, and delivery integration.
A self-hosted dispatcher provides the strongest control over data placement and change timing but may not be feasible if the team cannot staff the queue, database, and delivery integration as an on-call product.
Compliance evidence should be a key consideration. On-call load, lock-in boundaries, and the need for managed notification layers must be verified for export granularity and retention. While managed layers may reduce application burden, direct integration provides full control over evidence design, ensuring the controls meet auditor requirements. Self-hosted dispatchers offer the highest operational ownership but demand significant internal schema and infrastructure investment.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.