Urgent.News

What's breaking now, across thousands of outlets.

Tech

Node.js Feature Flag Safety: Kill Switch Control for Marketplace Incidents

Use a server-side flag to stop new notification delivery attempts, but do not call the incident contained until the queue has stopped producing failures. Short answer: a useful kill switch changes one narrow runtime decision, emits an audit event, and leaves failed jobs available for deliberate replay. For a marketplace notification service, that means disabling delivery while preserving the…

To prevent incident-related disruptions in a Node.js marketplace notification service, implement a server-side feature flag kill switch. This toggle controls the delivery of notifications without affecting the entire application. When activated, it stops new notification attempts, but the system continues processing existing jobs, preserving the order, seller, channel, and attempt context for later recovery.

The flag should be checked at the last safe point before the outbound call to avoid discarding important evidence needed for replay.

A decision rule based on a failure ratio, minimum sample count, and five-minute window provides a more reliable trigger than a single error counter. This approach ensures that the kill switch is activated only when the failure rate is significant enough to impact real users.

In production, the kill switch should gracefully stop the side effect at the last safe point before the outbound call, while still accepting valid marketplace events and durable jobs. Metrics and structured logs should reflect the deferred outcomes or the attempted deliveries, depending on the flag's state.

When operators change the flag through an authenticated control path, workers observe the change without requiring a new deployment. Rollback in this context means containing future attempts while maintaining a complete record of what has already occurred, rather than discarding any attempted deliveries.

Consider a scenario where a notification attempt fails at 14:02 due to a timeout, while a retry is valid as there's no positive acknowledgment. Even if several more jobs return the same retryable result, a single-error alarm would be premature without a sustained failure ratio in the rolling window. By 14:07, with enough attempts to show a sustained failure ratio, the operator can safely disable the revision.

Some workers finish requests that were already in flight, while others start recording deferred outcomes. The queue depth rises, proving the service is preserving work. The incident is contained only after metrics show no new outbound attempts, every live worker reports the new revision, and the number of deferred records agrees with newly accepted jobs.

To implement this in Node.js, create separate interfaces for the flag store, delivery gateway, and job ledger. Use an in-memory flag store for testing and demonstration purposes, but replace it with a durable, shared store in a production implementation. The key behavior is fail-closed for outbound delivery when the flag cannot be read, ensuring recoverable temporary losses and preventing accidental duplicate deliveries.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in Tech

More from Friday 2 October →