{
  "id": 12237357,
  "title": "Retrying Failed Jobs in Node.js: Choosing Queues for Idempotency and DLQs",
  "url": "https://urgent.news/2026/10/05/retrying-failed-jobs-in-node-js-choosing-queues-for-idempotency-and",
  "topic": "tech",
  "section": "Tech",
  "published": "2026-10-05T21:58:01.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/horatiofox1281/retrying-failed-jobs-in-nodejs-choosing-queues-for-idempotency-and-dlqs-3l6p"
  },
  "original_language": "en",
  "account": "When a media cleanup job extends beyond the lifespan of the initial HTTP request, operational constraints dictate a different approach. The optimal solution involves utilizing a managed queue for production retries, wherein workers remain operational even after web-process restarts. Idempotency keys should be recorded before acknowledging each message to maintain consistency. BullMQ is an excellent starting point for Node.js teams that already utilize Redis, while SQS proves to be a strong default when AWS is already in place. Small managed queue services are attractive when minimizing infrastructure maintenance is a priority. The key concept revolves around separating the request from the worker, with the queue serving as a buffer between traffic spikes and the cleanup process. When failed jobs are detached from web requests, they typically involve tasks such as removing expired thumbnails and orphaned captions on a scheduled basis. A cron trigger should be employed to publish this work, rather than maintaining a browser-facing request open while extensive libraries are scanned. Workers can then operate independently, effectively offloading retry traffic from the web process. BullMQ provides a quick and native API for Node.js development, leveraging Redis for familiar primitives. However, this comes with operational costs associated with Redis, including capacity, persistence, upgrades, and alerting. For teams with limited experience, SQS offers a more straightforward solution by eliminating Redis management and integrating seamlessly with AWS workers. Visibility timeouts, redrive policies, and IAM policies still need to be configured. Infrai's approach employs a REST API to expose the queue functionality, allowing any language capable of HTTPS communication to publish or consume messages. This approach also consolidates multiple backend capabilities, including scheduling and queues, under a single key and billing structure. To ensure idempotency and handle DLQs effectively, a worker should store a stable operation key in the application database prior to acknowledging the message. If the key already exists with a completed status, the worker should return success without repeating the side effect. The provided code snippet demonstrates a cleanup worker function that processes cleanup jobs, checks for existing entries, performs necessary actions, and marks the cleanup as complete. Retries should incorporate backoff strategies, such as exponential backoff with jitter, in response to 429 responses. If a failure occurs, the message should be routed to a DLQ for further inspection and redrive after addressing the underlying issues. When selecting a queue, factors like latency and cost must be carefully considered, as there is no one-size-fits-all solution. Measuring operator time and required latency is crucial, rather than solely focusing on the per-message pricing.",
  "summary": "When a media cleanup job outlives the HTTP request that started it, the operational constraint changes the answer. Short answer: use a managed queue for production retries when you want a worker that survives web-process restarts, then record an idempotency key before acknowledging each message. BullMQ is a good fast start for a Node.js team that already runs Redis; SQS is a strong default when…",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}