Move Slow LLM Calls Off the Request Path with BullMQ in Node.js
If an LLM call can take 20 seconds or more, keep it out of your HTTP handler. Put the work on a BullMQ queue backed by Redis, return a job ID right away with a 202 Accepted , and let a separate worker make the slow call while the client gets progress updates over Server-Sent Events. The rest of this post is working code for that setup. It covers retries with backoff, concurrency and rate limits,…
Long prompt completions from large language models (LLMs) can take 20 seconds or more to complete. Embedding these lengthy tasks directly into an HTTP handler can cause numerous issues in a real-world application. Timeouts stack up across load balancers, reverse proxies, and clients. Retries may lead to duplicate processing and exceed provider rate limits.
Restarting the API process could drop ongoing jobs. To mitigate these problems, move the slow LLM call to a separate queue with BullMQ, returning an immediate response while a dedicated worker handles the asynchronous processing. This approach allows a better retry policy, concurrency control, and failure handling without impacting the main request path.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.