Urgent.News

What's breaking now, across thousands of outlets.

AI

AI API call hangs in production, then 504: timeouts and budgets

We read 1,662 public posts from builders whose apps broke at or after launch and verified 215 recent cases. AI model and API failures make up 4% of them. Six of the verified cases share one pattern: an outbound or AI call with no timeout, no budget and no fallback. In most of the six, nobody noticed for a while, because the app didn't log model calls or their outcomes. From the builder's side it…

A recent analysis of 1,662 public posts from developers revealed 215 recent cases where AI API calls failed in production environments. Out of these verified cases, 4% were attributed to AI model and API failures. Six of these incidents demonstrated a consistent pattern: the absence of timeout, budget, and fallback mechanisms in outbound AI calls.

In most instances, the issue went unnoticed for some time due to the lack of logging for model calls and their outcomes. When it did surface, users encountered a spinning spinner that could last for minutes, followed by a 504 error from Vercel or a 524 error from Cloudflare. On localhost, the application functioned normally. Users were presented with polite fallback text while the spend dashboard remained unaffected, as the failed calls incurred minimal costs.

Eventually, the provider's spend limit was triggered, leading to a sudden cutoff of service for all users. The root cause of this issue lies in the fact that the API call lacks its own limits, causing other layers to impose restrictions. A simple Next.js App Router route using the OpenAI Node SDK with default settings exemplifies this problem.

The route attempts to complete a chat completion task using the gpt-5-mini model, handling any errors by returning a fallback message. However, when the SDK is instructed to connect to a stub provider that neither replies nor responds, the request hangs indefinitely, eventually being terminated by Vercel after 300 seconds, resulting in a 504 error.

The monitoring systems typically record this as a successful response, failing to capture the underlying issue. The underlying reason for this failure is that each layer between the browser and the AI model has its own timeout mechanism. In this setup, the code's timeout is the longest, making it the first to be reached. Factors such as the Cloudflare proxy's 125-second limit, Vercel's fluid compute function duration, Node's fetch timeout, and the SDK's maximum attempt time all contribute to the eventual failure.

The official documentation for these SDKs outlines the respective default timeout limits and retry mechanisms. The SDK's timeout applies to each attempt, and retries can significantly extend the overall duration. For non-streaming requests, the maximum duration for a single attempt is 10 minutes. The SDK automatically retries connection errors, 408, 409, 429, and 5xx status codes, along with timed-out requests.

This can result in a cumulative timeout of up to 15 minutes for a non-streamed request. To prevent such incidents, developers should implement appropriate timeout, budget, and fallback measures to ensure a smooth user experience and avoid unexpected service outages.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

Silicon Valley 101: Unpacking the Wild Growth of the AI‑Data Industry | 硅谷101:深度解析AI数据行业的野蛮生长

https://www.youtube.com/watch?v=I-rLxiIGf-4 谁在出题、卖题、判卷?深度解析AI数据行业的野蛮生长 说明:百分比按照文档token总长度12876做占比估算,代表内容在全文的位置占比。 第(一)部分 节目开篇预告:线下AI活动介绍 (0%‑7%) 1 主持人开场首先宣传《硅谷101》即将在湾区举办的黑客松活动,活动归属Alignment…

More from Monday 28 September →