What To Do When a Free LLM API Says 429
Every free LLM API says yes right up until the moment it says no. The no arrives as an HTTP 429, usually forty requests into a demo or halfway through a batch job, and most tutorials end the story at "add a retry loop." Here is what actually works when you are building on free tiers, provider by provider and layer by layer. First, find out which limit you hit A 429 is not one problem, it is four.…
When a free LLM API returns an HTTP 429 status code, it means the service is denying requests due to rate limits. This is common in free tiers, where providers meter requests on multiple axes such as requests per minute (RPM), tokens per minute (TPM), requests per day (RPD), and concurrency limits. The 429 response usually arrives after a few attempts and most tutorials only suggest adding a retry loop.
To handle 429 errors effectively, first determine which limit was reached. A 429 is not a single issue, but four. For RPM, smooth out the request rate as the provider cannot increase capacity. For TPM, shorten the prompt context or cap the max_tokens. For RPD, spread the work across days or switch providers. For concurrency limits, queue requests instead of parallelizing them.
Always read the response body and headers before modifying your code. Groq provides x-ratelimit-remaining-requests and x-ratelimit-remaining-tokens headers, along with retry-after. Gemini API returns 429 RESOURCE_EXHAUSTED and names the violated window in details[].quotaId, with retryDelay. OpenRouter free-model caps increase once purchased credits are added, with error text and x-ratelimit-reset-requests header indicating when the window resets.
If a provider sends no information with the 429, assume all windows are closed and back off to the hour. Retry using three rules: honor the retry-after header when present, and use exponential backoff with jitter otherwise. Retry a daily cap for only 3-5 attempts and then stop, as retrying will not increase the daily quota.
Remember, tokens matter more than requests. A retry loop optimized for RPM may still fail if TPM is the limiting factor, as retying consumes more tokens per minute. Treat free tier limitations as separate pools and switch to another provider when one quota is exhausted. Build a provider abstraction layer with a per-provider window tracker, queue, and health checks, or use a router like FreeLLMAPI which handles 429s, retries, backoff, fallbacks to other providers, and exposes an OpenAI-compatible endpoint.
Logging per-request details such as provider, model, latency, status, and rate-limit headers helps identify which window is being exhausted and guide optimization efforts. Most 429 issues stem from long prompts in loops, so read the 429 response carefully before rerunning requests.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.