Urgent.News

What's breaking now, across thousands of outlets.

Tech

Groq 429s: your retry loop is blind, read the headers instead

Getting 429s from Groq and just sleeping a few seconds before retrying? That works until it doesn't. import time , requests resp = requests . post ( " https://api.groq.com/openai/v1/chat/completions " , headers = { " Authorization " : f " Bearer { KEY } " }, json = payload ) if resp . status_code == 429 : time . sleep ( 5 ) # blind retry. this is the bug. resp = requests . post ( url , headers =…

Experiencing 429 errors from Groq and employing a simple retry loop that waits a few seconds before attempting the request again? This approach may seem effective at first, but it has its flaws. The issue lies in the fact that this method is "blind," ignoring crucial headers present in the response.

According to Groq's documentation, every response includes x-ratelimit-* headers. The retry-after header, which provides the authoritative wait time, appears only on 429 responses. The provided Python code snippet demonstrates this blind retry loop, where the script waits five seconds before reattempting the request.

However, this approach can lead to problems. Groq enforces limits at the organization level, not per user, meaning a single noisy retry loop can impact the entire organization by consuming all available resources. Additionally, accounts with input/output token splits (ITPM/OTPM) can experience unexpected slowdowns even when their total tokens per minute (TPM) appear sufficient.

Another factor to consider is cached tokens. High cache hit rates can reduce the effective pressure on limits, leading to misleading usage numbers. To avoid these issues and ensure proper retries, it is essential to read the retry-after header from the 429 response and back off to the exact window specified.

For further details and insights into the potential pitfalls of blind retries, as well as the organization-level limits and ITPM/OTPM split, refer to the full notes available at https://vectle.com/skills/skl_2mbinMsdstsavd9J2DD2jA.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in Tech

🌿WindQuest-Explore more. Scroll less.

This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass What I Built WildQuest — Explore more. Scroll less.

  • WildQuest is an open-source AI app transforming walks into personalized nature quests.
  • Users customize quests by selecting activity, duration, difficulty, and accessibility.
  • App runs locally with Gemma to ensure privacy and independence from cloud AI APIs.

India is Treating Starlink Like Pakistan

Elon Musk has ‘begged’ Indian billionaire Mukesh Ambani to allow Starlink to launch in India’s internet market. In a post … Read More The post India is Treating Starlink Like Pakistan appeared first…

Say exactly which tasks to run

Two questions decide every run: what does each task need first, and which projects should this run touch? vx answers the first in dependsOn and the second in --filter .

  • Tasks require determination of needs and projects
  • DependsOn field lists prerequisite tasks
  • --filter command specifies projects to include

The top 10 dark patterns of the web

It feels like we’ve been living in a dystopia for the last decade or so, with apps (and the companies behind them) competing to be the shadiest of them all.

  • Roach Motel pattern: Easy sign-up, but hard cancellation
  • Forced Continuity: Free trials auto-convert to paid subscriptions
  • Hidden Costs: Fees concealed until final checkout

More from Friday 9 October →