Route Once, Fail Over Among Equals: When NOT to Retry an LLM Call
Most LLM APIs I review have one IChatClient and a prayer. Hello goes to the same model as "why does this async code deadlock under load." When that provider returns 503 for twenty minutes, so does the product. Routing and failover both matter. Mixing them up is how you get a green uptime chart and worse answers. The full sample (Microsoft.Extensions.AI routing types, circuit breaker, tests,…
When designing an LLM routing system, it is crucial to follow a few hard rules. First and foremost, only fail over among models that are considered equals. This means that if one model fails, you should attempt to use another model of the same quality or reliability. Mixing up routing and failover can lead to a green uptime chart, but it may result in worse answers.
A recommended approach is to use the SemanticRoutingChatClient, which embeds the prompt and scores it against example utterances per route. The failover then walks this ordered model list. Inverting the layers can lead to outages on your reasoning model silently dumping hard debugging questions onto a small-talk model, which would negatively impact quality without raising any red flags.
Some situations should not trigger a failover. If the provider returns a 400, it indicates a malformed request or schema mistake, and retrying the same model will yield the same result. In such cases, you should not pay twice for the same bug. Rejecting content or policy filters should also be treated as a final answer for that request rather than an outage.
Similarly, if the API key is bad (401), failing over to another model while using the same broken credential will only waste time. Fixing the secret is the correct action. Additionally, if the connection is reset (429, 5xx), these are considered transient cases, and you should fail over to another model.
It's essential to commit only once when streaming the first token. Once the first update reaches the caller, you cannot switch providers, as failing over would mean sending half an answer. After commitment, a failure is final for that stream, and the user experience should be designed accordingly, either with a clear failed path or non-streaming for critical paths.
In cases where the same model times out multiple times consecutively, it is not recommended to keep trying it for every request. Instead, open a circuit for a break duration, skip that model, and probe later. This prevents every call from paying a full timeout tax while the provider is known to be bad. If the embeddings for routing are down, it is advisable to prefer a default route to ensure the chat still answers. Logging should be loud, but the system should not pretend that routing ran successfully.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.