Urgent.News

What's breaking now, across thousands of outlets.

AI

I stopped paying twice for the same LLM answer (and why embeddings alone can't dedupe prompts)

I kept paying for the same LLM answer twice. An agent retries, a user double-clicks, two workers ask the identical question a second apart, and each one is a fresh model call on the bill. The other thing that kept biting me: when the provider has a bad ten minutes, requests don't fail, they pile up until timeouts cascade through the whole app. So I put a small Rust server between my code and the…

The author was repeatedly paying for identical responses from the same large language model (LLM) due to various factors, such as agents retrying, users double-clicking, or multiple workers asking the same question in quick succession. This issue was exacerbated when the provider experienced performance issues, leading to a backlog of requests that caused timeouts across the entire application.

To address this problem, the author developed a Rust server called tokio-prompt-orchestrator, which can be easily integrated into existing OpenAI or Anthropic clients without any modifications to the client code. By dropping the server in front of the existing client, the same question asked twice would now be answered from the deduplication window instead of triggering an additional model call.

The server provides additional features such as a circuit breaker for handling provider downtime, timeout management, and a spend cap to prevent excessive usage (429 response when the limit is reached). It also includes a dead-letter queue to handle failed requests. The tool supports both OpenAI and Anthropic clients, allowing for a seamless switch with minimal changes to the existing codebase.

Semantic deduplication is another feature of the orchestrator, which involves embedding the prompt and comparing it to recent prompts. However, the author cautions that using embeddings alone is not sufficient for accurate deduplication, as exact matches are only captured. The author recommends further refinement by considering shared numbers and word order in the prompts.

A pure-similarity cache may unintentionally provide an answer to an opposite question, as demonstrated in the example where the similarity score between two paraphrased prompts exceeded 0.99, yet they were considered distinct.

The author stresses the importance of testing the threshold settings with real traffic before lowering them, as an overly aggressive approach may result in unnecessary model calls. They also mention an alternative approach of using the author's own documentation for deduplication, which involves indexing a folder of Markdown and text files with tantivy (BM25) and presenting the most relevant passages to each question.

If the search fails or takes over 2 seconds, the prompt is processed without context, ensuring that requests are never dropped.

The author provides instructions for installing the orchestrator on various platforms, including Linux, Windows, and macOS. They also offer the option to use it as a Rust library, with the source code available on GitHub. Additionally, the author shares a repl example that demonstrates the tool's functionality step by step, inviting feedback from anyone implementing semantic caching in production, specifically inquiring about the threshold settings used and any false hits encountered.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

More from Friday 9 October →